Category · 45 repos
Speech Recognition
Speech-to-text, transcription, diarization and voice typing. Ranked by star velocity over the last 24 hours.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.
The open-source AI voice studio. Clone, dictate, create.
A free, open source, and extensible speech-to-text application that works completely offline.
High-performance GPGPU inference of OpenAI's Whisper automatic speech recognition (ASR) model
Automatically cut long videos such as livestreams, courses, speeches, and talking-head recordings into highlights | Local speech recognition; audio stays on your device; LLM selects highlights; batch export.
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
Whisper Pro — personal macOS voice-to-text app
Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
Open source real-time translation app for Android that runs locally
local whisper.cpp voice transcription.
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket ser
A local-first video-to-structured-notes service: Bilibili/Douyin/local videos, platform subtitles + local offline transcription, LLM-generated timestamped notes; available as a desktop app / MCP / Agent Skill.
Robust Speech Recognition via Large-Scale Weak Supervision
mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local
Understand your personal finances. Forget Excels, try Whisper Money.
Open Feishu Recording Bean: self-hosted recording sync, transcription, and Feishu meeting notes without an official paid AI membership; HarmonyOS/Android/iOS, with user-selectable ASR/LLM.
Real-time voice translation and bilingual subtitles on macOS: translates calls, meetings, and audio from any app, powered by Qwen LiveTranslate · Real-time speech translation & subtitles for macOS
Word-accurate .lrc lyric sync. Paste your lyrics, point it at any audio file, get back an enhanced LRC with per-word timestamps. Powered by whisper.cpp + WhisperX (wav2vec2 forced alignment) + Demucs vocal isolation. Prebuilt for NVIDIA CUDA, Vulkan, and CPU. Manual editor for fine-tuning, live prev
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
The EarningsCall Python library provides convenient access to the EarningsCall API.
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Windows voice-assisted teleprompter built with Tauri, React, and local ASR.
ggml speech-to-text inference for 16+ model families
Outspoke is a free and open-source privacy focused speech-to-text keyboard for Android.
Native Go binding for the Whisper.cpp libraries
This Android application is meant for voice typing
Bangla Block-Online ASR with metrics
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X exclusive model access! 😉 - https://x.com/fluidvoiceapp
Add n-gram and large language model (LLM) support to Whisper models.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
A general digital human system: an intelligent interactive platform based on deep learning and WebRTC, integrating Azure Avatar digital-human rendering, speech recognition and synthesis, natural language processing, and other technologies. It supports real-time conversations, knowledge Q&A, and emotional interaction, delivering smooth rendering above 30 FPS and low-latency responses within 200…
Your AI assistant, built for the agentic era. Open source and local: talk to it, and it runs your agents, your coding CLIs, your browser and your apps. Windows, macOS, Linux.
Native local-first macOS speech-to-text, live captions, subtitle editing, offline translation, and on-device Gemma 4 transcript enhancement.
Stemdeck is an modern stem extraction platform for musicians,producers and hobbyists, designed to isolate vocals, drums, bass, piano and guitar for practice, transcription, remixing, and creative audio workflows through a modern and interactive interface
本地 FunASR ASR + 翻译统一网关(OpenAI /v1/audio/transcriptions + ferrum /transcribe + DeepL /v1/translate + LibreTranslate /translate),一键部署 setup.sh/run.sh + systemd
Port of OpenAI's Whisper model in C/C++
Open-source, local-first dictation for macOS. Hold a key, speak, release — text appears at your cursor. No audio leaves your machine.
A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness
Kalliope is a framework that will help you to create your own personal assistant.
Fast, private, local-first voice app for Apple Silicon Macs — dictation, file/media transcription, meeting recording, Transforms, and a public automation CLI. Free and open-source.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
David AI's ASR I18N Benchmark