Skip to main content

Category · 45 repos

Speech Recognition

Speech-to-text, transcription, diarization and voice typing. Ranked by star velocity over the last 24 hours.

1
debpalash/VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

PythonSpeech RecognitionText to Speech & Voice
2
OpenWhispr/openwhispr

Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.

JavaScriptSpeech Recognition
3
jamiepine/voicebox

The open-source AI voice studio. Clone, dictate, create.

TypeScriptSpeech RecognitionText to Speech & Voice
4
cjpais/Handy

A free, open source, and extensible speech-to-text application that works completely offline.

RustSpeech Recognition
5
Const-me/Whisper

High-performance GPGPU inference of OpenAI's Whisper automatic speech recognition (ASR) model

C++Speech Recognition
6
KnowClip-ai/knowclip

Automatically cut long videos such as livestreams, courses, speeches, and talking-head recordings into highlights | Local speech recognition; audio stays on your device; LLM selects highlights; batch export.

TypeScriptSpeech Recognition
7
bradautomates/claude-video

Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.

PythonSpeech Recognition
8
ZdenekCulik/whisper-pro

Whisper Pro — personal macOS voice-to-text app

SwiftSpeech Recognition
9
fermionresearch/phonon

Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images

PythonSpeech RecognitionLLM Inference & Serving
10
meizhong986/WhisperJAV

ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV

PythonSpeech RecognitionLocal LLMs
11
niedev/RTranslator

Open source real-time translation app for Android that runs locally

JavaSpeech RecognitionMobile & Desktop Apps
12
jbuck95/whisper.nvim

local whisper.cpp voice transcription.

LuaSpeech Recognition
13
k2-fsa/sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket ser

C++Speech RecognitionMobile & Desktop Apps
14
like-attract/video-to-note

A local-first video-to-structured-notes service: Bilibili/Douyin/local videos, platform subtitles + local offline transcription, LLM-generated timestamped notes; available as a desktop app / MCP / Agent Skill.

PythonSpeech RecognitionMCP Servers & Tools
15
openai/whisper

Robust Speech Recognition via Large-Scale Weak Supervision

PythonSpeech Recognition
16
tobi/qmd

mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local

TypeScriptSpeech RecognitionCLI & Developer Tools
17
whisper-money/whisper-money

Understand your personal finances. Forget Excels, try Whisper Money.

PHPSpeech RecognitionProductivity & Notes
18
Eliqi797/open-feishu-recording-bean

Open Feishu Recording Bean: self-hosted recording sync, transcription, and Feishu meeting notes without an official paid AI membership; HarmonyOS/Android/iOS, with user-selectable ASR/LLM.

PythonSpeech RecognitionMobile & Desktop Apps
19
hoobnn/livetranslate-bridge

Real-time voice translation and bilingual subtitles on macOS: translates calls, meetings, and audio from any app, powered by Qwen LiveTranslate · Real-time speech translation & subtitles for macOS

SwiftSpeech RecognitionNLP & Language Models
20
iamjrmh/usersync

Word-accurate .lrc lyric sync. Paste your lyrics, point it at any audio file, get back an enhanced LRC with per-word timestamps. Powered by whisper.cpp + WhisperX (wav2vec2 forced alignment) + Demucs vocal isolation. Prebuilt for NVIDIA CUDA, Vulkan, and CPU. Manual editor for fine-tuning, live prev

C++Speech RecognitionFine-tuning & Training
21
umlx5h/LLPlayer

The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!

C#Speech RecognitionLocal LLMs
22
0xShug0/audio.cpp

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.

C++Speech RecognitionText to Speech & Voice
23
EarningsCall/earningscall-python

The EarningsCall Python library provides convenient access to the EarningsCall API.

PythonSpeech Recognition
24
FluidInference/FluidAudio

Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.

SwiftSpeech RecognitionMobile & Desktop Apps
25
SallyZhou9527/prompter

Windows voice-assisted teleprompter built with Tauri, React, and local ASR.

RustSpeech RecognitionMobile & Desktop Apps
26
handy-computer/transcribe.cpp

ggml speech-to-text inference for 16+ model families

C++Speech RecognitionLocal LLMs
27
minburg/outspoke

Outspoke is a free and open-source privacy focused speech-to-text keyboard for Android.

KotlinSpeech RecognitionMobile & Desktop Apps
28
ardanlabs/bucky

Native Go binding for the Whisper.cpp libraries

GoSpeech Recognition
29
princeniithompson/VoxStream-version-2.

This Android application is meant for voice typing

KotlinSpeech RecognitionMobile & Desktop Apps
30
King-Rafat/Dynamic_Streaming_ASR

Bangla Block-Online ASR with metrics

PythonSpeech RecognitionMonitoring & Observability
31
altic-dev/FluidVoice

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X exclusive model access! 😉 - https://x.com/fluidvoiceapp

SwiftSpeech RecognitionMobile & Desktop Apps
32
hitz-zentroa/whisper-lm

Add n-gram and large language model (LLM) support to Whisper models.

Jupyter NotebookSpeech Recognition
33
sufyanaser/NASVoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

PythonSpeech RecognitionText to Speech & Voice
34
HansonJames/general_digital_human_system

A general digital human system: an intelligent interactive platform based on deep learning and WebRTC, integrating Azure Avatar digital-human rendering, speech recognition and synthesis, natural language processing, and other technologies. It supports real-time conversations, knowledge Q&A, and emotional interaction, delivering smooth rendering above 30 FPS and low-latency responses within 200…

PythonSpeech Recognition3D, Avatars & Digital Humans
35
PersonalJarvis/PersonalJarvis

Your AI assistant, built for the agentic era. Open source and local: talk to it, and it runs your agents, your coding CLIs, your browser and your apps. Windows, macOS, Linux.

PythonSpeech RecognitionAI Agents
36
maddylaneeee/ShengJi

Native local-first macOS speech-to-text, live captions, subtitle editing, offline translation, and on-device Gemma 4 transcript enhancement.

SwiftSpeech RecognitionMobile & Desktop Apps
37
stemdeckapp/stemdeck

Stemdeck is an modern stem extraction platform for musicians,producers and hobbyists, designed to isolate vocals, drums, bass, piano and guitar for practice, transcription, remixing, and creative audio workflows through a modern and interactive interface

JavaScriptSpeech Recognition
38
canxin121/subtitle-gateway

本地 FunASR ASR + 翻译统一网关(OpenAI /v1/audio/transcriptions + ferrum /transcribe + DeepL /v1/translate + LibreTranslate /translate),一键部署 setup.sh/run.sh + systemd

PythonSpeech Recognition
39
ggml-org/whisper.cpp

Port of OpenAI's Whisper model in C/C++

C++Speech Recognition
40
kishanhitk/suniye

Open-source, local-first dictation for macOS. Hold a key, speak, release — text appears at your cursor. No audio leaves your machine.

SwiftSpeech RecognitionMobile & Desktop Apps
41
OpenMOSS/MOSS-Transcribe-Diarize

A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness

PythonSpeech Recognition
42
kalliope-project/kalliope

Kalliope is a framework that will help you to create your own personal assistant.

PythonSpeech RecognitionChatbots & LLM Apps
43
moona3k/macparakeet

Fast, private, local-first voice app for Apple Silicon Macs — dictation, file/media transcription, meeting recording, Transforms, and a public automation CLI. Free and open-source.

SwiftSpeech RecognitionCLI & Developer Tools
44
modelscope/FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

PythonSpeech RecognitionMCP Servers & Tools
45
withdavid-ai/dai-asr-i18n

David AI's ASR I18N Benchmark

PythonSpeech Recognition