Skip to main content

Tag · 7 repos

asr

Repositories carrying the asr tag.

Exact tag match7 sample repositories
807
fermionresearch/phonon

Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images

PythonSpeech RecognitionLLM Inference & Serving
1092
k2-fsa/sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket ser

C++Mobile & Desktop AppsSpeech Recognition
1497
umlx5h/LLPlayer

The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!

C#Speech RecognitionLocal LLMs
2047
0xShug0/audio.cpp

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.

C++Text to Speech & VoiceAudio & Music Generation
2554
FluidInference/FluidAudio

Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.

SwiftSpeech RecognitionMobile & Desktop Apps
2660
handy-computer/transcribe.cpp

ggml speech-to-text inference for 16+ model families

C++Speech RecognitionLocal LLMs
5031
modelscope/FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

PythonSpeech RecognitionMCP Servers & Tools