Skip to main content

Category Β· 34 repos

Text to Speech & Voice

Speech synthesis, voice cloning, voice conversion and voice assistants. Ranked by star velocity over the last 24 hours.

1
debpalash/VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative β€” voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

PythonText to Speech & VoiceSpeech Recognition
2
harry0703/MoneyPrinterTurbo

Generate HD short videos from a topic or keyword with one click, using large AI models and automated workflows. Generate HD short videos from a topic or keyword with an automated AI workflow.

PythonText to Speech & VoiceVideo Generation & Editing
3
vincentsch/explainroo

Explainer videos and product demos made by your AI agent. Free and open source: a local voice (Kokoro), word timing (Whisper) and a canvas renderer turn a script into a narrated MP4.

JavaScriptText to Speech & VoiceVideo Generation & Editing
4
jamiepine/voicebox

The open-source AI voice studio. Clone, dictate, create.

TypeScriptText to Speech & VoiceLocal LLMs
5
jub0t/Concat

Beta: The truly free, and open-source cross-platform CapCut replacement (supports MCPs).

RustText to Speech & VoiceVideo Generation & Editing
6
livekit/agents

A framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

PythonText to Speech & VoiceAI Agents
7
Amoris202/qwen3-tts-webui

Gradio web UI and Windows one-click launcher for Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice.

PythonText to Speech & VoiceProductivity & Notes
8
SkyNotSilent/insightcut-jianying-image-video

InsightCut β€” an AI image/video and editable Jianying draft workspace: turn scripts into storyboards, images, voiceovers, and subtitles; supports Jianying draft / CapCut draft, MP4, and asset packages.

PythonText to Speech & VoiceVideo Generation & Editing
9
rishiskhare/parrot

A free, offline, private AI text-to-speech desktop app built on Rust 🦜

RustText to Speech & VoiceLocal LLMs
10
AhmadHassan-BTed/B

An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.

PythonText to Speech & VoiceComputer Vision
11
k2-fsa/OmniVoice

High-Quality Voice Cloning TTS for 600+ Languages

PythonText to Speech & Voice
12
fishaudio/fish-speech

SOTA Open Source TTS

PythonText to Speech & Voice
13
QwenLM/Qwen3-TTS

Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.

PythonText to Speech & Voice
14
0xShug0/audio.cpp

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.

C++Text to Speech & VoiceAudio & Music Generation
15
VividDub/VividDub

VividDub β€” product site for public-video voice translation; supports dubbing, subtitles, localization, and hard-subtitle removal

Text to Speech & VoiceNLP & Language Models
16
netease-youdao/EmotiVoice

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

PythonText to Speech & VoiceMachine Learning & Data Science
17
netease-youdao/Confucius4-TTS

Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine

PythonText to Speech & VoiceMachine Learning & Data Science
18
FluidInference/FluidAudio

Frontier CoreML audio models in your apps β€” text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.

SwiftText to Speech & VoiceSpeech Recognition
19
multimodal-art-projection/YuE

YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.

PythonText to Speech & VoiceAudio & Music Generation
20
LuwuDynamics/rig_omni

Open-source ESP32-S3 robot firmware supporting multiple forms (Puppy & Hover). Features voice AI, EAF emotion animation, IMU balance control, MCP remote tools, OTA updates, and a 240Γ—240 round LCD.

C++Text to Speech & VoiceSystems, Compilers & Build Tools
21
digitalsamba/claude-code-video-toolkit

AI-native video production toolkit for Claude Code

PythonText to Speech & VoiceVideo Generation & Editing
22
Kugelaudio/kugelaudio-open

Open-source text-to-speech for European languages with voice cloning

PythonText to Speech & Voice
23
xzf-thu/VoiceMem

Infrastructure for the next generation of voice agents, designed to provide universal memory. It is divided into a left brain and a right brain, storing information and emotions respectively, while a fully streaming architecture eliminates latency at the fundamental level.

PythonText to Speech & VoiceAI Agents
24
neural-maze/realtime-phone-agents-course

Build realtime AI voice agents using FastRTC for low-latency streaming, Superlinked for vector search, Twilio for live phone calls, and Runpod for scalable GPU deployment.

PythonText to Speech & VoiceVector Databases
25
sufyanaser/NASVoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative β€” voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

PythonText to Speech & VoiceSpeech Recognition
26
OpenBMB/VoxCPM

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

PythonText to Speech & VoiceMachine Learning & Data Science
27
PersonalJarvis/PersonalJarvis

Your AI assistant, built for the agentic era. Open source and local: talk to it, and it runs your agents, your coding CLIs, your browser and your apps. Windows, macOS, Linux.

PythonText to Speech & VoiceAI Agents
28
cosin2077/easyVoice

An open-source text-to-speech tool that supports very long text and multiple voice roles.

TypeScriptText to Speech & Voice
29
readest/readest

Readest is a modern, feature-rich ebook reader designed for avid readers offering seamless cross-platform access, powerful tools, and an intuitive interface to elevate your reading experience.

TypeScriptText to Speech & VoiceOCR & Document AI
30
ysharma3501/LuxTTS

A high-quality rapid TTS voice cloning model that reaches speeds of 150x realtime.

PythonText to Speech & Voice
31
kalliope-project/kalliope

Kalliope is a framework that will help you to create your own personal assistant.

PythonText to Speech & VoiceSpeech Recognition
32
microsoft/VibeVoice

Open-Source Frontier Voice AI

PythonText to Speech & Voice
33
dramaclaw/dramaclaw

A general-purpose AIGC video engine that takes scripts through a single pipeline to finished films, including animated dramas, ads, ecommerce videos, otome games, and more.

PythonText to Speech & VoiceVideo Generation & Editing
34
jxlarrea/kiosk-satellite

Turn any Android device into a beautiful, dedicated Home Assistant kiosk. Purpose-built for Home Assistant from the ground up.

DartText to Speech & VoiceMobile & Desktop Apps