Category Β· 34 repos
Text to Speech & Voice
Speech synthesis, voice cloning, voice conversion and voice assistants. Ranked by star velocity over the last 24 hours.
VoiceStudio is the open-source, fully-local ElevenLabs alternative β voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Generate HD short videos from a topic or keyword with one click, using large AI models and automated workflows. Generate HD short videos from a topic or keyword with an automated AI workflow.
Explainer videos and product demos made by your AI agent. Free and open source: a local voice (Kokoro), word timing (Whisper) and a canvas renderer turn a script into a narrated MP4.
The open-source AI voice studio. Clone, dictate, create.
Beta: The truly free, and open-source cross-platform CapCut replacement (supports MCPs).
A framework for building realtime voice AI agents π€ποΈπΉ
Gradio web UI and Windows one-click launcher for Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice.
InsightCut β an AI image/video and editable Jianying draft workspace: turn scripts into storyboards, images, voiceovers, and subtitles; supports Jianying draft / CapCut draft, MP4, and asset packages.
A free, offline, private AI text-to-speech desktop app built on Rust π¦
An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.
High-Quality Voice Cloning TTS for 600+ Languages
SOTA Open Source TTS
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
VividDub β product site for public-video voice translation; supports dubbing, subtitles, localization, and hard-subtitle removal
EmotiVoice π: a Multi-Voice and Prompt-Controlled TTS Engine
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
Frontier CoreML audio models in your apps β text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
Open-source ESP32-S3 robot firmware supporting multiple forms (Puppy & Hover). Features voice AI, EAF emotion animation, IMU balance control, MCP remote tools, OTA updates, and a 240Γ240 round LCD.
AI-native video production toolkit for Claude Code
Open-source text-to-speech for European languages with voice cloning
Infrastructure for the next generation of voice agents, designed to provide universal memory. It is divided into a left brain and a right brain, storing information and emotions respectively, while a fully streaming architecture eliminates latency at the fundamental level.
Build realtime AI voice agents using FastRTC for low-latency streaming, Superlinked for vector search, Twilio for live phone calls, and Runpod for scalable GPU deployment.
VoiceStudio is the open-source, fully-local ElevenLabs alternative β voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Your AI assistant, built for the agentic era. Open source and local: talk to it, and it runs your agents, your coding CLIs, your browser and your apps. Windows, macOS, Linux.
An open-source text-to-speech tool that supports very long text and multiple voice roles.
Readest is a modern, feature-rich ebook reader designed for avid readers offering seamless cross-platform access, powerful tools, and an intuitive interface to elevate your reading experience.
A high-quality rapid TTS voice cloning model that reaches speeds of 150x realtime.
Kalliope is a framework that will help you to create your own personal assistant.
Open-Source Frontier Voice AI
A general-purpose AIGC video engine that takes scripts through a single pipeline to finished films, including animated dramas, ads, ecommerce videos, otome games, and more.
Turn any Android device into a beautiful, dedicated Home Assistant kiosk. Purpose-built for Home Assistant from the ground up.