audio-ai-field-guide
A mental-model-first field guide to Audio AI — from raw waveforms to production voice systems. Covers ASR, TTS, audio classification, Transformer architectures (CTC/Seq2Seq), fine-tuning, and real-world pipelines like voice assistants and meeting transcription. Bilingual (EN/FA).
- asr
- audio-processing
- deep-learning
- huggingface
- meachine-learning
- speech-recognition
- speech-to-text
- transformers
- tts
- whisper
- audio-classification
- learning-notes
- mental-models
- nlp
- persian
- Stars
- 139
- Forks
- 2
- + today
- +1
- Created
- 2mo
Ranking data as of October 5, 2026 (UTC).
Star History
Today, hour by hour
01
Overview
A mental-model-first field guide to Audio AI — from raw waveforms to production voice systems. Covers ASR, TTS, audio classification, Transformer architectures (CTC/Seq2Seq), fine-tuning, and real-world pipelines like voice assistants and meeting transcription. Bilingual (EN/FA). It ranks #829 on GitTiger, gaining +1 star on October 5, 2026 (UTC).
It was created 2mo ago.
Installation
git clone https://github.com/Ali-hey-0/audio-ai-field-guide.git
cd audio-ai-field-guide
# see README for setup