All cursus
AIadvanced
How Voice Models Work: Tokens, Recognition, and Synthesis
One minute of talking is about 195 tokens as text and about 36,000 as audio. That ratio explains almost every design decision in speech systems, and this path starts there. How neural codecs turn a waveform into something a transformer can model, why speech recognition has three architectures rather than one and which of them can stream, how treating audio as tokens made cloning a voice from three seconds possible, and what the standard speech-to-text-to-speech pipeline throws away that it can never get back.
0 of 4 lessons complete
Sign in to track progress and earn a certificate.

