Interspeech: Where speech recognition and synthesis converge
Speech recognition and synthesis are converging—and it's reshaping how Amazon builds AI models.

Why it matters
Amazon scientist reveals how large language models and text-to-speech are sharing architectures, signaling a fundamental shift in how speech AI is engineered and could impact enterprise voice applications.
The key facts
10 to knowAmazon senior principal scientist Jasha Droppo on convergence of speech recognition and synthesis
Shared architectures between LLMs and spectrum quantization text-to-speech models
Focus on architectural convergence between two historically separate fields
Published on Amazon Science blog (August 2023)
Source: Amazon Science (credible internal research)
Speaker: Jasha Droppo, Senior Principal Scientist
Topic focus: Shared architectures between LLMs and spectrum quantization TTS models
Conference: Interspeech (academic/industry speech processing conference)
Published: August 2023
Cross-field convergence: speech recognition and synthesis architectural alignment
Go to the source
Amazon Scienceamazon.science
Publisher excerpt: Senior principal scientist Jasha Droppo on the shared architectures of large language models and spectrum quantization text-to-speech models — and other convergences between the two fields.
