20B-parameter Alexa model sets new marks in few-shot learning
20B parameters. Amazon's Alexa Teacher Model just outperformed major LLMs on few-shot learning tasks.

Why it matters
Amazon is positioning Alexa as a serious competitor in the LLM space with superior performance on summarization and translation tasks, challenging the decoder-only architecture dominance.
The key facts
5 to know20 billion parameters
Encoder-decoder architecture
Superior performance on few-shot summarization
Superior performance on machine translation
Outperforms other large language models
Go to the source
Amazon Scienceamazon.science
Publisher excerpt: With an encoder-decoder architecture — rather than decoder only — the Alexa Teacher Model excels other large language models on few-shot tasks such as summarization and machine translation.