FrontierThe story, in brief

Do large language models really need all those layers?

70% of attention heads are dead weight. Amazon Science just proved LLMs are massively over-engineered.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Amazon's research reveals that large language models contain significant redundancy—70% of attention heads and 20% of feed-forward networks can be removed without degrading performance. This challenges assumptions about LLM architecture efficiency and suggests models are undertrained, with major implications for inference cost and model compression.

The key facts

5 to know
  1. 70% of attention heads can be excised with minimal performance impact

  2. 20% of feed-forward networks can be removed without effect on in-context learning

  3. Suggests LLMs are undertrained and over-parameterized

  4. Published by Amazon Science

  5. Research focus: architectural redundancy in transformers

Go to the source

Amazon Scienceamazon.science

Publisher excerpt: Finding that 70% of attention heads and 20% of feed-forward networks can be excised with minimal effect on in-context learning suggests that large language models are undertrained.
Read original report
Back to today's editionMore frontier news

The wider picture

View all
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier01

OpenAI forms math advisory group as its AI resolves more than 100 open problems

OpenAI is demonstrating measurable AI capability breakthrough in mathematical problem-solving at scale, signaling a new frontier in reasoning and symbolic work. The advisory structure reveals the lab's confidence in autonomy and speed over external governance.

TechCrunch AI
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier02

Advancing AI for biology: Teaching models to design and characterize antibodies

AI capability in a high-value domain (drug discovery) is moving from theory to experimental validation. Practitioners building biotech AI tooling need to know what binding-prediction models are now reliable enough to trust in design loops.

Amazon Science
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier03

xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6

xAI's new model release underperforms Claude and GPT-6 on published benchmarks, but aggressive pricing could reshape how practitioners evaluate the capability-cost tradeoff in the lab race. A clear signal of competitive positioning and the emergence of a two-tier frontier.

The Decoder
Do large language models really need all those layers? | KeyNews.AI