Do large language models really need all those layers?
Amazon's research reveals that large language models contain significant redundancy—70% of attention heads and 20% of feed-forward networks can be removed without degrading performance. This challenges assumptions about LLM architecture efficiency and suggests models are undertrained, with major implications for inference cost and model compression.
Why it ranks · · 70% of attention heads can be excised with minimal performance impact · July 2023
Read full story