ChipsThe story, in brief

Sakana AI and NVIDIA Introduce TwELL with CUDA Kernels for 20.5% Inference and 21.9% Training Speedup in LLMs

20.5% faster inference. That's what Sakana AI and NVIDIA just unlocked with sparse CUDA kernels—turning 99% sparsity into real GPU throughput.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Sakana and NVIDIA demonstrate that inducing extreme sparsity (99%+) in LLM feedforward layers with L1 regularization translates into measurable inference and training speedups via optimized CUDA kernels. This is a practical infrastructure win for cost-conscious deployment at scale.

The key facts

7 to know
  1. 20.5% inference speedup achieved

  2. 21.9% training speedup achieved

  3. Over 99% sparsity induced in feedforward layers via L1 regularization

  4. Negligible downstream performance impact reported

  5. New sparse data formats and fused CUDA kernels developed

  6. Sakana AI and NVIDIA co-authored research

  7. Published May 11, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Sakana AI and NVIDIA Researchers demonstrate that simple L1 regularization can induce over 99% sparsity in feedforward layers with negligible downstream performance impact, and translate that sparsity into real GPU throughput gains using new sparse data formats and fused CUDA kernels.
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips