Sakana AI and NVIDIA Introduce TwELL with CUDA Kernels for 20.5% Inference and 21.9% Training Speedup in LLMs
20.5% faster inference. That's what Sakana AI and NVIDIA just unlocked with sparse CUDA kernels—turning 99% sparsity into real GPU throughput.

Why it matters
Sakana and NVIDIA demonstrate that inducing extreme sparsity (99%+) in LLM feedforward layers with L1 regularization translates into measurable inference and training speedups via optimized CUDA kernels. This is a practical infrastructure win for cost-conscious deployment at scale.
The key facts
7 to know20.5% inference speedup achieved
21.9% training speedup achieved
Over 99% sparsity induced in feedforward layers via L1 regularization
Negligible downstream performance impact reported
New sparse data formats and fused CUDA kernels developed
Sakana AI and NVIDIA co-authored research
Published May 11, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Sakana AI and NVIDIA Researchers demonstrate that simple L1 regularization can induce over 99% sparsity in feedforward layers with negligible downstream performance impact, and translate that sparsity into real GPU throughput gains using new sparse data formats and fused CUDA kernels.