Generative modeling with sparse transformers
OpenAI just processed sequences 30x longer. Here's why that changes everything.

Why it matters
Sparse Transformer represents a fundamental algorithmic breakthrough in attention mechanisms, enabling sequence modeling at previously impossible scales across text, images, and audio—a capability leap that redefined what large-scale neural networks could process.
The key facts
5 to know30x longer sequence length capability vs. prior approaches
Multimodal capability: text, images, sound
Algorithmic improvement to attention mechanism
Published April 2019 by OpenAI
Sets new prediction benchmarks for sequence modeling
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We’ve developed the Sparse Transformer, a deep neural network which sets new records at predicting what comes next in a sequence—whether text, images, or sound. It uses an algorithmic improvement of the attention mechanism to extract patterns from sequences 30x longer than possible previously.