FrontierThe story, in brief

There are no lossless transformations of natural-language text

A fundamental limit on language models: transformations always lose information. Here's what it means for reasoning and context windows.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Research reveals that no transformation of natural language can be truly lossless—a finding with direct implications for how LLMs process, compress, and reason over text. This reshapes thinking about tokenization, context windows, and the theoretical limits of language model capabilities.

The key facts

8 to know
  1. No lossless transformation of natural-language text exists

  2. Implication for tokenization strategies and context compression

  3. Affects reasoning chains and information preservation in long contexts

  4. Published by Simon Willison (Datasette creator, noted AI researcher)

  5. Core claim: all transformations of natural-language text lose information

  6. Implications for tokenization, embeddings, and model training

  7. Challenges the assumption that any encoding scheme is lossless

  8. Published August 11, 2026 on Simon Willison's blog (established AI thought leader)

Go to the source

Simon Willisonsimonwillison.net

Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier