AgentsThe story, in brief

Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy

Researchers just cracked prompt recovery: proprietary system prompts can now be reconstructed from output alone with near-perfect accuracy—no model weights needed.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A fundamental vulnerability in deployed LLM systems: adversaries can reverse-engineer proprietary prompts (jailbreaks, guardrails, agent instructions) from public outputs, threatening the security model that many enterprises depend on.

The key facts

6 to know
  1. IIT Bombay and Adobe Research developed 'Previous-Token Prediction' method

  2. Reconstructs original prompts from LLM output with near-perfect accuracy

  3. Works without access to model weights

  4. Model-agnostic (works across different models)

  5. Implications for proprietary system prompts and agent instructions

  6. Published August 12, 2026

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy. Their method, called "Previous-Token Prediction," doesn't need access to model weights and works across different models. For…
Read original report
Back to today's editionMore agents news

Keep reading

Related stories

More from Agents