Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
Google open-sourced RRSI: agents that rewrite their own prompts and tools without touching model weights. Terminal-Bench jumped 6 points; all held-out test splits improved.

Why it matters
A practitioner-relevant framework for agent self-improvement through harness optimization (prompts, tools, memory) with built-in safeguards against overfitting. Concrete gains on a standard benchmark suggest deployable technique for enterprises running agents on frozen models.
The key facts
7 to knowFramework: RRSI (open-source, Google Cloud AI Research)
Mechanism: agents rewrite prompts, tools, memory; model weights frozen
Safeguards: leakage critic, noise floor, cost rule, pruning
Benchmark: Terminal-Bench 2.1, Claude Opus 4.8
Baseline: 74.2% → 80.2% (+6 percentage points)
Generalization: all 6 held-out test splits improved
Status: open-source
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Google Cloud AI Research has open-sourced RRSI, a framework that lets LLM agents rewrite their own prompts, tools and memory while model weights stay frozen. It adds a leakage critic, a noise floor, a cost rule and pruning so gains carry over to new tasks. With Claude Opus 4.8, Terminal-Bench 2.1…