FrontierThe story, in brief

Multiagent AI for generating chain-of-thought training data

29%. That's the average benchmark improvement Amazon just unlocked using multiagent AI to generate training data.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Amazon Science demonstrates a scalable method for synthetic data generation that could reshape how AI models are trained, reducing dependence on expensive human annotation while improving model performance across multiple benchmarks.

The key facts

4 to know
  1. 29% average performance improvement across benchmarks

  2. Method: Multiagent ensembles generating chain-of-thought annotated interactions

  3. Source: Amazon Science (published July 31, 2025)

  4. Application: Training data generation and refinement

Go to the source

Amazon Scienceamazon.science

Publisher excerpt: Using ensembles of agents to generate and refine interactions annotated with chains of thought improves performance on a battery of benchmarks by an average of 29%.
Read original report
Back to today's editionMore frontier news

The wider picture

View all
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier01

Simulated students that make realistic mistakes help AI tutors learn faster

A novel approach to AI training using realistic synthetic feedback loops is accelerating tutor model development and reducing the cost of evaluation data. This represents a meaningful shift in how frontier labs can iterate on capability without massive labeled datasets.

The Decoder
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier02

Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages

Alibaba's new Qwen3.8-LiveTranslate represents a meaningful advance in multimodal capability (speech-to-speech interpretation with low latency, speaker diarization, and long-context understanding), shipped and available now. It's a model release that changes the frontier benchmark for real-time translation and interpretation.

MarkTechPost
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier03

Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks

A new multimodal model from Alibaba's Qwen achieves near-parity with Google's flagship on audio-video tasks while undercutting its pricing, intensifying competition in the agent-native model race and raising questions about Google's cost positioning.

The Decoder