FrontierThe story, in brief

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Apple's new training method lets a single LLM agent learn across wildly different environments without getting stuck on easy wins.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

RISED addresses a real training bottleneck: how to teach generalist agents across diverse interactive environments when some tasks succeed trivially while others fail consistently. The work proposes rubrics for curriculum selection and self-distillation that consider cross-environment rollout relationships, not just local rewards. Relevant to anyone building multi-task agents or evaluating agent training frameworks.

The key facts

11 to know
  1. Apple ML research paper on multi-environment agent training

  2. Problem: existing curriculum strategies allocate at environment level without modeling relationships between rollouts

  3. All-failure and all-success rollouts coexist in batches, leaving data without group-relative reward signals

  4. Solution: rubrics for prompt-group selection and self-distillation across environments

  5. Published October 2026 on Apple's ML research site

  6. Apple Research (machinelearning.apple.com)

  7. Published October 6, 2026

  8. Addresses curriculum/data-selection for multi-environment agent training

  9. Proposes prompt-group selection based on cross-environment rollout relationships

  10. Tackles batch heterogeneity problem: all-failure and all-success groups lacking relative reward signals

  11. No benchmark numbers or deployment outcomes disclosed in abstract

Go to the source

Apple Machine Learningmachinelearning.apple.com

Publisher excerpt: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier