RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
Apple's new training method lets a single LLM agent learn across wildly different environments without getting stuck on easy wins.

Why it matters
RISED addresses a real training bottleneck: how to teach generalist agents across diverse interactive environments when some tasks succeed trivially while others fail consistently. The work proposes rubrics for curriculum selection and self-distillation that consider cross-environment rollout relationships, not just local rewards. Relevant to anyone building multi-task agents or evaluating agent training frameworks.
The key facts
11 to knowApple ML research paper on multi-environment agent training
Problem: existing curriculum strategies allocate at environment level without modeling relationships between rollouts
All-failure and all-success rollouts coexist in batches, leaving data without group-relative reward signals
Solution: rubrics for prompt-group selection and self-distillation across environments
Published October 2026 on Apple's ML research site
Apple Research (machinelearning.apple.com)
Published October 6, 2026
Addresses curriculum/data-selection for multi-environment agent training
Proposes prompt-group selection based on cross-environment rollout relationships
Tackles batch heterogeneity problem: all-failure and all-success groups lacking relative reward signals
No benchmark numbers or deployment outcomes disclosed in abstract
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without…