Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
Google's Android Bench 2.0 now grades AI agents on long-horizon tasks. Here's what that changes for evaluating real-world autonomy.

Why it matters
Android Bench 2.0 introduces long-horizon task evaluation and agent-based scoring, enabling measurement of multi-step AI development workflows rather than isolated capabilities. This matters because agent reliability in production depends on how well benchmarks capture sequential failure modes and task completion over time.
The key facts
6 to knowLong-horizon tasks (LHTs) enable evaluation of complex, multi-step development workflows
Agent-based evaluation framework replaces single-step task assessment
Continuous scoring mechanism introduced (methodology details not disclosed in source)
Released October 2026
Benchmark focus: Android development tasks and model/agent performance
Update termed 'major' and 'Android Bench 2.0' (version numbering suggests prior 1.x baseline)
Go to the source
InfoQ AI/MLinfoq.com
Publisher excerpt: Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluation, and continuous scoring to better assess performance on complex, multi-step…