Google updates Android Bench with new LLMs, but Gemini still lags behind - Ars Technica
Google just changed how it grades AI models for Android coding—and Gemini still doesn't make the cut.

Why it matters
Google's Android Bench update introduces new evaluation methodology and expanded model coverage, but reveals a competitive gap where Gemini underperforms against rival LLMs on real-world Android development tasks. This signals pressure on Google to close capability gaps in a high-stakes developer tool category.
The key facts
5 to knowAndroid Bench upgraded to Harbor Framework
8 new LLM models added to leaderboard
Gemini underperforms vs. competing models on Android coding benchmarks
New evaluation methodology for measuring LLM performance on Android-specific tasks
Benchmark now measures real Android development work, not generic capability
Go to the source
Reuters Technologynews.google.com
Publisher excerpt: Google updates Android Bench with new LLMs, but Gemini still lags behind Ars Technica Google just changed how it grades the AI models you use for Android coding Digital Trends Android Bench Upgraded to Harbor Framework, 8 Models Added to Leaderboard Droid Life Evolving how LLMs are measured for…