FrontierThe story, in brief

Google updates Android Bench with new LLMs, but Gemini still lags behind - Ars Technica

Google just changed how it grades AI models for Android coding—and Gemini still doesn't make the cut.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Google's Android Bench update introduces new evaluation methodology and expanded model coverage, but reveals a competitive gap where Gemini underperforms against rival LLMs on real-world Android development tasks. This signals pressure on Google to close capability gaps in a high-stakes developer tool category.

The key facts

5 to know
  1. Android Bench upgraded to Harbor Framework

  2. 8 new LLM models added to leaderboard

  3. Gemini underperforms vs. competing models on Android coding benchmarks

  4. New evaluation methodology for measuring LLM performance on Android-specific tasks

  5. Benchmark now measures real Android development work, not generic capability

Go to the source

Reuters Technologynews.google.com

Publisher excerpt: Google updates Android Bench with new LLMs, but Gemini still lags behind Ars Technica Google just changed how it grades the AI models you use for Android coding Digital Trends Android Bench Upgraded to Harbor Framework, 8 Models Added to Leaderboard Droid Life Evolving how LLMs are measured for…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier