AI newsThe story, in brief

Are better models better?

Your AI model is 'better.' But better at what? The gap between model improvements and real-world answers is wider than you think.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

As AI models improve incrementally each week, leaders are confusing capability gains with business value. This piece challenges the assumption that better model performance translates to better outcomes, especially for factual/deterministic questions—a critical blind spot for enterprises betting on AI ROI.

The key facts

3 to know
  1. Weekly model improvements are not translating to better answers on deterministic questions

  2. Models struggle with 'right answers' vs. 'better answers' distinction

  3. Implications for enterprise AI deployment and expectation-setting

Go to the source

Benedict Evansben-evans.com

Publisher excerpt: Every week there’s a better AI model that gives better answers. But a lot of questions don’t have better answers, only ‘right’ answers, and these models can’t do that. So what does ‘better’ mean, how do we manage these things, and should we change what we expect from computers?
Read original report
Back to today's editionMore AI news

The wider picture

View all
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier01

Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters

A capable open-weight image model at 7B parameters challenges the closed-model dominance in generation and editing, expanding practitioner options for on-device and cost-efficient image workflows.

The Decoder
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Frontier02

Tencent's Gander aims to keep talking while it works in the background

A novel architecture for multimodal agents that separates conversational continuity from task execution. Demonstrates a real capability tradeoff: smoother UX vs. task reliability. Relevant to how frontier labs are rethinking agent design.

The Decoder
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work03

Trump now says he wants to form an ‘AI Force’

A major political signal on AI governance: the administration is positioning itself to accelerate rather than constrain AI development, with formal institutional backing (czar + task force). Practitioners and policy-watchers need to know the regulatory stance is shifting toward facilitation.

The Verge AI