WorkThe story, in brief

500 investment bankers review AI outputs and find none ready for client delivery

0 out of 500. Investment bankers tested GPT-5.4 and Claude Opus 4.6 on real workflows—none passed client-ready standards.

Paper-cut illustration of amber paths carrying capital toward a small coral research venture between larger buildings.
Capital and the next generation of AI ventures.AI illustration by KeyNews
The KeyNews take

Why it matters

Real-world benchmark data reveals a critical gap between headline model capabilities and enterprise deployment readiness. This challenges the narrative that frontier models are ready for high-stakes professional workflows, and signals what risk tolerance looks like in regulated industries.

The key facts

6 to know
  1. 500 investment bankers benchmarked GPT-5.4 and Claude Opus 4.6

  2. 0% of outputs rated client-ready

  3. Failures attributed to imprecision and factual errors

  4. 51% of bankers would use outputs as starting point (implicit trust gap)

  5. Real-world professional task evaluation (not academic benchmark)

  6. Published April 26, 2026

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: A new benchmark puts top models like GPT-5.4 and Claude Opus 4.6 to work on the kinds of tasks junior investment bankers handle every day. Not a single AI output was rated ready to send to a client; the results are too imprecise or flat-out wrong. Still, more than half of the bankers said they'd…
Read original report
Back to today's editionMore work news

The wider picture

View all
Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI illustration by KeyNews
Work01

The Emerging M&A Map For AI Agent Security

As agents move from pilots to production with real system access, enterprise security models are breaking. The M&A map is forming around who controls agent permissions, monitoring, and governance — a new class of identity management problem that practitioners need to architect for now.

Crunchbase News
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work02

AI privacy budgets: Ask for the calculation, not the claim

Enterprise AI buyers are accepting privacy budget numbers without verification. This deep dive explains what questions to ask vendors about federated learning privacy claims, and why the gap between contractual promises and operational evidence is where real exposure lives.

CIO
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work03

Andrew Kelley Interview: Why He Built Zig, Banned AI Contributions, and Moved Zig off GitHub

Open-source governance is shifting in response to AI-generated contributions. Zig's formal ban and migration off GitHub signals broader industry concern about code quality, maintainer burden, and the cultural impact of automated submissions — a flashpoint for how AI changes the work of software development.

InfoQ AI/ML