500 investment bankers review AI outputs and find none ready for client delivery
0 out of 500. Investment bankers tested GPT-5.4 and Claude Opus 4.6 on real workflows—none passed client-ready standards.

Why it matters
Real-world benchmark data reveals a critical gap between headline model capabilities and enterprise deployment readiness. This challenges the narrative that frontier models are ready for high-stakes professional workflows, and signals what risk tolerance looks like in regulated industries.
The key facts
6 to know500 investment bankers benchmarked GPT-5.4 and Claude Opus 4.6
0% of outputs rated client-ready
Failures attributed to imprecision and factual errors
51% of bankers would use outputs as starting point (implicit trust gap)
Real-world professional task evaluation (not academic benchmark)
Published April 26, 2026
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: A new benchmark puts top models like GPT-5.4 and Claude Opus 4.6 to work on the kinds of tasks junior investment bankers handle every day. Not a single AI output was rated ready to send to a client; the results are too imprecise or flat-out wrong. Still, more than half of the bankers said they'd…

