The Agent RaceJuly 3, 2026via The Decoder
GPT and Claude failed Bridgewater's finance tests because the right answers were never public
Why it matters
Closed-model dominance is cracking in specialized domains. When evaluation data is domain-specific and unavailable publicly, fine-tuned open models can outperform frontier models—and that changes ROI calculus for enterprises.
Key signals
- Fine-tuned open-weight model outperformed GPT and Claude on Bridgewater's financial document evaluation
- Performance achieved at a fraction of the cost of frontier models
- Evaluation based on domain-specific financial data not in public training sets
- Bridgewater and Thinking Machines Lab conducted the analysis
- Published July 3, 2026
The hook
GPT and Claude just got beaten at their own game. A fine-tuned open model outperformed them on Bridgewater's finance benchmarks—at a fraction of the cost.
The hedge fund Bridgewater and Thinking Machines Lab report that a finely tuned open-weight model outperforms the most powerful AI models in the evaluation of financial documents, at a fraction of the cost. The figures come from their own analysis.