FrontierThe story, in brief

Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data

Custom benchmarks just became actionable. Optima lets you test models against your own data—and compare on cost and latency, not just token price.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

AI benchmarking has relied on published datasets and metrics that don't reflect real-world workloads. Optima shifts evaluation toward practitioner-owned data and task-specific metrics, making model selection decisions sharper and more ROI-driven for enterprise deployments.

The key facts

5 to know
  1. Artificial Analysis launches Optima platform

  2. Custom benchmarks built from user's own data and workflows

  3. Models compared on quality, cost, and time per task

  4. Agent applications benefit most from cost/latency metrics over raw token pricing

  5. Addresses gap between published benchmarks and real-world performance

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks from their own data and workflows. Models can be compared not just on quality but also on cost and time per task. For agent-based applications, those metrics often tell you more than raw token pricing.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier