Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data
Custom benchmarks just became actionable. Optima lets you test models against your own data—and compare on cost and latency, not just token price.

Why it matters
AI benchmarking has relied on published datasets and metrics that don't reflect real-world workloads. Optima shifts evaluation toward practitioner-owned data and task-specific metrics, making model selection decisions sharper and more ROI-driven for enterprise deployments.
The key facts
5 to knowArtificial Analysis launches Optima platform
Custom benchmarks built from user's own data and workflows
Models compared on quality, cost, and time per task
Agent applications benefit most from cost/latency metrics over raw token pricing
Addresses gap between published benchmarks and real-world performance
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks from their own data and workflows. Models can be compared not just on quality but also on cost and time per task. For agent-based applications, those metrics often tell you more than raw token pricing.