Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products
Elastic built a production eval framework that bridges data science and TypeScript—here's how they catch agent regressions in RAG and security workflows without siloing.

Why it matters
A practitioner-focused walkthrough of production-grade agent evaluation architecture: moving from ad-hoc LLM-as-a-judge to deterministic rule-based hybrid evals, deep tracing for regression detection, and preserving domain context across complex multi-step agentic workloads.
The key facts
11 to knowElastic transitioned from siloed, ad-hoc evaluations to unified framework
Hybrid approach: LLM-as-a-judge balances with deterministic rules
Bridges Python data science evals with TypeScript production code
Deep tracing for regression detection across RAG and cybersecurity workloads
Focus on domain context preservation in complex agentic systems
Presentation format; no independent outcome data disclosed
Elastic transitioned from siloed, ad-hoc AI agent evaluations to unified framework
Balanced LLM-as-a-judge with deterministic rules for evaluation
Bridged Python data science evals with TypeScript production code
Implemented deep tracing to catch regressions across complex RAG and cybersecurity workloads
Framework preserves domain context in evaluation
Go to the source
InfoQ AI/MLinfoq.com
Publisher excerpt: Susan Chang explains how Elastic transitioned from siloed, ad-hoc AI agent evaluations to a unified, production-grade framework. She discusses balancing LLM-as-a-judge with deterministic rules, bridging Python data science evals with TypeScript production code, and implementing deep tracing to…