ToolsThe story, in brief

Evaluating Deep Agents using LangSmith on AWS

Deep agents hitting production. Here's how to evaluate them before they break.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

AWS and LangChain are publishing practical tooling for evaluating agentic AI systems in production—a critical gap as enterprises move agents from research to real workflows. This signals growing maturity in the agent deployment lifecycle.

The key facts

12 to know
  1. Five evaluation patterns for deep agents documented

  2. Offline evaluation using pytest and LangSmith

  3. Online monitoring for production agents

  4. Text-to-SQL deep agent walkthrough on Amazon Bedrock

  5. Full dev-to-production lifecycle guidance

  6. AWS, LangChain, and Anthropic collaboration on evals framework

  7. Offline evaluation approach using pytest and LangSmith

  8. Online production monitoring configuration included

  9. Text-to-SQL deep agent use case demonstrated

  10. Amazon Bedrock integration for inference

  11. Full development-to-production lifecycle covered

  12. Combines learnings from LangChain and Anthropic

Go to the source

AWS Machine Learning Blogaws.amazon.com

Publisher excerpt: This post combines learnings from LangChain’s work on evaluating deep agents and Anthropic’s guide to demystifying evals for AI agents into a practical guide. In this post, you will learn how to: 1) apply five evaluation patterns for deep agents, 2) build offline evaluations using pytest and…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools