Evals foundations
13 August 202610:00 AM PT

Jess WangDeveloper Advocate
Agents fail differently than traditional software. They drift and regress silently, so you need a new kind of observability to monitor and improve quality.
This workshop will cover the core foundations of active observability, evals, and how to get started with Braintrust.
What you’ll learn
- How Braintrust groups production traces into named failure patterns you can act on
- How to filter a failure cluster into a labeled dataset
- How to write an eval that targets a specific failure pattern and validate the fix
- How to run a repeatable diagnose-to-eval workflow in Braintrust
Recent on-demand workshops
How financial services teams ship AI agents to production
26 August 20267:30 AM PT
Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.
Build evals from real production data
10 July 202610:00 AM PT
Learn how to take the patterns Braintrust surfaces automatically, turn them into a labeled eval dataset, and run the same workflow every time a new pattern shows up.
Put the workshop into practice
Start using Braintrust to trace what ships, score it against real data, and turn every regression into an eval you can rerun.