Build evals from real production data
10 July 202610:00 AM PT

Amanda GilbertSolutions Engineer
Production traces capture where your AI falls short and what users are trying to do. Building evals from that data is how you catch failures earlier and make better calls about what ships next.
In this session, Amanda Gilbert shows how to take the patterns Braintrust surfaces automatically, turn them into a labeled eval dataset, and run the same workflow every time a new pattern shows up.
Recent on-demand workshops
How financial services teams ship AI agents to production
26 August 20267:30 AM PT
Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.
Evals foundations
13 August 202610:00 AM PT
Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.
Put the workshop into practice
Start using Braintrust to trace what ships, score it against real data, and turn every regression into an eval you can rerun.