Workshops

Evals foundations: Answering questions about your agent

4 November 202610:00 AM PT

Jess Wang, Developer Advocate

Jess WangDeveloper Advocate

A useful eval starts with a specific question about your agent. The dataset, scorers, and experiment design should help you answer that question, and the results should indicate ways to improve your agent's performance.

In this live workshop, Jess will share example evals developed at Braintrust for three different use cases. Testing an agent workflow, choosing a model judge, and comparing models.

Leave with a practical framework for designing an eval for your own agent, plus examples and skills to help you get started in Braintrust.

What you’ll learn

  • How to turn a question about your agent into an eval
  • How to build datasets that represent the behaviors and failure modes you care about
  • When to use deterministic scorers, model judges, or both
  • How to eval agent workflows for correctness, safety, and reliability
  • How to compare models against meaningful measures of quality
  • How to interpret results across accuracy, cost, and latency

Recent on-demand workshops

From reactive debugging to active observability

16 September 202610:00 AM PT

See how Braintrust’s latest release surfaces and investigates recurring production issues, then guides your team toward improvement.

Evals foundations

9 September 202610:00 AM PT

Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.

How financial services teams ship AI agents to production

26 August 20267:30 AM PT

A practical framework for moving AI agents and applications from pilot to production in regulated environments.

Share

Put the workshop into practice

Start using Braintrust to trace what ships, score it against real data, and turn every regression into an eval you can rerun.

Trace everything