Braintrust recently launched Patterns and Debugger, along with an enhanced Loop and Topics experience, to create one connected place for instrumenting, investigating, and measuring agents, enhanced with intelligence.
Braintrust's Hossein Niazmandi and Rahil Sondhi held an online workshop about these new releases. They covered how to connect production traces, evals, and iteration in a continuous feedback process, how Loop and Patterns help teams identify issues with their agents, shared customer stories from teams putting active observability into practice, and how to get started.
They also answered questions from the audience.
Signal can arrive from agents in three ways. First, the evals you define, which execute on every trace and look for the specific failures you told them to look for. Second, inference that runs across your traces to classify them more holistically, which is how you catch the unknown unknowns. That's Topics.
Third, an automation that inspects your traces, does deep analysis, and surfaces the failures with supporting traces. That's Patterns.
From there, you can set up Loop automations to push those findings to a webhook or to Slack, and you can always ask Loop directly to find issues in your traces.
Not necessarily. If all you've done is instrument your traces into Braintrust, with no evaluators at all, Patterns can still uncover failures automatically and then walk you through building evals from what it found.
No. Topics is one optional input to Patterns, not a prerequisite. Plenty of customers get a lot out of Patterns without using Topics. A common workflow is to find failures with Patterns, then use Loop to generate new facets for Topics so you start measuring what Patterns uncovered.
Patterns samples traces in the window you configure. Once it finds something interesting in that sample, it runs a larger investigation across a broader set of traces to establish whether it's a recurring theme.
Patterns is designed to give you the original insight. Then you can build an eval from that failure to determine the true blast radius. One customer found a bug affecting eight traces out of five million this way, starting from Patterns and then sizing the impact across all their traces.
Yes, through the open source Agent Behavior Spec, which we built in partnership with Basis. You define what your agent is supposed to do, the trajectory it should take, and the outcomes you expect, then pass that spec to Patterns so its output is tailored to your agent.
Not at all. Take an example of a customer service chatbot whose job is to process refunds and answer product questions. Defining that trajectory and those expected outcomes tells Patterns what the agent is for. You can set designated outcomes without giving up flexibility.
This is very common with our customers. Under Settings → Automations there's a default pattern discovery that runs generically across all traces, and you can define your own additional pattern automations on any criteria you like, including editing the instructions if your trace structure is unusual or if several agents log to the same project.
To scope a run to one agent, name that agent or its associated metadata ID in the pattern instructions and it will filter down. We're also working on an agents abstraction that sits one level below a project, so you can analyze per agent directly.
Nothing on the Braintrust side for now, since it is in public preview. Patterns currently runs on OpenAI models on a bring-your-own-key basis, so you connect your own key and pay only for inference when Loop runs the analysis. We encourage you to try it out while it's free and share your feedback.
Suggested fixes come primarily from investigating failing traces against traces that don't show the failure. We plan to let you connect external systems in the near future, so fixes can carry context from outside your tracing data.
Topics clusters run in real time, and every trace that lands is clustered against the available clusters. Because your agent changes over time, the cluster maps are regenerated every evening, and you can trigger a regeneration ad hoc. That nightly pass looks across the corpus for clusters that didn't exist the day before.
This is where Topics, Patterns, and Debugger work together. Topics surfaces the high level failure, then Patterns runs the deep analysis that gets you to that level of specificity. And if you want to investigate a single trace, run Debugger on it to find out its specific failure mode.
Either shape works. Topics classifies one trace at a time by default, which can fragment a multi-turn conversation. Set Scope to Group and point Group by at a field like metadata.session_id, and Topics classifies the whole conversation as one unit.
Evals can be scoped to an individual span, such as one failing tool execution, to the whole trace, or to the session. That last one allows for a long multi-turn conversation to be evaled end to end rather than turn by turn, and those sessions can be pulled into a dataset in one click. You can also eval the trajectory the agent took, not just its final answer.
Mostly this is a difference in vocabulary between platforms. Some customers nest an entire multi-turn conversation under a single trace, others send each turn as its own trace. Our view is that it doesn't matter. Use whichever is more readable for you. A thread is just tracing data rendered in a more human-readable format.
You can also group a set of traces representing one conversation with a Group by key, whether that's a session ID, a conversation ID, or any metadata key, then surface and score that group in the platform.
This varies by team and can be handled flexibly depending on how you work. Using Patterns as the example, one of the artifacts it produces is a chart you can monitor. So once Braintrust uncovers a failure, you build an eval, and you ship a fix, you can track whether that specific problem actually declined over time or whether it's still there.
Yes, and we're seeing more teams building evals for their MCP servers, where customers reach them through a coding agent. Topics is useful here because you want to know what kinds of interactions developers are actually having. One caveat is that there's a translation layer between developer intent, the coding agent, and the MCP call, and coding agents don't expose as much telemetry as we'd like. It's still an evolving area.
If you want opinionated, end to end production setups for different use cases, our cookbooks cover a wide range. And if something in Patterns or Loop isn't working the way you expect, tell us. We're iterating on both.