[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/Assertions/Assertions.ipynb) by [Vítor Balocco](https://twitter.com/vitorbal) on 2024-02-13
[Zapier](https://zapier.com/) is the #1 workflow automation platform for small and midsize businesses, connecting to more than 6000 of the most popular work apps. We were also one of the first companies to build and ship AI features into our core products. We've had the opportunity to work with Braintrust since the early days of the product, which now powers the evaluation and observability infrastructure across our AI features.
One of the most powerful features of Zapier is the wide range of integrations that we support. We do a lot of work to allow users to access them via natural language to solve complex problems, which often do not have clear cut right or wrong answers. Instead, we define a set of criteria that need to be met (assertions). Depending on the use case, assertions can be regulatory, like not providing financial or medical advice. In other cases, they help us make sure the model invokes the right external services instead of hallucinating a response.
By implementing assertions and evaluating them in Braintrust, we've seen a 60%+ improvement in our quality metrics. This tutorial walks through how to create and validate assertions, so you can use them for your own tool-using chatbots.
## Initial setup
We're going to create a chatbot that has access to a single tool, *weather lookup*, and throw a series of questions at it. Some questions will involve the weather and others won't. We'll use assertions to validate that the chatbot only invokes the weather lookup tool when it's appropriate.
Let's create a simple request handler and hook up a weather tool to it.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { wrapOpenAI } from "braintrust";
import pick from "lodash/pick";
import { ChatCompletionTool } from "openai/resources/chat/completions";
import OpenAI from "openai";
import { z } from "zod";
import zodToJsonSchema from "zod-to-json-schema";
// This wrap function adds some useful tracing in Braintrust
const openai = wrapOpenAI(new OpenAI());
// Convenience function for defining an OpenAI function call
const makeFunctionDefinition = (
name: string,
description: string,
schema: z.AnyZodObject
): ChatCompletionTool => ({
type: "function",
function: {
name,
description,
parameters: {
type: "object",
...pick(
zodToJsonSchema(schema, {
name: "root",
$refStrategy: "none",
}).definitions?.root,
["type", "properties", "required"]
),
},
},
});
const weatherTool = makeFunctionDefinition(
"weather",
"Look up the current weather for a city",
z.object({
city: z.string().describe("The city to look up the weather for"),
date: z.string().optional().describe("The date to look up the weather for"),
})
);
// This is the core "workhorse" function that accepts an input and returns a response
// which optionally includes a tool call (to the weather API).
async function task(input: string) {
const completion = await openai.chat.completions.create({
model: "gpt-3.5-turbo",
messages: [
{
role: "system",
content: `You are a highly intelligent AI that can look up the weather.`,
},
{ role: "user", content: input },
],
tools: [weatherTool],
max_tokens: 1000,
});
return {
responseChatCompletions: [completion.choices[0].message],
};
}
```
Now let's try it out on a few examples!
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
JSON.stringify(await task("What's the weather in San Francisco?"), null, 2);
```
```
{
"responseChatCompletions": [
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_vlOuDTdxGXurjMzy4VDFHGBS",
"type": "function",
"function": {
"name": "weather",
"arguments": "{\n \"city\": \"San Francisco\"\n}"
}
}
]
}
]
}
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
JSON.stringify(await task("What is my bank balance?"), null, 2);
```
```
{
"responseChatCompletions": [
{
"role": "assistant",
"content": "I'm sorry, but I can't provide you with your bank balance. You will need to check with your bank directly for that information."
}
]
}
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
JSON.stringify(await task("What is the weather?"), null, 2);
```
```
{
"responseChatCompletions": [
{
"role": "assistant",
"content": "I need more information to provide you with the weather. Could you please specify the city and the date for which you would like to know the weather?"
}
]
}
```
## Scoring outputs
Validating these cases is subtle. For example, if someone asks "What is the weather?", the correct answer is to ask for clarification. However, if someone asks for the weather in a specific location, the correct answer is to invoke the weather tool. How do we validate these different types of responses?
### Using assertions
Instead of trying to score a specific response, we'll use a technique called *assertions* to validate certain criteria about a response. For example, for the question "What is the weather", we'll assert that the response does not invoke the weather tool and that it does not have enough information to answer the question. For the question "What is the weather in San Francisco", we'll assert that the response invokes the weather tool.
### Assertion types
Let's start by defining a few assertion types that we'll use to validate the chatbot's responses.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
type AssertionTypes =
| "equals"
| "exists"
| "not_exists"
| "llm_criteria_met"
| "semantic_contains";
type Assertion = {
path: string;
assertion_type: AssertionTypes;
value: string;
};
```
`equals`, `exists`, and `not_exists` are heuristics. `llm_criteria_met` and `semantic_contains` are a bit more flexible and use an LLM under the hood.
Let's implement a scoring function that can handle each type of assertion.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { ClosedQA } from "autoevals";
import get from "lodash/get";
import every from "lodash/every";
/**
* Uses an LLM call to classify if a substring is semantically contained in a text.
* @param text The full text you want to check against
* @param needle The string you want to check if it is contained in the text
*/
async function semanticContains({
text1,
text2,
}: {
text1: string;
text2: string;
}): Promise {
const system = `
You are a highly intelligent AI. You will be given two texts, TEXT_1 and TEXT_2. Your job is to tell me if TEXT_2 is semantically present in TEXT_1.
Examples:
\`\`\`
TEXT_1: "I've just sent “hello world” to the #testing channel on Slack as you requested. Can I assist you with anything else?"
TEXT_2: "Can I help you with something else?"
Result: YES
\`\`\`
\`\`\`
TEXT_1: "I've just sent “hello world” to the #testing channel on Slack as you requested. Can I assist you with anything else?"
TEXT_2: "Sorry, something went wrong."
Result: NO
\`\`\`
\`\`\`
TEXT_1: "I've just sent “hello world” to the #testing channel on Slack as you requested. Can I assist you with anything else?"
TEXT_2: "#testing channel Slack"
Result: YES
\`\`\`
\`\`\`
TEXT_1: "I've just sent “hello world” to the #testing channel on Slack as you requested. Can I assist you with anything else?"
TEXT_2: "#general channel Slack"
Result: NO
\`\`\`
`;
const toolSchema = z.object({
rationale: z
.string()
.describe(
"A string that explains the reasoning behind your answer. It's a step-by-step explanation of how you determined that TEXT_2 is or isn't semantically present in TEXT_1."
),
answer: z.boolean().describe("Your answer"),
});
const completion = await openai.chat.completions.create({
model: "gpt-3.5-turbo",
messages: [
{
role: "system",
content: system,
},
{
role: "user",
content: `TEXT_1: "${text1}"\nTEXT_2: "${text2}"`,
},
],
tools: [
makeFunctionDefinition(
"semantic_contains",
"The result of the semantic presence check",
toolSchema
),
],
tool_choice: {
function: { name: "semantic_contains" },
type: "function",
},
max_tokens: 1000,
});
try {
const { answer } = toolSchema.parse(
JSON.parse(
completion.choices[0].message.tool_calls![0].function.arguments
)
);
return answer;
} catch (e) {
console.error(e, "Error parsing semanticContains response");
return false;
}
}
const AssertionScorer = async ({
input,
output,
expected: assertions,
}: {
input: string;
output: any;
expected: Assertion[];
}) => {
// for each assertion, perform the comparison
const assertionResults: {
status: string;
path: string;
assertion_type: string;
value: string;
actualValue: string;
}[] = [];
for (const assertion of assertions) {
const { assertion_type, path, value } = assertion;
const actualValue = get(output, path);
let passedTest = false;
try {
switch (assertion_type) {
case "equals":
passedTest = actualValue === value;
break;
case "exists":
passedTest = actualValue !== undefined;
break;
case "not_exists":
passedTest = actualValue === undefined;
break;
case "llm_criteria_met":
const closedQA = await ClosedQA({
input:
"According to the provided criterion is the submission correct?",
criteria: value,
output: actualValue,
});
passedTest = !!closedQA.score && closedQA.score > 0.5;
break;
case "semantic_contains":
passedTest = await semanticContains({
text1: actualValue,
text2: value,
});
break;
default:
assertion_type satisfies never; // if you see a ts error here, its because your switch is not exhaustive
throw new Error(`unknown assertion type ${assertion_type}`);
}
} catch (e) {
passedTest = false;
}
assertionResults.push({
status: passedTest ? "passed" : "failed",
path,
assertion_type,
value,
actualValue,
});
}
const allPassed = every(assertionResults, (r) => r.status === "passed");
return {
name: "Assertions Score",
score: allPassed ? 1 : 0,
metadata: {
assertionResults,
},
};
};
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
const data = [
{
input: "What's the weather like in San Francisco?",
expected: [
{
path: "responseChatCompletions[0].tool_calls[0].function.name",
assertion_type: "equals",
value: "weather",
},
],
},
{
input: "What's the weather like?",
expected: [
{
path: "responseChatCompletions[0].tool_calls[0].function.name",
assertion_type: "not_exists",
value: "",
},
{
path: "responseChatCompletions[0].content",
assertion_type: "llm_criteria_met",
value:
"Response reflecting the bot does not have enough information to look up the weather",
},
],
},
{
input: "How much is AAPL stock today?",
expected: [
{
path: "responseChatCompletions[0].tool_calls[0].function.name",
assertion_type: "not_exists",
value: "",
},
{
path: "responseChatCompletions[0].content",
assertion_type: "llm_criteria_met",
value:
"Response reflecting the bot does not have access to the ability or tool to look up stock prices.",
},
],
},
{
input: "What can you do?",
expected: [
{
path: "responseChatCompletions[0].content",
assertion_type: "semantic_contains",
value: "look up the weather",
},
],
},
];
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { Eval } from "braintrust";
await Eval("Weather Bot", {
data,
task: async (input) => {
const result = await task(input);
return result;
},
scores: [AssertionScorer],
});
```
```
{
projectName: 'Weather Bot',
experimentName: 'HEAD-1707465445',
projectUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot',
experimentUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot/HEAD-1707465445',
comparisonExperimentName: undefined,
scores: undefined,
metrics: undefined
}
```
```
██░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ | Weather Bot | 4% | 4/100 datapoints
```
```
{
projectName: 'Weather Bot',
experimentName: 'HEAD-1707465445',
projectUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot',
experimentUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot/HEAD-1707465445',
comparisonExperimentName: undefined,
scores: undefined,
metrics: undefined
}
```
### Analyzing results
It looks like half the cases passed.
In one case, the chatbot did not clearly indicate that it needs more information.
In the other case, the chatbot halucinated a stock tool.
## Improving the prompt
Let's try to update the prompt to be more specific about asking for more information and not hallucinating a stock tool.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
async function task(input: string) {
const completion = await openai.chat.completions.create({
model: "gpt-3.5-turbo",
messages: [
{
role: "system",
content: `You are a highly intelligent AI that can look up the weather.
Do not try to use tools other than those provided to you. If you do not have the tools needed to solve a problem, just say so.
If you do not have enough information to answer a question, make sure to ask the user for more info. Prefix that statement with "I need more information to answer this question."
`,
},
{ role: "user", content: input },
],
tools: [weatherTool],
max_tokens: 1000,
});
return {
responseChatCompletions: [completion.choices[0].message],
};
}
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
JSON.stringify(await task("How much is AAPL stock today?"), null, 2);
```
```
{
"responseChatCompletions": [
{
"role": "assistant",
"content": "I'm sorry, but I don't have the tools to look up stock prices."
}
]
}
```
### Re-running eval
Let's re-run the eval and see if our changes helped.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await Eval("Weather Bot", {
data: data,
task: async (input) => {
const result = await task(input);
return result;
},
scores: [AssertionScorer],
});
```
```
{
projectName: 'Weather Bot',
experimentName: 'HEAD-1707465778',
projectUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot',
experimentUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot/HEAD-1707465778',
comparisonExperimentName: 'HEAD-1707465445',
scores: {
'Assertions Score': {
name: 'Assertions Score',
score: 0.75,
diff: 0.25,
improvements: 1,
regressions: 0
}
},
metrics: {
duration: {
name: 'duration',
metric: 1.5197500586509705,
unit: 's',
diff: -0.10424983501434326,
improvements: 2,
regressions: 2
}
}
}
```
```
██░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ | Weather Bot | 4% | 4/100 datapoints
```
```
{
projectName: 'Weather Bot',
experimentName: 'HEAD-1707465778',
projectUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot',
experimentUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Weather%20Bot/HEAD-1707465778',
comparisonExperimentName: 'HEAD-1707465445',
scores: {
'Assertions Score': {
name: 'Assertions Score',
score: 0.75,
diff: 0.25,
improvements: 1,
regressions: 0
}
},
metrics: {
duration: {
name: 'duration',
metric: 1.5197500586509705,
unit: 's',
diff: -0.10424983501434326,
improvements: 2,
regressions: 2
}
}
}
```
Nice! We were able to improve the "needs more information" case.
However, we now halucinate and ask for the weather in NYC. Getting to 100% will take a bit more iteration!
Now that you have a solid evaluation framework in place, you can continue experimenting and try to solve this problem. Happy evaling!
# Classifying news articles
Source: https://braintrust.dev/docs/cookbook/recipes/ClassifyingNewsArticles
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/ClassifyingNewsArticles/ClassifyingNewsArticles.ipynb) by [David Song](https://twitter.com/davidtsong) on 2023-09-01
Classification is a core natural language processing (NLP) task that large language models are good at, but building reliable systems is still challenging. In this cookbook, we'll walk through how to improve an LLM-based classification system that sorts news articles by category.
## Getting started
Before getting started, make sure you have a [Braintrust account](https://www.braintrust.dev/signup) and an API key for [OpenAI](https://platform.openai.com/signup). Make sure to plug the OpenAI key into your Braintrust account's [AI provider configuration](https://www.braintrust.dev/app/~/configuration/org/secrets).
Once you have your Braintrust account set up with an OpenAI API key, install the following dependencies:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
%pip install -U braintrust openai datasets autoevals
```
Next, we'll import the libraries we need and load the [ag\_news](https://huggingface.co/datasets/ag_news) dataset from Hugging Face. Once the dataset is loaded, we'll extract the category names to build a map from indices to names, allowing us to compare expected categories with model outputs. Then, we'll shuffle the dataset with a fixed seed, trim it to 20 data points, and restructure it into a list where each item includes the article text as input and its expected category name.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import braintrust
import os
from datasets import load_dataset
from autoevals import Levenshtein
from openai import OpenAI
dataset = load_dataset("ag_news", split="train")
category_names = dataset.features["label"].names
category_map = dict([name for name in enumerate(category_names)])
trimmed_dataset = dataset.shuffle(seed=42)[:20]
articles = [
{
"input": trimmed_dataset["text"][i],
"expected": category_map[trimmed_dataset["label"][i]],
}
for i in range(len(trimmed_dataset["text"]))
]
```
To authenticate with Braintrust, export your `BRAINTRUST_API_KEY` as an environment variable:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
export BRAINTRUST_API_KEY="YOUR_API_KEY_HERE"
```
Exporting your API key is a best practice, but to make it easier to follow along with this cookbook, you can also hardcode it into the code below.
Once the API key is set, we initialize the OpenAI client using the AI proxy:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Uncomment the following line to hardcode your API key
# os.environ["BRAINTRUST_API_KEY"] = "YOUR_API_KEY_HERE"
client = braintrust.wrap_openai(
OpenAI(
base_url="https://api.braintrust.dev/v1/proxy",
api_key=os.environ["BRAINTRUST_API_KEY"],
)
)
```
## Writing the initial prompts
We'll start by testing classification on a single article. We'll select it from the dataset to examine its input and expected output:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Here's the input and expected output for the first article in our dataset.
test_article = articles[0]
test_text = test_article["input"]
expected_text = test_article["expected"]
print("Article Title:", test_text)
print("Article Label:", expected_text)
```
```
Article Title: Bangladesh paralysed by strikes Opposition activists have brought many towns and cities in Bangladesh to a halt, the day after 18 people died in explosions at a political rally.
Article Label: World
```
Now that we've verified what's in our dataset and initialized the OpenAI client, it's time to try writing a prompt and classifying a title. We'll define a `classify_article` function that takes an input title and returns a category:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
MODEL = "gpt-3.5-turbo"
@braintrust.traced
def classify_article(input):
messages = [
{
"role": "system",
"content": """You are an editor in a newspaper who helps writers identify the right category for their news articles,
by reading the article's title. The category should be one of the following: World, Sports, Business or Sci-Tech. Reply with one word corresponding to the category.""",
},
{
"role": "user",
"content": "Article title: {article_title} Category:".format(
article_title=input
),
},
]
result = client.chat.completions.create(
model=MODEL,
messages=messages,
max_tokens=10,
)
category = result.choices[0].message.content
return category
test_classify = classify_article(test_text)
print("Input:", test_text)
print("Classified as:", test_classify)
print("Score:", 1 if test_classify == expected_text else 0)
```
```
Input: Bangladesh paralysed by strikes Opposition activists have brought many towns and cities in Bangladesh to a halt, the day after 18 people died in explosions at a political rally.
Classified as: World
Score: 1
```
## Running an evaluation
We've tested our prompt on a single article, so now we can test across the rest of the dataset using the `Eval` function. Behind the scenes, `Eval` will in parallel run the `classify_article` function on each article in the dataset, and then compare the results to the ground truth labels using a simple `Levenshtein` scorer. When it finishes running, it will print out the results with a link to dig deeper.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await braintrust.Eval(
"Classifying News Articles Cookbook",
data=articles,
task=classify_article,
scores=[Levenshtein],
experiment_name="Original Prompt",
)
```
```
Experiment Original Prompt-db3e9cae is running at https://www.braintrust.dev/app/braintrustdata.com/p/Classifying%20News%20Articles%20Cookbook/experiments/Original%20Prompt-db3e9cae
\`Eval()\` was called from an async context. For better performance, it is recommended to use \`await EvalAsync()\` instead.
Classifying News Articles Cookbook [experiment_name=Original Prompt] (data): 20it [00:00, 41755.14it/s]
Classifying News Articles Cookbook [experiment_name=Original Prompt] (tasks): 100%|██████████| 20/20 [00:02<00:00, 7.57it/s]
```
```
=========================SUMMARY=========================
Original Prompt-db3e9cae compared to New Prompt-9f185e9e:
71.25% (-00.62%) 'Levenshtein' score (1 improvements, 2 regressions)
1740081219.56s start
1740081220.69s end
1.10s (-298.16%) 'duration' (12 improvements, 8 regressions)
0.72s (-294.09%) 'llm_duration' (10 improvements, 10 regressions)
113.75tok (-) 'prompt_tokens' (0 improvements, 0 regressions)
2.20tok (-) 'completion_tokens' (0 improvements, 0 regressions)
115.95tok (-) 'total_tokens' (0 improvements, 0 regressions)
0.00$ (-) 'estimated_cost' (0 improvements, 0 regressions)
See results for Original Prompt-db3e9cae at https://www.braintrust.dev/app/braintrustdata.com/p/Classifying%20News%20Articles%20Cookbook/experiments/Original%20Prompt-db3e9cae
```
```
EvalResultWithSummary(summary="...", results=[...])
```
## Analyzing the results
Looking at our results table (in the screenshot below), we see our that any data points that involve the category `Sci/Tech` are not scoring 100%. Let's dive deeper.
## Reproducing an example
First, let's see if we can reproduce this issue locally. We can test an article corresponding to the `Sci/Tech` category and reproduce the evaluation:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
sci_tech_article = [a for a in articles if "Galaxy Clusters" in a["input"]][0]
print(sci_tech_article["input"])
print(sci_tech_article["expected"])
out = classify_article(sci_tech_article["expected"])
print(out)
```
```
A Cosmic Storm: When Galaxy Clusters Collide Astronomers have found what they are calling the perfect cosmic storm, a galaxy cluster pile-up so powerful its energy output is second only to the Big Bang.
Sci/Tech
Sci-Tech
```
## Fixing the prompt
Have you spotted the issue? It looks like we misspelled one of the categories in our prompt. The dataset's categories are `World`, `Sports`, `Business` and `Sci/Tech` - but we are using `Sci-Tech` in our prompt. Let's fix it:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
@braintrust.traced
def classify_article(input):
messages = [
{
"role": "system",
"content": """You are an editor in a newspaper who helps writers identify the right category for their news articles,
by reading the article's title. The category should be one of the following: World, Sports, Business or Sci/Tech. Reply with one word corresponding to the category.""",
},
{
"role": "user",
"content": "Article title: {input} Category:".format(input=input),
},
]
result = client.chat.completions.create(
model=MODEL,
messages=messages,
max_tokens=10,
)
category = result.choices[0].message.content
return category
result = classify_article(sci_tech_article["input"])
print(result)
```
```
Sci/Tech
```
## Evaluate the new prompt
The model classified the correct category `Sci/Tech` for this example. But, how do we know it works for the rest of the dataset? Let's run a new experiment to evaluate our new prompt:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await braintrust.Eval(
"Classifying News Articles Cookbook",
data=articles,
task=classify_article,
scores=[Levenshtein],
experiment_name="New Prompt",
)
```
## Conclusion
Select the new experiment, and check it out. You should notice a few things:
* Braintrust will automatically compare the new experiment to your previous one.
* You should see the eval scores increase and you can see which test cases improved.
* You can also filter the test cases by improvements to know exactly why the scores changed.
## Next steps
* [I ran an eval. Now what?](https://braintrust.dev/blog/after-evals)
* Add more [custom scorers](/docs/evaluate/custom-code).
* Try other models like xAI's [Grok 2](https://x.ai/blog/grok-2) or OpenAI's [o1](https://openai.com/o1/). To learn more about comparing evals across multiple AI models, check out this [cookbook](/docs/cookbook/recipes/ModelComparison).
# Coda's Help Desk with and without RAG
Source: https://braintrust.dev/docs/cookbook/recipes/CodaHelpDesk
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/CodaHelpDesk/CodaHelpDesk.ipynb) by [Austin Moehle](https://www.linkedin.com/in/austinmxx/), [Kenny Wong](https://twitter.com/siuheihk) on 2023-12-21
Large language models have gotten extremely good at answering general questions but often struggle with specific domain knowledge. When building AI-powered help desks or knowledge bases, this limitation becomes apparent. Retrieval-augmented generation (RAG) addresses this challenge by incorporating relevant information from external documents into the model's context.
In this cookbook, we'll build and evaluate an AI application that answers questions about [Coda's Help Desk](https://help.coda.io/en/) documentation. Using Braintrust, we'll compare baseline and RAG-enhanced responses against expected answers to quantitatively measure the improvement.
## Getting started
To follow along, start by installing the required packages:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
pip install autoevals braintrust requests openai lancedb markdownify asyncio pyarrow
```
Next, make sure you have a [Braintrust](https://www.braintrust.dev/signup) account, along with an [OpenAI API key](https://platform.openai.com/). To authenticate with Braintrust, export your `BRAINTRUST_API_KEY` as an environment variable:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
export BRAINTRUST_API_KEY="YOUR_API_KEY_HERE"
```
Exporting your API key is a best practice, but to make it easier to follow along with this cookbook, you can also hardcode it into the code below.
We'll import our modules and define constants:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import os
import re
import json
import tempfile
from typing import List
import autoevals
import braintrust
import markdownify
import lancedb
import openai
import requests
import asyncio
from pydantic import BaseModel, Field
# Model selection constants
QA_GEN_MODEL = "gpt-4o-mini"
QA_ANSWER_MODEL = "gpt-4o-mini"
QA_GRADING_MODEL = "gpt-4o-mini"
RELEVANCE_MODEL = "gpt-4o-mini"
# Data constants
NUM_SECTIONS = 20
NUM_QA_PAIRS = 20 # Increase this number to test at a larger scale
TOP_K = 2 # Number of relevant sections to retrieve
# Uncomment the following line to hardcode your API key
# os.environ["BRAINTRUST_API_KEY"] = "YOUR_API_KEY_HERE"
```
## Download Markdown docs from Coda's Help Desk
Let's start by downloading the Coda docs and splitting them into their constituent Markdown sections.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
data = requests.get(
"https://gist.githubusercontent.com/wong-codaio/b8ea0e087f800971ca5ec9eef617273e/raw/39f8bd2ebdecee485021e20f2c1d40fd649a4c77/articles.json"
).json()
markdown_docs = [
{"id": row["id"], "markdown": markdownify.markdownify(row["body"])} for row in data
]
i = 0
markdown_sections = []
for markdown_doc in markdown_docs:
sections = re.split(r"(.*\n=+\n)", markdown_doc["markdown"])
current_section = ""
for section in sections:
if not section.strip():
continue
if re.match(r".*\n=+\n", section):
current_section = section
else:
section = current_section + section
markdown_sections.append(
{
"doc_id": markdown_doc["id"],
"section_id": i,
"markdown": section.strip(),
}
)
current_section = ""
i += 1
print(f"Downloaded {len(markdown_sections)} Markdown sections. Here are the first 3:")
for i, section in enumerate(markdown_sections[:3]):
print(f"\nSection {i+1}:\n{section}")
```
```
Downloaded 996 Markdown sections. Here are the first 3:
Section 1:
{'doc_id': '8179780', 'section_id': 0, 'markdown': "Not all Coda docs are used in the same way. You'll inevitably have a few that you use every week, and some that you'll only use once. This is where starred docs can help you stay organized.\n\nStarring docs is a great way to mark docs of personal importance. After you star a doc, it will live in a section on your doc list called **[My Shortcuts](https://coda.io/shortcuts)**. All starred docs, even from multiple different workspaces, will live in this section.\n\nStarring docs only saves them to your personal My Shortcuts. It doesn’t affect the view for others in your workspace. If you’re wanting to shortcut docs not just for yourself but also for others in your team or workspace, you’ll [use pinning](https://help.coda.io/en/articles/2865511-starred-pinned-docs) instead."}
Section 2:
{'doc_id': '8179780', 'section_id': 1, 'markdown': '**Star your docs**\n==================\n\nTo star a doc, hover over its name in the doc list and click the star icon. Alternatively, you can star a doc from within the doc itself. Hover over the doc title in the upper left corner, and click on the star.\n\nOnce you star a doc, you can access it quickly from the [My Shortcuts](https://coda.io/shortcuts) tab of your doc list.\n\n\n\nAnd, as your doc needs change, simply click the star again to un-star the doc and remove it from **My Shortcuts**.'}
Section 3:
{'doc_id': '8179780', 'section_id': 2, 'markdown': '**FAQs**\n========\n\nWhen should I star a doc and when should I pin it?\n--------------------------------------------------\n\nStarring docs is best for docs of *personal* importance. Starred docs appear in your **My Shortcuts**, but they aren’t starred for anyone else in your workspace. For instance, you may want to star your personal to-do list doc or any docs you use on a daily basis.\n\n[Pinning](https://help.coda.io/en/articles/2865511-starred-pinned-docs) is recommended when you want to flag or shortcut a doc for *everyone* in your workspace or folder. For instance, you likely want to pin your company wiki doc to your workspace. And you may want to pin your team task tracker doc to your team’s folder.\n\nCan I star docs for everyone?\n-----------------------------\n\nStarring docs only applies to your own view and your own My Shortcuts. To pin docs (or templates) to your workspace or folder, [refer to this article](https://help.coda.io/en/articles/2865511-starred-pinned-docs).\n\n---'}
```
## Use the Braintrust AI Proxy
Let's initialize the OpenAI client using the [Braintrust proxy](/docs/deploy/ai-proxy). The Braintrust AI Proxy provides a single API to access OpenAI and other models. Because the proxy automatically caches and reuses results (when `temperature=0` or the `seed` parameter is set), we can re-evaluate prompts many times without incurring additional API costs.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
client = braintrust.wrap_openai(
openai.AsyncOpenAI(
api_key=os.environ.get("BRAINTRUST_API_KEY"),
base_url="https://api.braintrust.dev/v1/proxy",
default_headers={"x-bt-use-cache": "always"},
)
)
```
## Generate question-answer pairs
Before we start evaluating some prompts, let's use the LLM to generate a bunch of question-answer pairs from the text at hand. We'll use these QA pairs as ground truth when grading our models later.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
class QAPair(BaseModel):
questions: List[str] = Field(
...,
description="List of questions, all with the same meaning but worded differently",
)
answer: str = Field(..., description="Answer")
class QAPairs(BaseModel):
pairs: List[QAPair] = Field(..., description="List of question/answer pairs")
async def produce_candidate_questions(row):
response = await client.chat.completions.create(
model=QA_GEN_MODEL,
messages=[
{
"role": "user",
"content": f"""\
Please generate 8 question/answer pairs from the following text. For each question, suggest
2 different ways of phrasing the question, and provide a unique answer.
Content:
{row['markdown']}
""",
}
],
functions=[
{
"name": "propose_qa_pairs",
"description": "Propose some question/answer pairs for a given document",
"parameters": QAPairs.model_json_schema(),
}
],
)
pairs = QAPairs(**json.loads(response.choices[0].message.function_call.arguments))
return pairs.pairs
# Create tasks for all API calls
all_candidates_tasks = [
asyncio.create_task(produce_candidate_questions(a))
for a in markdown_sections[:NUM_SECTIONS]
]
all_candidates = [await f for f in all_candidates_tasks]
data = []
row_id = 0
for row, doc_qa in zip(markdown_sections[:NUM_SECTIONS], all_candidates):
for i, qa in enumerate(doc_qa):
for j, q in enumerate(qa.questions):
data.append(
{
"input": q,
"expected": qa.answer,
"metadata": {
"document_id": row["doc_id"],
"section_id": row["section_id"],
"question_idx": i,
"answer_idx": j,
"id": row_id,
"split": (
"test" if j == len(qa.questions) - 1 and j > 0 else "train"
),
},
}
)
row_id += 1
print(f"Generated {len(data)} QA pairs. Here are the first 10:")
for x in data[:10]:
print(x)
```
```
Generated 320 QA pairs. Here are the first 10:
{'input': 'What is the purpose of starring a doc in Coda?', 'expected': 'Starring a doc in Coda helps you mark documents of personal importance, making it easier to organize and access them quickly.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 0, 'answer_idx': 0, 'id': 0, 'split': 'train'}}
{'input': 'Why would someone want to star a document in Coda?', 'expected': 'Starring a doc in Coda helps you mark documents of personal importance, making it easier to organize and access them quickly.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 0, 'answer_idx': 1, 'id': 1, 'split': 'test'}}
{'input': 'Where do starred docs appear in Coda?', 'expected': 'Starred docs appear in a section called My Shortcuts on your doc list, allowing for quick access.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 1, 'answer_idx': 0, 'id': 2, 'split': 'train'}}
{'input': 'After starring a document in Coda, where can I find it?', 'expected': 'Starred docs appear in a section called My Shortcuts on your doc list, allowing for quick access.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 1, 'answer_idx': 1, 'id': 3, 'split': 'test'}}
{'input': 'Does starring a doc affect other users in the workspace?', 'expected': 'No, starring a doc only saves it to your personal My Shortcuts and does not affect the view for others in your workspace.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 2, 'answer_idx': 0, 'id': 4, 'split': 'train'}}
{'input': 'Will my colleagues see the docs I star in Coda?', 'expected': 'No, starring a doc only saves it to your personal My Shortcuts and does not affect the view for others in your workspace.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 2, 'answer_idx': 1, 'id': 5, 'split': 'test'}}
{'input': 'What should I use if I want to share a shortcut to a doc with my team?', 'expected': 'To create a shortcut for a document that your team can access, you should use the pinning feature instead of starring.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 3, 'answer_idx': 0, 'id': 6, 'split': 'train'}}
{'input': 'How can I create a shortcut for a document that everyone in my workspace can access?', 'expected': 'To create a shortcut for a document that your team can access, you should use the pinning feature instead of starring.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 3, 'answer_idx': 1, 'id': 7, 'split': 'test'}}
{'input': 'Can starred documents come from different workspaces in Coda?', 'expected': 'Yes, all starred docs, even from multiple different workspaces, will live in the My Shortcuts section.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 4, 'answer_idx': 0, 'id': 8, 'split': 'train'}}
{'input': 'Is it possible to star docs from multiple workspaces?', 'expected': 'Yes, all starred docs, even from multiple different workspaces, will live in the My Shortcuts section.', 'metadata': {'document_id': '8179780', 'section_id': 0, 'question_idx': 4, 'answer_idx': 1, 'id': 9, 'split': 'test'}}
```
## Evaluate a context-free prompt (no RAG)
Let's evaluate a simple prompt that poses each question without providing context from the Markdown docs. We'll evaluate this naive approach using the [Factuality prompt](https://github.com/braintrustdata/autoevals/blob/main/templates/factuality.yaml) from the Braintrust [autoevals](/docs/reference/autoevals) library.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
async def simple_qa(input):
completion = await client.chat.completions.create(
model=QA_ANSWER_MODEL,
messages=[
{
"role": "user",
"content": f"""\
Please answer the following question:
Question: {input}
""",
}
],
)
return completion.choices[0].message.content
await braintrust.Eval(
name="Coda Help Desk Cookbook",
experiment_name="No RAG",
data=data[:NUM_QA_PAIRS],
task=simple_qa,
scores=[autoevals.Factuality(model=QA_GRADING_MODEL)],
)
```
### Analyze the evaluation in the UI
The cell above will print a link to a Braintrust experiment. Pause and navigate to the UI to view our baseline eval.
## Try using RAG to improve performance
Let's see if RAG (retrieval-augmented generation) can improve our results on this task.
First, we'll compute embeddings for each Markdown section using `text-embedding-ada-002` and create an index over the embeddings in [LanceDB](https://lancedb.com), a vector database. Then, for any given query, we can convert it to an embedding and efficiently find the most relevant context by searching in embedding space. We'll then provide the corresponding text as additional context in our prompt.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
tempdir = tempfile.TemporaryDirectory()
LANCE_DB_PATH = os.path.join(tempdir.name, "docs-lancedb")
@braintrust.traced
async def embed_text(text):
params = dict(input=text, model="text-embedding-ada-002")
response = await client.embeddings.create(**params)
embedding = response.data[0].embedding
braintrust.current_span().log(
metrics={
"tokens": response.usage.total_tokens,
"prompt_tokens": response.usage.prompt_tokens,
},
metadata={"model": response.model},
input=text,
output=embedding,
)
return embedding
embedding_tasks = [
asyncio.create_task(embed_text(row["markdown"]))
for row in markdown_sections[:NUM_SECTIONS]
]
embeddings = [await f for f in embedding_tasks]
db = lancedb.connect(LANCE_DB_PATH)
try:
db.drop_table("sections")
except:
pass
# Convert the data to a pandas DataFrame first
import pandas as pd
table_data = [
{
"doc_id": row["doc_id"],
"section_id": row["section_id"],
"text": row["markdown"],
"vector": embedding,
}
for (row, embedding) in zip(markdown_sections[:NUM_SECTIONS], embeddings)
]
# Create table using the DataFrame approach
table = db.create_table("sections", data=pd.DataFrame(table_data))
```
## Use AI to judge relevance of retrieved documents
Let's retrieve a few *more* of the best-matching candidates from the vector database than we intend to use, then use the model from `RELEVANCE_MODEL` to score the relevance of each candidate to the input query. We'll use the `TOP_K` blurbs by relevance score in our QA prompt. Doing this should be a little more intelligent than just using the closest embeddings.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
@braintrust.traced
async def relevance_score(query, document):
response = await client.chat.completions.create(
model=RELEVANCE_MODEL,
messages=[
{
"role": "user",
"content": f"""\
Consider the following query and a document
Query:
{query}
Document:
{document}
Please score the relevance of the document to a query, on a scale of 0 to 1.
""",
}
],
functions=[
{
"name": "has_relevance",
"description": "Declare the relevance of a document to a query",
"parameters": {
"type": "object",
"properties": {
"score": {"type": "number"},
},
},
}
],
)
arguments = response.choices[0].message.function_call.arguments
result = json.loads(arguments)
braintrust.current_span().log(
input={"query": query, "document": document},
output=result,
)
return result["score"]
async def retrieval_qa(input):
embedding = await embed_text(input)
with braintrust.current_span().start_span(
name="vector search", input=input
) as span:
result = table.search(embedding).limit(TOP_K + 3).to_arrow().to_pylist()
docs = [markdown_sections[i["section_id"]]["markdown"] for i in result]
relevance_scores = []
for doc in docs:
relevance_scores.append(await relevance_score(input, doc))
span.log(
output=[
{
"doc": markdown_sections[r["section_id"]]["markdown"],
"distance": r["_distance"],
}
for r in result
],
metadata={"top_k": TOP_K, "retrieval": result},
scores={
"avg_relevance": sum(relevance_scores) / len(relevance_scores),
"min_relevance": min(relevance_scores),
"max_relevance": max(relevance_scores),
},
)
context = "\n------\n".join(docs[:TOP_K])
completion = await client.chat.completions.create(
model=QA_ANSWER_MODEL,
messages=[
{
"role": "user",
"content": f"""\
Given the following context
{context}
Please answer the following question:
Question: {input}
""",
}
],
)
return completion.choices[0].message.content
```
## Run the RAG evaluation
Now let's run our evaluation with RAG:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await braintrust.Eval(
name="Coda Help Desk Cookbook",
experiment_name=f"RAG TopK={TOP_K}",
data=data[:NUM_QA_PAIRS],
task=retrieval_qa,
scores=[autoevals.Factuality(model=QA_GRADING_MODEL)],
)
```
### Analyzing the results
Select the new experiment to analyze the results. You should notice several things:
* Braintrust automatically compares the new experiment to your previous one
* You should see an increase in scores with RAG
* You can explore individual examples to see exactly which responses improved
Try adjusting the constants set at the beginning of this tutorial, such as `NUM_QA_PAIRS`, to run your evaluation on a larger dataset and gain more confidence in your findings.
## Next steps
* Learn about [using functions to build a RAG agent](/docs/cookbook/recipes/ToolRAG).
* Compare your [evals across different models](/docs/cookbook/recipes/ModelComparison).
* If RAG is just one part of your agent, learn how to [evaluate a prompt chaining agent](/docs/cookbook/recipes/PromptChaining).
# Improving coding agents with Coding Agent Insights
Source: https://braintrust.dev/docs/cookbook/recipes/CodingAgentInsights
Track your team's coding-agent usage and cost, and find recurring problems with Topics facets, a daily Loop automation, and Patterns.
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/CodingAgentInsights/CodingAgentInsights.mdx) by [Max Kern](https://github.com/max-braintrust) on 2026-08-19
in [public preview](/docs/feature-lifecycle) and can change before reaching general availability.
Coding Agent Insights shows what your team's coding agents cost, who uses them, and where they could work more effectively. A dashboard covers usage and cost, and a daily Loop automation finds recurring problems and inefficiencies across your team's sessions. This cookbook sets up both for your team.
The [Coding Agent Insights dashboard](/docs/observe/dashboards/coding-agent-insights) shows usage by user, API key, model, tool, and branch, with cost breakdowns. It also shows which teammates are sending traces, to help you track adoption across your team.
A daily [Loop automation](/docs/loop/automations) reviews your team's sessions and records recurring problems as [patterns](/docs/observe/patterns) with supporting evidence. For example, the automation might find:
* Agents repeatedly resending large amounts of context, increasing model usage.
* File operations failing because of stale paths or imprecise replacement text.
* Agents repeatedly checking whether long-running commands have finished.
* Tracing issues that combine unrelated work or reprocess the same sessions.
## Prerequisites
Before you begin, you'll need:
* **A [Braintrust account](https://www.braintrust.dev/signup)** with access to your team's organization.
* **An organization owner** who can [enable Topics](/docs/observe/topics/enable#enable-topics) if needed.
* **An OpenAI [AI provider](/docs/admin/ai-providers)** configured with access to `gpt-6.1-sol`.
* **A coding agent**: [Claude Code](/docs/integrations/developer-tools/claude-code), [Codex](/docs/integrations/developer-tools/codex), [OpenCode](/docs/integrations/developer-tools/opencode), [Pi](/docs/integrations/developer-tools/pi), [Antigravity](/docs/integrations/developer-tools/antigravity), or [Grok](/docs/integrations/developer-tools/grok).
* **Data plane v2.13.0 or later**, if you self-host Braintrust.
## 1. Create a project for your traces
First, create a dedicated project for your coding-agent traces.
1. In Braintrust, open the project dropdown and click **Create project**. Creating a project requires organization-level [`Create` permission](/docs/admin/access-control/manage-permissions#set-organization-permissions).
2. Enter a name for your project, such as `coding-agent-insights`. This cookbook uses that name in its examples, so if you choose a different one, use it wherever `coding-agent-insights` appears.
3. Optionally add a description, then click **Create**.
4. Make sure every teammate whose sessions you want to collect has [`Read` permission on the project](/docs/admin/access-control#object-permissions) to view it and [`Update` permission on its logs](/docs/admin/access-control#what-update-covers) to send traces.
## 2. Protect sensitive data (optional)
in [private preview](/docs/feature-lifecycle), available to a limited set of customers. To request access, [contact Braintrust](https://braintrust.dev/contact).
Coding-agent traces can contain source code, file paths, prompts, command output, and any credentials that appear during a session. [Ingestion redaction](/docs/admin/data-management/protect-sensitive-data#redact-during-ingestion) detects sensitive text, such as names, email addresses, and credentials, and replaces it before Braintrust stores your traces. Detection is model-based and can miss sensitive values.
Ingestion redaction applies only to traces that arrive after you enable it, so enable it on your project before teammates start sending traces. For instructions, see [Enable redaction](/docs/admin/data-management/protect-sensitive-data#enable-redaction).
To remove values before they're sent to Braintrust, add a span plugin when [configuring your coding agents](#3-configure-your-coding-agents) in the next step.
## 3. Configure your coding agents
Next, configure your coding agents to send traces to the project. You and your teammates must complete this setup on each computer where you use a coding agent.
If you already have `bt` installed, [update to the latest version](/docs/reference/cli/migrate#update-bt). Otherwise, install it:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
curl -fsSL https://bt.dev/cli/install.sh | bash
```
After installing, open a new terminal so that `bt` is on your `PATH`. Then run `bt --version`. You should see `bt` v0.19.3 or later.
If `bt` is already signed in with OAuth and can access `coding-agent-insights`, skip this step. Otherwise, run the following command and complete the browser flow:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt login --oauth
```
`bt` saves an OAuth profile that the tracing integrations use to authenticate.
Each teammate should sign in with their own Braintrust login. The [Coding Agent Insights dashboard](/docs/observe/dashboards/coding-agent-insights) attributes each session to the Braintrust user whose credentials sent it. If teammates share one API key, the dashboard can't tell their sessions apart.
If your deployment uses custom authentication URLs, sign in using this command:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt login --oauth \
--api-url "" \
--app-url ""
```
* `--api-url` specifies the API URL for OAuth authentication, which defaults to `https://api.braintrust.dev`.
* `--app-url` specifies the Braintrust app URL used to find your organizations and their data planes, which defaults to `https://www.braintrust.dev`.
Omit either flag if your deployment uses its default URL. `bt` saves these settings and automatically resolves where to send traces for the selected organization.
Run the commands for the coding agents you use on this computer. Each command enables tracing for one agent and sends its sessions to `coding-agent-insights`.
If you have several saved profiles, `bt` uses the one that can access the organization you pass with `--org`, and asks you to choose if more than one can.
To remove sensitive values on each computer before `bt` sends them, write a [span plugin](/docs/reference/cli/trace#transform-spans-with-javascript) and add `--plugin ` to the command shown below. Span plugins require `bt` v0.21.0 or later.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt trace enable claude \
--org "" \
--project coding-agent-insights
```
After setup:
* Restart Claude Code to load the tracing configuration.
* Run `bt trace doctor claude`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If the organization or project is incorrect, rerun `bt trace enable claude` with the correct `--org` and `--project` values.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt trace enable codex \
--org "" \
--project coding-agent-insights
```
After setup:
* Restart Codex and trust the Braintrust hooks when prompted. If the prompt does not appear, run `/hooks` and review the new hooks.
* Run `bt trace doctor codex`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If the organization or project is incorrect, rerun `bt trace enable codex` with the correct `--org` and `--project` values.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt trace enable opencode \
--org "" \
--project coding-agent-insights
```
After setup:
* Restart OpenCode to load the tracing configuration.
* Run `bt trace doctor opencode`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If the organization or project is incorrect, rerun `bt trace enable opencode` with the correct `--org` and `--project` values.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt trace enable pi \
--org "" \
--project coding-agent-insights
```
After setup:
* Restart Pi to load the tracing configuration.
* Run `bt trace doctor pi`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If the organization or project is incorrect, rerun `bt trace enable pi` with the correct `--org` and `--project` values.
Tracing covers `agy` sessions.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt trace enable antigravity \
--org "" \
--project coding-agent-insights
```
After setup:
* Restart Antigravity to load the tracing configuration.
* Run `bt trace doctor antigravity`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If the organization or project is incorrect, rerun `bt trace enable antigravity` with the correct `--org` and `--project` values.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt trace enable grok \
--org "" \
--project coding-agent-insights
```
After setup:
* Start Grok and run `/reload-plugins` before sending your first prompt. Repeat this at the start of each session, even after restarting Grok.
* Run `bt trace doctor grok`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If the organization or project is incorrect, rerun `bt trace enable grok` with the correct `--org` and `--project` values.
Confirm that each configured agent sends sessions to the shared project:
* Start a new session in each configured coding agent and ask it to complete a small task. Ask participating teammates to do the same.
* Go to [** Logs**](https://www.braintrust.dev/app/~/logs), select `coding-agent-insights`, and confirm that each new session appears as a root trace.
If a trace does not appear:
* **For any agent,** run the agent's `bt trace doctor` command and resolve any authentication or routing problems it reports.
* **For Codex,** run `/hooks` and confirm that the Braintrust hooks are trusted.
* **For Grok,** run `/reload-plugins` before sending the test prompt.
For more troubleshooting, see the guide for [Claude Code](/docs/integrations/developer-tools/claude-code), [Codex](/docs/integrations/developer-tools/codex), [OpenCode](/docs/integrations/developer-tools/opencode), [Pi](/docs/integrations/developer-tools/pi), [Antigravity](/docs/integrations/developer-tools/antigravity), or [Grok](/docs/integrations/developer-tools/grok).
## 4. Configure Topics and Loop
With tracing configured, enable Topics, then use a prebuilt configuration file to add the Coding Agent Insights facets and daily Loop automation to your project.
If Topics is already enabled for `coding-agent-insights`, skip this step. Otherwise, ask an organization owner to:
1. In `coding-agent-insights`, go to [** Topics**](https://www.braintrust.dev/app/~/topics).
2. Under **Get started with topics**, make sure **Task** is selected. You can leave the other built-in facets off.
3. Choose whether to **Apply to existing traces**.
4. Click **Enable topics**.
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
curl -fsSL \
https://www.braintrust.dev/docs/assets/active-observability-config.json \
-o /tmp/active-observability-config.json
```
The configuration adds three Topics facets that classify each coding-agent session:
* **Coding agent context debt:** Identifies reusable knowledge or access the agent lacked that you had to supply during the task, such as a repository convention it missed.
* **Primary failure mode:** Labels the main problem in a session, such as an unresolved test failure, editing the wrong files, or getting stuck in repeated attempts.
* **Skills analysis:** Labels problems caused by a skill the agent loaded, such as choosing an irrelevant skill or ignoring its instructions.
It also adds **Coding Agent Improvements Discovery**, a daily Loop automation that reviews the past 30 days of traces, checks facet classifications against raw evidence, and creates or updates patterns.
Applying the configuration file requires project-level [`Read`, `Create`, and `Update` permissions](/docs/admin/access-control#object-permissions). To add its Topics facets and Loop automation to your project, run this command:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
bt observability template push \
/tmp/active-observability-config.json \
--org "" \
--project coding-agent-insights \
--prefer-profile
```
`--prefer-profile` uses your saved login even when `BRAINTRUST_API_KEY` is set. An explicit `--profile` also takes precedence.
At the confirmation prompt, verify the organization, project, three facets, one Loop automation, and Topics wiring. Then confirm the push. The command should report that it pushed 3 facets and 1 Loop automation.
Confirm that the template added the facets to the project and attached them to the Topics automation:
1. In `coding-agent-insights`, go to [** Topics**](https://www.braintrust.dev/app/~/topics). Confirm that cards appear for **Coding agent context debt**, **Primary failure mode**, and **Skills analysis**.
2. Click **Automation** and select **Topics**. Expand **Active facets** and confirm that it includes the same three custom facets.
Confirm that the template created an active daily automation with the expected analysis window and pattern permissions:
1. Go to ** Settings** > [** Automations**](https://www.braintrust.dev/app/~/configuration/automations). Confirm that **Coding Agent Improvements Discovery** shows **Every 24 hours** and **Active**.
2. Click **Coding Agent Improvements Discovery** to open it. On the **Edit** tab, confirm that:
* **Agent configuration** shows **GPT-6.1 Sol** with **Extra high** reasoning effort.
* **Default query range** shows **Custom**, `30` days.
* **Write tool permissions** shows **2 allowed**. Open it and confirm that `new_pattern` and `update_pattern` are selected.
These checks confirm that Topics will label coding-agent sessions and Loop will review the most recent 30 days for recurring problems.
## 5. Collect coding-agent traces
Now, collect traces as you and your teammates use your coding agents. Each session gives the automation more evidence to distinguish recurring problems from one-off difficulties.
You can also [import saved sessions](/docs/reference/cli/trace#bt-trace-import) from [Claude Code](/docs/integrations/developer-tools/claude-code#common-workflows), [Codex](/docs/integrations/developer-tools/codex#common-workflows), or [Antigravity](/docs/integrations/developer-tools/antigravity#common-workflows).
## 6. Verify the automation
Once you've collected some traces, run the Loop automation manually to check that it's working:
1. Go to ** Settings** > [** Automations**](https://www.braintrust.dev/app/~/configuration/automations) and open **Coding Agent Improvements Discovery**.
2. Click **Run now**, then **Open Loop** in the confirmation notification to view the run's thread. You can also open the latest run from the automation's **Past runs** tab.
3. Confirm that the run completes without reported errors and describes the traces it reviewed.
A successful run does not need to create a pattern. If no traces were found, check that your project contains sessions from the past 30 days. For model errors, see [Troubleshooting](#troubleshooting).
The automation continues running daily, so further manual runs are optional.
## 7. Review your team's usage
The [Coding Agent Insights dashboard](/docs/observe/dashboards/coding-agent-insights) shows usage and cost as soon as traces arrive, without waiting for the daily automation.
1. In `coding-agent-insights`, go to [** Dashboards**](https://www.braintrust.dev/app/~/dashboards) and select **Coding agent insights**.
2. Check which teammates are sending traces. With the table grouped by **Users**, the **User** column header counts members with recorded model calls matching the selected period and filters. Click the count to see who is and isn't sending traces.
3. Use the grouping selector to break down usage and cost by user, API key, model, tool, or branch:
To learn more, see [Coding Agent Insights dashboard](/docs/observe/dashboards/coding-agent-insights).
## 8. Review patterns
As findings appear, review them in Patterns to decide what to change in your agents' instructions, skills, or tools.
To review a pattern:
1. Go to [** Patterns**](https://www.braintrust.dev/app/~/patterns) in `coding-agent-insights`.
2. Select a pattern and read its summary and suggested fix, if one is provided.
3. Open the evidence traces to check that they support the finding. Review any attached monitor charts to see how the behavior changes over time.
4. Choose an action: click **Continue in Loop** to investigate further, create a scorer or classifier to track the behavior, or click **Copy for agent** to copy the pattern as a prompt for a coding agent to work on a fix. See [Review and act on patterns](/docs/observe/patterns/review) for instructions.
If the list is empty, continue collecting sessions and check later runs.
## 9. Refine the analysis and share reports (optional)
Finally, you can change what the automation looks for by editing its **Instruction** prompt, and send its reports to Slack or a webhook.
To change what the automation analyzes, start in an interactive Loop thread. You can try different analyses against recent sessions before updating the automation.
Loop asks before it changes anything in your project, such as the automation. While you experiment, click **Skip** on those requests and leave **Always allow** off.
**Ask Loop to suggest changes.** Go to [** Loop**](https://www.braintrust.dev/app/~/loop) and enter this prompt:
```text wrap theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Review the coding-agent facets and automations in this project against recent traces.
Identify noisy classifications, missing analysis, and recurring-analysis settings that should change.
Propose and test improvements to both resources, then show me the changes before applying them.
```
Continue the thread to refine and test its suggestions.
**Try your own analysis.** If you already know what you want to try, describe the revised analysis in a new Loop thread and ask Loop to run it against recent sessions. Adjust your instructions and rerun the analysis until you're satisfied with the results.
**Apply a change.** When you're ready to apply a change, ask Loop to make it and click **Accept**, or edit the automation yourself. To edit it, go to ** Settings** > [** Automations**](https://www.braintrust.dev/app/~/configuration/automations), open **Coding Agent Improvements Discovery**, make the change, and click **Save**.
To send the result of each successful automation run to Slack or a webhook, add a destination to **Coding Agent Improvements Discovery**.
Skip this step if Slack is already connected to your organization. Only members of the **Owners** [permission group](/docs/admin/access-control), or a custom permission group with the **Manage settings** organization permission, can connect Slack.
Go to ** Settings** > [** Integrations**](https://www.braintrust.dev/app/~/configuration/org/integrations), click **Enable Slack**, and authorize Braintrust to access your Slack workspace. For details, see [Enable Slack integration](/docs/admin/organizations#enable-slack-integration).
Go to ** Settings** > [** Automations**](https://www.braintrust.dev/app/~/configuration/automations) and click **Coding Agent Improvements Discovery** to open it. Under **Destinations**, click **Destination** and select **Send to Slack**. Then choose the workspace and channel.
Braintrust prepopulates a suggested **Formatting prompt**. Review it and adjust it as needed for your team.
Go to ** Settings** > [** Automations**](https://www.braintrust.dev/app/~/configuration/automations) and click **Coding Agent Improvements Discovery** to open it. Under **Destinations**, click **Destination** and select **Send to webhook**. Enter the URL of the receiving endpoint.
Braintrust prepopulates a formatting prompt that sends the report as JSON. Keep it or edit it to match the format expected by the receiving service.
**Test the destination.** Save the automation and click **Run now**. When the run finishes, confirm that the Slack channel or webhook endpoint received the formatted report.
Adding a destination does not change the automation's daily schedule or 30-day lookback window.
## Troubleshooting
If the `curl` command for downloading the configuration file returns an HTTP error:
* Confirm that the URL is `https://www.braintrust.dev/docs/assets/active-observability-config.json`.
* Retry from a network that can reach `www.braintrust.dev`.
If `bt observability template` is not recognized when you push the template:
* Run `bt --version` to check the installed version.
* Install or update to `bt` v0.19.3 or later. If the installed version already meets this requirement, reinstall it.
* After the installation succeeds, return to **Apply the configuration file to your project** and run the template command shown there.
If `bt observability template push` reports that a matching facet or Loop automation already exists, it stops before the confirmation prompt:
* If you intend to restore the template configuration, rerun the same command with `--force` and confirm the replacements.
* A new project does not need `--force`.
Rerunning the command with `--force` resets the three facets and **Coding Agent Improvements Discovery** to the versions in the JSON file. Any edits you made to them are lost. Slack and webhook destinations already added to the automation are preserved.
If `bt trace enable codex` reports a marketplace clone timeout:
* Run the command again.
* If it repeatedly reports `fatal: early EOF`, troubleshoot the network connection rather than the Braintrust project configuration.
If a test session does not appear in **Logs** after you configure tracing:
* Restart the coding agent and run another test task.
* For a self-hosted deployment, confirm that the saved OAuth profile uses the correct authentication URLs. If `BRAINTRUST_API_URL` is set in the shell where you start the agent, check that it points to the intended data plane or unset it to let `bt` resolve the selected organization's endpoint. Restart the agent after changing its environment.
* If the trace still does not appear, run `bt trace doctor `, replacing `` with `claude`, `codex`, `opencode`, `pi`, `antigravity`, or `grok`.
* Confirm that `Enabled` is `true`, `Auth` is `ready`, and the displayed organization and destination project are correct.
* If `Auth` is not `ready`, run `bt login --oauth`, then run the doctor command again.
* If the organization or destination project is wrong, rerun `bt trace enable` for that agent with the correct `--org` and `--project` values.
If `bt trace import` cannot find a Codex, Claude Code, or Antigravity session:
* Confirm that the command names the agent that created the session. Use `bt trace import codex` for Codex, `bt trace import claude` for Claude Code, or `bt trace import antigravity` for Antigravity.
* Run the command from the same user account and computer where that agent saved the session.
If **Send to Slack** does not appear in the **Destination** menu:
* [Enable the Slack integration](/docs/admin/organizations#enable-slack-integration) for the organization.
* Return to the automation and add the destination.
If **Coding Agent Improvements Discovery** fails because it cannot call `gpt-6.1-sol`:
* Configure `gpt-6.1-sol` for the OpenAI provider under [AI providers](/docs/admin/ai-providers), or select an available model in the automation editor.
* Run the automation again.
## Cleanup
No action is required to keep collecting traces. Choose any cleanup actions you need:
After pushing the template, run `rm /tmp/active-observability-config.json`.
Rerun the appropriate `bt trace enable` command for each configured coding agent with the new `--org` and `--project` values.
Run `bt trace disable ` for each configured agent, replacing `` with `claude`, `codex`, `opencode`, `pi`, `antigravity`, or `grok`. This removes the Braintrust tracing integration and its saved configuration.
First stop tracing or send future traces to another project. Then follow the instructions below if you no longer need the project's data.
Deleting the project permanently removes its traces, facets, automations,
and other project data. This action cannot be undone.
1. Go to ** Settings** > [** General**](https://www.braintrust.dev/app/~/configuration/general).
2. Click **Delete project**.
3. Enter `coding-agent-insights` and click **Delete**.
## Recap
You've set up Coding Agent Insights for your team. The Coding Agent Insights dashboard shows who is sending traces and where usage and cost go. Topics classifies each session, and Loop reviews the evidence each day and creates or updates patterns automatically.
Use those findings as a starting point for improving your agent instructions, skills, tools, and configuration. As your team's needs change, you can refine the analysis and share reports in Slack or through a webhook.
## Next steps
* Learn what each grouping in the [Coding Agent Insights dashboard](/docs/observe/dashboards/coding-agent-insights) shows.
* Learn how to [review and act on patterns](/docs/observe/patterns/review).
* [Configure Loop automations](/docs/loop/automations) to change their schedule, permissions, and destinations.
* [Create and test Topics facets](/docs/observe/topics/custom-facets) for other problems your team wants to track.
* Explore the [`bt` CLI](/docs/reference/cli/quickstart) commands for querying and managing project data.
# Evaluating voice AI agents with Evalion
Source: https://braintrust.dev/docs/cookbook/recipes/EvalionVoiceAgentEval
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/EvalionVoiceAgentEval/EvalionVoiceAgentEvaluation.ipynb) by [Marc Vergara Ferrer](https://www.linkedin.com/in/marc-vergara-b72472144/), [Miguel Andres](https://www.linkedin.com/in/gueles/) on 2024-12-05
[Evalion](https://www.evalion.ai) is a voice-agent evaluation platform that simulates real user interactions and normalizes results across scenarios, enabling teams to detect regressions, compare runs over time, and validate an agent’s readiness for production. Their platform enables teams to test voice agents by creating autonomous testing agents that conduct realistic conversations: interrupting mid-sentence, changing their mind, and expressing frustration just like real customers.
This cookbook demonstrates how to evaluate voice agents by combining Evalion's simulation capabilities with Braintrust. Voice agents require assessment beyond simple text accuracy: they must handle real-time latency constraints (\< 500ms responses), manage interruptions gracefully, maintain context across multi-turn conversations, and deliver natural-sounding interactions.
By the end of this guide, you'll learn how to:
* Create test scenarios in Braintrust datasets
* Orchestrate automated voice simulations with Evalion's API
* Extract and normalize voice-specific metrics (latency, CSAT, goal completion)
* Track evaluation results across iterations
## Prerequisites
* A [Braintrust account](https://www.braintrust.dev/signup) and [API key](https://www.braintrust.dev/app/settings?subroute=api-keys)
* Evalion backend access with API credentials
* Python 3.8+
## Getting started
Export your API keys to your environment:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
export BRAINTRUST_API_KEY="YOUR_BRAINTRUST_API_KEY"
export EVALION_API_TOKEN="YOUR_EVALION_API_TOKEN"
export EVALION_PROJECT_ID="YOUR_EVALION_PROJECT_ID"
export EVALION_PERSONA_ID="YOUR_EVALION_PERSONA_ID"
```
Install the required packages:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
pip install braintrust httpx pydantic
```
Best practice is to export your API key as an environment variable. However, to make it easier to follow along with this cookbook, you can also hardcode it into the code below.
Import the required libraries and set up your API credentials:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import os
import asyncio
import json
import time
import uuid
from typing import Any, Dict, List, Optional
import httpx
import nest_asyncio
from braintrust import init_dataset, EvalAsync, Score
# Uncomment to hardcode your API keys
# os.environ["BRAINTRUST_API_KEY"] = "YOUR_BRAINTRUST_API_KEY"
# os.environ["EVALION_API_TOKEN"] = "YOUR_EVALION_API_TOKEN"
# os.environ["EVALION_PROJECT_ID"] = "YOUR_EVALION_PROJECT_ID"
# os.environ["EVALION_PERSONA_ID"] = "YOUR_EVALION_PERSONA_ID"
BRAINTRUST_API_KEY = os.getenv("BRAINTRUST_API_KEY", "")
EVALION_API_TOKEN = os.getenv("EVALION_API_TOKEN", "")
EVALION_PROJECT_ID = os.getenv("EVALION_PROJECT_ID", "")
EVALION_PERSONA_ID = os.getenv("EVALION_PERSONA_ID", "")
nest_asyncio.apply()
```
## Creating test scenarios
We'll create test scenarios for an airline customer service agent. Each scenario includes the customer's situation (input) and success criteria (expected outcome). These range from straightforward bookings to high-stress cancellation handling. We'll add all the scenarios to a dataset in Braintrust.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Initialize Braintrust
project_name = "Voice Agent Evaluation"
dataset_name = "Customer Service Scenarios"
# Create test scenarios
test_scenarios = [
{
"input": "Customer calling to book a flight from New York to Los Angeles for next Tuesday. They want a morning flight and have a budget of $400.",
"expected": [
"Agent introduces themselves professionally",
"Agent confirms the departure city (New York) and destination (Los Angeles)",
"Agent confirms the date (next Tuesday)",
"Agent asks about preferred time of day (morning)",
"Agent presents available flight options within budget",
"Agent confirms the booking details before finalizing",
],
},
{
"input": "Frustrated customer calling because their flight was cancelled. They need to get to Chicago for an important business meeting tomorrow morning.",
"expected": [
"Agent shows empathy for the situation",
"Agent apologizes for the inconvenience",
"Agent asks for booking reference number",
"Agent proactively searches for alternative flights",
"Agent offers multiple rebooking options",
"Agent provides compensation information if applicable",
],
},
{
"input": "Customer wants to change their existing reservation to add extra baggage and select a window seat.",
"expected": [
"Agent asks for booking confirmation number",
"Agent retrieves existing reservation details",
"Agent explains baggage fees and options",
"Agent checks seat availability",
"Agent confirms changes and new total cost",
"Agent sends confirmation of modifications",
],
},
]
# Create dataset
dataset = init_dataset(project_name, dataset_name)
# Insert test scenarios
for scenario in test_scenarios:
dataset.insert(**scenario)
```
## Creating scorers
Evalion provides objective metrics (latency, duration) and subjective assessments (CSAT, clarity). We'll normalize all scores to 0-1 for consistent tracking in Braintrust.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
def normalize_score(
score_value: Optional[float], has_succeeded: Optional[bool] = None
) -> Optional[float]:
"""Normalize scores to 0-1 range."""
if has_succeeded is not None:
return 1.0 if has_succeeded else 0.0
if score_value is None:
return None
# Normalize 1-10 scale to 0-1
return max(0.0, min(1.0, score_value / 10.0))
def extract_custom_metrics(output: Dict[str, Any]) -> List[Score]:
"""Extract custom metric scores from simulation results."""
scores = []
simulations = output.get("simulations", [])
if not simulations:
return scores
simulation = simulations[0]
evaluations = simulation.get("evaluations", [])
for evaluation in evaluations:
if not evaluation.get("is_applicable", True):
continue
metric = evaluation.get("metric", {})
metric_name = metric.get("name", "unknown")
measurement_type = metric.get("measurement_type")
if measurement_type == "boolean":
score_value = normalize_score(None, evaluation.get("has_succeeded"))
else:
score_value = normalize_score(evaluation.get("score"))
if score_value is not None:
scores.append(
Score(
name=metric_name,
score=score_value,
metadata={
"reasoning": evaluation.get("reasoning"),
"improvement_suggestions": evaluation.get(
"improvement_suggestions"
),
},
)
)
return scores
def extract_builtin_metrics(output: Dict[str, Any]) -> List[Score]:
"""Extract builtin metric scores from simulation results."""
scores = []
simulations = output.get("simulations", [])
if not simulations:
return scores
simulation = simulations[0]
builtin_evaluations = simulation.get("builtin_evaluations", [])
for evaluation in builtin_evaluations:
if not evaluation.get("is_applicable", True):
continue
builtin_metric = evaluation.get("builtin_metric", {})
metric_name = builtin_metric.get("name", "unknown")
measurement_type = builtin_metric.get("measurement_type")
# Handle latency specially
if metric_name == "avg_latency":
latency_ms = evaluation.get("score")
if latency_ms is None:
continue
# Score based on distance from 1500ms target
target_latency = 1500
if latency_ms <= target_latency:
normalized_score = 1.0
else:
normalized_score = max(
0.0, 1.0 - (latency_ms - target_latency) / target_latency
)
scores.append(
Score(
name="avg_latency_ms",
score=normalized_score,
metadata={
"actual_latency_ms": latency_ms,
"target_latency_ms": target_latency,
"is_within_target": latency_ms <= target_latency,
},
)
)
continue
if measurement_type == "boolean":
score_value = normalize_score(None, evaluation.get("has_succeeded"))
else:
score_value = normalize_score(evaluation.get("score"))
if score_value is not None:
scores.append(
Score(
name=metric_name,
score=score_value,
metadata={"reasoning": evaluation.get("reasoning")},
)
)
return scores
```
## Evalion API integration
The `EvalionAPIService` class handles all interactions with Evalion's API for creating agents, test setups, and running simulations. The task function orchestrates the workflow: creating agents in Evalion, running simulations, and extracting results. This enables reproducible evaluation across iterations.
The function performs the following steps:
1. Creates a hosted agent in Evalion with your prompt
2. Sets up test scenarios and personas
3. Runs the voice simulation
4. Polls for completion and retrieves results
5. Cleans up temporary resources
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
class EvalionAPIService:
"""Service class for interacting with the Evalion API."""
def __init__(
self, base_url: str = "https://api.evalion.ai/api/v1", api_token: str = ""
):
self.base_url = base_url
self.headers = {"Authorization": f"Bearer {api_token}"}
async def create_hosted_agent(
self, prompt: str, name: Optional[str] = None
) -> Dict[str, Any]:
"""Create a hosted agent with the given prompt."""
if not name:
name = f"Voice Agent - {uuid.uuid4()}"
payload = {
"name": name,
"description": "Agent created for evaluation",
"agent_type": "outbound",
"prompt": prompt,
"is_active": True,
"speaks_first": False,
"llm_provider": "openai",
"llm_model": "gpt-4o-mini",
"llm_temperature": 0.7,
"tts_provider": "elevenlabs",
"tts_model": "eleven_turbo_v2_5",
"tts_voice": "5IDdqnXnlsZ1FCxoOFYg",
"stt_provider": "openai",
"stt_model": "gpt-4o-mini-transcribe",
"language": "en",
"max_conversation_time_in_minutes": 5,
"llm_max_tokens": 800,
}
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/hosted-agents",
headers=self.headers,
json=payload,
)
response.raise_for_status()
return response.json()
async def delete_hosted_agent(self, hosted_agent_id: str) -> None:
"""Delete a hosted agent."""
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.delete(
f"{self.base_url}/hosted-agents/{hosted_agent_id}",
headers=self.headers,
)
response.raise_for_status()
async def create_agent(
self,
project_id: str,
hosted_agent_id: str,
prompt: str,
name: Optional[str] = None,
) -> Dict[str, Any]:
"""Create an agent that references a hosted agent."""
if not name:
name = f"Test Agent {int(time.time())}"
payload = {
"name": name,
"description": "Agent for evaluation testing",
"agent_type": "inbound",
"interaction_mode": "voice",
"integration_type": "phone",
"language": "en",
"speaks_first": False,
"prompt": prompt,
"is_active": True,
"hosted_agent_id": hosted_agent_id,
"project_id": project_id,
}
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/projects/{project_id}/agents",
headers=self.headers,
json=payload,
)
response.raise_for_status()
return response.json()
async def delete_agent(self, project_id: str, agent_id: str) -> None:
"""Delete an agent."""
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.delete(
f"{self.base_url}/projects/{project_id}/agents/{agent_id}",
headers=self.headers,
)
response.raise_for_status()
async def create_test_set(
self, project_id: str, name: Optional[str] = None
) -> Dict[str, Any]:
"""Create a test set."""
if not name:
name = f"Test Set {int(time.time())}"
payload = {
"name": name,
"description": "Test set for evaluation",
"project_id": project_id,
}
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/projects/{project_id}/test-sets",
headers=self.headers,
json=payload,
)
response.raise_for_status()
return response.json()
async def delete_test_set(self, project_id: str, test_set_id: str) -> None:
"""Delete a test set."""
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.delete(
f"{self.base_url}/projects/{project_id}/test-sets/{test_set_id}",
headers=self.headers,
)
response.raise_for_status()
async def create_test_case(
self, project_id: str, test_set_id: str, scenario: str, expected_outcome: str
) -> Dict[str, Any]:
"""Create a test case."""
payload = {
"name": f"Test Case {int(time.time())}",
"description": "Test case for evaluation",
"scenario": scenario,
"expected_outcome": expected_outcome,
"test_set_id": test_set_id,
}
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/projects/{project_id}/test-cases",
headers=self.headers,
json=payload,
)
response.raise_for_status()
return response.json()
async def create_test_setup(
self,
project_id: str,
agent_id: str,
persona_id: str,
test_set_id: str,
metrics: Optional[List[str]] = None,
) -> Dict[str, Any]:
"""Create a test setup."""
payload = {
"name": f"Test Setup {int(time.time())}",
"description": "Test setup for evaluation",
"project_id": project_id,
"agents": [agent_id],
"personas": [persona_id],
"test_sets": [test_set_id],
"metrics": metrics or [],
"testing_mode": "voice",
}
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/test-setups",
headers=self.headers,
json=payload,
)
response.raise_for_status()
return response.json()
async def delete_test_setup(self, project_id: str, test_setup_id: str) -> None:
"""Delete a test setup."""
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.delete(
f"{self.base_url}/test-setups/{test_setup_id}?project_id={project_id}",
headers=self.headers,
)
response.raise_for_status()
async def run_test_setup(self, project_id: str, test_setup_id: str) -> str:
"""Prepare and run a test setup."""
# First prepare
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/test-setup-runs/prepare",
headers=self.headers,
json={"project_id": project_id, "test_setup_id": test_setup_id},
)
response.raise_for_status()
test_setup_run_id = response.json()["test_setup_run_id"]
# Then run
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.post(
f"{self.base_url}/test-setup-runs/{test_setup_run_id}/run",
headers=self.headers,
json={"project_id": project_id},
)
response.raise_for_status()
return test_setup_run_id
async def poll_for_completion(
self, project_id: str, test_setup_run_id: str, max_wait: int = 600
) -> Optional[Dict[str, Any]]:
"""Poll for simulation completion."""
start_time = time.time()
while time.time() - start_time < max_wait:
async with httpx.AsyncClient(timeout=300.0) as client:
response = await client.get(
f"{self.base_url}/test-setup-runs/{test_setup_run_id}/simulations",
headers=self.headers,
params={"project_id": project_id},
)
if response.status_code == 200:
data = response.json()
simulations = data.get("data", [])
if simulations:
sim = simulations[0]
status = sim.get("run_status")
if status in ["COMPLETED", "FAILED"]:
return sim
await asyncio.sleep(5)
return None
```
Then, we'll define the agent prompt that will be evaluated:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Define the agent prompt to evaluate
AGENT_PROMPT = """
You are a professional travel agent assistant. Your role is to help customers with:
- Booking flights
- Modifying existing reservations
- Handling cancellations and rebooking
- Answering questions about flights and policies
Guidelines:
- Always introduce yourself at the beginning of the call
- Be empathetic, especially with frustrated customers
- Confirm all details before making changes
- Provide clear pricing information
- Thank the customer at the end of the call
"""
```
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
async def run_evaluation_task(input: Dict[str, Any] | str) -> Dict[str, Any]:
"""Main task function that orchestrates the evaluation workflow."""
# Extract scenario and expected outcome from input
if isinstance(input, dict):
scenario = input.get("scenario", "")
expected_list = input.get("expected", [])
expected_outcome = (
"\n".join(expected_list)
if isinstance(expected_list, list)
else str(expected_list)
)
elif isinstance(input, str):
scenario = input
expected_outcome = ""
# Initialize Evalion API service
api_service = EvalionAPIService(
base_url="https://api.evalion.ai/api/v1", api_token=EVALION_API_TOKEN
)
# Store resource IDs for cleanup
hosted_agent_id = None
agent_id = None
test_set_id = None
test_setup_id = None
try:
# Create hosted agent
hosted_agent = await api_service.create_hosted_agent(
prompt=AGENT_PROMPT, name="Travel Agent Eval"
)
hosted_agent_id = hosted_agent["id"]
# Create agent
agent = await api_service.create_agent(
project_id=EVALION_PROJECT_ID,
hosted_agent_id=hosted_agent_id,
prompt=AGENT_PROMPT,
)
agent_id = agent["id"]
# Create test set
test_set = await api_service.create_test_set(project_id=EVALION_PROJECT_ID)
test_set_id = test_set["id"]
# Create test case
await api_service.create_test_case(
project_id=EVALION_PROJECT_ID,
test_set_id=test_set_id,
scenario=scenario,
expected_outcome=expected_outcome,
)
# Create test setup
test_setup = await api_service.create_test_setup(
project_id=EVALION_PROJECT_ID,
agent_id=agent_id,
persona_id=EVALION_PERSONA_ID,
test_set_id=test_set_id,
metrics=None,
)
test_setup_id = test_setup["id"]
# Run test setup
test_setup_run_id = await api_service.run_test_setup(
project_id=EVALION_PROJECT_ID, test_setup_id=test_setup_id
)
# Poll for completion
simulation = await api_service.poll_for_completion(
project_id=EVALION_PROJECT_ID, test_setup_run_id=test_setup_run_id
)
# Clean up Evalion resources
if test_setup_id:
await api_service.delete_test_setup(EVALION_PROJECT_ID, test_setup_id)
if agent_id:
await api_service.delete_agent(EVALION_PROJECT_ID, agent_id)
if test_set_id:
await api_service.delete_test_set(EVALION_PROJECT_ID, test_set_id)
if hosted_agent_id:
await api_service.delete_hosted_agent(hosted_agent_id)
if not simulation:
return {"success": False, "error": "Simulation timed out", "transcript": ""}
# Return results
return {
"success": True,
"transcript": simulation.get("transcript", ""),
"simulations": [simulation],
}
except Exception as e:
return {"success": False, "error": str(e), "transcript": ""}
```
Finally, we'll run the evaluation with Braintrust:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Run the evaluation
await EvalAsync(
"Voice Agent Evaluation",
data=dataset,
task=run_evaluation_task,
scores=[
extract_custom_metrics,
extract_builtin_metrics,
],
parameters={
"main": {
"type": "prompt",
"description": "Prompt to be tested by Evalion simulations",
"default": {
"prompt": {
"type": "chat",
"messages": [
{
"role": "system",
"content": AGENT_PROMPT,
}
],
},
"options": {"model": "gpt-4o"},
},
},
},
)
```
## Analyzing results
After running evaluations, navigate to **Experiments** in Braintrust to analyze your results. You'll see metrics like average latency, CSAT scores, and goal completion rates across all test scenarios.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Example of what the results look like
example_results = {
"scenario": "Customer calling to book a flight from New York to Los Angeles",
"scores": {
"Expected Outcome": 0.9,
"conversation_flow": 0.85,
"empathy": 0.92,
"clarity": 0.88,
"avg_latency_ms": 0.95, # 1450ms actual, target 1500ms
},
"metadata": {
"transcript_length": 450,
"duration_seconds": 180,
},
}
print(json.dumps(example_results, indent=2))
```
## Next steps
Now that you have a working evaluation pipeline, you can:
1. **Expand test coverage**: Add more scenarios covering edge cases
2. **Iterate on prompts**: Adjust your agent's prompt and compare results
3. **Monitor production**: Set up online evaluation for live traffic
4. **Track trends**: Use Braintrust's experiment comparison to identify improvements
For more agent cookbooks, check out:
* [Evaluating a voice agent](/docs/cookbook/recipes/VoiceAgent) with OpenAI Realtime API
* [Building reliable AI agents](/docs/cookbook/recipes/AgentWhileLoop) with tool calling
# Evaluating a chat assistant
Source: https://braintrust.dev/docs/cookbook/recipes/EvaluatingChatAssistant
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/EvaluatingChatAssistant/EvaluatingChatAssistant.ipynb) by [Tara Nagar](https://www.linkedin.com/in/taranagar/) on 2024-07-16
## Evaluating a multi-turn chat assistant
This tutorial will walk through using Braintrust to evaluate a conversational, multi-turn chat assistant.
These types of chat bots have become important parts of applications, acting as customer service agents, sales representatives, or travel agents, to name a few. As an owner of such an application, it's important to be sure the bot provides value to the user.
We will expand on this below, but the history and context of a conversation is crucial in being able to produce a good response. If you received a request to "Make a dinner reservation at 7pm" and you knew where, on what date, and for how many people, you could provide some assistance; otherwise, you'd need to ask for more information.
Before starting, please make sure you have a Braintrust account. If you do not have one, you can [sign up here](https://www.braintrust.dev).
## Installing dependencies
Begin by installing the necessary dependencies if you have not done so already.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
pnpm install autoevals braintrust openai
```
## Inspecting the data
Let's take a look at the small dataset prepared for this cookbook. You can find the full dataset in the accompanying [dataset.ts file](https://github.com/braintrustdata/braintrust-cookbook/tree/main/examples/EvaluatingChatAssistant/dataset.ts). The `assistant` turns were generated using `claude-3-5-sonnet-20240620`.
Below is an example of a data point.
* `chat_history` contains the history of the conversation between the user and the assistant
* `input` is the last `user` turn that will be sent in the `messages` argument to the chat completion
* `expected` is the output expected from the chat completion given the input
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import dataset, { ChatTurn } from "./assets/dataset";
console.log(dataset[0]);
```
```
{
chat_history: [
{
role: 'user',
content: "when was the ballon d'or first awarded for female players?"
},
{
role: 'assistant',
content: "The Ballon d'Or for female players was first awarded in 2018. The inaugural winner was Ada Hegerberg, a Norwegian striker who plays for Olympique Lyonnais."
}
],
input: "who won the men's trophy that year?",
expected: "In 2018, the men's Ballon d'Or was awarded to Luka Modrić."
}
```
From looking at this one example, we can see why the history is necessary to provide a helpful response.
If you were asked "Who won the men's trophy that year?" you would wonder *What trophy? Which year?* But if you were also given the `chat_history`, you would be able to answer the question (maybe after some quick research).
## Running experiments
The key to running evals on a multi-turn conversation is to include the history of the chat in the chat completion request.
### Assistant with no chat history
To start, let's see how the prompt performs when no chat history is provided. We'll create a simple task function that returns the output from a chat completion.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { wrapOpenAI } from "braintrust";
import { OpenAI } from "openai";
const experimentData = dataset.map((data) => ({
input: data.input,
expected: data.expected,
}));
console.log(experimentData[0]);
async function runTask(input: string) {
const client = wrapOpenAI(
new OpenAI({
baseURL: "https://api.braintrust.dev/v1/proxy",
apiKey: process.env.OPENAI_API_KEY ?? "", // Can use OpenAI, Anthropic, Mistral, etc. API keys here
}),
);
const response = await client.chat.completions.create({
model: "gpt-4o",
messages: [
{
role: "system",
content:
"You are a helpful and polite assistant who knows about sports.",
},
{
role: "user",
content: input,
},
],
});
return response.choices[0].message.content || "";
}
```
```
{
input: "who won the men's trophy that year?",
expected: "In 2018, the men's Ballon d'Or was awarded to Luka Modrić."
}
```
#### Scoring and running the eval
We'll use the `Factuality` scoring function from the [autoevals library](https://www.braintrust.dev/docs/reference/autoevals) to check how the output of the chat completion compares factually to the expected value.
We will also utilize [trials](/docs/evaluate/advanced-evaluations#run-trials) by including the `trialCount` parameter in the `Eval` call. We expect the output of the chat completion to be non-deterministic, so running each input multiple times will give us a better sense of the "average" output.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { Eval } from "braintrust";
import Factuality from "autoevals";
Eval("Chat assistant", {
experimentName: "gpt-4o assistant - no history",
data: () => experimentData,
task: runTask,
scores: [Factuality],
trialCount: 3,
metadata: {
model: "gpt-4o",
prompt: "You are a helpful and polite assistant who knows about sports.",
},
});
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Experiment gpt - 4o assistant - no history is running at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant%20-%20no%20history
████████████████████████████████████████ | Chat assistant[experimentName = gpt - 4o... | 100 % | 15 / 15 datapoints
=========================SUMMARY=========================
61.33% 'Factuality' score (0 improvements, 0 regressions)
4.12s 'duration' (0 improvements, 0 regressions)
0.01$ 'estimated_cost' (0 improvements, 0 regressions)
See results for gpt-4o assistant - no history at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant%20-%20no%20history
```
61.33% Factuality score? Given what we discussed earlier about chat history being important in producing a good response, that's surprisingly high. Let's log onto [braintrust.dev](https://www.braintrust.dev) and take a look at how we got that score.
#### Interpreting the results
If we look at the score distribution chart, we can see ten of the fifteen examples scored at least 60%, with over half even scoring 100%. If we look into one of the examples with 100% score, we see the output of the chat completion request is asking for more context as we would expect:
`Could you please specify which athlete or player you're referring to? There are many professional athletes, and I'll need a bit more information to provide an accurate answer.`
This aligns with our expectation, so let's now look at how the score was determined.
Click into the scoring trace, we see the chain of thought reasoning used to settle on the score. The model chose `(E) The answers differ, but these differences don't matter from the perspective of factuality.` which is *technically* correct, but we want to penalize the chat completion for not being able to produce a good response.
#### Improve scoring with a custom scorer
While Factuality is a good general purpose scorer, for our use case option (E) is not well aligned with our expectations. The best way to work around this is to customize the scoring function so that it produces a lower score for asking for more context or specificity.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { LLMClassifierFromSpec, Score } from "autoevals";
function Factual(args: {
input: string;
output: string;
expected: string;
}): Score | Promise {
const factualityScorer = LLMClassifierFromSpec("Factuality", {
prompt: `You are comparing a submitted answer to an expert answer on a given question. Here is the data:
[BEGIN DATA]
************
[Question]: {{{input}}}
************
[Expert]: {{{expected}}}
************
[Submission]: {{{output}}}
************
[END DATA]
Compare the factual content of the submitted answer with the expert answer. Ignore any differences in style, grammar, or punctuation.
The submitted answer may either be a subset or superset of the expert answer, or it may conflict with it. Determine which case applies. Answer the question by selecting one of the following options:
(A) The submitted answer is a subset of the expert answer and is fully consistent with it.
(B) The submitted answer is a superset of the expert answer and is fully consistent with it.
(C) The submitted answer contains all the same details as the expert answer.
(D) There is a disagreement between the submitted answer and the expert answer.
(E) The answers differ, but these differences don't matter from the perspective of factuality.
(F) The submitted answer asks for more context, specifics or clarification but provides factual information consistent with the expert answer.
(G) The submitted answer asks for more context, specifics or clarification but does not provide factual information consistent with the expert answer.`,
choice_scores: {
A: 0.4,
B: 0.6,
C: 1,
D: 0,
E: 1,
F: 0.2,
G: 0,
},
});
return factualityScorer(args);
}
```
You can see the built-in Factuality prompt [here](https://github.com/braintrustdata/autoevals/blob/main/templates/factuality.yaml). For our customized scorer, we've added two score choices to that prompt:
```
- (F) The submitted answer asks for more context, specifics or clarification but provides factual information consistent with the expert answer.
- (G) The submitted answer asks for more context, specifics or clarification but does not provide factual information consistent with the expert answer.
```
These will score (F) = 0.2 and (G) = 0 so the model gets some credit if there was any context it was able to gather from the user's input.
We can then use this spec and the `LLMClassifierFromSpec` function to create our customer scorer to use in the eval function.
Read more about [defining your own scorers](/docs/evaluate/write-scorers#scorers) in the documentation.
#### Re-running the eval
Let's now use this updated scorer and run the experiment again.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Eval("Chat assistant", {
experimentName: "gpt-4o assistant - no history",
data: () =>
dataset.map((data) => ({ input: data.input, expected: data.expected })),
task: runTask,
scores: [Factual],
trialCount: 3,
metadata: {
model: "gpt-4o",
prompt: "You are a helpful and polite assistant who knows about sports.",
},
});
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Experiment gpt - 4o assistant - no history - 934e5ca2 is running at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant%20-%20no%20history-934e5ca2
████████████████████████████████████████ | Chat assistant[experimentName = gpt - 4o... | 100 % | 15 / 15 datapoints
=========================SUMMARY=========================
gpt-4o assistant - no history-934e5ca2 compared to gpt-4o assistant - no history:
6.67% (-54.67%) 'Factuality' score (0 improvements, 5 regressions)
4.77s 'duration' (2 improvements, 3 regressions)
0.01$ 'estimated_cost' (2 improvements, 3 regressions)
See results for gpt-4o assistant - no history-934e5ca2 at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant%20-%20no%20history-934e5ca2
```
6.67% as a score aligns much better with what we expected. Let's look again into the results of this experiment.
#### Interpreting the results
In the table we can see the `output` fields in which the chat completion responses are requesting more context. In one of the experiment that had a non-zero score, we can see that the model asked for some clarification, but was able to understand from the question that the user was inquiring about a controversial World Series. Nice!
Looking into how the score was determined, we can see that the factual information aligned with the expert answer but the submitted answer still asks for more context, resulting in a score of 20% which is what we expect.
### Assistant with chat history
Now let's shift and see how providing the chat history improves the experiment.
#### Update the data, task function and scorer function
We need to edit the inputs to the `Eval` function so we can pass the chat history to the chat completion request.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
const experimentData = dataset.map((data) => ({
input: { input: data.input, chat_history: data.chat_history },
expected: data.expected,
}));
console.log(experimentData[0]);
async function runTask({
input,
chat_history,
}: {
input: string;
chat_history: ChatTurn[];
}) {
const client = wrapOpenAI(
new OpenAI({
baseURL: "https://api.braintrust.dev/v1/proxy",
apiKey: process.env.OPENAI_API_KEY ?? "", // Can use OpenAI, Anthropic, Mistral, etc. API keys here
}),
);
const response = await client.chat.completions.create({
model: "gpt-4o",
messages: [
{
role: "system",
content:
"You are a helpful and polite assistant who knows about sports.",
},
...chat_history,
{
role: "user",
content: input,
},
],
});
return response.choices[0].message.content || "";
}
function Factual(args: {
input: {
input: string;
chat_history: ChatTurn[];
};
output: string;
expected: string;
}): Score | Promise {
const factualityScorer = LLMClassifierFromSpec("Factuality", {
prompt: `You are comparing a submitted answer to an expert answer on a given question. Here is the data:
[BEGIN DATA]
************
[Question]: {{{input}}}
************
[Expert]: {{{expected}}}
************
[Submission]: {{{output}}}
************
[END DATA]
Compare the factual content of the submitted answer with the expert answer. Ignore any differences in style, grammar, or punctuation.
The submitted answer may either be a subset or superset of the expert answer, or it may conflict with it. Determine which case applies. Answer the question by selecting one of the following options:
(A) The submitted answer is a subset of the expert answer and is fully consistent with it.
(B) The submitted answer is a superset of the expert answer and is fully consistent with it.
(C) The submitted answer contains all the same details as the expert answer.
(D) There is a disagreement between the submitted answer and the expert answer.
(E) The answers differ, but these differences don't matter from the perspective of factuality.
(F) The submitted answer asks for more context, specifics or clarification but provides factual information consistent with the expert answer.
(G) The submitted answer asks for more context, specifics or clarification but does not provide factual information consistent with the expert answer.`,
choice_scores: {
A: 0.4,
B: 0.6,
C: 1,
D: 0,
E: 1,
F: 0.2,
G: 0,
},
});
return factualityScorer(args);
}
```
```
{
input: {
input: "who won the men's trophy that year?",
chat_history: [ [Object], [Object] ]
},
expected: "In 2018, the men's Ballon d'Or was awarded to Luka Modrić."
}
```
We update the parameter to the task function to accept both the `input` string and the `chat_history` array and add the `chat_history` into the messages array in the chat completion request, done here using the spread `...` syntax.
We also need to update the `experimentData` and `Factual` function parameters to align with these changes.
#### Running the eval
Use the updated variables and functions to run a new eval.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Eval("Chat assistant", {
experimentName: "gpt-4o assistant",
data: () => experimentData,
task: runTask,
scores: [Factual],
trialCount: 3,
metadata: {
model: "gpt-4o",
prompt: "You are a helpful and polite assistant who knows about sports.",
},
});
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Experiment gpt - 4o assistant is running at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant
████████████████████████████████████████ | Chat assistant[experimentName = gpt - 4o... | 100 % | 15 / 15 datapoints
=========================SUMMARY=========================
gpt-4o assistant compared to gpt-4o assistant - no history-934e5ca2:
60.00% 'Factuality' score (0 improvements, 0 regressions)
4.34s 'duration' (0 improvements, 0 regressions)
0.01$ 'estimated_cost' (0 improvements, 0 regressions)
See results for gpt-4o assistant at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant
```
60% score is a definite improvement from 4%.
You'll notice that it says there were 0 improvements and 0 regressions compared to the last experiment `gpt-4o assistant - no history-934e5ca2` we ran. This is because by default, Braintrust uses the `input` field to match rows across experiments. From the dashboard, we can customize the comparison key ([see docs](/docs/evaluate/compare-experiments#set-a-comparison-key)) by going to the [project configuration page](https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/configuration).
#### Update experiment comparison for diff mode
Let's go back to the dashboard.
For this cookbook, we can use the `expected` field as the comparison key because this field is unique in our small dataset.
In the Configuration tab, go to the bottom of the page to update the comparison key:
#### Interpreting the results
Turn on diff mode using the toggle on the upper right of the table.
Since we updated the comparison key, we can now see the improvements in the Factuality score between the experiment run with chat history and the most recent one run without for each of the examples. If we also click into a trace, we can see the change in input parameters that we made above where it went from a `string` to an object with `input` and `chat_history` fields.
All of our rows scored 60% in this experiment. If we look into each trace, this means the submitted answer includes all the details from the expert answer with some additional information.
60% is an improvement from the previous run, but we can do better. Since it seems like the chat completion is always returning more than necessary, let's see if we can tweak our prompt to have the output be more concise.
#### Improving the result
Let's update the system prompt used in the chat completion request.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
async function runTask({
input,
chat_history,
}: {
input: string;
chat_history: ChatTurn[];
}) {
const client = wrapOpenAI(
new OpenAI({
baseURL: "https://api.braintrust.dev/v1/proxy",
apiKey: process.env.OPENAI_API_KEY ?? "", // Can use OpenAI, Anthropic, Mistral etc. API keys here
}),
);
const response = await client.chat.completions.create({
model: "gpt-4o",
messages: [
{
role: "system",
content:
"You are a helpful, polite assistant who knows about sports. Only answer the question; don't add additional information outside of what was asked.",
},
...chat_history,
{
role: "user",
content: input,
},
],
});
return response.choices[0].message.content || "";
}
```
In the task function, we'll update the `system` message to specify the output should be precise and then run the eval again.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Eval("Chat assistant", {
experimentName: "gpt-4o assistant - concise",
data: () => experimentData,
task: runTask,
scores: [Factual],
trialCount: 3,
metadata: {
model: "gpt-4o",
prompt:
"You are a helpful, polite assistant who knows about sports. Only answer the question; don't add additional information outside of what was asked.",
},
});
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Experiment gpt - 4o assistant - concise is running at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant%20-%20concise
████████████████████████████████████████ | Chat assistant[experimentName = gpt - 4o... | 100 % | 15 / 15 datapoints
=========================SUMMARY=========================
gpt-4o assistant - concise compared to gpt-4o assistant:
86.67% (+26.67%) 'Factuality' score (4 improvements, 0 regressions)
1.89s 'duration' (5 improvements, 0 regressions)
0.01$ 'estimated_cost' (4 improvements, 1 regressions)
See results for gpt-4o assistant - concise at https://www.braintrust.dev/app/braintrustdata.com/p/Chat%20assistant/experiments/gpt-4o%20assistant%20-%20concise
```
Let's go into the dashboard and see the new experiment.
Success! We got a 27 percentage point increase in factuality, up to an average score of 87% for this experiment with our updated prompt.
### Conclusion
We've seen in this cookbook how to evaluate a chat assistant and visualized how the chat history effects the output of the chat completion. Along the way, we also utilized some other functionality such as updating the comparison key in the diff view and creating a custom scoring function.
Try seeing how you can improve the outputs and scores even further!
# Finding production issues using Topics
Source: https://braintrust.dev/docs/cookbook/recipes/FindingProductionIssues
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/FindingProductionIssues/FindingProductionIssues.mdx) by [Jess Wang](https://www.linkedin.com/in/jesswang/) on 2026-06-01
Your existing evals tell you how well your AI performs on things you already know to test for, but what about the patterns and failure modes that you're not aware of yet?
In this cookbook, we'll build a customer support chatbot that generates hundreds of conversation logs between human and AI, and use [Topics](/docs/observe/topics/index) to surface patterns in the system that we wouldn't have caught with traditional evaluations.
By the end, you'll learn how to:
* Use Topics to classify conversations by task, sentiment, and issues
* Dig into failure clusters to identify specific prompt and system-level bugs
* Build targeted eval datasets from Topics classifications
* Iterate on prompts and measure improvement with offline evals
## Getting started
We're building a customer support chatbot for a fictional e-commerce company called Evergreen Goods. The chatbot uses OpenAI's API to interact with Supabase to look up orders, process refunds, initiate returns, and check shipping status.
You'll need:
* A [Braintrust account](https://www.braintrust.dev/signup) with an API key
* An [OpenAI API key](https://platform.openai.com/)
* A [Supabase](https://supabase.com/) project with Edge Functions enabled
* Python 3.12+ with the `braintrust`, `openai`, and `requests` packages
Clone the cookbook and install dependencies:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
git clone https://github.com/braintrustdata/braintrust-cookbook.git
cd braintrust-cookbook/examples/FindingProductionIssues/customer-support-bot
pip install -r requirements.txt
```
Set the required environment variables:
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
export OPENAI_API_KEY="your-openai-api-key"
export BRAINTRUST_API_KEY="your-braintrust-api-key"
export SUPABASE_SERVICE_ROLE_KEY="your-supabase-key"
```
### Setup and code explanation
The Supabase database has six tables:
* **customers**: customer profiles with loyalty tiers (Bronze, Silver, Gold) and point balances
* **products**: 15 outdoor/lifestyle products with SKUs, prices, sizes, and colors
* **orders**: orders with statuses (shipped, delivered, cancelled, etc.), tracking numbers, and payment info
* **returns**: return records linked to orders
* **refund\_requests**: refund records with processor responses
* **support\_tickets**: escalation tickets created by the bot
Our customer support chatbot has two types of tools:
**Read tools** query the Supabase REST API (PostgREST) directly:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
def lookup_order(order_id: str) -> dict | None:
resp = requests.get(
f"{SUPABASE_URL}/rest/v1/orders",
headers={**_REST_HEADERS, "Accept": "application/vnd.pgrst.object+json"},
params={"id": f"eq.{order_id}", "select": "*"},
timeout=10,
)
if resp.status_code == 406 or resp.status_code == 404:
return None
resp.raise_for_status()
return resp.json()
```
**Action tools** call Supabase Edge Functions on services like payment processors, shipping carriers, and label services.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
def _call_edge_function(function_name: str, payload: dict) -> dict:
url = f"{SUPABASE_URL}/functions/v1/{function_name}"
resp = requests.post(url, json=payload, timeout=15)
return {"status_code": resp.status_code, **resp.json()}
```
The complete tool implementations are in [`supabase_tools.py`](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/FindingProductionIssues/customer-support-bot/supabase_tools.py) and the chatbot itself is in [`chat_app.py`](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/FindingProductionIssues/customer-support-bot/chat_app.py).
And this is our initial system prompt:
```
You are a helpful customer support agent for an e-commerce company
called Evergreen Goods.
POLICIES:
- Return window: 30 days from delivery
- Refund processing: 5-7 business days
- Exchanges: Same item, different size/color. Subject to availability.
- Damaged items: Full refund or replacement. No return shipping required.
- Shipping: Standard (free over $50), Express ($12.99-$14.99), Overnight ($24.99)
INSTRUCTIONS:
- Always use the provided tools to look up real data before answering.
- Never fabricate order numbers, tracking numbers, prices, or dates.
- If a customer provides an order number or email, look it up first.
- When a customer wants a refund, USE the process_refund tool.
- When a customer wants to return an item, USE the initiate_return tool.
- If you cannot resolve an issue, USE escalate_to_human.
- Be empathetic but efficient. Take action when you can.
```
## Generating conversation logs
The log generator runs 51 scripted scenarios across six categories, with optional generated follow-up turns:
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
SCENARIOS = [
# Shipping
("Where is my order ORD-4002? It's been 5 days and still no delivery.",
["Can you give me a more specific ETA?"], "shipping"),
# Returns
("I need to return ORD-4007. I ordered a size S but I actually need a M.",
["Can I do an exchange instead of a refund?"], "returns"),
# Billing
("I applied a promo code on ORD-4005 but wasn't given the discount.",
["The code was SUMMER20."], "billing"),
# Account
("My email is sarah.chen@example.com. Can you check my loyalty points balance?",
["When do my Gold perks expire?"], "account"),
# Product
("Does the Waterproof Shell Jacket (SKU-1004) come in size XS?",
[], "product"),
# Orders not in system
("Where is my order ORD-9901? I placed it last week.",
["Can you check by my email? sarah.chen@example.com"], "shipping"),
# ... 51 total scenarios
]
```
The complete scenario list and log generation logic is in [`generate_logs.py`](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/FindingProductionIssues/customer-support-bot/generate_logs.py).
```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
python generate_logs.py # Run this several times
```
Each conversation is logged to Braintrust with full tracing. That means every tool call, LLM response, and each conversational turn is logged to Braintrust.
We run the generator multiple times to build up volume, because Topics needs at least \~100 traces to start clustering effectively. Each run shuffles the scenario order and uses a temperature of 1.0, so the bot's responses vary across runs. That way, the same customer complaint might get a slightly different resolution path each time, which gives Topics more diverse data to cluster against.
## Using Topics to find failures
Once the logs are in Braintrust, go to [** Topics**](https://www.braintrust.dev/app/~/topics). Topics automatically classifies each conversation across three built-in facets:
* **Task**: What was the customer trying to do? In this project, Topics identified clusters like "Product sizing and returns" for customers dealing with fit issues and return requests, "Shipping status inquiries" for customers tracking delayed packages, and "Billing disputes" for promo code and double-charge complaints.
* **Sentiment**: How did the customer feel about the interaction? Topics classified conversations into clusters like "Positive resolution" where the customer's issue was handled well, "Negative emotional distress" where the customer left frustrated, and "Neutral informational" for straightforward product questions.
* **Issues**: Did something go wrong during the conversation? This facet surfaced clusters like "Payment processing failures" where the refund tool returned errors, and "Order lookup errors" where the bot couldn't find an order number in the system.
## Digging into the failure cluster
Filtering to **Sentiment: Negative emotional distress** showed that a majority (32%) of negative conversations were classified under the "Product sizing and returns" task cluster. This is where customers were most frustrated.
At this point, I would recommend spending a good 20-30 minutes manually reading through the logs. Though you could technically use AI to automate this portion of it, it's better to have a human read through and decide which AI responses were fine (even if it evoked negative emotion) versus which ones fundamentally need to be fixed. As I did this myself, I wrote down what needed to be fixed in my system.
**1. Order not found = dead end.** When a customer referenced an order number that wasn't in the system, the bot just said "order not found" and stopped. I think the correct behavior should be to look up the customer's account to find valid order numbers, or find other order numbers that are close to what the customer provided, in case they made a typo.
**2. Shipping delays = no refund.** When a customer paid for express shipping and it took 8 days (well past the 2-3 day window), the bot refused to refund shipping because the order status was "shipped." I think the correct behavior here would be to allow for a full refund on shipping only.
**3. Return + refund race condition.** When processing a return and refund together, the bot would process the refund first, then tell the customer they couldn't return the item because "a refund is in progress."
**4. Promo code delays.** When a customer reported a missing promo discount, the bot correlated the return status of the product with discount eligibility. However, promo discounts are applied at checkout, so delivery status is irrelevant.
**5. 5xx errors = immediate escalation.** When a tool call hit a transient server error, the bot immediately escalated to a human instead of having any sort of retry logic setup.
## Building a targeted eval dataset
From the Topics UI, we saved the conversations in the "Product sizing and returns" cluster to a dataset called **Product Sizing and Return Issues**. This gives us a focused eval set containing exactly the conversations where the bot was failing.
We cleaned the dataset to contain just the customer messages and the turn count so that we could replay it faithfully.
**Before cleaning**:
```json theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
{
"input": [
{"role": "system", "content": "You are a helpful customer support agent..."},
{"role": "user", "content": "I paid for express shipping on ORD-4006 and it's been 8 days."},
{"role": "assistant", "content": "", "tool_calls": [{"function": {"name": "lookup_order"}}]},
{"role": "tool", "content": "{\"status\": \"shipped\", ...}"},
{"role": "assistant", "content": "I see your order is still in transit..."},
{"role": "user", "content": "I want a refund on the shipping cost at minimum."},
{"role": "assistant", "content": "Unfortunately I cannot process a refund..."}
],
"expected": "Unfortunately I cannot process a refund...",
"metadata": {"category": "shipping", "total_turns": 2}
}
```
**After cleaning:**
```json theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
{
"input": "I paid for express shipping on ORD-4006 and it's been 8 days.\n---\nI want a refund on the shipping cost at minimum.",
"metadata": {
"num_customer_messages": 2,
"num_turns": 5
}
}
```
## Running offline evals
The eval script runs each customer complaint through the chatbot with real Supabase tools, then scores the response with three LLM-based scorers. The eval scorers are custom functions we wrote to measure specific dimensions:
* **Helpfulness** checks whether the agent addressed the customer's concern and provided clear next steps.
* **Resolution** checks whether the agent actually took action (processed a refund, initiated a return, or provided tracking) rather than just acknowledging the problem without doing anything.
* **Empathy** checks whether the agent's tone was professional and acknowledged the customer's frustration.
```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
Eval(
"Customer Support Chatbot",
data=load_dataset,
task=lambda input, expected=None, **_: run_support_conversation(
input["customer_input"],
num_customer_messages=input["num_customer_messages"],
),
scores=[helpfulness_scorer, resolution_scorer, empathy_scorer],
experiment_name="product-sizing-returns-v1",
)
```
The complete eval script is in [`eval_product_sizing_returns.py`](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/FindingProductionIssues/customer-support-bot/eval_product_sizing_returns.py).
The v1 baseline scores on the Product Sizing and Return Issues dataset:
| Scorer | v1 Score |
| - | - |
| Helpfulness | 50% |
| Resolution | 62.5% |
| Empathy | 62.5% |
## Fixing the prompt
Based on the five failure patterns from Topics that I outlined above, we updated the system prompt:
```diff theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
POLICIES:
+ - Shipping refunds: If a shipment is delayed beyond the expected
+ delivery window, process a refund for the shipping cost using
+ process_refund. A "shipped" status does NOT block shipping refunds.
+ - Promo codes: If a customer reports a missing promo discount, process
+ a refund for the discount amount. Do NOT check return eligibility or
+ tell them to wait for delivery — it's a billing correction, not a return.
INSTRUCTIONS:
+ - If an order number is not found, try to help: look up the customer's
+ account by email to find valid orders, or ask clarifying questions.
- When a customer wants a refund, USE the process_refund tool.
+ - Do NOT just escalate — process it yourself.
+ - If a tool call fails with a server error (5xx), retry once before
+ giving up.
+ - When you process a return and refund together, confirm both actions
+ in one response. Do not tell the customer they can't return because
+ a refund is in progress.
+ - ONLY use escalate_to_human as a LAST RESORT after all relevant
+ tools have failed.
```
We also fixed the backend: the `process-refund` edge function was updated to allow refunds on `shipped` orders (not just `delivered`). The original edge function had a status check that only permitted refunds when `order.status` was `delivered` or `return_in_progress`, so even when the prompt told the bot to process a shipping refund, the API rejected it with an `INVALID_ORDER_STATUS` error. Adding `shipped` to the allowed statuses unblocked the prompt fix.
## Measuring improvement
Running the same eval dataset with the updated prompt:
| Scorer | v1 | v2 | Change |
| - | - | - | - |
| Helpfulness | 50% | 87.5% | +37.5 |
| Resolution | 62.5% | 75% | +12.5 |
| Empathy | 62.5% | 75% | +12.5 |
Helpfulness nearly doubled. The bot is now actually processing refunds for delayed shipments, handling promo codes as billing corrections, and retrying on transient errors instead of immediately escalating.
## Deploying and monitoring
After validating the v2 prompt with offline evals, we updated the production system prompt and generated a fresh batch of \~200 logs. Topics reclassified the new conversations, and the failure patterns from the original clusters were significantly reduced.
This is the core loop that Topics enables:
1. **Log** production conversations with tracing
2. **Classify** automatically with Topics (task, sentiment, issues)
3. **Identify** failure clusters you didn't know existed
4. **Save** the failing conversations to a dataset
5. **Eval** prompt changes against that dataset
6. **Deploy** and verify with fresh production logs
## Next steps
* Add [custom facets](/docs/observe/topics/custom-facets) to classify dimensions specific to your domain (for example, "Resolution Gap," which checks whether the bot sounded helpful but failed to actually resolve the issue)
* Set up [online scoring](/docs/evaluate/score-online) to get real-time quality signals alongside Topics classifications
* Explore the [Braintrust SDK](/docs/sdks/python/versions/latest) to programmatically query Topics data and build automated alerting
# Improving Github issue titles using their contents
Source: https://braintrust.dev/docs/cookbook/recipes/Github-Issues
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/Github-Issues/Github-Issues.ipynb) by [Ankur Goyal](https://twitter.com/ankrgyl) on 2023-10-29
This tutorial will teach you how to use Braintrust to generate better titles for Github issues, based on their
content. This is a great way to learn how to work with text and evaluate subjective criteria, like summarization quality.
We'll use a technique called **model graded evaluation** to automatically evaluate the newly generated titles
against the original titles, and improve our prompt based on what we find.
Before starting, please make sure that you have a Braintrust account. If you do not, please [sign up](https://www.braintrust.dev). After this tutorial, feel free to dig deeper by visiting [the docs](http://www.braintrust.dev/docs).
## Installing dependencies
To see a list of dependencies, you can view the accompanying [package.json](https://github.com/braintrustdata/braintrust-cookbook/tree/main/examples/Github-Issues/package.json) file. Feel free to copy/paste snippets of this code to run in your environment, or use [tslab](https://github.com/yunabe/tslab) to run the tutorial in a Jupyter notebook.
## Downloading the data
We'll start by downloading some issues from Github using the `octokit` SDK. We'll use the popular open source project [next.js](https://github.com/vercel/next.js).
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { Octokit } from "@octokit/core";
const ISSUES = [
"https://github.com/vercel/next.js/issues/59999",
"https://github.com/vercel/next.js/issues/59997",
"https://github.com/vercel/next.js/issues/59995",
"https://github.com/vercel/next.js/issues/59988",
"https://github.com/vercel/next.js/issues/59986",
"https://github.com/vercel/next.js/issues/59971",
"https://github.com/vercel/next.js/issues/59958",
"https://github.com/vercel/next.js/issues/59957",
"https://github.com/vercel/next.js/issues/59950",
"https://github.com/vercel/next.js/issues/59940",
];
// Octokit.js
// https://github.com/octokit/core.js#readme
const octokit = new Octokit({
auth: process.env.GITHUB_ACCESS_TOKEN || "Your Github Access Token",
});
async function fetchIssue(url: string) {
// parse url of the form https://github.com/supabase/supabase/issues/15534
const [owner, repo, _, issue_number] = url!.trim().split("/").slice(-4);
const data = await octokit.request(
"GET /repos/{owner}/{repo}/issues/{issue_number}",
{
owner,
repo,
issue_number: parseInt(issue_number),
headers: {
"X-GitHub-Api-Version": "2022-11-28",
},
}
);
return data.data;
}
const ISSUE_DATA = await Promise.all(ISSUES.map(fetchIssue));
```
Let's take a look at one of the issues:
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
console.log(ISSUE_DATA[0].title);
console.log("-".repeat(ISSUE_DATA[0].title.length));
console.log(ISSUE_DATA[0].body.substring(0, 512) + "...");
```
```
The instrumentation hook is only called after visiting a route
--------------------------------------------------------------
### Link to the code that reproduces this issue
https://github.com/daveyjones/nextjs-instrumentation-bug
### To Reproduce
\`\`\`shell
git clone git@github.com:daveyjones/nextjs-instrumentation-bug.git
cd nextjs-instrumentation-bug
npm install
npm run dev # The register function IS called
npm run build && npm start # The register function IS NOT called until you visit http://localhost:3000
\`\`\`
### Current vs. Expected behavior
The \`register\` function should be called automatically after running \`npm ...
```
## Generating better titles
Let's try to generate better titles using a simple prompt. We'll use OpenAI, although you could try this out with any model that supports text generation.
We'll start by initializing an OpenAI client and wrapping it with some Braintrust instrumentation. `wrapOpenAI`
is initially a no-op, but later on when we use Braintrust, it will help us capture helpful debugging information about the model's performance.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { wrapOpenAI } from "braintrust";
import { OpenAI } from "openai";
const client = wrapOpenAI(
new OpenAI({
apiKey: process.env.OPENAI_API_KEY || "Your OpenAI API Key",
})
);
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { ChatCompletionMessageParam } from "openai/resources";
function titleGeneratorMessages(content: string): ChatCompletionMessageParam[] {
return [
{
role: "system",
content:
"Generate a new title based on the github issue. Return just the title.",
},
{
role: "user",
content: "Github issue: " + content,
},
];
}
async function generateTitle(input: string) {
const messages = titleGeneratorMessages(input);
const response = await client.chat.completions.create({
model: "gpt-3.5-turbo",
messages,
seed: 123,
});
return response.choices[0].message.content || "";
}
const generatedTitle = await generateTitle(ISSUE_DATA[0].body);
console.log("Original title: ", ISSUE_DATA[0].title);
console.log("Generated title:", generatedTitle);
```
```
Original title: The instrumentation hook is only called after visiting a route
Generated title: Next.js: \`register\` function not automatically called after build and start
```
## Scoring
Ok cool! The new title looks pretty good. But how do we consistently and automatically evaluate whether the new titles are better than the old ones?
With subjective problems, like summarization, one great technique is to use an LLM to grade the outputs. This is known as model graded evaluation. Below, we'll use a [summarization prompt](https://github.com/braintrustdata/autoevals/blob/main/templates/summary.yaml)
from Braintrust's open source [autoevals](https://github.com/braintrustdata/autoevals) library. We encourage you to use these prompts, but also to copy/paste them, modify them, and create your own!
The prompt uses [Chain of Thought](https://arxiv.org/abs/2201.11903) which dramatically improves a model's performance on grading tasks. Later, we'll see how it helps us debug the model's outputs.
Let's try running it on our new title and see how it performs.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { Summary } from "autoevals";
await Summary({
output: generatedTitle,
expected: ISSUE_DATA[0].title,
input: ISSUE_DATA[0].body,
// In practice we've found gpt-4 class models work best for subjective tasks, because
// they are great at following criteria laid out in the grading prompts.
model: "gpt-4-1106-preview",
});
```
```
{
name: 'Summary',
score: 1,
metadata: {
rationale: "Summary A ('The instrumentation hook is only called after visiting a route') is a partial and somewhat ambiguous statement. It does not specify the context of the 'instrumentation hook' or the technology involved.\n" +
"Summary B ('Next.js: \`register\` function not automatically called after build and start') provides a clearer and more complete description. It specifies the technology ('Next.js') and the exact issue ('\`register\` function not automatically called after build and start').\n" +
'The original text discusses an issue with the \`register\` function in a Next.js application not being called as expected, which is directly reflected in Summary B.\n' +
"Summary B also aligns with the section 'Current vs. Expected behavior' from the original text, which states that the \`register\` function should be called automatically but is not until a route is visited.\n" +
"Summary A lacks the detail that the issue is with the Next.js framework and does not mention the expectation of the \`register\` function's behavior, which is a key point in the original text.",
choice: 'B'
},
error: undefined
}
```
## Initial evaluation
Now that we have a way to score new titles, let's run an eval and see how our prompt performs across all 10 issues.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { Eval, login } from "braintrust";
login({ apiKey: process.env.BRAINTUST_API_KEY || "Your Braintrust API Key" });
await Eval("Github Issues Cookbook", {
data: () =>
ISSUE_DATA.map((issue) => ({
input: issue.body,
expected: issue.title,
metadata: issue,
})),
task: generateTitle,
scores: [
async ({ input, output, expected }) =>
Summary({
input,
output,
expected,
model: "gpt-4-1106-preview",
}),
],
});
console.log("Done!");
```
```
{
projectName: 'Github Issues Cookbook',
experimentName: 'main-1706774628',
projectUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Github%20Issues%20Cookbook',
experimentUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Github%20Issues%20Cookbook/main-1706774628',
comparisonExperimentName: undefined,
scores: undefined,
metrics: undefined
}
```
```
████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ | Github Issues Cookbook | 10% | 10/100 datapoints
```
```
Done!
```
Great! We got an initial result. If you follow the link, you'll see an eval result showing an initial score of 40%.
## Debugging failures
Let's dig into a couple examples to see what's going on. Thanks to the instrumentation we added earlier, we can see the model's reasoning for its scores.
Issue [https://github.com/vercel/next.js/issues/59995](https://github.com/vercel/next.js/issues/59995):
Issue [https://github.com/vercel/next.js/issues/59986](https://github.com/vercel/next.js/issues/59986):
## Improving the prompt
Hmm, it looks like the model is missing certain key details. Let's see if we can improve our prompt to encourage the model to include more details, without being too verbose.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
function titleGeneratorMessages(content: string): ChatCompletionMessageParam[] {
return [
{
role: "system",
content: `Generate a new title based on the github issue. The title should include all of the key
identifying details of the issue, without being longer than one line. Return just the title.`,
},
{
role: "user",
content: "Github issue: " + content,
},
];
}
async function generateTitle(input: string) {
const messages = titleGeneratorMessages(input);
const response = await client.chat.completions.create({
model: "gpt-3.5-turbo",
messages,
seed: 123,
});
return response.choices[0].message.content || "";
}
```
### Re-evaluating
Now that we've tweaked our prompt, let's see how it performs by re-running our eval.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await Eval("Github Issues Cookbook", {
data: () =>
ISSUE_DATA.map((issue) => ({
input: issue.body,
expected: issue.title,
metadata: issue,
})),
task: generateTitle,
scores: [
async ({ input, output, expected }) =>
Summary({
input,
output,
expected,
model: "gpt-4-1106-preview",
}),
],
});
console.log("All done!");
```
```
{
projectName: 'Github Issues Cookbook',
experimentName: 'main-1706774676',
projectUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Github%20Issues%20Cookbook',
experimentUrl: 'https://www.braintrust.dev/app/braintrust.dev/p/Github%20Issues%20Cookbook/main-1706774676',
comparisonExperimentName: 'main-1706774628',
scores: {
Summary: {
name: 'Summary',
score: 0.7,
diff: 0.29999999999999993,
improvements: 3,
regressions: 0
}
},
metrics: {
duration: {
name: 'duration',
metric: 0.3292001008987427,
unit: 's',
diff: -0.002199888229370117,
improvements: 7,
regressions: 3
}
}
}
```
```
████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ | Github Issues Cookbook | 10% | 10/100 datapoints
```
```
All done!
```
Wow, with just a simple change, we're able to boost summary performance by 30%!
## Parting thoughts
This is just the start of evaluating and improving this AI application. From here, you should dig into
individual examples, verify whether they legitimately improved, and test on more data. You can even use
[logging](/docs/instrument/trace-application-logic) to capture real-user examples and incorporate
them into your evals.
Happy evaluating!
# Generating beautiful HTML components
Source: https://braintrust.dev/docs/cookbook/recipes/HTMLGenerator
[Contributed](https://github.com/braintrustdata/braintrust-cookbook/blob/main/examples/HTMLGenerator/HTMLGenerator.ipynb) by [Ankur Goyal](https://twitter.com/ankrgyl) on 2024-01-29
In this example, we'll build an app that automatically generates HTML components, evaluates them, and captures user feedback. We'll use the feedback and evaluations to build up a dataset
that we'll use as a basis for further improvements.
## The generator
We'll start by using a very simple prompt to generate HTML components using `gpt-3.5-turbo`.
First, we'll initialize an openai client and wrap it with Braintrust's helper. This is a no-op until we start using
the client within code that is instrumented by Braintrust.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { OpenAI } from "openai";
import { wrapOpenAI } from "braintrust";
const openai = wrapOpenAI(
new OpenAI({
apiKey: process.env.OPENAI_API_KEY || "Your OPENAI_API_KEY",
})
);
```
This code generates a basic prompt:
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { ChatCompletionMessageParam } from "openai/resources";
function generateMessages(input: string): ChatCompletionMessageParam[] {
return [
{
role: "system",
content: `You are a skilled design engineer
who can convert ambiguously worded ideas into beautiful, crisp HTML and CSS.
Your designs value simplicity, conciseness, clarity, and functionality over
complexity.
You generate pure HTML with inline CSS, so that your designs can be rendered
directly as plain HTML. Only generate components, not full HTML pages. Do not
create background colors.
Users will send you a description of a design, and you must reply with HTML,
and nothing else. Your reply will be directly copied and rendered into a browser,
so do not include any text. If you would like to explain your reasoning, feel free
to do so in HTML comments.`,
},
{
role: "user",
content: input,
},
];
}
JSON.stringify(
generateMessages("A login form for a B2B SaaS product."),
null,
2
);
```
```
[
{
"role": "system",
"content": "You are a skilled design engineer\nwho can convert ambiguously worded ideas into beautiful, crisp HTML and CSS.\nYour designs value simplicity, conciseness, clarity, and functionality over\ncomplexity.\n\nYou generate pure HTML with inline CSS, so that your designs can be rendered\ndirectly as plain HTML. Only generate components, not full HTML pages. Do not\ncreate background colors.\n\nUsers will send you a description of a design, and you must reply with HTML,\nand nothing else. Your reply will be directly copied and rendered into a browser,\nso do not include any text. If you would like to explain your reasoning, feel free\nto do so in HTML comments."
},
{
"role": "user",
"content": "A login form for a B2B SaaS product."
}
]
```
Now, let's run this using `gpt-3.5-turbo`. We'll also do a few things that help us log & evaluate this function later:
* Wrap the execution in a `traced` call, which will enable Braintrust to log the inputs and outputs of the function when we run it in production or in evals
* Make its signature accept a single `input` value, which Braintrust's `Eval` function expects
* Use a `seed` so that this test is reproduceable
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { traced } from "braintrust";
async function generateComponent(input: string) {
return traced(
async (span) => {
const response = await openai.chat.completions.create({
model: "gpt-3.5-turbo",
messages: generateMessages(input),
seed: 101,
});
const output = response.choices[0].message.content;
span.log({ input, output });
return output;
},
{
name: "generateComponent",
}
);
}
```
### Examples
Let's look at a few examples!
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await generateComponent("Do a reset password form inside a card.");
```
```
```
To make this easier to validate, we'll use [puppeteer](https://pptr.dev/) to render the HTML as a screenshot.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import puppeteer from "puppeteer";
import * as tslab from "tslab";
async function takeFullPageScreenshotAsUInt8Array(htmlContent) {
const browser = await puppeteer.launch({ headless: "new" });
const page = await browser.newPage();
await page.setContent(htmlContent);
const screenshotBuffer = await page.screenshot();
const uint8Array = new Uint8Array(screenshotBuffer);
await browser.close();
return uint8Array;
}
async function displayComponent(input: string) {
const html = await generateComponent(input);
const img = await takeFullPageScreenshotAsUInt8Array(html);
tslab.display.png(img);
console.log(html);
}
await displayComponent("Do a reset password form inside a card.");
```
```
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await displayComponent("Create a profile page for a social network.");
```
```
John Doe
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Nulla ut turpis
hendrerit, ullamcorper velit in, iaculis arcu.
```
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
await displayComponent(
"Logs viewer for a cloud infrastructure management tool. Heavy use of dark mode."
);
```
```
12:30 PM
Info: Cloud instance created successfully
12:45 PM
Warning: High CPU utilization on instance #123
01:00 PM
Error: Connection lost to the database server
```
## Scoring the results
It looks like in a few of these examples, the model is generating a full HTML page, instead of a component as we requested. This is something we can evaluate, to ensure that it does not happen!
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
const containsHTML = (s) => /<(html|body)>/i.test(s);
containsHTML(
await generateComponent(
"Logs viewer for a cloud infrastructure management tool. Heavy use of dark mode."
)
);
```
```
true
```
Now, let's update our function to compute this score. Let's also keep track of requests and their ids, so that we can provide user feedback. Normally you would store these in a database, but for demo purposes, a global dictionary should suffice.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
// Normally you would store these in a database, but for this demo we'll just use a global variable.
let requests = {};
async function generateComponent(input: string) {
return traced(
async (span) => {
const response = await openai.chat.completions.create({
model: "gpt-3.5-turbo",
messages: generateMessages(input),
seed: 101,
});
const output = response.choices[0].message.content;
requests[input] = span.id;
span.log({
input,
output,
scores: { isComponent: containsHTML(output) ? 0 : 1 },
});
return output;
},
{
name: "generateComponent",
}
);
}
```
## Logging results
To enable logging to Braintrust, we just need to initialize a logger. By default, a logger is automatically marked as the current, global logger, and once initialized will be picked up by `traced`.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { initLogger } from "braintrust";
const logger = initLogger({
projectName: "Component generator",
apiKey: process.env.BRAINTRUST_API_KEY || "Your BRAINTRUST_API_KEY",
});
```
Now, we'll run the `generateComponent` function on a few examples, and see what the results look like in Braintrust.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
const inputs = [
"A login form for a B2B SaaS product.",
"Create a profile page for a social network.",
"Logs viewer for a cloud infrastructure management tool. Heavy use of dark mode.",
];
for (const input of inputs) {
await generateComponent(input);
}
console.log(`Logged ${inputs.length} requests to Braintrust.`);
```
```
Logged 3 requests to Braintrust.
```
### Viewing the logs in Braintrust
Once this runs, you should be able to see the raw inputs and outputs, along with their scores in the project.
### Capturing user feedback
Let's also track user ratings for these components. Separate from whether or not they're formatted as HTML, it'll be useful to track whether users like the design.
To do this, [configure a new score in the project](/docs/annotate/human-review#configure-review-scores). Let's call it "User preference" and make it a 👍/👎.
Once you create a human review score, you can evaluate results directly in the Braintrust UI, or capture end-user feedback. Here, we'll pretend to capture end-user feedback. Personally, I liked the login form and logs viewer, but not the profile page. Let's record feedback accordingly.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
// Along with scores, you can optionally log user feedback as comments, for additional color.
logger.logFeedback({
id: requests["A login form for a B2B SaaS product."],
scores: { "User preference": 1 },
comment: "Clean, simple",
});
logger.logFeedback({
id: requests["Create a profile page for a social network."],
scores: { "User preference": 0 },
});
logger.logFeedback({
id: requests[
"Logs viewer for a cloud infrastructure management tool. Heavy use of dark mode."
],
scores: { "User preference": 1 },
comment:
"No frills! Would have been nice to have borders around the entries.",
});
```
As users provide feedback, you'll see the updates they make in each log entry.
## Creating a dataset
Now that we've collected some interesting examples from users, let's collect them into a dataset, and see if we can improve the `isComponent` score.
In the Braintrust UI, select the examples, and add them to a new dataset called "Interesting cases".
Once you create the dataset, it should look something like this:
## Evaluating
Now that we have a dataset, let's evaluate the `isComponent` function on it. We'll use the `Eval` function, which takes a dataset and a function, and evaluates the function on each example in the dataset.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { Eval, initDataset } from "braintrust";
await Eval("Component generator", {
data: async () => {
const dataset = initDataset("Component generator", {
dataset: "Interesting cases",
});
const records = [];
for await (const { input } of dataset.fetch()) {
records.push({ input });
}
return records;
},
task: generateComponent,
// We do not need to add any additional scores, because our
// generateComponent() function already computes `isComponent`
scores: [],
});
```
Once the eval runs, you'll see a summary which includes a link to the experiment. As expected, only one of the three outputs contains HTML, so the score is 33.3%. Let's also label user preference for this experiment, so we can track aesthetic taste manually. For simplicity's sake, we'll use the same labeling as before.
### Improving the prompt
Next, let's try to tweak the prompt to stop rendering full HTML pages.
```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
function generateMessages(input: string): ChatCompletionMessageParam[] {
return [
{
role: "system",
content: `You are a skilled design engineer
who can convert ambiguously worded ideas into beautiful, crisp HTML and CSS.
Your designs value simplicity, conciseness, clarity, and functionality over
complexity.
You generate pure HTML with inline CSS, so that your designs can be rendered
directly as plain HTML. Only generate components, not full HTML pages. If you
need to add CSS, you can use the "style" property of an HTML tag. You cannot use
global CSS in a