Requires TypeScript SDK v3.33.0 or later. Import from the top-level
braintrust package. The API is marked experimental in the SDK.How workflow evaluations work
A workflow evaluation has three stages:- Define an evaluation: Call
defineWorkflowEval()with your data, task, scorers, and a store. It works likeEval(), except that the task or scorers can send their work to a provider’s asynchronous API. It returns an object with methods that you use to start and resume runs. - Start the run: Call
start()to submit the requests. It returns arunIdright away, without waiting for the provider. - Resume until completion: Pass the
runIdtopoll()on a schedule, or toprocessSubmissionResult()from a webhook handler. Each call collects the results that are ready and scores them. The run is done when its status is"completed".
Example: OpenAI Batch API
This example sends questions about capital cities through the OpenAI Batch API and scores each answer against the expected city. Here’s what the script does:- Submits one OpenAI batch for each case.
- Polls OpenAI until each batch finishes, then collects the answer.
- Scores each answer with
exactMatch(), a plain scorer function, and logs the results to Braintrust.
The SDK submits each case as its own request. With a batch API, that means one batch per case, not one batch for the whole evaluation. Some providers limit how many batches you can create, and a large evaluation might reach that limit.
braintrust and openai packages, and set your BRAINTRUST_API_KEY and OPENAI_API_KEY environment variables:
evaluation.ts. They live in their own module so the Redis and webhook examples later on this page can import them:
evaluation.ts
workflow-eval.ts:
workflow-eval.ts
npx tsx workflow-eval.ts. OpenAI batches can take minutes or hours, so the script can run for a while. After each poll(), it prints the run’s status and how many requests are still pending. When it finishes, the Braintrust experiment contains two rows, each with an exactMatch score.
The rest of this section walks through the example’s three stages.
Define an evaluation
First,evaluation.ts defines the evaluation in createEvaluation(). In outline, the call looks like this:
defineWorkflowEval().
Project and experiment
Project and experiment
The first argument is the project name,
"capital-cities-eval" in the example. Braintrust creates the project if it doesn’t exist. Each run logs to a new experiment unless you set experimentName.Data
Data
The
data argument accepts the same cases as Eval(). Each case needs a unique id, so the SDK can match results to cases when the run resumes. Cases can also set metadata, tags, and trialCount.Case data must be JSON-serializable, so the store can save it.Task
Task
The
task argument accepts a plain task function or a WorkflowTask. Use a WorkflowTask when the provider accepts a request now and returns its result later. The example’s task, with polling completion, looks like this:submit: Sends one case’s request to the provider. Returns data, such as the provider’s request ID. The SDK saves it and passes it to thecompletionandcollectcallbacks when it calls them.completion: How the SDK learns the request is done: by polling with a function such as the example’spollSubmission(), or from a webhook. See completion options.collect: Fetches the finished result from the provider and returns an object with the task’soutput.
submit and collect return, so both must be JSON-serializable. In the example, those are { batchId: ... } and the model’s answer.The example passes type arguments to defineWorkflowEval<string, string, string>() for the input, output, and expected output. TypeScript uses them to check the data and to infer the parameter types of callbacks written inline, so input is a string and submission is a Submission without annotations. For the fields of the submitted item and the collected result, see WorkflowTask.Scores
Scores
The It takes the same callbacks as a Then change
scores argument accepts plain scorer functions, like the example’s exactMatch(), and WorkflowScorer instances. Use a WorkflowScorer when scoring needs a request that finishes later, such as an LLM-as-a-judge call through a batch API. It has this shape:WorkflowTask, plus a name. Its submit receives the task’s output along with the case’s fields, and its collect returns an object with a score. See WorkflowScorer.Example: Score with a batch requestThe following llmJudge scorer asks a model to grade each answer through the OpenAI Batch API. It reuses the example’s submitRequest(), collectText(), and pollSubmission(). To try it, add WorkflowScorer to the braintrust import in evaluation.ts, then add this code right before createEvaluation():scores: [exactMatch] to scores: [exactMatch, llmJudge] and run workflow-eval.ts again. Your exactMatch() scorer runs as soon as each answer arrives, and a later poll() collects the grade.Store
Store
The
store argument saves the run’s progress between calls:new WorkflowEvalMemoryStore(): Keeps progress in memory. Use it when the whole run happens in one process, as in the example.new WorkflowEvalRedisStore({ client }): Keeps progress in Redis, using a connectedredis,ioredis, or@upstash/redisclient. Use it when the run must survive process exits or be resumed by another process. Records expire after seven days by default.
WorkflowEvalStore. Every process that works on a run must use the same store and the same evaluation definition.Start the run
Next,main() in workflow-eval.ts starts a run by calling start(), which submits the first requests and returns a result with the runId, without waiting for the provider.
Save the runId. You pass it to poll(), processSubmissionResult(), and status(). The example also saves it in each OpenAI batch’s metadata, so a webhook handler can find the run.
Resume until completion
Finally, the example callspoll() in a loop until every case is scored. A waiting run can resume in either of two ways, or both:
- Polling: Call
evaluation.poll({ runId })on a schedule. - Webhooks: Call
evaluation.processSubmissionResult()when the provider sends a completion event.
poll({ runId }) to collect the scorer results.
Check progress. Each call returns a result whose status stays "waiting" until the run completes, and whose pending field has poll and webhook counts of the requests not yet collected. To check without advancing the run, call evaluation.status({ runId }).
Get the results. When the status is "completed", result.summary has the scores and a link to the experiment.
Resume with polling
Resume with polling
The main example runs in one process that waits until the run finishes, and it keeps progress in memory.For runs that take hours, an alternative is a separate script, such as the following To use it:
resume-eval.ts, that runs in place of workflow-eval.ts. Each time it runs, it does one step and exits, and it saves progress in Redis so the next run continues where the last one stopped. It imports createEvaluation() from evaluation.ts so it doesn’t repeat the evaluation. It requires the redis package (npm install redis) and a REDIS_URL environment variable, such as redis://localhost:6379:resume-eval.ts
- Running
npx tsx resume-eval.tswith no arguments submits the requests and prints the run ID. - Running
npx tsx resume-eval.ts <runId>advances the run once. A scheduled job can run it periodically until it printscompleted.
Resume from provider webhooks
Resume from provider webhooks
Use webhook completion when your provider sends an event when a request finishes. For OpenAI, configure a webhook endpoint for Set
batch.completed events, and use a shared Redis store. The following webhook-handler.ts reuses createEvaluation() and client from evaluation.ts:webhook-handler.ts
OPENAI_WEBHOOK_SECRET to the endpoint’s signing secret from OpenAI. Pass the incoming Request from your HTTP endpoint, such as a Next.js route handler, to handleWebhook(). It reads the unparsed body, which signature verification requires. Return an error status if it throws, so OpenAI redelivers the event. Start the run with startRun().Keep in mind:- Failed batches: The example’s
handleWebhook()ignoresbatch.failed,batch.expired, andbatch.cancelledevents, so a failed batch stays pending and the run never completes. The SDK can’t mark a webhook submission as failed, so subscribe to those events to detect failures, and start a new run to recover. - Security:
processSubmissionResult()doesn’t check that the event came from the provider. In the example,unwrap()does. - Repeated events: The same event can arrive more than once. Make your
collectcallbacks safe to run more than once.
processSubmissionResult().Handle failures and concurrency
Handle failures and concurrency
- Concurrency:
maxConcurrencylimits how many of yoursubmit,getExternalId, polling, andcollectcallbacks the SDK runs at once during each call. It defaults to10. SeedefineWorkflowEval(). - Provider failures: When your polling callback returns
{ status: "failed", error },evaluation.poll()throws the error. The SDK checks the request again on the nextpoll(), so a permanent failure makes every laterpoll()throw, and the run never completes. Start a new run to recover. The SDK doesn’t resubmit or cancel requests at the provider, so retries and cancellation are up to your integration. - Errors in your functions: For polling, the SDK retries your polling callback and your
collectfunction on the nextpoll(). For webhooks,collectruns again only whenprocessSubmissionResult()is called again, such as when the provider redelivers the event. It runs yoursubmitfunction and your plain task, scorer, and classifier functions only once per case, so if one of them fails, that case never finishes. Fix the problem and start a new run.
Next steps
- Review
defineWorkflowEval()and the workflow evaluation methods. - Learn how to run experiments in code.
- Add custom scorers to measure task quality.