Requires Python SDK v0.39.0 or later. Import from
braintrust.workflow_eval, not the top-level braintrust package.How workflow evaluations work
A workflow evaluation has three stages:- Define an evaluation: Call
define_workflow_eval()with your data, task, scorers, and a store. It works likeEval(), except that the task or scorers can send their work to a provider’s asynchronous API. It returns aWorkflowEval, which you use to start and resume runs. - Start the run: Call
start()to submit the requests. It returns arun_idright away, without waiting for the provider. - Resume until completion: Pass the
run_idtopoll()on a schedule, or toprocess_submission_result()from a webhook handler. Each call collects the results that are ready and scores them. The run is done when its status is"completed".
Example: OpenAI Batch API
This example sends questions about capital cities through the OpenAI Batch API and scores each answer against the expected city. Here’s what the script does:- Submits one OpenAI batch for each case.
- Polls OpenAI until each batch finishes, then collects the answer.
- Scores each answer with
exact_match(), a plain scorer function, and logs the results to Braintrust.
The SDK submits each case as its own request. With a batch API, that means one batch per case, not one batch for the whole evaluation. Some providers limit how many batches you can create, and a large evaluation might reach that limit.
braintrust and openai packages, and set your BRAINTRUST_API_KEY and OPENAI_API_KEY environment variables:
workflow_eval.py:
workflow_eval.py
python workflow_eval.py. OpenAI batches can take minutes or hours, so the script can run for a while. After each poll(), it prints the run’s status and how many requests are still pending. When it finishes, the Braintrust experiment contains two rows, each with an exact_match score.
The rest of this section walks through the example’s three stages.
Define an evaluation
First, the example defines the evaluation inmake_evaluation(). In outline, the call looks like this:
define_workflow_eval().
Project and experiment
Project and experiment
The first argument is the project name,
"capital-cities-eval" in the example. Braintrust creates the project if it doesn’t exist. Each run logs to a new experiment unless you set experiment_name.Data
Data
The
data argument accepts the same cases as Eval(). Each case needs a unique id, so the SDK can match results to cases when the run resumes. Cases can also set metadata, tags, and trial_count.Case data must be JSON-serializable, so the store can save it.Task
Task
The
task argument accepts a plain task function or a WorkflowTask. Use a WorkflowTask when the provider accepts a request now and returns its result later. The example’s task looks like this:submit(submit_task()in the example): Sends one case’s request to the provider. Returns data, such as the provider’s request ID. The SDK saves it and passes it to yourpoll_submission()andcollect_task()functions when it calls them.completion: How the SDK learns the request is done: by polling with a function such as the example’spoll_submission(), or from a webhook. See completion options.collect(collect_task()in the example): Fetches the finished result from the provider and returns aWorkflowTaskResult.
submit_task() and collect_task() functions return, so both must be JSON-serializable. In the example, those are {"batch_id": ...} and the model’s answer.For the fields of item and the result, see WorkflowTask.Scores
Scores
The It takes the same functions as a Then change
scores argument accepts plain scorer functions, like the example’s exact_match(), and WorkflowScorer instances. Use a WorkflowScorer when scoring needs a request that finishes later, such as an LLM-as-a-judge call through a batch API. It has this shape:WorkflowTask, plus a name. Its submit receives the task’s output as item.output, and its collect returns a WorkflowScorerResult. See WorkflowScorer.Example: Score with a batch requestThe following llm_judge scorer asks a model to grade each answer through the OpenAI Batch API. It reuses the example’s Submission type, submit_request(), collect_text(), and poll_submission(). To try it, add this code to workflow_eval.py, right before make_evaluation():scores=[exact_match] to scores=[exact_match, llm_judge] and run the script again. Your exact_match() scorer runs as soon as each answer arrives, and a later poll() collects the grade.Store
Store
The
store argument saves the run’s progress between calls:WorkflowEvalMemoryStore(): Keeps progress in memory. Use it when the whole run happens in one process, as in the example.WorkflowEvalRedisStore(client): Keeps progress in Redis. Use it when the run must survive process exits or be resumed by another process. Records expire after seven days by default.
WorkflowEvalStore. Every process that works on a run must use the same store and the same evaluation definition.Start the run
Next,main() starts a run by calling start(), which submits the first requests and returns a result with the run_id, without waiting for the provider.
Save the run_id. You pass it to poll(), process_submission_result(), and status(). The example also saves it in each OpenAI batch’s metadata, so a webhook handler can find the run.
Resume until completion
Finally, the example callspoll() in a loop until every case is scored. A waiting run can resume in either of two ways, or both:
- Polling: Call
evaluation.poll(run_id)on a schedule. - Webhooks: Call
evaluation.process_submission_result()when the provider sends a completion event.
poll(run_id) to collect the scorer results.
Check progress. Each call returns a result whose status stays "waiting" until the run completes, and whose pending field counts the requests still out. To check without advancing the run, call evaluation.status(run_id).
Get the results. When the status is "completed", result.summary has the scores and a link to the experiment.
Resume with polling
Resume with polling
The main example runs in one process that waits until the run finishes, and it keeps progress in memory.For runs that take hours, an alternative is a separate script, such as the following To use it:
resume_eval.py, that runs in place of workflow_eval.py. Each time it runs, it does one step and exits, and it saves progress in Redis so the next run continues where the last one stopped. It imports make_evaluation() and the callbacks from workflow_eval.py so it doesn’t repeat them. It requires the redis package and a REDIS_URL environment variable:resume_eval.py
- Running
python resume_eval.pywith no arguments submits the requests and prints the run ID. - Running
python resume_eval.py <run_id>advances the run once. A scheduled job can run it periodically until it printscompleted.
Resume from provider webhooks
Resume from provider webhooks
Use Pass the raw request body and headers from your HTTP endpoint to
WorkflowSubmissionCompletionWebhook when your provider sends an event when a request finishes. For OpenAI, configure a webhook endpoint for batch.completed events, and use a shared Redis store. The following webhook_handler.py reuses make_evaluation() and client from workflow_eval.py:webhook_handler.py
handle_webhook(), and start the run from a script that imports evaluation from webhook_handler.py.Keep in mind:- Failed batches: The example’s
handle_webhook()ignores failure events, so a failed batch stays pending and the run never completes. Handle failure events in your application. - Security:
process_submission_result()doesn’t check that the event came from the provider. In the example,unwrap()does. - Repeated events: The same event can arrive more than once. Make your
collectcallbacks safe to run more than once.
process_submission_result().Handle failures and concurrency
Handle failures and concurrency
- Concurrency:
max_concurrencylimits how many of your functions the SDK runs at once during each call. It defaults to10. Seedefine_workflow_eval(). - Provider failures: When your polling callback returns
WorkflowSubmissionPoll("failed", error=...),evaluation.poll()raises the error. The SDK checks the request again on the nextpoll(), so if the failure is permanent, start a new run. The SDK doesn’t resubmit or cancel requests at the provider, so retries and cancellation are up to your integration. - Errors in your functions: The SDK retries your polling callback and your
collectfunction on the nextpoll(). It runs yoursubmitfunction and your plain task, scorer, and classifier functions only once per case, so if one of them fails, that case never finishes. Fix the problem and start a new run.
Next steps
- Review
define_workflow_eval()andWorkflowEval. - Learn how to run experiments in code.
- Add custom scorers to measure task quality.