Skip to main content
Python workflow evaluations are in public preview and can change before reaching general availability.
Requires Python SDK v0.39.0 or later. Import from braintrust.workflow_eval, not the top-level braintrust package.
Workflow evaluations let you submit tasks or scorers to an asynchronous provider API, such as the OpenAI Batch API, and collect the results later. Use them for evaluations that run through a batch API, which providers often offer at a discount, or that need to continue across process exits. You write functions that submit requests, check whether they’re done, and fetch the results. The SDK calls them, saves progress in a store, and records the outputs and scores in one Braintrust experiment.

How workflow evaluations work

A workflow evaluation has three stages:
  1. Define an evaluation: Call define_workflow_eval() with your data, task, scorers, and a store. It works like Eval(), except that the task or scorers can send their work to a provider’s asynchronous API. It returns a WorkflowEval, which you use to start and resume runs.
  2. Start the run: Call start() to submit the requests. It returns a run_id right away, without waiting for the provider.
  3. Resume until completion: Pass the run_id to poll() on a schedule, or to process_submission_result() from a webhook handler. Each call collects the results that are ready and scores them. The run is done when its status is "completed".
The following example walks through each of these stages.

Example: OpenAI Batch API

This example sends questions about capital cities through the OpenAI Batch API and scores each answer against the expected city. Here’s what the script does:
  • Submits one OpenAI batch for each case.
  • Polls OpenAI until each batch finishes, then collects the answer.
  • Scores each answer with exact_match(), a plain scorer function, and logs the results to Braintrust.
The SDK submits each case as its own request. With a batch API, that means one batch per case, not one batch for the whole evaluation. Some providers limit how many batches you can create, and a large evaluation might reach that limit.
To get started, install the braintrust and openai packages, and set your BRAINTRUST_API_KEY and OPENAI_API_KEY environment variables:
Then, save the following code in workflow_eval.py:
workflow_eval.py
Run it with python workflow_eval.py. OpenAI batches can take minutes or hours, so the script can run for a while. After each poll(), it prints the run’s status and how many requests are still pending. When it finishes, the Braintrust experiment contains two rows, each with an exact_match score. The rest of this section walks through the example’s three stages.

Define an evaluation

First, the example defines the evaluation in make_evaluation(). In outline, the call looks like this:
The following sections describe each argument. For every option, see define_workflow_eval().
The first argument is the project name, "capital-cities-eval" in the example. Braintrust creates the project if it doesn’t exist. Each run logs to a new experiment unless you set experiment_name.
The data argument accepts the same cases as Eval(). Each case needs a unique id, so the SDK can match results to cases when the run resumes. Cases can also set metadata, tags, and trial_count.Case data must be JSON-serializable, so the store can save it.
The task argument accepts a plain task function or a WorkflowTask. Use a WorkflowTask when the provider accepts a request now and returns its result later. The example’s task looks like this:
  • submit (submit_task() in the example): Sends one case’s request to the provider. Returns data, such as the provider’s request ID. The SDK saves it and passes it to your poll_submission() and collect_task() functions when it calls them.
  • completion: How the SDK learns the request is done: by polling with a function such as the example’s poll_submission(), or from a webhook. See completion options.
  • collect (collect_task() in the example): Fetches the finished result from the provider and returns a WorkflowTaskResult.
The store saves what your submit_task() and collect_task() functions return, so both must be JSON-serializable. In the example, those are {"batch_id": ...} and the model’s answer.For the fields of item and the result, see WorkflowTask.
The scores argument accepts plain scorer functions, like the example’s exact_match(), and WorkflowScorer instances. Use a WorkflowScorer when scoring needs a request that finishes later, such as an LLM-as-a-judge call through a batch API. It has this shape:
It takes the same functions as a WorkflowTask, plus a name. Its submit receives the task’s output as item.output, and its collect returns a WorkflowScorerResult. See WorkflowScorer.Example: Score with a batch requestThe following llm_judge scorer asks a model to grade each answer through the OpenAI Batch API. It reuses the example’s Submission type, submit_request(), collect_text(), and poll_submission(). To try it, add this code to workflow_eval.py, right before make_evaluation():
Then change scores=[exact_match] to scores=[exact_match, llm_judge] and run the script again. Your exact_match() scorer runs as soon as each answer arrives, and a later poll() collects the grade.
The store argument saves the run’s progress between calls:
  • WorkflowEvalMemoryStore(): Keeps progress in memory. Use it when the whole run happens in one process, as in the example.
  • WorkflowEvalRedisStore(client): Keeps progress in Redis. Use it when the run must survive process exits or be resumed by another process. Records expire after seven days by default.
For Redis options and custom stores, see WorkflowEvalStore. Every process that works on a run must use the same store and the same evaluation definition.

Start the run

Next, main() starts a run by calling start(), which submits the first requests and returns a result with the run_id, without waiting for the provider. Save the run_id. You pass it to poll(), process_submission_result(), and status(). The example also saves it in each OpenAI batch’s metadata, so a webhook handler can find the run.
Calling start() again creates a new run and submits every request to the provider again. To continue an existing run, pass its run_id to poll() or process_submission_result().

Resume until completion

Finally, the example calls poll() in a loop until every case is scored. A waiting run can resume in either of two ways, or both:
  • Polling: Call evaluation.poll(run_id) on a schedule.
  • Webhooks: Call evaluation.process_submission_result() when the provider sends a completion event.
Each task and scorer chooses its own method. If the task uses webhooks and a scorer uses polling, you still need to schedule poll(run_id) to collect the scorer results. Check progress. Each call returns a result whose status stays "waiting" until the run completes, and whose pending field counts the requests still out. To check without advancing the run, call evaluation.status(run_id). Get the results. When the status is "completed", result.summary has the scores and a link to the experiment.
The main example runs in one process that waits until the run finishes, and it keeps progress in memory.For runs that take hours, an alternative is a separate script, such as the following resume_eval.py, that runs in place of workflow_eval.py. Each time it runs, it does one step and exits, and it saves progress in Redis so the next run continues where the last one stopped. It imports make_evaluation() and the callbacks from workflow_eval.py so it doesn’t repeat them. It requires the redis package and a REDIS_URL environment variable:
resume_eval.py
To use it:
  • Running python resume_eval.py with no arguments submits the requests and prints the run ID.
  • Running python resume_eval.py <run_id> advances the run once. A scheduled job can run it periodically until it prints completed.
Use WorkflowSubmissionCompletionWebhook when your provider sends an event when a request finishes. For OpenAI, configure a webhook endpoint for batch.completed events, and use a shared Redis store. The following webhook_handler.py reuses make_evaluation() and client from workflow_eval.py:
webhook_handler.py
Pass the raw request body and headers from your HTTP endpoint to handle_webhook(), and start the run from a script that imports evaluation from webhook_handler.py.Keep in mind:
  • Failed batches: The example’s handle_webhook() ignores failure events, so a failed batch stays pending and the run never completes. Handle failure events in your application.
  • Security: process_submission_result() doesn’t check that the event came from the provider. In the example, unwrap() does.
  • Repeated events: The same event can arrive more than once. Make your collect callbacks safe to run more than once.
For how the SDK matches an event to a request, see process_submission_result().
  • Concurrency: max_concurrency limits how many of your functions the SDK runs at once during each call. It defaults to 10. See define_workflow_eval().
  • Provider failures: When your polling callback returns WorkflowSubmissionPoll("failed", error=...), evaluation.poll() raises the error. The SDK checks the request again on the next poll(), so if the failure is permanent, start a new run. The SDK doesn’t resubmit or cancel requests at the provider, so retries and cancellation are up to your integration.
  • Errors in your functions: The SDK retries your polling callback and your collect function on the next poll(). It runs your submit function and your plain task, scorer, and classifier functions only once per case, so if one of them fails, that case never finishes. Fix the problem and start a new run.

Next steps