Skip to main content
TypeScript workflow evaluations are in public preview and can change before reaching general availability.
Requires TypeScript SDK v3.33.0 or later. Import from the top-level braintrust package. The API is marked experimental in the SDK.
Workflow evaluations let you submit tasks or scorers to an asynchronous provider API, such as the OpenAI Batch API, and collect the results later. Use them for evaluations that run through a batch API, which providers often offer at a discount, or that need to continue across process exits. You write functions that submit requests, check whether they’re done, and fetch the results. The SDK calls them, saves progress in a store, and records the outputs and scores in one Braintrust experiment.

How workflow evaluations work

A workflow evaluation has three stages:
  1. Define an evaluation: Call defineWorkflowEval() with your data, task, scorers, and a store. It works like Eval(), except that the task or scorers can send their work to a provider’s asynchronous API. It returns an object with methods that you use to start and resume runs.
  2. Start the run: Call start() to submit the requests. It returns a runId right away, without waiting for the provider.
  3. Resume until completion: Pass the runId to poll() on a schedule, or to processSubmissionResult() from a webhook handler. Each call collects the results that are ready and scores them. The run is done when its status is "completed".
The following example walks through each of these stages.

Example: OpenAI Batch API

This example sends questions about capital cities through the OpenAI Batch API and scores each answer against the expected city. Here’s what the script does:
  • Submits one OpenAI batch for each case.
  • Polls OpenAI until each batch finishes, then collects the answer.
  • Scores each answer with exactMatch(), a plain scorer function, and logs the results to Braintrust.
The SDK submits each case as its own request. With a batch API, that means one batch per case, not one batch for the whole evaluation. Some providers limit how many batches you can create, and a large evaluation might reach that limit.
To get started, install the braintrust and openai packages, and set your BRAINTRUST_API_KEY and OPENAI_API_KEY environment variables:
Then, save the evaluation and its callbacks in evaluation.ts. They live in their own module so the Redis and webhook examples later on this page can import them:
evaluation.ts
Next, save the script that runs the evaluation in workflow-eval.ts:
workflow-eval.ts
Run it with npx tsx workflow-eval.ts. OpenAI batches can take minutes or hours, so the script can run for a while. After each poll(), it prints the run’s status and how many requests are still pending. When it finishes, the Braintrust experiment contains two rows, each with an exactMatch score. The rest of this section walks through the example’s three stages.

Define an evaluation

First, evaluation.ts defines the evaluation in createEvaluation(). In outline, the call looks like this:
The following sections describe each argument. For every option, see defineWorkflowEval().
The first argument is the project name, "capital-cities-eval" in the example. Braintrust creates the project if it doesn’t exist. Each run logs to a new experiment unless you set experimentName.
The data argument accepts the same cases as Eval(). Each case needs a unique id, so the SDK can match results to cases when the run resumes. Cases can also set metadata, tags, and trialCount.Case data must be JSON-serializable, so the store can save it.
The task argument accepts a plain task function or a WorkflowTask. Use a WorkflowTask when the provider accepts a request now and returns its result later. The example’s task, with polling completion, looks like this:
  • submit: Sends one case’s request to the provider. Returns data, such as the provider’s request ID. The SDK saves it and passes it to the completion and collect callbacks when it calls them.
  • completion: How the SDK learns the request is done: by polling with a function such as the example’s pollSubmission(), or from a webhook. See completion options.
  • collect: Fetches the finished result from the provider and returns an object with the task’s output.
The store saves what submit and collect return, so both must be JSON-serializable. In the example, those are { batchId: ... } and the model’s answer.The example passes type arguments to defineWorkflowEval<string, string, string>() for the input, output, and expected output. TypeScript uses them to check the data and to infer the parameter types of callbacks written inline, so input is a string and submission is a Submission without annotations. For the fields of the submitted item and the collected result, see WorkflowTask.
The scores argument accepts plain scorer functions, like the example’s exactMatch(), and WorkflowScorer instances. Use a WorkflowScorer when scoring needs a request that finishes later, such as an LLM-as-a-judge call through a batch API. It has this shape:
It takes the same callbacks as a WorkflowTask, plus a name. Its submit receives the task’s output along with the case’s fields, and its collect returns an object with a score. See WorkflowScorer.Example: Score with a batch requestThe following llmJudge scorer asks a model to grade each answer through the OpenAI Batch API. It reuses the example’s submitRequest(), collectText(), and pollSubmission(). To try it, add WorkflowScorer to the braintrust import in evaluation.ts, then add this code right before createEvaluation():
Then change scores: [exactMatch] to scores: [exactMatch, llmJudge] and run workflow-eval.ts again. Your exactMatch() scorer runs as soon as each answer arrives, and a later poll() collects the grade.
The store argument saves the run’s progress between calls:
  • new WorkflowEvalMemoryStore(): Keeps progress in memory. Use it when the whole run happens in one process, as in the example.
  • new WorkflowEvalRedisStore({ client }): Keeps progress in Redis, using a connected redis, ioredis, or @upstash/redis client. Use it when the run must survive process exits or be resumed by another process. Records expire after seven days by default.
For Redis options and custom stores, see WorkflowEvalStore. Every process that works on a run must use the same store and the same evaluation definition.

Start the run

Next, main() in workflow-eval.ts starts a run by calling start(), which submits the first requests and returns a result with the runId, without waiting for the provider. Save the runId. You pass it to poll(), processSubmissionResult(), and status(). The example also saves it in each OpenAI batch’s metadata, so a webhook handler can find the run.
Calling start() again creates a new run and submits every request to the provider again. To continue an existing run, pass its runId to poll() or processSubmissionResult().

Resume until completion

Finally, the example calls poll() in a loop until every case is scored. A waiting run can resume in either of two ways, or both:
  • Polling: Call evaluation.poll({ runId }) on a schedule.
  • Webhooks: Call evaluation.processSubmissionResult() when the provider sends a completion event.
Each task and scorer chooses its own method. If the task uses webhooks and a scorer uses polling, you still need to schedule poll({ runId }) to collect the scorer results. Check progress. Each call returns a result whose status stays "waiting" until the run completes, and whose pending field has poll and webhook counts of the requests not yet collected. To check without advancing the run, call evaluation.status({ runId }). Get the results. When the status is "completed", result.summary has the scores and a link to the experiment.
The main example runs in one process that waits until the run finishes, and it keeps progress in memory.For runs that take hours, an alternative is a separate script, such as the following resume-eval.ts, that runs in place of workflow-eval.ts. Each time it runs, it does one step and exits, and it saves progress in Redis so the next run continues where the last one stopped. It imports createEvaluation() from evaluation.ts so it doesn’t repeat the evaluation. It requires the redis package (npm install redis) and a REDIS_URL environment variable, such as redis://localhost:6379:
resume-eval.ts
To use it:
  • Running npx tsx resume-eval.ts with no arguments submits the requests and prints the run ID.
  • Running npx tsx resume-eval.ts <runId> advances the run once. A scheduled job can run it periodically until it prints completed.
Use webhook completion when your provider sends an event when a request finishes. For OpenAI, configure a webhook endpoint for batch.completed events, and use a shared Redis store. The following webhook-handler.ts reuses createEvaluation() and client from evaluation.ts:
webhook-handler.ts
Set OPENAI_WEBHOOK_SECRET to the endpoint’s signing secret from OpenAI. Pass the incoming Request from your HTTP endpoint, such as a Next.js route handler, to handleWebhook(). It reads the unparsed body, which signature verification requires. Return an error status if it throws, so OpenAI redelivers the event. Start the run with startRun().Keep in mind:
  • Failed batches: The example’s handleWebhook() ignores batch.failed, batch.expired, and batch.cancelled events, so a failed batch stays pending and the run never completes. The SDK can’t mark a webhook submission as failed, so subscribe to those events to detect failures, and start a new run to recover.
  • Security: processSubmissionResult() doesn’t check that the event came from the provider. In the example, unwrap() does.
  • Repeated events: The same event can arrive more than once. Make your collect callbacks safe to run more than once.
For how the SDK matches an event to a request, see processSubmissionResult().
  • Concurrency: maxConcurrency limits how many of your submit, getExternalId, polling, and collect callbacks the SDK runs at once during each call. It defaults to 10. See defineWorkflowEval().
  • Provider failures: When your polling callback returns { status: "failed", error }, evaluation.poll() throws the error. The SDK checks the request again on the next poll(), so a permanent failure makes every later poll() throw, and the run never completes. Start a new run to recover. The SDK doesn’t resubmit or cancel requests at the provider, so retries and cancellation are up to your integration.
  • Errors in your functions: For polling, the SDK retries your polling callback and your collect function on the next poll(). For webhooks, collect runs again only when processSubmissionResult() is called again, such as when the provider redelivers the event. It runs your submit function and your plain task, scorer, and classifier functions only once per case, so if one of them fails, that case never finishes. Fix the problem and start a new run.

Next steps