> ## Documentation Index
> Fetch the complete documentation index at: https://braintrust.dev/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Workflow evaluations

> Submit evaluation tasks and scorers to asynchronous provider APIs, then collect results through polling or webhooks in one Braintrust experiment.

export const feature_0 = "TypeScript workflow evaluations"

export const verb_0 = "are"

<Warning>
  {feature_0} {verb_0} in [public preview](/docs/feature-lifecycle) and can change before reaching general availability.
</Warning>

<Note>
  Requires TypeScript SDK v3.33.0 or later. Import from the top-level `braintrust` package. The API is marked experimental in the SDK.
</Note>

Workflow evaluations let you submit tasks or scorers to an asynchronous provider API, such as the [OpenAI Batch API](https://developers.openai.com/api/docs/guides/batch), and collect the results later. Use them for evaluations that run through a batch API, which providers often offer at a discount, or that need to continue across process exits.

You write functions that submit requests, check whether they're done, and fetch the results. The SDK calls them, saves progress in a store, and records the outputs and scores in one Braintrust experiment.

## How workflow evaluations work

A workflow evaluation has three stages:

1. **[Define an evaluation](#define-an-evaluation)**: Call [`defineWorkflowEval()`](/docs/sdks/typescript/api-reference#defineworkfloweval) with your data, task, scorers, and a store. It works like [`Eval()`](/docs/sdks/typescript/api-reference#eval), except that the task or scorers can send their work to a provider's asynchronous API. It returns an object with [methods](/docs/sdks/typescript/api-reference#workflow-evaluation-methods) that you use to start and resume runs.
2. **[Start the run](#start-the-run)**: Call [`start()`](/docs/sdks/typescript/api-reference#workflow-evaluation-methods) to submit the requests. It returns a `runId` right away, without waiting for the provider.
3. **[Resume until completion](#resume-until-completion)**: Pass the `runId` to [`poll()`](/docs/sdks/typescript/api-reference#workflow-evaluation-methods) on a schedule, or to [`processSubmissionResult()`](/docs/sdks/typescript/api-reference#workflow-evaluation-methods) from a webhook handler. Each call collects the results that are ready and scores them. The run is done when its status is `"completed"`.

The following example walks through each of these stages.

## Example: OpenAI Batch API

This example sends questions about capital cities through the [OpenAI Batch API](https://developers.openai.com/api/docs/guides/batch) and scores each answer against the expected city. Here's what the script does:

* Submits one OpenAI batch for each case.
* Polls OpenAI until each batch finishes, then collects the answer.
* Scores each answer with `exactMatch()`, a plain scorer function, and logs the results to Braintrust.

<Note>
  The SDK submits each case as its own request. With a batch API, that means one batch per case, not one batch for the whole evaluation. Some providers limit how many batches you can create, and a large evaluation might reach that limit.
</Note>

To get started, install the `braintrust` and `openai` packages, and set your `BRAINTRUST_API_KEY` and `OPENAI_API_KEY` environment variables:

```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
npm install "braintrust@>=3.33.0" openai
export BRAINTRUST_API_KEY="your-braintrust-api-key"
export OPENAI_API_KEY="your-openai-api-key"
```

Then, save the evaluation and its callbacks in `evaluation.ts`. They live in their own module so the Redis and webhook examples later on this page can import them:

```typescript evaluation.ts expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import OpenAI, { toFile } from "openai";
import {
  WorkflowTask,
  defineWorkflowEval,
  type EvalScorerArgs,
  type WorkflowEvalStore,
} from "braintrust";

export const client = new OpenAI();

// Data your submit callback returns and the SDK saves between calls.
// The SDK passes it back to the polling and collection callbacks.
// Here it holds the OpenAI batch ID. You choose the fields. Use a type
// alias rather than an interface, because TypeScript can't confirm that an
// interface is JSON-serializable.
export type Submission = { batchId: string };

// Sends one prompt to OpenAI as a single-request batch. context holds the
// SDK's run ID and submission ID.
export async function submitRequest(
  prompt: string,
  context: { runId: string; submissionId: string },
): Promise<Submission> {
  const request = {
    // Ties the batch request to this SDK submission.
    custom_id: context.submissionId,
    method: "POST",
    url: "/v1/responses",
    body: { model: "gpt-5-mini", input: prompt },
  };
  // The Batch API reads its requests from an uploaded JSONL file.
  const batchFile = await client.files.create({
    file: await toFile(
      Buffer.from(JSON.stringify(request) + "\n"),
      "requests.jsonl",
    ),
    purpose: "batch",
  });
  const batch = await client.batches.create({
    input_file_id: batchFile.id,
    endpoint: "/v1/responses",
    completion_window: "24h",
    // Lets a webhook handler find the run.
    metadata: { braintrust_run_id: context.runId },
  });
  // The SDK saves this object and passes it to the polling and collect
  // callbacks.
  return { batchId: batch.id };
}

// The polling callback. Called during each poll() until the batch
// reports complete.
export async function pollSubmission({ batchId }: Submission) {
  const batch = await client.batches.retrieve(batchId);
  switch (batch.status) {
    case "completed":
      return { status: "complete" } as const;
    case "failed":
    case "expired":
    case "cancelled": {
      const errors = JSON.stringify(batch.errors);
      const error = new Error(`Batch ${batch.id}: ${batch.status}, ${errors}`);
      return { status: "failed", error } as const;
    }
    default:
      return { status: "pending" } as const;
  }
}

// Reads the model's answer from the batch output file. The file has one
// JSON line because each batch holds one request.
export async function collectText({ batchId }: Submission): Promise<string> {
  const batch = await client.batches.retrieve(batchId);
  if (batch.error_file_id) {
    const errors = await client.files.content(batch.error_file_id);
    throw new Error(await errors.text());
  }
  if (!batch.output_file_id) {
    throw new Error(`Batch ${batch.id} has no output file`);
  }
  const content = await client.files.content(batch.output_file_id);
  const result = JSON.parse(await content.text());
  const response: OpenAI.Responses.Response = result.response.body;
  return response.output
    .flatMap((item) => (item.type === "message" ? item.content : []))
    .flatMap((part) => (part.type === "output_text" ? [part.text] : []))
    .join("");
}

// A plain scorer function, like the ones you pass to Eval(). The SDK runs
// it on each case's collected output.
function exactMatch({
  output,
  expected,
}: EvalScorerArgs<string, string, string>) {
  return output.trim().toLowerCase() === expected.toLowerCase() ? 1 : 0;
}

// The store and completion mode are parameters so the Redis and webhook
// examples later on this page can reuse this function.
export function createEvaluation(
  store: WorkflowEvalStore,
  completionMode: "poll" | "webhook",
) {
  // The type arguments are the input, output, and expected output types.
  // TypeScript uses them to infer the parameter types of the callbacks below.
  return defineWorkflowEval<string, string, string>("capital-cities-eval", {
    store,
    // Each case needs a stable, unique id so the run can resume.
    data: [
      { id: "france", input: "Capital of France?", expected: "Paris" },
      { id: "japan", input: "Capital of Japan?", expected: "Tokyo" },
    ],
    task: new WorkflowTask({
      // Called once for each case in data, or once per trial if you set
      // trialCount.
      submit: ({ input }, context) =>
        submitRequest(`${input} Reply with only the city name.`, context),
      // How the SDK learns that a batch finished: by polling, or from a
      // webhook that the SDK matches to the batch ID getExternalId returns.
      completion:
        completionMode === "poll"
          ? { mode: "poll", poll: pollSubmission }
          : { mode: "webhook", getExternalId: ({ batchId }) => batchId },
      // Called after the batch completes.
      collect: async (submission) => ({
        output: await collectText(submission),
      }),
    }),
    // Plain scorer functions and WorkflowScorer instances can go here.
    scores: [exactMatch],
  });
}
```

Next, save the script that runs the evaluation in `workflow-eval.ts`:

```typescript workflow-eval.ts theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import { setTimeout as sleep } from "node:timers/promises";
import { WorkflowEvalMemoryStore } from "braintrust";
import { createEvaluation } from "./evaluation";

async function main() {
  // The memory store works in this example because the whole run
  // happens in one process.
  const evaluation = createEvaluation(new WorkflowEvalMemoryStore(), "poll");
  // Submits one batch per case and returns right away with status
  // "waiting".
  let result = await evaluation.start();
  console.log(result.runId, result.status);
  // In an application, schedule poll() instead of sleeping in one process.
  while (result.status === "waiting") {
    await sleep(60_000);
    result = await evaluation.poll({ runId: result.runId });
    // pending counts submitted requests that haven't been collected.
    console.log(result.status, result.pending);
  }
  console.log(result.summary);
}

main();
```

Run it with `npx tsx workflow-eval.ts`. OpenAI batches can take minutes or hours, so the script can run for a while. After each `poll()`, it prints the run's status and how many requests are still pending. When it finishes, the Braintrust experiment contains two rows, each with an `exactMatch` score.

The rest of this section walks through the example's three stages.

### Define an evaluation

First, `evaluation.ts` defines the evaluation in `createEvaluation()`. In outline, the call looks like this:

```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
const evaluation = defineWorkflowEval("capital-cities-eval", {
  data: [/* ... */],
  task: new WorkflowTask({ submit, completion, collect }),
  scores: [/* ... */],
  store: new WorkflowEvalMemoryStore(),
});
```

The following sections describe each argument. For every option, see [`defineWorkflowEval()`](/docs/sdks/typescript/api-reference#defineworkfloweval).

<AccordionGroup>
  <Accordion title="Project and experiment">
    The first argument is the project name, `"capital-cities-eval"` in the example. Braintrust creates the project if it doesn't exist. Each run logs to a new experiment unless you set `experimentName`.
  </Accordion>

  <Accordion title="Data">
    The `data` argument accepts the same cases as `Eval()`. Each case needs a unique `id`, so the SDK can match results to cases when the run resumes. Cases can also set `metadata`, `tags`, and `trialCount`.

    Case data must be JSON-serializable, so the store can save it.
  </Accordion>

  <Accordion title="Task">
    The `task` argument accepts a plain task function or a `WorkflowTask`. Use a `WorkflowTask` when the provider accepts a request now and returns its result later. The example's task, with polling completion, looks like this:

    ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    new WorkflowTask({
      submit: ({ input }, context) =>
        submitRequest(`${input} Reply with only the city name.`, context),
      completion: { mode: "poll", poll: pollSubmission },
      collect: async (submission) => ({ output: await collectText(submission) }),
    })
    ```

    * **`submit`**: Sends one case's request to the provider. Returns data, such as the provider's request ID. The SDK saves it and passes it to the `completion` and `collect` callbacks when it calls them.
    * **`completion`**: How the SDK learns the request is done: by polling with a function such as the example's `pollSubmission()`, or from a webhook. See [completion options](/docs/sdks/typescript/api-reference#completion-configuration).
    * **`collect`**: Fetches the finished result from the provider and returns an object with the task's `output`.

    The store saves what `submit` and `collect` return, so both must be JSON-serializable. In the example, those are `{ batchId: ... }` and the model's answer.

    The example passes type arguments to `defineWorkflowEval<string, string, string>()` for the input, output, and expected output. TypeScript uses them to check the data and to infer the parameter types of callbacks written inline, so `input` is a `string` and `submission` is a `Submission` without annotations. For the fields of the submitted item and the collected result, see [`WorkflowTask`](/docs/sdks/typescript/api-reference#workflowtask-and-workflowscorer).
  </Accordion>

  <Accordion title="Scores">
    The `scores` argument accepts [plain scorer functions](/docs/evaluate/write-scorers), like the example's `exactMatch()`, and `WorkflowScorer` instances. Use a `WorkflowScorer` when scoring needs a request that finishes later, such as an LLM-as-a-judge call through a batch API. It has this shape:

    ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    new WorkflowScorer({
      name: "llmJudge",
      submit: async ({ input, output, expected }, context) => {/* ... */},
      completion: { mode: "poll", poll: pollSubmission },
      collect: async (submission) => {/* ... */},
    })
    ```

    It takes the same callbacks as a `WorkflowTask`, plus a `name`. Its `submit` receives the task's `output` along with the case's fields, and its `collect` returns an object with a `score`. See [`WorkflowScorer`](/docs/sdks/typescript/api-reference#workflowtask-and-workflowscorer).

    **Example: Score with a batch request**

    The following `llmJudge` scorer asks a model to grade each answer through the OpenAI Batch API. It reuses the example's `submitRequest()`, `collectText()`, and `pollSubmission()`. To try it, add `WorkflowScorer` to the `braintrust` import in `evaluation.ts`, then add this code right before `createEvaluation()`:

    ```typescript expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    const llmJudge = new WorkflowScorer({
      // Names the score in Braintrust. Keep it the same across resumes.
      name: "llmJudge",
      // Called once for each case in data, or once per trial if you set
      // trialCount, after that case's task output is ready. A scorer defined
      // outside defineWorkflowEval() has no type arguments to infer from, so
      // annotate the item with EvalScorerArgs.
      submit: (
        { input, output, expected }: EvalScorerArgs<string, string, string>,
        context,
      ) =>
        submitRequest(
          [
            "Does the answer name the expected city? Reply with only 1 for yes or 0 for no.",
            `Question: ${input}`,
            `Expected: ${expected}`,
            `Answer: ${output}`,
          ].join("\n"),
          context,
        ),
      // Reuses the task's polling callback, since both are OpenAI batches.
      completion: { mode: "poll", poll: pollSubmission },
      // Called after pollSubmission() reports the grading batch complete.
      collect: async (submission) => {
        const verdict = (await collectText(submission)).trim();
        // Throwing here blocks the run: the SDK retries collect on every poll().
        if (verdict !== "0" && verdict !== "1") {
          throw new Error(`Unexpected judge response: ${verdict}`);
        }
        return { score: Number(verdict) };
      },
    });
    ```

    Then change `scores: [exactMatch]` to `scores: [exactMatch, llmJudge]` and run `workflow-eval.ts` again. Your `exactMatch()` scorer runs as soon as each answer arrives, and a later `poll()` collects the grade.
  </Accordion>

  <Accordion title="Store">
    The `store` argument saves the run's progress between calls:

    * **[`new WorkflowEvalMemoryStore()`](/docs/sdks/typescript/api-reference#workflowevalstore)**: Keeps progress in memory. Use it when the whole run happens in one process, as in the example.
    * **[`new WorkflowEvalRedisStore({ client })`](/docs/sdks/typescript/api-reference#workflowevalstore)**: Keeps progress in Redis, using a connected `redis`, `ioredis`, or `@upstash/redis` client. Use it when the run must survive process exits or be resumed by another process. Records expire after seven days by default.

    For Redis options and custom stores, see [`WorkflowEvalStore`](/docs/sdks/typescript/api-reference#workflowevalstore). Every process that works on a run must use the same store and the same evaluation definition.
  </Accordion>
</AccordionGroup>

### Start the run

Next, `main()` in `workflow-eval.ts` starts a run by calling `start()`, which submits the first requests and returns a result with the `runId`, without waiting for the provider.

**Save the `runId`.** You pass it to `poll()`, `processSubmissionResult()`, and [`status()`](/docs/sdks/typescript/api-reference#workflow-evaluation-methods). The example also saves it in each OpenAI batch's metadata, so a webhook handler can find the run.

<Warning>
  Calling `start()` again creates a new run and submits every request to the provider again. To continue an existing run, pass its `runId` to `poll()` or `processSubmissionResult()`.
</Warning>

### Resume until completion

Finally, the example calls `poll()` in a loop until every case is scored. A waiting run can resume in either of two ways, or both:

* **Polling**: Call `evaluation.poll({ runId })` on a schedule.
* **Webhooks**: Call `evaluation.processSubmissionResult()` when the provider sends a completion event.

Each task and scorer chooses its own method. If the task uses webhooks and a scorer uses polling, you still need to schedule `poll({ runId })` to collect the scorer results.

**Check progress.** Each call returns a result whose `status` stays `"waiting"` until the run completes, and whose `pending` field has `poll` and `webhook` counts of the requests not yet collected. To check without advancing the run, call `evaluation.status({ runId })`.

**Get the results.** When the status is `"completed"`, `result.summary` has the scores and a link to the experiment.

<AccordionGroup>
  <Accordion title="Resume with polling">
    The main example runs in one process that waits until the run finishes, and it keeps progress in memory.

    For runs that take hours, an alternative is a separate script, such as the following `resume-eval.ts`, that runs in place of `workflow-eval.ts`. Each time it runs, it does one step and exits, and it saves progress in Redis so the next run continues where the last one stopped. It imports `createEvaluation()` from `evaluation.ts` so it doesn't repeat the evaluation. It requires the `redis` package (`npm install redis`) and a `REDIS_URL` environment variable, such as `redis://localhost:6379`:

    ```typescript resume-eval.ts expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    import { createClient } from "redis";
    import { WorkflowEvalRedisStore } from "braintrust";
    import { createEvaluation } from "./evaluation";

    async function main() {
      const redis = await createClient({ url: process.env.REDIS_URL }).connect();
      try {
        // Saves run state in Redis, so any process with the same Redis URL
        // and evaluation definition can resume the run.
        const evaluation = createEvaluation(
          new WorkflowEvalRedisStore({ client: redis }),
          "poll",
        );
        const runId = process.argv[2];
        // With no argument, starts a new run and prints its run ID.
        // With a run ID, checks the run's pending batches once and advances
        // any cases whose results are ready.
        const result = runId
          ? await evaluation.poll({ runId })
          : await evaluation.start();
        console.log(result.runId, result.status);
      } finally {
        await redis.close();
      }
    }

    main();
    ```

    To use it:

    * Running `npx tsx resume-eval.ts` with no arguments submits the requests and prints the run ID.
    * Running `npx tsx resume-eval.ts <runId>` advances the run once. A scheduled job can run it periodically until it prints `completed`.
  </Accordion>

  <Accordion title="Resume from provider webhooks">
    Use [webhook completion](/docs/sdks/typescript/api-reference#completion-configuration) when your provider sends an event when a request finishes. For OpenAI, configure a [webhook endpoint](https://developers.openai.com/api/docs/guides/webhooks) for `batch.completed` events, and use a shared Redis store. The following `webhook-handler.ts` reuses `createEvaluation()` and `client` from `evaluation.ts`:

    ```typescript webhook-handler.ts expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    import { createClient } from "redis";
    import { WorkflowEvalRedisStore } from "braintrust";
    import { client, createEvaluation } from "./evaluation";

    const redis = createClient({ url: process.env.REDIS_URL });
    // Uses webhook completion. After submit, the SDK saves the batch ID, so a
    // later webhook can find the submission.
    const evaluation = createEvaluation(
      new WorkflowEvalRedisStore({ client: redis }),
      "webhook",
    );

    async function connectRedis() {
      if (!redis.isOpen) {
        await redis.connect();
      }
    }

    // Starts a run that completes through webhooks.
    export async function startRun() {
      await connectRedis();
      return evaluation.start();
    }

    // Call this from your HTTP endpoint with the incoming request.
    export async function handleWebhook(request: Request) {
      await connectRedis();
      // Checks that the request came from OpenAI, using the signing secret in
      // OPENAI_WEBHOOK_SECRET, and parses the event.
      const event = await client.webhooks.unwrap(
        await request.text(),
        request.headers,
      );
      // Ignores failure events. Handle those in your application.
      if (event.type !== "batch.completed") {
        return;
      }
      const batch = await client.batches.retrieve(event.data.id);
      // The run ID comes from the batch metadata that submitRequest() set.
      const runId = batch.metadata?.braintrust_run_id;
      // Skips batches that this evaluation didn't submit.
      if (!runId) {
        return;
      }
      try {
        // The SDK finds the submission by batch ID, then calls the task's
        // collect callback.
        return await evaluation.processSubmissionResult({
          runId,
          externalId: batch.id,
        });
      } catch (error) {
        // Skips batches that use polling, such as llmJudge's grading batches.
        if (
          error instanceof Error &&
          error.message === "No submission matches this result"
        ) {
          return;
        }
        throw error;
      }
    }
    ```

    Set `OPENAI_WEBHOOK_SECRET` to the endpoint's [signing secret](https://developers.openai.com/api/docs/guides/webhooks#verifying-webhook-signatures) from OpenAI. Pass the incoming [`Request`](https://developer.mozilla.org/en-US/docs/Web/API/Request) from your HTTP endpoint, such as a Next.js route handler, to `handleWebhook()`. It reads the unparsed body, which signature verification requires. Return an error status if it throws, so OpenAI redelivers the event. Start the run with `startRun()`.

    Keep in mind:

    * **Failed batches**: The example's `handleWebhook()` ignores `batch.failed`, `batch.expired`, and `batch.cancelled` events, so a failed batch stays pending and the run never completes. The SDK can't mark a webhook submission as failed, so subscribe to those events to detect failures, and start a new run to recover.
    * **Security**: `processSubmissionResult()` doesn't check that the event came from the provider. In the example, `unwrap()` does.
    * **Repeated events**: The same event can arrive more than once. Make your `collect` callbacks safe to run more than once.

    For how the SDK matches an event to a request, see [`processSubmissionResult()`](/docs/sdks/typescript/api-reference#workflow-evaluation-methods).
  </Accordion>

  <Accordion title="Handle failures and concurrency">
    * **Concurrency**: `maxConcurrency` limits how many of your `submit`, `getExternalId`, polling, and `collect` callbacks the SDK runs at once during each call. It defaults to `10`. See [`defineWorkflowEval()`](/docs/sdks/typescript/api-reference#defineworkfloweval).
    * **Provider failures**: When your polling callback returns [`{ status: "failed", error }`](/docs/sdks/typescript/api-reference#completion-configuration), `evaluation.poll()` throws the error. The SDK checks the request again on the next `poll()`, so a permanent failure makes every later `poll()` throw, and the run never completes. Start a new run to recover. The SDK doesn't resubmit or cancel requests at the provider, so retries and cancellation are up to your integration.
    * **Errors in your functions**: For polling, the SDK retries your polling callback and your `collect` function on the next `poll()`. For webhooks, `collect` runs again only when `processSubmissionResult()` is called again, such as when the provider redelivers the event. It runs your `submit` function and your plain task, scorer, and classifier functions only once per case, so if one of them fails, that case never finishes. Fix the problem and start a new run.
  </Accordion>
</AccordionGroup>

## Next steps

* Review [`defineWorkflowEval()`](/docs/sdks/typescript/api-reference#defineworkfloweval) and the [workflow evaluation methods](/docs/sdks/typescript/api-reference#workflow-evaluation-methods).
* Learn how to [run experiments in code](/docs/evaluate/run-in-code).
* Add [custom scorers](/docs/evaluate/write-scorers) to measure task quality.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.