> ## Documentation Index
> Fetch the complete documentation index at: https://braintrust.dev/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Workflow evaluations

> Submit evaluation tasks and scorers to asynchronous provider APIs, then collect results through polling or webhooks in one Braintrust experiment.

export const feature_0 = "Python workflow evaluations"

export const verb_0 = "are"

<Warning>
  {feature_0} {verb_0} in [public preview](/docs/feature-lifecycle) and can change before reaching general availability.
</Warning>

<Note>
  Requires Python SDK v0.39.0 or later. Import from `braintrust.workflow_eval`, not the top-level `braintrust` package.
</Note>

Workflow evaluations let you submit tasks or scorers to an asynchronous provider API, such as the [OpenAI Batch API](https://developers.openai.com/api/docs/guides/batch), and collect the results later. Use them for evaluations that run through a batch API, which providers often offer at a discount, or that need to continue across process exits.

You write functions that submit requests, check whether they're done, and fetch the results. The SDK calls them, saves progress in a store, and records the outputs and scores in one Braintrust experiment.

## How workflow evaluations work

A workflow evaluation has three stages:

1. **[Define an evaluation](#define-an-evaluation)**: Call [`define_workflow_eval()`](/docs/sdks/python/api-reference#define_workflow_eval) with your data, task, scorers, and a store. It works like [`Eval()`](/docs/sdks/python/api-reference#eval), except that the task or scorers can send their work to a provider's asynchronous API. It returns a [`WorkflowEval`](/docs/sdks/python/api-reference#workfloweval), which you use to start and resume runs.
2. **[Start the run](#start-the-run)**: Call [`start()`](/docs/sdks/python/api-reference#workfloweval) to submit the requests. It returns a `run_id` right away, without waiting for the provider.
3. **[Resume until completion](#resume-until-completion)**: Pass the `run_id` to [`poll()`](/docs/sdks/python/api-reference#workfloweval) on a schedule, or to [`process_submission_result()`](/docs/sdks/python/api-reference#workfloweval) from a webhook handler. Each call collects the results that are ready and scores them. The run is done when its status is `"completed"`.

The following example walks through each of these stages.

## Example: OpenAI Batch API

This example sends questions about capital cities through the [OpenAI Batch API](https://developers.openai.com/api/docs/guides/batch) and scores each answer against the expected city. Here's what the script does:

* Submits one OpenAI batch for each case.
* Polls OpenAI until each batch finishes, then collects the answer.
* Scores each answer with `exact_match()`, a plain scorer function, and logs the results to Braintrust.

<Note>
  The SDK submits each case as its own request. With a batch API, that means one batch per case, not one batch for the whole evaluation. Some providers limit how many batches you can create, and a large evaluation might reach that limit.
</Note>

To get started, install the `braintrust` and `openai` packages, and set your `BRAINTRUST_API_KEY` and `OPENAI_API_KEY` environment variables:

```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
pip install "braintrust>=0.39.0" openai
export BRAINTRUST_API_KEY="your-braintrust-api-key"
export OPENAI_API_KEY="your-openai-api-key"
```

Then, save the following code in `workflow_eval.py`:

```python workflow_eval.py expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import asyncio
import json
from typing import TypedDict

from openai import AsyncOpenAI
from openai.types.responses import Response

from braintrust.workflow_eval import (
    WorkflowEvalMemoryStore,
    WorkflowEvalStore,
    WorkflowSubmissionCompletion,
    WorkflowSubmissionCompletionPoll,
    WorkflowSubmissionContext,
    WorkflowSubmissionPoll,
    WorkflowTask,
    WorkflowTaskItem,
    WorkflowTaskResult,
    define_workflow_eval,
)

client = AsyncOpenAI()

# Data your submit callback returns and the SDK saves between calls.
# The SDK passes it back to the polling and collection callbacks.
# Here it holds the OpenAI batch ID. You choose the fields.
class Submission(TypedDict):
    batch_id: str

# Sends one prompt to OpenAI as a single-request batch.
async def submit_request(
    prompt: str, context: WorkflowSubmissionContext
) -> Submission:
    request = {
        # Ties the batch request to this SDK submission.
        "custom_id": context.submission_id,
        "method": "POST",
        "url": "/v1/responses",
        "body": {"model": "gpt-5-mini", "input": prompt},
    }
    # The Batch API reads its requests from an uploaded JSONL file.
    batch_file = await client.files.create(
        file=("requests.jsonl", (json.dumps(request) + "\n").encode()),
        purpose="batch",
    )
    batch = await client.batches.create(
        input_file_id=batch_file.id,
        endpoint="/v1/responses",
        completion_window="24h",
        # Lets a webhook handler find the run.
        metadata={"braintrust_run_id": context.run_id},
    )
    # The SDK saves this dictionary and passes it to
    # poll_submission() and collect_task().
    return {"batch_id": batch.id}

# The task's submit callback. The SDK calls it once for each case in data,
# or once per trial if you set trial_count. item holds the case's input and
# expected output. context holds the SDK's run ID and submission ID.
async def submit_task(
    item: WorkflowTaskItem[str, str], context: WorkflowSubmissionContext
) -> Submission:
    prompt = f"{item.input} Reply with only the city name."
    return await submit_request(prompt, context)

# The polling callback. Called during each poll() until the batch
# reports complete.
async def poll_submission(
    submission: Submission, _context: WorkflowSubmissionContext
) -> WorkflowSubmissionPoll:
    batch = await client.batches.retrieve(submission["batch_id"])
    if batch.status == "completed":
        return WorkflowSubmissionPoll("complete")
    if batch.status in {"failed", "expired", "cancelled"}:
        message = f"Batch {batch.id}: {batch.status}, {batch.errors}"
        return WorkflowSubmissionPoll("failed", error=RuntimeError(message))
    return WorkflowSubmissionPoll("pending")

# Reads the model's answer from the batch output file. The file has one
# JSON line because each batch holds one request.
async def collect_text(submission: Submission) -> str:
    batch = await client.batches.retrieve(submission["batch_id"])
    if batch.error_file_id:
        errors = await client.files.content(batch.error_file_id)
        raise RuntimeError(errors.text)
    assert batch.output_file_id is not None
    content = await client.files.content(batch.output_file_id)
    result = json.loads(content.text)
    response = Response.model_validate(result["response"]["body"])
    return response.output_text

# The task's collect callback. Called after poll_submission() reports
# the batch complete.
async def collect_task(
    submission: Submission, _context: WorkflowSubmissionContext
) -> WorkflowTaskResult[str]:
    return WorkflowTaskResult(output=await collect_text(submission))

# A plain scorer function, like the ones you pass to Eval(). The SDK runs
# it on each case's collected output.
def exact_match(input: str, output: str, expected: str) -> float:
    return float(output.strip().casefold() == expected.casefold())

# The store and completion method are parameters so the Redis and
# webhook examples later on this page can reuse this function.
def make_evaluation(
    store: WorkflowEvalStore,
    completion: WorkflowSubmissionCompletion[Submission],
):
    return define_workflow_eval(
        "capital-cities-eval",
        store=store,
        # Each case needs a stable, unique id so the run can resume.
        data=[
            {
                "id": "france",
                "input": "Capital of France?",
                "expected": "Paris",
            },
            {
                "id": "japan",
                "input": "Capital of Japan?",
                "expected": "Tokyo",
            },
        ],
        # Bundles the task's submit and collect callbacks with its
        # completion method.
        task=WorkflowTask(
            submit=submit_task,
            completion=completion,
            collect=collect_task,
        ),
        # Plain scorer functions and WorkflowScorer instances can go here.
        scores=[exact_match],
    )

async def main():
    async with client:
        # The memory store works in this example because the whole run
        # happens in one process.
        evaluation = make_evaluation(
            WorkflowEvalMemoryStore(),
            WorkflowSubmissionCompletionPoll(poll=poll_submission),
        )
        # Submits one batch per case and returns right away with status
        # "waiting".
        result = await evaluation.start()
        print(result.run_id, result.status)
        # In an application, schedule poll() instead of sleeping in one process.
        while result.status == "waiting":
            await asyncio.sleep(60)
            result = await evaluation.poll(result.run_id)
            # pending counts submitted requests that haven't been collected.
            print(result.status, result.pending)
        print(result.summary)

if __name__ == "__main__":
    asyncio.run(main())
```

Run it with `python workflow_eval.py`. OpenAI batches can take minutes or hours, so the script can run for a while. After each `poll()`, it prints the run's status and how many requests are still pending. When it finishes, the Braintrust experiment contains two rows, each with an `exact_match` score.

The rest of this section walks through the example's three stages.

### Define an evaluation

First, the example defines the evaluation in `make_evaluation()`. In outline, the call looks like this:

```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
evaluation = define_workflow_eval(
    "capital-cities-eval",
    data=[...],
    task=WorkflowTask(submit=..., completion=..., collect=...),
    scores=[...],
    store=WorkflowEvalMemoryStore(),
)
```

The following sections describe each argument. For every option, see [`define_workflow_eval()`](/docs/sdks/python/api-reference#define_workflow_eval).

<AccordionGroup>
  <Accordion title="Project and experiment">
    The first argument is the project name, `"capital-cities-eval"` in the example. Braintrust creates the project if it doesn't exist. Each run logs to a new experiment unless you set `experiment_name`.
  </Accordion>

  <Accordion title="Data">
    The `data` argument accepts the same cases as `Eval()`. Each case needs a unique `id`, so the SDK can match results to cases when the run resumes. Cases can also set `metadata`, `tags`, and `trial_count`.

    Case data must be JSON-serializable, so the store can save it.
  </Accordion>

  <Accordion title="Task">
    The `task` argument accepts a plain task function or a `WorkflowTask`. Use a `WorkflowTask` when the provider accepts a request now and returns its result later. The example's task looks like this:

    ```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    WorkflowTask(
        submit=submit_task,
        completion=WorkflowSubmissionCompletionPoll(poll=poll_submission),
        collect=collect_task,
    )
    ```

    * **`submit`** (`submit_task()` in the example): Sends one case's request to the provider. Returns data, such as the provider's request ID. The SDK saves it and passes it to your `poll_submission()` and `collect_task()` functions when it calls them.
    * **`completion`**: How the SDK learns the request is done: by polling with a function such as the example's `poll_submission()`, or from a webhook. See [completion options](/docs/sdks/python/api-reference#workflowsubmissioncompletionpoll-and-workflowsubmissioncompletionwebhook).
    * **`collect`** (`collect_task()` in the example): Fetches the finished result from the provider and returns a [`WorkflowTaskResult`](/docs/sdks/python/api-reference#workflowtask-and-workflowscorer).

    The store saves what your `submit_task()` and `collect_task()` functions return, so both must be JSON-serializable. In the example, those are `{"batch_id": ...}` and the model's answer.

    For the fields of `item` and the result, see [`WorkflowTask`](/docs/sdks/python/api-reference#workflowtask-and-workflowscorer).
  </Accordion>

  <Accordion title="Scores">
    The `scores` argument accepts [plain scorer functions](/docs/evaluate/write-scorers), like the example's `exact_match()`, and `WorkflowScorer` instances. Use a `WorkflowScorer` when scoring needs a request that finishes later, such as an LLM-as-a-judge call through a batch API. It has this shape:

    ```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    WorkflowScorer(
        name="llm_judge",
        submit=submit_score,
        completion=WorkflowSubmissionCompletionPoll(poll=poll_submission),
        collect=collect_score,
    )
    ```

    It takes the same functions as a `WorkflowTask`, plus a `name`. Its `submit` receives the task's output as `item.output`, and its `collect` returns a [`WorkflowScorerResult`](/docs/sdks/python/api-reference#workflowtask-and-workflowscorer). See [`WorkflowScorer`](/docs/sdks/python/api-reference#workflowtask-and-workflowscorer).

    **Example: Score with a batch request**

    The following `llm_judge` scorer asks a model to grade each answer through the OpenAI Batch API. It reuses the example's `Submission` type, `submit_request()`, `collect_text()`, and `poll_submission()`. To try it, add this code to `workflow_eval.py`, right before `make_evaluation()`:

    ```python expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    from braintrust.workflow_eval import (
        WorkflowScorer,
        WorkflowScorerItem,
        WorkflowScorerResult,
    )

    # The scorer's submit callback. The SDK calls it once for each case in
    # data, or once per trial if you set trial_count, after that case's task
    # output is ready. item holds the case's input, expected output, and the
    # task's output.
    async def submit_score(
        item: WorkflowScorerItem[str, str, str], context: WorkflowSubmissionContext
    ) -> Submission:
        return await submit_request(
            "Does the answer name the expected city? "
            "Reply with only 1 for yes or 0 for no.\n"
            f"Question: {item.input}\n"
            f"Expected: {item.expected}\n"
            f"Answer: {item.output}",
            context,
        )

    # The scorer's collect callback. Called after poll_submission() reports
    # the grading batch complete.
    async def collect_score(
        submission: Submission, _context: WorkflowSubmissionContext
    ) -> WorkflowScorerResult:
        verdict = (await collect_text(submission)).strip()
        # Raising here blocks the run: the SDK retries collect on every poll().
        if verdict not in {"0", "1"}:
            raise ValueError(f"Unexpected judge response: {verdict!r}")
        return WorkflowScorerResult(score=float(verdict))

    # Bundles the scorer's submit and collect callbacks with its completion
    # method.
    llm_judge = WorkflowScorer(
        # Names the score in Braintrust. Keep it the same across resumes.
        name="llm_judge",
        submit=submit_score,
        # Reuses the task's polling callback, since both are OpenAI batches.
        completion=WorkflowSubmissionCompletionPoll(poll=poll_submission),
        collect=collect_score,
    )
    ```

    Then change `scores=[exact_match]` to `scores=[exact_match, llm_judge]` and run the script again. Your `exact_match()` scorer runs as soon as each answer arrives, and a later `poll()` collects the grade.
  </Accordion>

  <Accordion title="Store">
    The `store` argument saves the run's progress between calls:

    * **[`WorkflowEvalMemoryStore()`](/docs/sdks/python/api-reference#workflowevalstore)**: Keeps progress in memory. Use it when the whole run happens in one process, as in the example.
    * **[`WorkflowEvalRedisStore(client)`](/docs/sdks/python/api-reference#workflowevalstore)**: Keeps progress in Redis. Use it when the run must survive process exits or be resumed by another process. Records expire after seven days by default.

    For Redis options and custom stores, see [`WorkflowEvalStore`](/docs/sdks/python/api-reference#workflowevalstore). Every process that works on a run must use the same store and the same evaluation definition.
  </Accordion>
</AccordionGroup>

### Start the run

Next, `main()` starts a run by calling `start()`, which submits the first requests and returns a result with the `run_id`, without waiting for the provider.

**Save the `run_id`.** You pass it to `poll()`, `process_submission_result()`, and [`status()`](/docs/sdks/python/api-reference#workfloweval). The example also saves it in each OpenAI batch's metadata, so a webhook handler can find the run.

<Warning>
  Calling `start()` again creates a new run and submits every request to the provider again. To continue an existing run, pass its `run_id` to `poll()` or `process_submission_result()`.
</Warning>

### Resume until completion

Finally, the example calls `poll()` in a loop until every case is scored. A waiting run can resume in either of two ways, or both:

* **Polling**: Call `evaluation.poll(run_id)` on a schedule.
* **Webhooks**: Call `evaluation.process_submission_result()` when the provider sends a completion event.

Each task and scorer chooses its own method. If the task uses webhooks and a scorer uses polling, you still need to schedule `poll(run_id)` to collect the scorer results.

**Check progress.** Each call returns a result whose `status` stays `"waiting"` until the run completes, and whose `pending` field counts the requests still out. To check without advancing the run, call `evaluation.status(run_id)`.

**Get the results.** When the status is `"completed"`, `result.summary` has the scores and a link to the experiment.

<AccordionGroup>
  <Accordion title="Resume with polling">
    The main example runs in one process that waits until the run finishes, and it keeps progress in memory.

    For runs that take hours, an alternative is a separate script, such as the following `resume_eval.py`, that runs in place of `workflow_eval.py`. Each time it runs, it does one step and exits, and it saves progress in Redis so the next run continues where the last one stopped. It imports `make_evaluation()` and the callbacks from `workflow_eval.py` so it doesn't repeat them. It requires the `redis` package and a `REDIS_URL` environment variable:

    ```python resume_eval.py expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    import asyncio
    import os
    import sys

    from redis.asyncio import Redis

    from braintrust.workflow_eval import (
        WorkflowEvalRedisStore,
        WorkflowSubmissionCompletionPoll,
    )
    # Reuses the evaluation and callbacks from the main example.
    from workflow_eval import client, make_evaluation, poll_submission

    async def main():
        async with Redis.from_url(os.environ["REDIS_URL"]) as redis, client:
            # Saves run state in Redis, so any process with the same Redis URL
            # and evaluation definition can resume the run.
            evaluation = make_evaluation(
                WorkflowEvalRedisStore(redis),
                WorkflowSubmissionCompletionPoll(poll=poll_submission),
            )
            # With no argument, starts a new run and prints its run ID.
            if len(sys.argv) == 1:
                result = await evaluation.start()
            # With a run ID, checks the run's pending batches once and advances
            # any cases whose results are ready.
            else:
                result = await evaluation.poll(sys.argv[1])
            print(result.run_id, result.status)

    if __name__ == "__main__":
        asyncio.run(main())
    ```

    To use it:

    * Running `python resume_eval.py` with no arguments submits the requests and prints the run ID.
    * Running `python resume_eval.py <run_id>` advances the run once. A scheduled job can run it periodically until it prints `completed`.
  </Accordion>

  <Accordion title="Resume from provider webhooks">
    Use [`WorkflowSubmissionCompletionWebhook`](/docs/sdks/python/api-reference#workflowsubmissioncompletionpoll-and-workflowsubmissioncompletionwebhook) when your provider sends an event when a request finishes. For OpenAI, configure a [webhook endpoint](https://developers.openai.com/api/docs/guides/webhooks) for `batch.completed` events, and use a shared Redis store. The following `webhook_handler.py` reuses `make_evaluation()` and `client` from `workflow_eval.py`:

    ```python webhook_handler.py expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
    import os
    from collections.abc import Mapping

    from redis.asyncio import Redis

    from braintrust.workflow_eval import (
        WorkflowEvalRedisStore,
        WorkflowSubmissionCompletionWebhook,
    )
    # Reuses the evaluation and OpenAI client from the main example.
    from workflow_eval import client, make_evaluation

    redis = Redis.from_url(os.environ["REDIS_URL"])
    # Uses webhook completion. After submit, the SDK saves the batch ID that
    # get_external_id returns, so a later webhook can find the submission.
    evaluation = make_evaluation(
        WorkflowEvalRedisStore(redis),
        WorkflowSubmissionCompletionWebhook(
            get_external_id=lambda submission, _context: submission["batch_id"],
        ),
    )

    # Call this from your HTTP endpoint with the raw request body and headers.
    async def handle_webhook(body: str, headers: Mapping[str, str]):
        # Checks that the request came from OpenAI, using the signing secret in
        # OPENAI_WEBHOOK_SECRET, and parses the event.
        event = client.webhooks.unwrap(body, headers)
        # Ignores failure events. Handle those in your application.
        if event.type != "batch.completed":
            return
        batch = await client.batches.retrieve(event.data.id)
        # The run ID comes from the batch metadata that submit_request() set.
        run_id = (batch.metadata or {}).get("braintrust_run_id")
        # Skips batches that this evaluation didn't submit.
        if run_id is None:
            return
        try:
            # The SDK finds the submission by batch ID, then calls collect_task().
            return await evaluation.process_submission_result(
                run_id, external_id=batch.id
            )
        except ValueError:
            # Skips batches that use polling, such as llm_judge's grading batches.
            return
    ```

    Pass the raw request body and headers from your HTTP endpoint to `handle_webhook()`, and start the run from a script that imports `evaluation` from `webhook_handler.py`.

    Keep in mind:

    * **Failed batches**: The example's `handle_webhook()` ignores failure events, so a failed batch stays pending and the run never completes. Handle failure events in your application.
    * **Security**: `process_submission_result()` doesn't check that the event came from the provider. In the example, `unwrap()` does.
    * **Repeated events**: The same event can arrive more than once. Make your `collect` callbacks safe to run more than once.

    For how the SDK matches an event to a request, see [`process_submission_result()`](/docs/sdks/python/api-reference#workfloweval).
  </Accordion>

  <Accordion title="Handle failures and concurrency">
    * **Concurrency**: `max_concurrency` limits how many of your functions the SDK runs at once during each call. It defaults to `10`. See [`define_workflow_eval()`](/docs/sdks/python/api-reference#define_workflow_eval).
    * **Provider failures**: When your polling callback returns [`WorkflowSubmissionPoll("failed", error=...)`](/docs/sdks/python/api-reference#workflowsubmissioncompletionpoll-and-workflowsubmissioncompletionwebhook), `evaluation.poll()` raises the error. The SDK checks the request again on the next `poll()`, so if the failure is permanent, start a new run. The SDK doesn't resubmit or cancel requests at the provider, so retries and cancellation are up to your integration.
    * **Errors in your functions**: The SDK retries your polling callback and your `collect` function on the next `poll()`. It runs your `submit` function and your plain task, scorer, and classifier functions only once per case, so if one of them fails, that case never finishes. Fix the problem and start a new run.
  </Accordion>
</AccordionGroup>

## Next steps

* Review [`define_workflow_eval()`](/docs/sdks/python/api-reference#define_workflow_eval) and [`WorkflowEval`](/docs/sdks/python/api-reference#workfloweval).
* Learn how to [run experiments in code](/docs/evaluate/run-in-code).
* Add [custom scorers](/docs/evaluate/write-scorers) to measure task quality.
