# Access control Source: https://braintrust.dev/docs/admin/access-control/index Control who can access your projects, experiments, and data using permission groups scoped to organizations, project groups, projects, and objects ## How permissions work Braintrust uses a hierarchical permission model built around *permission groups*, which are collections of users (or [service accounts](/docs/admin/access-control/manage-permissions#use-service-accounts)) that share a set of permissions. Permissions cascade down the following hierarchy: * **Organization**: The top level. Permissions granted at the organization level apply to every project and every object within those projects, including projects created later. * **Project group**: A named collection of projects. Permissions granted to a project group apply to every project in the group and the objects inside those projects. * **Project**: Permissions granted at the project level apply to a single project and the objects inside it. * **Object**: The most granular level. Permissions granted at the object level apply to that specific object (for example, experiment, dataset, log, prompt, or playground). ### Rules to consider Braintrust permissions are subject to the following rules: * **Permissions are additive only.** Permissions granted at a higher scope cannot be removed at a lower scope. If a user has organization-level `Read` permissions, you cannot use project-level groups to restrict their access to specific projects. Instead, remove them from the broader group and grant project-level access. Project groups follow the same rule, so they cannot narrow access a project already has. * **Group permissions are unioned.** A user can belong to multiple groups. Their effective permissions are the union of every group they're in. A project can belong to multiple project groups, and its permissions are the union of every project group it's in. * **Project group membership takes effect immediately.** Braintrust resolves a project's group memberships when it checks a permission, rather than copying permissions onto the project. Adding a project to a group, or removing it, changes effective access right away. * **The same model applies to service accounts.** Service accounts can be members of permission groups, just like users. The hierarchy and cascade rules apply identically. * **Creators get a direct owner grant by default.** When a user creates an object, they receive owner access to that object directly, in addition to any inherited organization and project permissions. Organization owners can turn this off for newly created objects with the **Create direct ownership grants** toggle at the top of ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). When an organization owner turns off this toggle, they also have the option to remove all *existing* direct owner grants. ### Inheritance labels on the project permissions page On a project's ** Settings** > [** Project permissions**](https://www.braintrust.dev/app/~/configuration/permissions) page, each entry with access on the **Permission groups**, **Members**, and **Service accounts** tabs shows **Direct access** or **Inherited access**. An **Inherited access** label lists the sources of that access in parentheses, for example **Inherited access (Engineers, Organization Project)**. When more than three permission groups contribute, the label shows the first three and the number of others. Each source is one of two kinds. **Organization**, **Organization Project**, and **Project group** say where a grant was made. A permission group name says which permission group the access comes through. | Source | What it means | | - | - | | **Organization** | A grant on the organization itself, from membership in **Owners** or from the **Organization** column of a permission group's, member's, or service account's permissions. Not every grant with this label reaches the project. **Owners** membership grants full access to the project, and **Manage access** lets the holder manage the project's permissions. Other **Organization** permissions, such as **Manage settings**, also show this label but don't grant access to the project's data. | | **Organization Project** | A grant from the **All projects** column. It covers every project in the organization, including projects created later. | | **Project group** | A grant on a [project group](/docs/admin/access-control/manage-permissions#manage-project-groups) that contains this project. The grant goes to a permission group, member, or service account, and applies to every project in the project group. | | A permission group name, such as **Engineers** | Access comes through that permission group. The member or service account belongs to it, either directly or through another permission group. See [Manage group membership](/docs/admin/access-control/manage-permissions#manage-group-membership). | | **Organization**, **Organization Project**, or **Project group** with no permission group name | On the **Members** and **Service accounts** tabs, access is granted directly to that member or service account, not through a permission group. On the **Permission groups** tab, access is granted to the permission group itself. | Permission group names and the other sources are listed separately, so the label doesn't say which permission group holds which grant. For example, **Inherited access (Engineers, Organization Project)** can mean that the Engineers permission group has an **All projects** grant, that the member has their own **All projects** grant and also gets access through Engineers, or both. To tell them apart, check the member's [direct permissions](/docs/admin/access-control/manage-permissions#set-direct-permissions). Because permissions are additive, project-level settings cannot remove inherited access. To change who can reach a project, change the grant where it comes from: * **A permission group's grant:** [Edit the permission group's permissions](/docs/admin/access-control/manage-permissions#set-organization-permissions). To remove only one member's access, [remove them from the permission group](/docs/admin/access-control/manage-permissions#manage-group-membership). * **A member's direct grant:** [Edit the member's direct permissions](/docs/admin/access-control/manage-permissions#set-direct-permissions). * **A service account's direct grant:** [Edit the service account's direct permissions](/docs/admin/access-control/manage-permissions#use-service-accounts). ### Plan-specific access control features The level of access control available depends on your [plan](/docs/plans-and-limits#plans): | Capability | Starter | Pro | Enterprise | | - | :-: | :-: | :-: | | **Owners** built-in group | ✓ | ✓ | ✓ | | **Engineers**, **Viewers**, and **All AI Provider Access** built-in groups | — | ✓ | ✓ | | Custom permission groups | — | — | ✓ | | Project groups | — | — | ✓ | | Project-level permissions | — | — | ✓ | | Object-level permissions (AI providers, experiments, datasets, logs, prompts, playgrounds) | — | — | ✓ | To grant permissions at the project or object level, use custom permission groups (Enterprise only). Built-in permission groups (all plans) apply to the entire organization. ## Built-in permission groups Braintrust provides built-in permission groups for managing team access. **These groups are scoped to the entire organization.** Permissions granted through a built-in group cascade to all projects in the organization, including projects created later. | Group | Access | Plan availability | | - | - | - | | **Owners** | Full access to organization, projects, data, and settings. Can invite/remove members, manage permissions, and delete resources. | All plans | | **Engineers** | Create, read, update, and delete projects and resources across the organization. Cannot manage members, access controls, or organization settings. | Pro and Enterprise | | **Viewers** | Read-only access to all projects and resources across the organization. Cannot create, update, or delete anything. | Pro and Enterprise | | **All AI Provider Access** | Read access to every organization AI provider, and no other permissions. See [Organization AI providers](/docs/admin/ai-providers#organization-ai-providers). | Pro and Enterprise | To assign a user to a built-in group, invite them from ** Settings** > [** Members**](https://www.braintrust.dev/app/~/configuration/org/team), add them to the group's member list from ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups), or use that member's **Manage permissions** menu on the Members page. See [Manage group membership](/docs/admin/access-control/manage-permissions#manage-group-membership) for details. Built-in groups are groups with default names and permissions. An Owner can *add* permissions to a built-in group but cannot *remove* any default permissions. ### When to use custom permission groups Use custom permission groups (Enterprise) when you need to: * Grant access to a subset of projects rather than the entire organization. * Restrict access to specific organization AI providers or objects within a project (experiments, datasets, logs, prompts, or playgrounds). * Assemble a permission set that doesn't match Owners, Engineers, Viewers, or All AI Provider Access (for example, a group that can read data across all projects but only update prompts in one). Project-level read is often all you need to share resources across projects. For example, granting read access to a central scorer project makes its scorers available in dropdowns across the organization, without organization-wide read access. See [Use scorers from another project](/docs/evaluate/score-online#use-scorers-from-another-project). For setup instructions, see [Create custom permission groups](/docs/admin/access-control/manage-permissions#create-custom-permission-groups). ## Permissions reference Each permission applies to a scope (organization, project, or object) or controls administrative actions at the organization or group level. When granted at the organization level, scope-based permissions cascade to all projects and the objects within them. ### Object permissions These permissions apply to organizations, projects, organization AI providers, and individual objects within projects (experiments, datasets, logs, prompts, and playgrounds). | Permission | What it allows | | - | - | | **Read** | View the object and its contents. At the project level, lets users see the project and its data. Includes data export via the API and SDK. No separate export-only permission exists. In the UI, download and export controls are shown only to users who also have `Update` or `Delete`. | | **Create** | Create new resources within the scope. At the project level, includes creating experiments, datasets, logs, prompts, and playgrounds. | | **Update** | Modify existing resources within the scope. For the full list, see [What `Update` covers](#what-update-covers). | | **Delete** | Remove resources. Granted independently of `Update`; having `Update` does not let you delete. | | **Manage access** | Grant and revoke permissions on this object. A super-user permission: a user with `Manage access` on a scope can grant themselves any other permission on that scope. Assign carefully. | ### Organization-only permissions These permissions exist only at the organization level and control administrative actions on the organization itself: | Permission | What it allows | | - | - | | **Manage settings** | Change organization-level configuration, including the API URL and billing. | | **Invite members** | Invite new users to the organization. | | **Remove members** | Remove users from the organization. (At least one member must remain.) | | **Read audit logs** | Read audit log entries for the organization, such as permission changes, member invites, and API key creation. | ### What `Update` covers The `Update` permission controls modifications to existing resources. At the project scope, it includes: **Project-level resources:** * Project name and description * Project scores (human review scores; not automated evaluation scores). Per-score visibility is a soft filter, not an access boundary. See [Restrict score visibility](/docs/annotate/human-review#restrict-score-visibility). * Project tags * Span iframes * MCP servers * Project automations, including retention policies * Project-level AI provider credentials. See [Configure AI providers](/docs/admin/ai-providers#permissions-and-access). **Resources inside the project:** * Experiment metadata and event data * Dataset metadata and records * Logs `Update` does **not** include `Create`, `Delete`, or `Manage access` — each is granted separately. ### Permissions vs. Manage access `Permissions` and `Manage access` control different things: * **Permissions** define what group members can do across all projects (organization-level) or within a single project (project-level). When you click **Permissions** on a group, you are setting what that group's members can do in Braintrust. * **Manage access** defines who can administer the group itself: who can invite new members, rename it, edit its permissions, or grant access to it. When you grant `Manage access` on a group, you are deciding who controls the group — not what its members can do. The same distinction applies to projects and objects: `Manage access` on a project controls who can change the project's permissions, while the other permissions control what those people can do with the project's data. ## Next steps * [Manage permissions](/docs/admin/access-control/manage-permissions) to create permission groups, set organization and project permissions, and provision service accounts. * [Manage organizations](/docs/admin/organizations) to invite members and assign groups. * [Manage projects](/docs/admin/projects) to configure project-level permissions. # Manage permissions Source: https://braintrust.dev/docs/admin/access-control/manage-permissions Create permission groups, assign permissions, and provision service accounts Set up permission groups, assign members, set organization and project permissions, and provision service accounts for system integrations. For the permission model, see the [Access control overview](/docs/admin/access-control). ## Create custom permission groups Build groups with specific permissions: only available on the [Enterprise plan](/docs/plans-and-limits#plans). 1. Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). 2. Click **Create permission group**. 3. Enter a name and description. 4. Set the group's permissions inline. Configure organization-level permissions for the **Organization** and **All projects** columns, plus project-specific and object-level permissions in the **Project-specific permissions** section. 5. Click **Create**. ## Manage access to a permission group Control who can administer a permission group itself: who can view it, edit its permissions, rename it, or grant others access to it. This is separate from the permissions the group grants its members. For the distinction, see [Permissions vs. Manage access](/docs/admin/access-control#permissions-vs-manage-access). only available on the [Enterprise plan](/docs/plans-and-limits#plans). 1. Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). 2. Find the group in the permission groups list, then click the more options menu () on its row. 3. Select ** Manage access**. 4. In the **Object permissions** dialog, select the tab for who you want to grant access to: **Permission groups**, **Members**, or **Service accounts**. 5. Search for the user, group, or service account, then click the edit icon next to it. 6. Select the permissions to grant on the group: * **Read**: View the group and its permissions. * **Update**: Edit the group's name, description, and permissions. * **Delete**: Delete the group. * **Manage access**: Grant and revoke access to the group (super-user ability). 7. Click **Save**. ## Set organization permissions Grant organization-level permissions to custom groups: 1. Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). 2. Find the group in the permission groups list, then click **Permissions** on its row. 3. Select organization-level permissions: * **Manage settings**: Change organization configuration. * **Invite members**: Invite users. * **Remove members**: Remove users (organizations must have at least one member). * **Manage access**: Grant and revoke permissions (super-user ability). * **Read audit logs**: Read organization audit log entries. 4. Select permissions for all projects: * **Read**: View all projects and their resources. * **Create**: Create projects, and create experiments, logs, and datasets in all projects. * **Update**: Modify existing resources in all projects. * **Delete**: Remove resources from all projects. * **Manage access**: Grant permissions on all projects. 5. (Optional) Select project-specific and object-level permissions in the **Project-specific permissions** section. This section lets you set project-specific and object-level permissions directly from the permission group dialog, without going to each project's **Project permissions** settings. 6. Click **Save**. **Manage access** is a super-user permission. Users with this permission can grant themselves any other permission. Assign it carefully. **Manage settings** grants users the ability to change organization-level settings, like the API URL. ## Set project permissions Specify a group's permissions for a particular project and its objects: 1. [Create a custom permission group](#create-custom-permission-groups). 2. In your project, go to ** Settings** > [** Project permissions**](https://www.braintrust.dev/app/~/configuration/permissions). 3. Search for your group. 4. Click the pencil icon next to the group. 5. Select project permissions: * **Read**: View project and its resources. * **Create**: Create experiments, logs, datasets. * **Update**: Modify existing resources. * **Delete**: Remove resources. * **Manage access**: Grant permissions on this project. 6. Select object-level permissions for experiments, datasets, logs, prompts, playgrounds, functions, scorers, and classifiers: * **Create**: Create the object. * **Read**: View the object. * **Update**: Modify the object. * **Delete**: Remove the object. * **Manage access**: Grant permissions on this object. 7. Click **Save**. Users must have Read permission on a project to see it in the UI. ## Manage project groups A project group is a named collection of projects. Grant a permission group access to the project group once, and every project in the group inherits that access. Without project groups, giving a team the same access across 100 projects means creating and maintaining 100 separate project-level grants. For how project-group access works, see [How permissions work](/docs/admin/access-control#how-permissions-work). only available on the [Enterprise plan](/docs/plans-and-limits#plans). Use a project group when the same team, role, or service account needs the same access across multiple projects. You create a group, grant access to it, and choose its projects in a single sheet: 1. Go to ** Settings** > [** Project groups**](https://www.braintrust.dev/app/~/configuration/org/project-groups). 2. Click **Create project group**. 3. Enter a **Name** and an optional **Description**. 4. In the **Permissions** section, choose which permission groups, members, and service accounts can access the group. Select the **Permission groups**, **Members**, or **Service accounts** tab, click the edit icon () next to a group, member, or service account, select the permissions to grant, then click **Apply**. You can also grant access later from the permission group itself. See [Set project group permissions](#set-project-group-permissions). 5. In the **Projects** section, select the projects to include. A project can belong to more than one project group. 6. Click **Create project group**. Grant access to permission groups rather than to individual members whenever you can. Group-based ownership is easier to audit, and access stays correct as people join and leave teams. Use members for one-off exceptions, and service accounts for automations and integrations. To change a group's name, description, permissions, or projects, click its row in the project groups list, or click the edit icon () on the row. Make your changes, then click **Save**. Membership is edited in this sheet, not from the **Projects** column in the list. Adding or removing a project requires **Manage access** on that project, so you can only change membership for projects you administer. Removing a project from a group revokes only the access that group granted. Permissions the project has from other sources stay in effect. A project group can contain up to 10,000 projects. A single save can change up to 1,000 project assignments, counting additions and removals together. To assign more projects than that, save in batches. To remove a group, click the delete icon () on its row. You need **Manage access** on every project in a group to delete it. Deleting a project group removes the permissions its member projects inherited from the group. The projects themselves are not deleted. Deleting a group frees its name for reuse. ## Set project group permissions Grant a permission group access to a project group from the permission group side. To grant the same access while creating the project group, use the **Permissions** section described in [Manage project groups](#manage-project-groups). 1. Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). 2. Find the group in the permission groups list, then click **Permissions** on its row. 3. In the **Project and project-group permissions** section, find **Project groups** and click **Project group**. 4. Select a project group, then set its permissions: * **Read**: View the projects in the group and their resources. * **Create**: Create experiments, logs, and datasets in those projects. * **Update**: Modify existing resources in those projects. * **Delete**: Remove resources from those projects. * **Manage access**: Grant permissions on those projects. 5. (Optional) Set object-level permissions for experiments, datasets, logs, prompts, playgrounds, functions, scorers, and classifiers. These work the same as they do at the project level, and apply to those objects in every project in the group. See [Set project permissions](#set-project-permissions). 6. Click **Save**. ## Manage group membership Users can belong to multiple permission groups, either directly or through a group that is itself a member of another group. Their effective permissions are the union of all group permissions. To change the groups a specific user belongs to: 1. Go to ** Settings** > [** Members**](https://www.braintrust.dev/app/~/configuration/org/team). 2. Find the member, then click **Manage permissions** on their row. 3. Select **Edit permission groups**. 4. The dialog lists the groups you can manage under **Member of**, **Member of via inheritance** (groups the user joins through another group's membership), and **Not a member of**. 5. To add: Click **+** next to a group under **Not a member of**. 6. To remove: Click the **x** next to a group under **Member of**. To view and edit all members of a specific group: 1. Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). 2. Find the group in the permission groups list. 3. Click **Members**. 4. To add: Search for users and click **+**. 5. To remove: Click the **x** next to a user's name. Both paths are disabled for permission groups managed by [SCIM provisioning](/docs/admin/scim). To change membership, edit the mapped group in your identity provider instead. Groups without a SCIM mapping are unaffected. ## Set direct permissions A direct permission is granted to one member or service account instead of through a permission group, so changing a permission group's permissions or membership doesn't affect it. To find where a member's access to a project comes from, see [Inheritance labels on the project permissions page](/docs/admin/access-control#inheritance-labels-on-the-project-permissions-page). To set a service account's direct permissions, see [Use service accounts](#use-service-accounts). only available on the [Enterprise plan](/docs/plans-and-limits#plans). To set a member's direct permissions: 1. Go to ** Settings** > [** Members**](https://www.braintrust.dev/app/~/configuration/org/team). 2. Find the member, then click **Manage permissions** on their row. 3. Select **Edit direct permissions**. 4. Select the permissions to grant or remove. The dialog has the same columns and **Project-specific permissions** section as a permission group's permissions. See [Set organization permissions](#set-organization-permissions). Permissions the member inherits from a permission group are shown but can't be edited here. 5. Click **Save**. ## Use service accounts A service account is an identity for system integrations, authenticated with a service token that you use like an API key. Unlike a personal API key, which inherits the full permissions of the user who created it, a service account is a separate identity whose permissions come from its own permission group memberships and isn't tied to any individual. Service tokens also support [user impersonation](/docs/api-reference#impersonate-users). The service account must have the `Owner` role in every organization the target user belongs to, and the target user must belong to at least one organization. Braintrust checks the service account's permissions, not those of the token's creator. Use service accounts to: * **Authenticate integrations and [automations](/docs/admin/data-management/export)** that shouldn't depend on one person's account. A service account keeps working even after team members leave. * **Grant least-privilege, project-scoped access.** Because its permissions come from its own groups, you can scope a service account more restrictively than your own access, for example to a single project or a subset of projects. A personal API key can't be scoped below your own permissions. * **Separate environments.** Create distinct tokens for development, staging, and production, each assigned to a permission group scoped to the right projects. To create a service account: 1. Go to ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens). 2. Click **+ Service token**. 3. Enter a service account name. 4. Assign permission groups or grant specific permissions. To scope the account to specific projects, assign it to a permission group limited to those projects. 5. Click **Create**. 6. Copy and save the auto-generated service token somewhere safe and accessible. For security reasons, you will not be able to view it again. If you lose the service token, you must create a new one. 7. Use the token like an API key in SDK or API calls. Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can edit a service account's permission groups or direct permissions after creation. To edit its permission groups: 1. Go to ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens). 2. Find the service account, then click the more options menu () on its card. 3. Select **Edit permission groups**. 4. The dialog lists the groups you can manage under **Member of**, **Member of via inheritance**, and **Not a member of**. 5. To add: Click **+** next to a group under **Not a member of**. 6. To remove: Click the **x** next to a group under **Member of**. To set its direct permissions: 1. Go to ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens). 2. Find the service account, then click the more options menu () on its card. 3. Select **Edit direct permissions**. 4. Select the permissions to grant or remove. Permissions the service account inherits from its permission groups are shown but can't be edited here. To change them, update the service account's group membership instead. 5. Click **Save**. Only organization owners can create service tokens, at ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens) in the Braintrust UI or by calling [`POST /v1/service_token`](/docs/api-reference/servicetokens/create-service_token) with a service token that has organization-owner permissions. User API keys cannot be used to create service tokens. Users with permission to add organization members can create service accounts by calling [`PATCH /v1/organization/members`](/docs/api-reference/organizations/modify-organization-membership). To also create an initial service token, include `token_name` (this requires authenticating with a service token that has organization-owner permissions). For self-hosted deployments, you must configure a service token for the data plane to enable features like data retention. See [Data plane manager](/docs/admin/self-hosting/configure/telemetry#data-retention) for more details. ## Programmatic access control To automate the creation of permission groups and their access control rules, use the Braintrust API. See the API reference for [groups](/docs/api-reference/groups/list-groups) and [permissions](/docs/api-reference/acls/list-acls). ## Next steps * Review the [permissions reference](/docs/admin/access-control#permissions-reference) to understand what each permission grants. * [Set up automations](/docs/admin/data-management/export) with service accounts. * [API reference](/docs/api-reference/groups/list-groups) for programmatic access control. # Configure AI providers Source: https://braintrust.dev/docs/admin/ai-providers Manage organization-level and project-level AI provider credentials for playgrounds, experiments, and the gateway. Braintrust manages AI provider credentials on the ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets) page, so you can share access without each team member needing an individual API key. Most configured providers are available in playgrounds, experiments, and the gateway. AI providers can be configured at two scopes: * **Organization AI providers** are the defaults available across every project in the organization. * **Project AI providers** are configured for the selected project. They can override matching organization providers or provide additional options. Each row shows a **Last updated** timestamp that tracks when the key value itself was last changed, along with the user who made the change. If you have permission to edit the section, each row also shows a redacted preview of the key (for example, `abc...xyz`). Renaming a provider or editing other metadata does not bump this timestamp. Keys that have not been rotated in over six months display a warning indicator. Braintrust recommends disabling and rotating AI provider secrets periodically. ## Add a provider For provider-specific configuration (authentication methods, regions, model registries), see the [AI provider integrations](/docs/integrations/ai-providers). ### Organization providers Organization providers make credentials available across projects, subject to [provider permissions](#permissions-and-access). 1. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 2. Under **Organization AI providers**, click **Organization provider** and choose the provider you want to configure. 3. On **Setup**, enter your API key or configure another supported [authentication method](#authentication). Complete any required provider settings, such as the endpoint or region. 4. Open **More settings** to configure request settings, additional headers, and [model restrictions](#restrict-available-models). 5. Click **Create**. The sheet opens on **Models**, where you can [browse models and test a request](#browse-and-test-models). ### Project providers Project providers let you configure credentials for a specific project. Use them when you need separate billing or rate limits, isolated API usage, different provider accounts or credentials, or a specific regional endpoint (for example, US-specific OpenAI keys to keep traffic in-region). To add a provider for your project: 1. In your project, navigate to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 2. Under **Project AI providers**, click **Project provider** and choose the provider you want to configure. 3. On **Setup**, enter your credentials and complete any required settings for the provider. See the [AI provider integrations](/docs/integrations/ai-providers) for provider-specific instructions. 4. Open **More settings** to configure request settings, additional headers, and [model restrictions](#restrict-available-models). 5. Click **Create**. The sheet opens on **Models**, where you can [browse models and test a request](#browse-and-test-models). You can also add a project-level provider inline from playgrounds within that project. When you attempt to run a playground without a configured provider, you'll see an option to add your API key without leaving the page. ### Custom providers Braintrust supports custom AI providers at both the organization and project level. Add them from the same **Organization provider** or **Project provider** picker. See [Custom providers](/docs/integrations/ai-providers/custom) for endpoint configuration, headers, streaming, and cost metadata. ### Authentication Most providers authenticate with a long-lived API key. Some also support alternatives that avoid storing a long-lived provider credential in Braintrust: * **Workload identity federation**: Braintrust exchanges a short-lived, Braintrust-signed OIDC token for a provider access token at request time, so no long-lived key is stored. Available for [OpenAI](/docs/integrations/ai-providers/openai), [Anthropic](/docs/integrations/ai-providers/anthropic), [Google Vertex AI](/docs/integrations/ai-providers/google), and [Azure AI Foundry](/docs/integrations/ai-providers/azure), for organization-level providers on Braintrust-hosted organizations with the gateway enabled. * **Cloud-native role assumption**: [Bedrock](/docs/integrations/ai-providers/bedrock) supports AWS `AssumeRole` instead of storing long-lived access keys. See each provider's integration page for setup details. ## Update a provider To change a configured provider's API key or other settings: 1. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 2. In the row for the provider you want to change, click the edit icon, or click the row itself. 3. On the **Setup** tab, update the API key or other settings. 4. Open **More settings**, if collapsed, to configure request settings, additional headers, and [model restrictions](#restrict-available-models). 5. Click **Update**. The sheet opens on **Models**. Configured providers also appear in the **Organization provider** and **Project provider** pickers with a edit icon. Click a configured tile to edit it without leaving the picker. Custom providers appear only in the table, not the picker. ## Delete a provider 1. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 2. In the row for the provider you want to remove, click the delete icon. You can also open the provider's **Setup** tab and click **Delete** in the footer. 3. Confirm the deletion in the dialog that appears. Deleting a provider removes its credentials and configuration. Requests that rely on the provider fail unless another eligible provider can handle them. ## Browse and test models Open a configured provider to see its models and test a request through the [Gateway](/docs/deploy/gateway). 1. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 2. Click the provider's row, then select **Models**. If you don't have permission to edit the provider, the sheet opens on **Models** and **Setup** is disabled. 3. Under **Available models**, search by model name or ID. 4. Choose a model from the selector above the request example. 5. Use an existing Braintrust API key in place of `$BRAINTRUST_API_KEY`, or click **Generate API key**, enter a name, and click **Create API key**. The generated key is a Braintrust API key and is inserted into the command. 6. Copy the `curl` command and run it in your terminal. The command requests the selected model through this provider. For a project provider, it also includes the project's ID. A successful request returns a model response. The list includes models Braintrust supports for this provider, filtered by the [model restrictions](#restrict-available-models) you've configured, plus any custom models you've added. It excludes deprecated and embedding models, and does not verify which models your upstream provider account can access. Run a request to confirm access to a specific model. If no models are listed, the request example is unavailable. If you can edit the provider, check its [model restrictions](#restrict-available-models) or add a **Custom model** on **Setup**, then save the configuration. Gateway requests use the permissions of the user or service account that owns the Braintrust API key. Generating a key does not grant additional provider access. See [Permissions and access](#permissions-and-access). ## Request routing Multiple organization and project providers can be configured to handle requests that target the same model. When a request arrives, it specifies a particular model. Braintrust determines which providers are eligible to handle the request, then selects one to handle it. ### Provider eligibility First, Braintrust determines which organization and project providers are eligible to handle the incoming request. To be eligible, a provider must support the requested model, allow it under its [model restrictions](#restrict-available-models), and be accessible to the caller through [provider permissions](#permissions-and-access). Project providers can also make matching organization providers ineligible, as described below. Consider the following scenarios: * **Matching organization and project providers both support and allow the requested model** The project provider remains eligible, and the organization provider becomes ineligible. * **An organization provider supports and allows the requested model, but a matching project provider does not** Whether the organization provider remains eligible depends on your deployment: * **Braintrust-hosted:** The project provider is ineligible because it does not support the requested model or excludes it under model restrictions. The organization provider remains eligible. (Therefore, excluding a model through a project provider's [model restrictions](#restrict-available-models) does not block access to that model through an organization provider the caller has permission to use.) * **Self-hosted:** The project provider does not support the requested model or excludes it under model restrictions. It still overrides the organization provider, making both ineligible. The request fails unless another eligible provider can handle the request. The legacy [AI proxy](/docs/deploy/ai-proxy) excludes a matching organization provider only when the project provider supports and allows the requested model. In this scenario, the organization provider remains eligible. * **Nonmatching organization and project providers both support and allow the requested model** Both providers remain eligible for the request. Neither overrides the other, and the project provider does not automatically take priority. For example, a project-level custom provider named `Staging OpenAI` does not override an organization-level custom provider named `Production OpenAI`. If both support and allow `gpt-5-mini` and the caller can use both, both remain eligible for the request. When an organization provider has a matching project provider, the organization provider's row in ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets) displays an **Overridden** badge. [Model restrictions](#restrict-available-models) do not affect the badge, so the badge does not indicate which provider handles a particular request. ### Provider selection Braintrust selects a provider from those eligible for the request. If multiple providers are eligible, Braintrust selects one automatically. The legacy [AI proxy](/docs/deploy/ai-proxy) randomly selects an eligible provider for the first attempt. If that attempt fails, the proxy can retry with other eligible providers. **Specify a particular provider** To specify which provider should handle a given request, set the [`x-bt-endpoint-name`](/docs/deploy/gateway#advanced-configuration) header to the provider's configuration name, as shown on the provider's row in ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets): | Provider | Header value | | - | - | | Built-in OpenAI | `OPENAI_API_KEY` | | Built-in Anthropic | `ANTHROPIC_API_KEY` | | Custom provider named `Staging OpenAI` | `Staging OpenAI` | For built-in providers, use the identifier displayed beneath the provider's friendly name. For custom providers, use the exact name, including capitalization. The command in the provider's **Models** tab includes the correct header value. This header selects a provider by name, not by organization or project scope. If a project provider overrides an organization provider with that name, the routing rules above still apply. Specifying a provider does not bypass permissions or model restrictions. To confirm which provider handled a request, check the `x-bt-used-endpoint` response header, which contains the name of the provider that handled it. ## Restrict available models Model restrictions let you choose which models a provider allows. Combined with [provider permissions](#permissions-and-access), these restrictions let you control which models members and service accounts can use. 1. If you're restricting a project-level provider, make sure you're in that project. 2. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 3. Click the provider's row, then select **Setup**. 4. Expand **More settings**. 5. Under **Models**, set **Registry models**: * **All models**: Include all models in Braintrust's registry for this provider. * **Selected models**: Include only the registry models you select. Search the list and check each model you want to allow. * **None**: Exclude registry models. Only models you add with **Custom model** remain available. 6. Click **Update** to save the configuration, or **Create** if you are adding a provider. When configuring model restrictions, keep the following in mind: To edit an organization provider in the UI, you need the organization's **Manage settings** permission. To edit a project provider, you need the project's **Update** permission. The **Update** permission on an organization provider allows API changes but does not grant editing access in this UI. See [Permissions and access](#permissions-and-access). Organization providers support model restrictions on both Braintrust-hosted and self-hosted deployments. For project providers, **Selected models** is available on Braintrust-hosted deployments. Self-hosted deployments support **All models** and **None**, along with custom models. Custom models remain available in all three modes. To remove a custom model, remove its entry from the provider's configuration. **None** does not disable the provider or remove its custom models. Before excluding a model, check whether any playgrounds or LLM scorers use it, including scorers used for online scoring. Their model requests may fail unless another eligible provider can handle those requests. ### Combine model restrictions with permissions Model restrictions determine which models a provider allows, while provider permissions determine who can use that provider. On Enterprise, you can combine them to give teams access to different model selections. To set this up: 1. Create separate organization provider configurations for the model selections each team needs. If teams need different model selections from the same upstream provider, create separate [custom providers](/docs/integrations/ai-providers/custom) with distinct names. 2. Set each provider's [model restrictions](#restrict-available-models) to allow only the models that team should use. 3. In [organization provider permissions](#organization-ai-providers), grant each team's permission group **Read** access to that team's provider. This lets members use the models allowed by that provider. Permissions are additive. To ensure each team can access only the intended models, also check: Restricting a team's provider does not prevent members from using a model through another organization provider they can access. For example, membership in **All AI Provider Access** grants **Read** access to every organization provider, including providers that allow models outside the team's intended set. Restricting an organization provider does not restrict access through project providers. If a project provider allows a model excluded from the team's organization provider, members with [**Read** access to that project](#project-ai-providers) can still use the model through the project provider. [Braintrust built-in models](#manage-built-in-models) are available without configuring your own provider credentials. Restrictions on your configured providers do not control access to these models. Disable [**Allow built-in models**](#enable-or-disable-built-in-models) in ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets) if members should use only your configured providers. ## Permissions and access Use permissions to control who can use and manage your AI providers. At the organization level, permissions can apply to all providers or to a specific provider. At the project level, permissions control who can use and manage all project providers. ### Organization AI providers Permissions can apply to all organization AI providers or to an individual provider: Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups), then select a group. Under **Organization**, the **Manage settings** permission lets group members add providers and edit them from the **AI providers** page. Under **AI providers**, set the following permissions as needed: * **Read** lets group members view and use providers anywhere Braintrust calls them, including playgrounds, experiments, scorers, and the [Gateway](/docs/deploy/gateway). * **Update** lets group members rename a provider or change its configuration through the [API](/docs/api-reference). * **Delete** lets group members delete providers. * **Manage access** lets group members grant other permission groups access to providers. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets), then select the provider's **Provider permissions** icon. On the **Permission groups**, **Members**, or **Service accounts** tab, select who receives access, then set the following permissions as needed: * **Read** lets recipients view and use the provider anywhere Braintrust calls it, including playgrounds, experiments, scorers, and the [Gateway](/docs/deploy/gateway). * **Update** lets recipients rename the provider or change its configuration through the [API](/docs/api-reference). * **Delete** lets recipients delete the provider. * **Manage access** lets recipients grant others access to the provider. **Built-in permission groups** To ensure continued provider access for existing accounts during the rollout of organization AI provider permissions, Braintrust created the **All AI Provider Access** group in every organization and added all existing members and service accounts to it. The group grants **Read** on every organization AI provider and no other permissions. New members are not added automatically. Your plan determines which [permission groups](/docs/admin/access-control#built-in-permission-groups) you can assign accounts to: * **Starter**: You can assign accounts only to **Owners**. Owners have all organization AI provider permissions. * **Pro**: You can assign accounts to any built-in group. **Engineers**, **Viewers**, and **All AI Provider Access** grant **Read** on every organization AI provider. * **Enterprise**: All built-in groups, plus custom groups. Custom groups can grant permissions across every organization AI provider or on individual providers. Permissions are additive: a user or service account receives the combined access of every group they belong to. An account may therefore appear in **All AI Provider Access** even when another group also grants provider access. Before removing an account from **All AI Provider Access**, confirm it has **Read** access to each provider it needs through another permission group or a provider-specific permission. **Service accounts** Assign service accounts you create to permission groups that grant access to the providers they need. Each service token inherits the permissions of its service account. Braintrust adds service accounts it provisions to **All AI Provider Access** so features such as online scoring and [Topics](/docs/observe/topics) can call organization-level AI providers. Removing these accounts from the group can prevent those features from working. ### Project AI providers Project permissions apply to every provider in the project and do not affect access to organization-level providers. You cannot configure permissions for a specific project provider. To configure a permission group's ability to use or manage project-level AI providers: * Go to the project's ** Settings** > [** Project permissions**](https://www.braintrust.dev/app/~/configuration/permissions), then select a permission group. * Under **Project**: * **Read** lets group members view the project and use all its AI providers. * **Update** lets group members add, edit, and delete project-level AI providers. To allow provider use without allowing changes, grant **Read** but not **Update**. The **Update** permission controls project resources other than AI providers. Review [What `Update` covers](/docs/admin/access-control#what-update-covers) before removing it. ## Manage built-in models Braintrust provides a set of models that your organization can use without configuring its own AI provider. They back the built-in model choices in playgrounds, prompts, and scorers, the models that [Loop](/docs/loop) and [Patterns](/docs/observe/patterns) run on, [message translation](/docs/observe/examine-traces#translate-message-content) in traces, and the facet summarization, embeddings, and cluster naming behind [Topics](/docs/observe/topics). ### Available models In playgrounds, prompts, and scorers, you can select these models from the **Braintrust** provider. When your organization has no AI providers configured, GLM-5.2 is selected by default. On Pro and Enterprise plans, you can also call each model directly through the [gateway](/docs/deploy/gateway) by requesting the model ID below. See [Requirements](#requirements). | Model | Model ID | | - | - | | GLM-5.2 | `glm-5.2` | | GLM-5.3 Flash | `glm-5.3-flash` | | Kimi K3 | `kimi-k3` | | DeepSeek V4 Flash 0731 | `deepseek-v4-flash-0731` | | DeepSeek V4.1 Flash | `deepseek-v4.1-flash` | Loop and Patterns run on GPT-6 Astra, GPT-6 Luna, GPT-6.1 Sol, GPT-6 Sol, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna as built-in models. You select them from Loop's model picker rather than from the **Braintrust** provider, and they aren't available in playgrounds, prompts, or scorers. See [Choose a provider and model](/docs/loop#choose-a-provider-and-model). Topics runs on a family of `brain-*` models that Braintrust selects for you. They aren't selectable anywhere in the product. ### Requirements Using a built-in model requires both of the following: * **Built-in models are allowed for your organization.** Braintrust-hosted organizations have them on by default. Self-hosted organizations have them off by default, so that no trace data leaves your network boundary, and must turn them on first. * **On the Starter plan, an eligible owner or a payment method.** Your organization needs at least one owner with a work email address, or a payment method on file. Braintrust checks the email domain of each [organization owner](/docs/admin/access-control#built-in-permission-groups) against a maintained list of personal email providers. The Starter requirement covers every built-in model you can select: the open-source models in playgrounds, prompts, and scorers, and the models that Loop and Patterns run on. Topics' `brain-*` models are exempt. Until your organization qualifies, the affected models don't appear in any model picker, and gateway requests for them return HTTP 403. To qualify, [add a payment method](/docs/admin/billing/change-plan#enable-on-demand-usage), or add an organization owner whose email uses a work domain. Access resumes as soon as ownership changes. Calling a built-in model directly through the [gateway](/docs/deploy/gateway) or the [AI proxy](/docs/deploy/ai-proxy) with a Braintrust API key or service token requires a Pro or Enterprise plan. On the Starter plan, these requests return HTTP 403 with the message `Native inference requires a Pro or Enterprise plan`, even when your organization meets the Starter requirement above. Built-in models still work everywhere Braintrust runs them for you: playgrounds, prompts, scorers, Loop, Patterns, and the `invoke()` function for saved prompts and scorers. To call built-in models directly, [upgrade your plan](/docs/admin/billing/change-plan#upgrade-your-plan). ### Enable or disable built-in models 1. Go to ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets). 2. Under **Built-in models**, turn **Allow built-in models** on or off. [Jev](/docs/integrations/ai-providers/typesafe#llm-as-a-judge) needs a second opt-in of its own. With built-in models allowed, click **Enable Jev** and authorize Braintrust to engage TypeSafe as a subprocessor. The same permission requirement below applies to that setting. Disabling built-in models forces Loop to run on your own AI providers. **Braintrust** does not appear as a provider option in Loop until you re-enable built-in models. See [Models and providers](/docs/loop/manage#models-and-providers). Disabling built-in models automatically pauses your topic automations. Re-enabling built-in models does not automatically resume them. You must [resume](/docs/observe/topics/manage#pause-resume-automation) each automation manually. Only members of the **Owners** [permission group](/docs/admin/access-control), or a custom permission group with the **Manage settings** organization permission, can enable or disable built-in models. ### Cost and credits Built-in Jev is free to use and does not consume model credits. The following billing rules apply to other built-in models. On Starter and Pro plans, usage of built-in models draws down your [model credits](/docs/plans-and-limits#usage-limits), shared with Topics, and continues at the [on-demand token rates](/docs/plans-and-limits#model-credits) once your credits are exhausted. On Starter, they become unavailable once you use up the credit, until you [enable on-demand usage](/docs/admin/billing/change-plan#enable-on-demand-usage) or upgrade to Pro. Pro and on-demand usage continue at the same rates beyond the credit. ### Data handling Braintrust hosts GLM-5.2, GLM-5.3 Flash, Kimi K3, DeepSeek V4 Flash 0731, DeepSeek V4.1 Flash, and the `brain-*` models behind Topics on [Baseten](https://www.baseten.co/), which is included in the Braintrust [DPA](https://www.braintrust.dev/legal/dpa) as a subprocessor. These endpoints run with Zero Data Retention (ZDR) for every organization, whether Braintrust-hosted (SaaS), [BYOC](/docs/admin/deployment/byoc), or self-hosted. ZDR is on by default and requires no configuration, so inference inputs and outputs are not stored by the model host. See Baseten's [data privacy documentation](https://docs.baseten.co/observability/security#data-privacy). [Jev](/docs/integrations/ai-providers/typesafe#llm-as-a-judge) is the exception. It runs on TypeSafe rather than Baseten, so the terms above do not describe how that data is handled, and your organization has to opt in before anyone can select it. Enabling Jev requires authorizing Braintrust to engage TypeSafe as a subprocessor. Data submitted to Jev is sent to TypeSafe for processing, and TypeSafe does not retain it after processing. When you run a model through one of your own [AI providers](#organization-providers) instead, the request goes to that provider on your own key. Your agreement with that provider, including any ZDR agreement, governs the prompt and the response. ## Next steps * Browse [supported AI providers](/docs/integrations/ai-providers) for provider-specific configuration. * [Manage permissions](/docs/admin/access-control/manage-permissions) to control who can add or modify project-level AI providers. # Audit logging Source: https://braintrust.dev/docs/admin/audit-logs Track and review administrative actions in your Braintrust organization, such as permission changes, member management, and API key creation The [** Audit log**](https://www.braintrust.dev/app/~/configuration/org/audit-log) records administrative actions in your Braintrust organization, such as creating projects, changing settings, granting permissions, managing members, and creating API keys. Reads of your data can also be logged as an optional add-on. only available on the [Enterprise plan](/docs/plans-and-limits#plans). Audit logs are only recorded while your organization is on an Enterprise plan. If you upgrade from Pro to Enterprise, audit logs are recorded from the time of the upgrade. Self-hosted deployments require [data plane v2.6.0](/docs/data-plane-changelog) or later for audit logging. On earlier versions, administrative events are recorded but aren't queryable. ## Grant access Members of the [**Owners**](/docs/admin/access-control#built-in-permission-groups) permission group can read organization audit logs by default. To let anyone else read them, grant the organization-level **Read audit logs** permission in ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). The **Read audit logs** permission grants read access to audit log entries for the organization. It does not grant access to modify the audited resources. ## View the audit log To view recent organization activity, go to ** Settings** > [** Audit log**](https://www.braintrust.dev/app/~/configuration/org/audit-log). The table lists actions performed by members of your organization, with the most recent events first. Use the controls above the table to customize the view: * **Time range**: Select **Last 24 hours**, **Last 7 days**, or **Last 30 days**. The default is **Last 7 days**. * **Filters**: Narrow results by fields such as actor, event type, or resource. * **Columns**: Show or hide columns. ID, details, and before and after change columns are hidden by default. To run more complex queries or download results, select **Open SQL sandbox**. See [Audit data reads and modifications](#audit-data-reads-and-modifications) and the [SQL reference](/docs/reference/sql) for query details. ## Query the audit log Users with the **Read audit logs** permission can query audit logs with [SQL](/docs/reference/sql) using the `audit_logs()` data source, identified by your organization's ID (for example, `audit_logs('my-org-id')`). No additional configuration is required to query them. Run a query from the [** SQL sandbox**](https://www.braintrust.dev/app/~/sql), the [`bt sql`](/docs/reference/cli/sql) CLI, or the [API](/docs/reference/sql#api). **Examples:** ```sql Recent activity across the organization theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} SELECT created, actor_id, event_type, resource_type, resource_name FROM audit_logs('') -- Replace with your organization ID WHERE created > NOW() - INTERVAL 1 DAY ORDER BY created DESC LIMIT 100 ``` ```sql All actions taken by a specific member theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} SELECT created, event_type, resource_type, resource_name FROM audit_logs('') -- Replace with your organization ID WHERE actor_id = '' -- Replace with the member's user ID ORDER BY created DESC ``` ```sql Permission and access control changes in the last 30 days theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} SELECT created, actor_id, event_type, resource_name, after_changes FROM audit_logs('') -- Replace with your organization ID WHERE created > NOW() - INTERVAL 30 DAY AND (event_type LIKE 'acl.%' OR event_type LIKE 'group%' OR event_type LIKE 'role%') ORDER BY created DESC ``` To log reads of your data in the audit trail, see [Audit data reads and modifications](#audit-data-reads-and-modifications). ## What gets logged Each audit log entry records a single event: what happened, who performed it, and what changed. ### Fields Each organization audit log entry includes: | Field | Description | | - | - | | `created` | Event timestamp. | | `org_id` | Organization where the event occurred. | | `project_id` | Project associated with the event, when applicable. | | `actor_id` | User or service account that performed the action. | | `event_type` | Event name in `.` form, such as `project.updated`. | | `event_details` | Additional event-specific metadata. | | `resource_type` | Type of resource that changed. | | `resource_id` | ID of the resource that changed. | | `resource_name` | Human-readable resource name. | | `actor_details` | Request metadata, including IP address, user agent, request ID, and authentication token details. | | `before_changes` | Relevant resource fields before the event. Populated for update and delete events. | | `after_changes` | Relevant resource fields after the event. Populated for create and update events. | For create events, `before_changes` is `null`. For delete events, `after_changes` is `null`. For update events, both fields contain the changed resource values. Readonly events contain neither. ### Events Braintrust records organization audit log events for these resource categories: | Resource category | Resource types | Event types | | - | - | - | | Organizations | `organization` | `organization.created`, `organization.updated` | | Projects | `project` | `project.created`, `project.updated`, `project.deleted` | | Experiments | `experiment` | `experiment.created`, `experiment.updated`, `experiment.deleted` | | Datasets | `dataset` | `dataset.created`, `dataset.updated`, `dataset.deleted` | | Project automations | `project_automation` | `project_automation.created`, `project_automation.updated`, `project_automation.deleted` (configuration changes only, not automation runs) | | Organization automations | `org_automation` | `org_automation.created`, `org_automation.updated`, `org_automation.deleted` (configuration changes only, not automation runs) | | Token budgets | `token_budget_config` | `token_budget_config.created`, `token_budget_config.updated`, `token_budget_config.deleted` (`scope_type`, `user_id`, `api_key_id`, and `config` fields) | | Organization AI providers and secrets | `ai_secret` | `ai_secret.created`, `ai_secret.updated`, `ai_secret.deleted` | | API keys | `api_key` | `api_key.created`, `api_key.deleted` | | Data plane manager service tokens | `service_token` | `data_plane_service_token.created`, `data_plane_service_token.replaced` | | Permission groups | `group` | `group.created`, `group.updated`, `group.deleted` | | Permission group membership | `group_member` | `group_member.created`, `group_member.deleted` | | Roles | `role` | `role.created`, `role.updated`, `role.deleted` | | Role membership | `role_member` | `role_member.created`, `role_member.deleted` | | Role permissions | `role_permission` | `role_permission.created`, `role_permission.deleted` | | Organization members | `org_member` | `org_member.created`, `org_member.deleted` | | Access grants | `acl` | `acl.created`, `acl.deleted` | Some operations emit multiple audit log entries. For example, inviting a user can create an organization member entry, permission group membership entries, access grants, and an API key entry. Bulk operations create one audit log entry per changed resource. Audit logs can take a few minutes to show up after an action occurs. ### Sensitive values Braintrust excludes or redacts sensitive values in audit logs: * API key hashes and raw keys are not included. Audit entries include the API key preview name when available. * AI provider secrets are redacted. Audit entries include a secret preview and omit encrypted secret material and key names. * Webhook action URLs (`config.action.url`) in project automation configurations are redacted. When a webhook URL changes, `event_details.redacted_fields_changed` lists `config.action.url`. * Resource IDs, organization IDs, project IDs, creation timestamps, update timestamps, and deletion timestamps are omitted from `before_changes` and `after_changes` when they would add noise to the change diff. ## Audit data reads and modifications Administrative actions are logged automatically with no setup. Row-level auditing of data reads and modifications is different: these events can be high volume, so Braintrust records them only when you opt in. `query.read` events are particularly high volume. When enabled, `query.read` covers both SQL queries run manually and ones run implicitly by the Braintrust UI when you browse logs, experiments, and traces. Braintrust records the following events for data rows: | Resource category | Resource types | Event types | | - | - | - | | SQL queries | `query` | `query.read` | | Logs and spans | `project_log` | `project_log.deleted` | | Experiment rows | `experiment_log` | `experiment_log.deleted` | | Dataset rows | `dataset_log` | `dataset_log.deleted` | How you enable row-level audit logging depends on your deployment. Enabling `query.read` audit logging counts toward your processed data usage. Because these events can be high volume, this may result in additional charges. For Braintrust-hosted and BYOC deployments, contact [Braintrust support](mailto:support@braintrust.dev). For self-hosted deployments, enable `query.read` auditing per organization through Terraform or Helm. See [Enable read auditing](/docs/admin/self-hosting/configure/security#enable-read-auditing) for the strict and best-effort modes and the variables to set. ## Retention Audit logs are retained indefinitely for the lifetime of your organization's Enterprise subscription, giving you a permanent record of administrative activity for compliance and regulatory requirements. Unlike logs, experiments, and datasets, they aren't subject to your [plan's retention window](/docs/admin/data-management/retention) or to retention automations. If your organization leaves the Enterprise plan, Braintrust stops recording new audit log events. ## Next steps * [Access control](/docs/admin/access-control) to learn how organization permissions work. * [Manage permissions](/docs/admin/access-control/manage-permissions) to grant **Read audit logs** to a permission group. * [SQL reference](/docs/reference/sql) to learn about how to query audit logs with SQL. # Authentication Source: https://braintrust.dev/docs/admin/authentication Understand how users and services authenticate to Braintrust, and how Braintrust authenticates to model providers on your behalf. Braintrust authenticates in two directions: your users and services authenticate **to** Braintrust, and Braintrust authenticates **outbound** to model providers when it calls models on your behalf. ## Authenticating to Braintrust These are the ways users and services prove their identity to Braintrust. ### End-user authentication The most common form of authentication is end-user authentication to the Braintrust application. Users authenticate with your enterprise's identity provider (e.g. Google, Okta) and receive credentials directly to their browser. In a self-hosted deployment, your API endpoints and data live in your own cloud environment, and these credentials communicate directly with the Braintrust API endpoint deployed in your cloud. You could even run these endpoints in a VPN that Braintrust's servers can't access, and the application will work. #### Single sign-on (SSO) Braintrust supports single sign-on (SSO) with your organization's identity provider (powered by [Clerk](https://clerk.com/)): * **Social login**: Google. * **SAML**: Okta Workforce, Microsoft Entra ID, Google Workspace, or a custom SAML provider. * **OpenID Connect (OIDC)**: A custom OIDC provider. only available on the [Enterprise plan](/docs/plans-and-limits#plans). To get set up with SAML or OIDC SSO, contact Braintrust at [support@braintrust.dev](mailto:support@braintrust.dev) to exchange the appropriate configuration URLs. Once everything's configured, Braintrust will turn it on for your domain and your team can start signing in using their regular work credentials. #### Domain mappings Domain mappings use just-in-time (JIT) provisioning to add users from specific email domains to your organization. On a successful sign-in, Braintrust adds the user without an invitation if they are not already an organization member. Domain mappings do not create accounts for everyone in a configured domain ahead of time. Each domain mapping matches one email domain and can optionally require one SAML group value. A matching mapping adds the user to one Braintrust organization and can optionally add them to one permission group. To match on a SAML group, configure your IdP to send each group as a separate value in the `public_metadata_groups` attribute. Braintrust applies a domain mapping only when the user first joins the organization. Later IdP group changes do not automatically update their Braintrust group membership. Users added this way still sign in through end-user authentication. only available on the [Enterprise plan](/docs/plans-and-limits#plans). #### SCIM provisioning Provision organization members from your identity provider instead of adding them by hand or by domain. Braintrust derives organization membership, and optionally permission group membership, from the SCIM groups a user belongs to. For setup steps, see [SCIM provisioning](/docs/admin/scim). in [private preview](/docs/feature-lifecycle), available to a limited set of customers. To request access, [contact Braintrust](https://braintrust.dev/contact). only available on the [Enterprise plan](/docs/plans-and-limits#plans). ### API authentication You can authenticate on behalf of users in your experiments or services using an API key. Braintrust API keys inherit their user's permissions, and essentially are another way to authenticate as a user. To increase security, API keys are stored as one-way cryptographic hashes and cannot be recovered. The actual key is only displayed once upon creation. If you lose an API key, you will need to generate a new one (and can deactivate the old one). You can create an API key by going to ** Settings** > [** API keys**](https://www.braintrust.dev/app/~/configuration/org/api-keys). When creating an API key or service token, you can optionally set an expiration, which cannot be changed after creation. After it expires, it stops authenticating and cannot be renewed, so create a new one to replace it. Keys created without an expiration will never expire. ### MCP authentication The [Braintrust MCP (Model Context Protocol)](/docs/integrations/developer-tools/mcp) server uses API key or OAuth 2.0 authentication, depending on the AI tool used to access the server. When AI tools use OAuth 2.0 to authentication, they: 1. Initiate an OAuth authorization flow. 2. Redirect users to authenticate with their Braintrust account. 3. Receive access tokens for API requests. 4. Use refresh tokens to maintain long-lived sessions. This authentication method inherits your organization's security policies and SSO configuration. MCP OAuth tokens follow the same permission model as your user account, providing access only to projects and resources you can normally access. ## Authenticating to model providers To make *outbound* calls to model providers on your behalf, Braintrust must authenticate with those providers. This is unrelated to how your users or services authenticate to Braintrust. By default, Braintrust authenticates to a model provider with a long-lived API key you store on the ** Settings** > [** AI providers**](https://www.braintrust.dev/app/~/configuration/org/secrets) page. However, some providers allow Braintrust to obtain short-lived credentials at request time, eliminating the need to store long-lived credentials in Braintrust. Braintrust supports two such methods, depending on the provider: * **Workload identity federation**: Braintrust exchanges a short-lived, Braintrust-signed OIDC token for a provider access token. Available for [OpenAI](/docs/integrations/ai-providers/openai), [Anthropic](/docs/integrations/ai-providers/anthropic), [Google Vertex AI](/docs/integrations/ai-providers/google), and [Azure AI Foundry](/docs/integrations/ai-providers/azure). * **Assume role**: Braintrust assumes an IAM role in your AWS account via the AWS STS `AssumeRole`, receives temporary credentials, and uses them to authenticate. Available for [AWS Bedrock](/docs/integrations/ai-providers/bedrock#connect-bedrock-to-braintrust). Both methods avoid storing a long-lived provider credential in Braintrust, but they rely on different flows: workload identity federation uses OIDC token exchange, while assume role uses AWS STS role assumption. These short-lived credential methods are available only on Braintrust-hosted organizations, and workload identity federation additionally requires an [organization-level provider](/docs/admin/ai-providers). Depending on the provider, you can also authenticate using stored credentials, such as an API key. # Change your plan Source: https://braintrust.dev/docs/admin/billing/change-plan Upgrade or downgrade between Starter, Pro, and Enterprise Change your [Braintrust plan](/docs/plans-and-limits) to match your team's needs. You can upgrade from Starter to Pro directly in the application, or contact Braintrust for Enterprise plans. Downgrading is also supported with considerations for feature access and data retention. ## Enable on-demand usage On-demand usage lets you grow beyond [Starter plan limits](/docs/plans-and-limits#usage-limits) without upgrading to Pro. Add a credit card, and you'll automatically be charged for additional usage at the end of each calendar month, with billing on the 1st of the following month. On-demand usage rates on Starter are higher than Pro rates. If your usage consistently grows beyond the Starter plan limits, [upgrading to Pro](/docs/admin/billing/change-plan#upgrade-your-plan) provides better economics with lower rates and a higher included usage allowance. To add a credit card: 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **On-demand usage**, click **Add card**. 3. Enter your credit card information. 4. Click **Save**. When you add a card, Braintrust charges a \$1 card validation fee to verify your payment method. The receipt is available in the **Invoices** dropdown on the ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing) page. After on-demand usage is enabled, use **Edit billing information** under **On-demand usage** to update your card, billing address, billing email, purchase order number, or tax information. ## Disable on-demand usage To stop future charges for usage above Starter plan limits, disable on-demand billing from the **Edit billing information** dialog. Your card stays on file. This change takes effect immediately. Starter plan limits are enforced from the moment you confirm, and usage above those limits will be blocked. 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **On-demand usage**, click **Edit billing information**. 3. Under **On-demand billing**, click **Disable**. 4. Click **Disable on-demand billing** to confirm. The change takes effect immediately and Starter plan limits apply from that point forward. ## Upgrade your plan Upgrade from Starter to Pro to access advanced features, higher usage limits, and lower pay-as-you-go rates. 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **Current plan**, click **Adjust plan**. 3. In the plan selection dialog, click **Upgrade to Pro**. 4. Review the Pro plan features and pricing. 5. Enter your payment information. 6. Click **Subscribe**. You'll be charged a prorated amount for the remainder of the current month. For billing addresses in taxable US jurisdictions, the upgrade preview shows an **Estimated tax** line that updates as you enter your billing address, and the tax is collected together with the prorated charge. On the 1st of the following month, you'll be charged the full \$249 monthly platform fee plus any usage-based charges from the previous month. If you upgrade on the 15th of a 30-day month: * Days remaining: 15 (from the 15th to the end of the month) * Prorated charge: \$249 × (15/30) = \$124.50 (using the example \$249 monthly platform fee) * For billing addresses in taxable US jurisdictions, the estimated tax is added on top of the prorated charge. When you upgrade to Pro, you immediately gain access to: * Custom dashboards * Environments * Unlimited human review scorers (Starter limited to 1 per project) * Built-in permission groups (Starter allows new assignments only to Owners; Enterprise adds custom permission groups) * [Higher usage limits](/docs/plans-and-limits#usage-limits) and lower on-demand rates for processed data and scores * A larger [monthly Topics credit](/docs/plans-and-limits#model-credits), with the same overage rates as Starter * Longer data retention (30-day retention vs. 14-day retention), [extendable up to 180-day retention](/docs/admin/data-management/retention#extend-the-retention-window) for a storage charge * Priority email support **Alternative**: If you only need higher usage limits, consider [enabling on-demand usage](#enable-on-demand-usage) instead. On-demand usage lets you grow beyond Starter limits without the monthly Pro platform fee, though [on-demand rates](/docs/plans-and-limits#usage-limits) are higher than Pro rates. Enterprise plans include custom usage limits, advanced security and compliance features, and dedicated support for large organizations. To upgrade to Enterprise: 1. [Contact Braintrust](https://www.braintrust.dev/contact) to discuss your requirements. 2. Work with the Braintrust team to configure a custom plan. 3. Sign an Enterprise agreement. Enterprise plans include everything in Pro plus: * Custom usage limits for processed data and scores * Custom data retention policies * Custom permission groups (Pro limited to Owner/Engineer/Viewer) * SAML/OIDC SSO providers (Pro and Starter limited to Google only) * Domain mappings * Audit logging * Flexible [deployment options](/docs/admin/deployment), including BYOC and self-hosted * Compliance and security: * SOC2 attestation * Custom DPA (Pro has click-through DPA) * BAA * Dedicated Slack support channel * Guaranteed SLAs * Custom legal terms * Annual invoicing See [Plans and limits](/docs/plans-and-limits) for a complete feature comparison. ## Downgrade your plan Before downgrading, ensure your usage fits within [Starter plan limits](/docs/plans-and-limits#usage-limits). If your usage is higher than these limits when you downgrade, you'll be charged on-demand usage rates at the Starter tier pricing (which are higher than Pro rates). To downgrade from Pro to Starter: 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **Current plan**, click **Adjust plan**. 3. In the plan selection dialog, click **Downgrade to starter**. 4. Review what will change when you downgrade. 5. Select a downgrade reason. 6. Click **Downgrade**. When you downgrade from Pro to Starter, the change is scheduled for the end of your current billing period. You'll keep access to Pro features until then. You can cancel the scheduled downgrade any time before the billing period ends from the Billing page. | Feature | Pro → Starter change | | - | - | | Custom dashboards | Unlimited → Keep existing, can't edit, duplicate, or create new | | Environments | Unlimited → Keep and edit existing, can't create new | | Human review scorers | Unlimited per project → Limited to 1 per project | | Permission groups | Owner/Engineer/Viewer → Owner only for new projects | | Usage type | Pro → Starter change | | - | - | | Processed data | 5 GB/month, then \$3/GB → 1 GB/month, then \$4/GB | | Scores | 50k/month, then \$1.50/1k → 10k/month, then \$2.50/1k | | [Model credits](/docs/plans-and-limits#usage-limits) | \$100/month → \$10/month. Overage rates unchanged. | | Data retention | 30-day retention (extendable up to 180-day retention) → 14-day retention | Enterprise customers on annual contracts cannot downgrade until the contract term ends. To discuss your options: 1. [Contact Braintrust support](mailto:support@braintrust.dev). 2. Discuss your situation and potential solutions. 3. Review contract terms and renewal options. When your Enterprise contract ends, you can choose to: * Renew at a different tier (Pro or Enterprise) * Move to a month-to-month Pro plan * Move to the Starter plan Enterprise features and custom limits will remain active until the end of your contract term. Before downgrading, ensure your usage fits within [Pro or Starter plan limits](/docs/plans-and-limits#usage-limits). If your usage is higher than plan limits when you move to a new tier, you'll be charged on-demand usage rates. ## Redeem a coupon On Starter with on-demand billing and Pro plans, you can redeem coupon codes to apply discounts to your subscription. To redeem a coupon: 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **Current plan**, click **Redeem coupon**. 3. Enter your coupon code. 4. Click **Redeem**. The discount appears on your billing page after the coupon is applied. The **Redeem coupon** button is disabled when a plan change is already scheduled on your subscription. Contact [support@braintrust.dev](mailto:support@braintrust.dev) for assistance. ## Legacy plans If you're on a legacy plan (signed up before March 2026), see [Legacy plans FAQ](/docs/admin/billing/faq#legacy-plans) for details on how plan changes affect your existing features. ## Next steps * [Plans and limits](/docs/plans-and-limits) — Compare plans and understand usage limits * [Monitor usage](/docs/admin/billing/monitor-usage) — Track your usage and costs * [Billing FAQ](/docs/admin/billing/faq) — Common questions about pricing and billing * [Contact sales](https://www.braintrust.dev/contact) — Discuss Enterprise plans or custom requirements # Billing FAQ Source: https://braintrust.dev/docs/admin/billing/faq Common questions about Braintrust pricing, billing, and plans ## Plans and payment ### Which plan is right for me? See [Plans and limits](/docs/plans-and-limits) for a detailed comparison of features, usage limits, and pricing for each tier. To estimate your costs, visit the [Pricing page](https://www.braintrust.dev/pricing). For custom requirements or questions about which plan fits your needs, [contact our team](https://braintrust.dev/contact). ### How do I change my plan? For details on upgrading, downgrading, feature changes, and billing impacts, see [Change your plan](/docs/admin/billing/change-plan). ### How do I update my payment method? 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **On-demand usage**, click **Edit billing information**. 3. Under **Payment method**, click **Replace card**, then enter the new card details. ### How do I update my billing address? Update your billing address on its own, without replacing the card on file. **Edit address** appears only when a card is on file. If there is no card yet, the address you enter when you [add a card](/docs/admin/billing/change-plan#enable-on-demand-usage) becomes the billing address. 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **On-demand usage**, click **Edit billing information**. 3. Under **Billing address**, click **Edit address**. 4. Update your address and click **Save address**. Braintrust verifies US addresses against address records so that tax is calculated correctly. If an address can't be fully verified, a **Check your billing address** dialog shows the address you entered next to a suggested address when one is available. Click **Use suggested address** or **Use entered address** to continue, or close the dialog to keep editing. ### How do I get an invoice or receipt? 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Under **On-demand usage**, open the **Invoices** dropdown to view recent invoices and card validation fee receipts, or click **Detailed usage breakdown** to access the full invoice history in the billing portal. ### Are taxes included in my charges? For billing addresses in taxable US jurisdictions, sales tax is estimated when you upgrade to Pro and collected together with the prorated platform fee. The upgrade preview shows an **Estimated tax** line that updates as you enter your billing address, so you can review the total before confirming. Tax appears as a separate line on the resulting invoice. For billing addresses outside taxable jurisdictions, no tax is applied at checkout. ### What happens if my payment fails? If a payment fails, Braintrust will attempt to collect payment several times over a grace period. During this time: * Your access continues uninterrupted * You'll receive email notifications about the failed payment If payment cannot be collected after the grace period, your account will be downgraded to the Starter plan. You can restore Pro access at any time by updating your payment method and resubscribing. ### How does Vercel Marketplace billing work? If you installed Braintrust through the [Vercel Marketplace](https://vercel.com/integrations/braintrust), your billing is handled through Vercel. To manage your subscription, go to your [Vercel billing dashboard](https://vercel.com/dashboard/billing), find Braintrust in your integrations, and manage your subscription there. Usage limits and features work the same as direct Braintrust billing. ### Do you offer a startup program? Qualifying early-stage startups can apply for 6–12 months of free Pro plan access. To be eligible, you must be a new Braintrust customer at Series A or earlier and have raised at least \$100K in funding. Visit the [startups page](https://www.braintrust.dev/startups) to apply. ## Usage and limits See [Plans and limits](/docs/plans-and-limits#usage-limits) for included limits and on-demand usage rates. ### What does processed data mean? Processed data refers to the total bytes of data ingested across logs, experiments, and datasets. Includes inputs, outputs, prompts, metadata, traces and spans, datasets, attachments, and any other related information. Braintrust measures the size of each item when you send it and adds it to your total for that calendar month. ### Does deleting data reduce my usage? No. Processed data is measured at ingestion, based on what you send during the month, not on what's currently stored. Deleting logs, applying a [retention policy](/docs/admin/data-management/retention), or shortening your retention window doesn't reduce your processed data usage for the month. To lower processed data usage, reduce how much you ingest. [Retention storage](/docs/admin/billing/monitor-usage#retention-storage), a separate charge on the Pro plan that applies when you retain data beyond the base retention window, is measured the same way. It bills the data you ingest into each month you keep beyond the base window, not what's currently stored, so deleting data, applying a retention policy, or shortening your window doesn't reduce it. The only way to lower retention storage is to ingest less. ### What are scores? [Scores](/docs/evaluate/write-scorers) are used to measure the results of offline or online evaluations run in Braintrust. Each time you record a score, the total number of scores counted toward your monthly usage increases by one. Your monthly total is calculated cumulatively from the first to the last day of each calendar month. ### How do I track my usage? See [Monitor usage](/docs/admin/billing/monitor-usage) for instructions on viewing usage charts and accessing detailed reports. ### Can I set spending limits? Braintrust doesn't offer hard spending limits that prevent usage. However, Pro and Starter plans can [set up spend alerts](/docs/admin/billing/monitor-usage#set-up-spend-alerts) to get notified when monthly costs reach custom thresholds. Starter plan users also receive automated alerts at 80%, 90%, and 100% of their included usage limits. ### How does Topics billing work? [Topics](/docs/observe/topics) is metered separately from processed data and scores. Each plan includes a [monthly Topics credit](/docs/plans-and-limits#model-credits) applied to the input and output tokens consumed by the daily Topics pipeline. Once the credit is used up, uniform [overage rates](/docs/plans-and-limits#model-credits) apply across all plans. Credits do not roll over month-to-month. To track your Topics token usage, see [Monitor usage](/docs/admin/billing/monitor-usage). ### What happens when I run out of Topics credit? On Starter and Pro plans, Braintrust emails you automatically when you've consumed 60% and 100% of your monthly Topics credit, so you know before and when you reach the limit. What happens at the limit depends on your plan: * **Starter without on-demand usage**: Topics is paused for the rest of the billing cycle. [Add a card](/docs/admin/billing/change-plan#enable-on-demand-usage) to continue running Topics at the overage rates, or wait for the next monthly cycle. * **Starter with on-demand usage**: Token usage beyond the credit is billed at the overage rates and appears in your monthly on-demand invoice. * **Pro**: Token usage beyond the credit is billed at the overage rates and appears in your monthly invoice. ### Can I use built-in models with my credit? Yes. [Built-in models](/docs/admin/ai-providers#available-models) require no AI provider setup. On Starter and Pro plans, their usage draws down the same monthly credit as Topics. On Starter, they're unavailable once you use up the credit unless you enable on-demand usage. With on-demand usage on Starter, and on Pro, usage beyond the credit is billed at the overage rates. See [model pricing](/docs/plans-and-limits#model-credits). ### Why can't I select a built-in model? On the Starter plan, built-in models require at least one organization owner with a work email address, or a payment method on file. This covers the models in playgrounds, prompts, and scorers, as well as the models [Loop](/docs/loop) and [Patterns](/docs/observe/patterns) run on. Affected models don't appear in the model picker. To qualify, [add a payment method](/docs/admin/billing/change-plan#enable-on-demand-usage), or add an owner whose email uses a work domain. See [Requirements](/docs/admin/ai-providers#requirements) for how Braintrust determines eligibility. ### Why does the gateway return HTTP 403 for a built-in model? Calling a built-in model directly through the [gateway](/docs/deploy/gateway) with a Braintrust API key or service token requires a Pro or Enterprise plan. On the Starter plan, these requests return HTTP 403 with the message `Native inference requires a Pro or Enterprise plan`. On Starter, you can still use built-in models in playgrounds, prompts, scorers, Loop, and Patterns, and via the `invoke()` function for saved prompts and scorers. To call built-in models through the gateway, [upgrade your plan](/docs/admin/billing/change-plan#upgrade-your-plan). ## Legacy plans This section applies to self-serve Starter and Pro plans. Enterprise customers are governed by their contract and should contact their account representative for questions about plan features. ### Am I on a legacy plan? If you signed up before March 16, 2026, and haven't changed plans since, you're on a legacy plan. Legacy plans preserve access to features that were available when you signed up, allowing you to continue using features even if they're no longer available on your plan tier. For confirmation of your plan status, [contact support](mailto:support@braintrust.dev). ### What features do legacy plans include? Legacy plans preserve full access to features that were available when you signed up. This includes: * Custom and built-in permission groups * Data export automations * Custom charts, environments, and human review scorers * Any other features you were using before March 2026 The entire feature is grandfathered. You can create new instances, edit existing ones, and continue using them, even if they're no longer available on your current plan tier. ### What happens if I upgrade or downgrade from legacy? **Upgrading from a legacy plan:** You retain your legacy status when you upgrade, so you keep all the features you had before plus gain access to everything on the new plan. For example, if you upgrade from Starter to Pro and you were on a legacy plan, you'll have all Pro features in addition to any legacy features you previously had access to (such as full RBAC). **Downgrading from a legacy plan:** Downgrading permanently removes your legacy status. You will move to the current version of the lower plan and cannot return to your legacy plan. You lose the ability to create new instances of legacy features. Existing items (such as custom charts or environments you already created) remain accessible, but you won't be able to add new ones. Downgrading from a legacy plan is permanent — you cannot return to your legacy plan afterward. Before downgrading, make sure you understand the current plan limits and features. If you're unsure, [contact support](mailto:support@braintrust.dev) to discuss your options. ## Enterprise plans For Enterprise features, pricing, and contracts, see [Plans and limits](/docs/plans-and-limits) or [contact Braintrust](https://www.braintrust.dev/contact) to discuss your requirements. ## Need help? For billing questions, contact [support@braintrust.dev](mailto:support@braintrust.dev). For Enterprise plans or custom requirements, [contact sales](https://www.braintrust.dev/contact). # Manage billing Source: https://braintrust.dev/docs/admin/billing/index Change plans, monitor usage, and manage billing Manage your Braintrust subscription, track usage, and understand your costs. This section covers plan changes, usage monitoring, and common billing questions. ## Plans * **Starter** (\$0 platform fee) — Start building and evaluating AI applications for free with generous usage limits. No credit card required. * Best for individual developers and small teams. * On-demand usage pricing applies when your usage grows beyond the free-tier. * **Pro** (\$249/month) — Everything in Starter, plus advanced observability, security features, and higher usage limits. * Best for growing teams and production applications. * Lower on-demand usage rates than Starter. * **Enterprise** (Custom pricing) — Everything in Pro, plus custom usage limits, advanced security and compliance features, flexible deployment options (BYOC and self-hosted), and dedicated support. * Best for large organizations with enterprise requirements. * Annual invoice. See [Plans and limits](/docs/plans-and-limits) for a detailed comparison of features and usage limits. ## Change your plan [Upgrade or downgrade](/docs/admin/billing/change-plan) between Starter, Pro, and Enterprise plans. Learn about: * Upgrading from Starter to Pro or Enterprise * Downgrading to a lower tier * Enabling pay-as-you-go on Starter * How plan changes affect your features and billing * Redeeming coupon codes on Starter and Pro plans ## Monitor usage [Track your usage and costs](/docs/admin/billing/monitor-usage) throughout the billing period: * See your current plan and projected on-demand usage at a glance * View usage charts for trace spans, processed data, and scores * Access detailed usage reports (Pro and Enterprise) * Set up cost alerts using log alerts * Understand usage metrics and billing cycles ## Billing FAQ [Common questions](/docs/admin/billing/faq) about pricing, billing, and plans: * Plans and payment (which plan is right for me, how to change plans, payment methods, invoices) * Usage and limits (processed data, scores, tracking usage, spending limits) * Legacy plans (feature preservation, upgrade/downgrade impacts) * Enterprise plans and support ## Additional resources * [Plans and limits](/docs/plans-and-limits) — Detailed feature comparison and usage limits * [Pricing page](https://www.braintrust.dev/pricing) — Cost estimator * [Contact sales](https://www.braintrust.dev/contact) — Enterprise plans and custom requirements # Monitor usage Source: https://braintrust.dev/docs/admin/billing/monitor-usage Track usage and set up cost alerts Track your Braintrust usage and costs on the Billing page. **Usage analytics** is available on all plans. Pro plans and Starter plans with [on-demand usage enabled](/docs/admin/billing/change-plan#enable-on-demand-usage) also show invoice projections and a **Detailed usage breakdown** link to the Orb portal. For Enterprise invoice details, [contact Braintrust](https://www.braintrust.dev/contact). ## View usage 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. If available, review **On-demand usage** for your projected invoice, Logs and Scores breakdown, and billing cycle reset date. 3. Under **Usage analytics**, select the current month or a custom date range (up to 100 days). Use the scope dropdown to view **Entire organization** (default) or specific projects. * **Entire organization**: toggle between daily and cumulative views of logs (GB), scores, and Topics input and output tokens (when [Topics](/docs/observe/topics) is enabled). * **By project**: daily logs, scores, and Topics input and output tokens broken down by project. Project data is available from September 1, 2026, and selecting a start date before then prompts you to choose a later date. Requires data plane Usage analytics section with Entire organization selected, showing a CSV export button and daily logs, scores, and Topics input and output token charts In either scope, click **CSV** to download the data for the selected date range. 4. If available, click **Detailed usage breakdown** under **On-demand usage** to open the Orb portal, where you can: * View detailed breakdowns of trace spans, processed data, scores, Topics tokens, and retention storage. * Download usage reports for internal tracking or accounting. * Monitor costs and billing cycles. * Compare usage across multiple billing periods. ## Understanding usage metrics Your Braintrust usage is measured across four dimensions. Each contributes to your monthly costs if your usage grows beyond your plan's included limits. See [Plans and limits](/docs/plans-and-limits#usage-limits) for included limits and on-demand usage rates by plan, and [Model credits](/docs/plans-and-limits#model-credits) for the monthly model credit and overage rates. ### Processed data Processed data refers to the total bytes of data ingested by Braintrust when you create logs or experiments. This includes: * Inputs and outputs * Prompts and completions * Metadata and tags * Traces and spans * Datasets * Attachments (images, audio, files) * Any other related information **How it's measured**: Braintrust measures the size of each log, experiment, and dataset record when you send it and adds it to your total for that calendar month, which is invoiced at the end of the month. On Enterprise plans, monthly usage draws down from your committed volume. Processed data is measured at ingestion, based on what you send, not on what's currently stored. Deleting data or shortening your [retention window](/docs/admin/data-management/retention) doesn't reduce your processed data usage for the month. [Retention storage](#retention-storage), a separate charge on Pro and Enterprise plans, is measured the same way, so deletion and shorter windows don't reduce it either. ### Scores Scores are used to measure the results of offline or online evaluations run in Braintrust. Each time you record a score, it counts toward your monthly usage. Examples of scored operations: * Offline evaluation scores from experiments * Online evaluation scores from production logs * Human review scores * Custom scorer outputs **How it's counted**: Each score recorded (whether automatic or manual) increments your monthly score count by one. ### Topics tokens [Topics](/docs/observe/topics) is metered in input and output tokens consumed by the daily Topics pipeline (facet summarization, embedding, and cluster naming). Each plan includes a [monthly credit](/docs/plans-and-limits#usage-limits), shared with the built-in open-source models, that covers a portion of token usage at no extra cost. Once the credit is used up, uniform [overage rates](/docs/plans-and-limits#model-credits) apply across all plans. **How it's measured**: Input and output tokens are summed daily across all Topics pipeline calls and counted toward your monthly credit. The credit does not roll over to the next month. On the Starter plan, if you exhaust your monthly credit and haven't [enabled on-demand usage](/docs/admin/billing/change-plan#enable-on-demand-usage), Topics is paused for the rest of the billing cycle. Add a card to continue running Topics at the [overage rates](/docs/plans-and-limits#model-credits), or wait for the next monthly cycle. [Built-in model](/docs/admin/ai-providers#available-models) usage also draws from your model credits and appears under **Native inference** in your usage breakdown. ### Retention storage On the Pro plan, retaining data beyond [30-day retention](/docs/plans-and-limits#usage-limits) incurs a retention storage charge. Like processed data, retention storage is measured at ingestion: it bills the data you ingest into each month you keep beyond the base window, not what's currently stored. This covers logs and playground data, including each row's [activity history](/docs/observe/examine-traces#view-activity-history). Experiments are never deleted by your retention window and have up to 365-day retention on all plans, but the data you ingest into them still counts toward retention storage. On Enterprise plans, retention beyond your base window is arranged with your account team. To keep data longer, [extend your retention window](/docs/admin/data-management/retention#extend-the-retention-window). Because retention storage is measured at ingestion, deleting data, setting a [retention automation](/docs/admin/data-management/retention#automate-data-deletion), or shortening your window doesn't reduce it. The only way to lower retention storage is to ingest less. **How it's measured**: Braintrust measures the size of your retained data daily and bills the highest daily amount during the billing month. See [Plans and limits](/docs/plans-and-limits#usage-limits) and the [Pricing page](https://www.braintrust.dev/pricing) for rates. ## Set up spend alerts Spend alerts are available for plans with [on-demand usage](/docs/admin/billing/change-plan#enable-on-demand-usage) (Starter and Pro). Receive a notification when your total invoice amount reaches a threshold in the current billing cycle. You can set up to three spending thresholds and choose to notify via **Billing email**, a **Slack channel**, or both. 1. Go to ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing). 2. Click ** Spend alerts**. 3. Set your spending thresholds. 4. Under **Notify via**, select **Billing email**, **Slack channel**, or both. * To notify a Slack channel, [enable the Slack integration](/docs/admin/organizations#enable-slack-integration) first, then select a workspace and channel. 5. Click **Save**. Spend alerts are informational-only and do not interrupt or limit your usage. You can update or disable alerts at any time. Spend alerts are distinct from Topics credit notifications. Braintrust emails you automatically when you've consumed 60% and 100% of your [monthly Topics credit](/docs/plans-and-limits#model-credits). These notifications are sent automatically and aren't configurable, whereas spend alerts track your total invoice amount against thresholds you set. ## Next steps * [Set up alerts](/docs/observe/alerts) — Learn more about creating and managing alerts * [Configure data retention](/docs/admin/data-management/retention) — Configure retention policies to control costs * [Plans and limits](/docs/plans-and-limits) — Review your plan limits and features * [Billing FAQ](/docs/admin/billing/faq) — Common questions about billing and usage * [Write SQL filters](/docs/observe/alerts#common-alert-patterns) — Advanced SQL patterns for alert conditions # Export to cloud storage Source: https://braintrust.dev/docs/admin/data-management/export Deliver project logs to S3 or GCS in JSONL or Parquet format, with Hive partitioning for efficient downstream querying. Braintrust-hosted customers can export to AWS S3. Self-hosted customers can export to the cloud that hosts their data plane: AWS S3 or Google Cloud Storage (cross-cloud export is not supported). Data plane v2.1 introduces Google Cloud Storage as an export destination, a new file layout, and other changes that may affect existing integrations. See [Migrate exports to data plane v2.1](#migrate-exports-to-data-plane-v2-1). only available on the [Enterprise plan](/docs/plans-and-limits#plans). ## How exports work An export is a project-scoped automation that runs on a configured interval. The interval is a target cadence between runs (start-to-start), not a fixed wall-clock schedule. Choose an interval based on your data volume and optimal file sizes for downstream query performance. On each run, the export: * Queries recent log data. * Writes the result to your cloud storage bucket as either JSON Lines or Parquet, organized in a Hive-partitioned layout that warehouses and query engines can read. * Advances its cursor, so that the same data isn't re-exported on the next run. * Schedules the next run: immediately if more data remains, otherwise roughly one interval from when the current run finished. For one-off exports, or to export data other than logs, use [SQL](/docs/reference/sql) or the [REST API](/docs/api-reference/query#export-data) instead of exports. ## Create an export Configure exports per project, in the UI or via the API. Rewind and start-from date can only be configured through the UI. To create an export, follow these steps: 1. Go to ** Settings** > **Project** > [** Data management**](https://www.braintrust.dev/app/~/configuration/data-management). 2. Click **Create automation**. If another automation already exists, click **Automation**. 3. Configure export settings: * **Automation name**: Identify the export. * **Type**: Export * **Data to export**: Logs (traces), Logs (spans), or Custom SQL query. See [export data types](#export-data-types). * **Export path**: Target bucket and prefix (e.g., `s3://my-bucket/braintrust/logs` or `gs://my-bucket/braintrust/logs`). Once the automation is created, this path cannot be changed. * **Cloud auth**: On AWS data planes, provide a **Role ARN**. On GCP data planes, provide a **Service account email**. See [identity and access management](#iam). * **Format**: JSON Lines or Parquet. * **Interval**: How often to export (5 min to 1 day). * (Optional) **Start from** (requires data plane v2.1+): Export data starting from this time. The picker uses your browser's local timezone, but `date=` folders partition by UTC. Leave empty to export from the beginning of your data. Logs already deleted under your [retention settings](/docs/admin/data-management/retention) will not be included. 4. Expand **Setup instructions** in the dialog to see the IAM steps for your data plane's cloud. 5. Click **Test automation** to verify access. The test writes a small file to your export path and immediately deletes it. 6. Click **Create automation**. The first export interval starts immediately. On data plane v2.0 and earlier, some labels differ (**S3 export**, **S3 path**, **Role creation instructions**) and the **Start from** field isn't available. The setup flow is otherwise the same. To create an export automation using the Braintrust API: Call [`POST /v1/project_automation`](/docs/api-reference/projectautomations/create-project_automation) with the export configuration. Capture the `id` returned in the response for the next step. ```bash curl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} # The example below exports trace summaries. To export spans or a custom # SQL (or BTQL) query, swap export_definition for one of: # { "type": "log_spans" } # { "type": "btql_query", # "btql_query": "select: id, input, output from: project_logs('PROJECT_ID') filter: scores.correctness < 0.5" } # The 'PROJECT_ID' literal inside btql_query must match the top-level project_id. curl --request POST \ --url https://api.braintrust.dev/v1/project_automation \ --header "Authorization: Bearer $BRAINTRUST_API_KEY" \ --header "Content-Type: application/json" \ --data '{ "project_id": "PROJECT_ID", "name": "Hourly logs export", "config": { "event_type": "btql_export", "export_definition": { "type": "log_traces" }, "export_path": "s3://my-bucket/braintrust/logs", "format": "parquet", "interval_seconds": 3600, "credentials": { "type": "aws_iam", "role_arn": "arn:aws:iam::123456789012:role/braintrust-export", "external_id": "EXTERNAL_ID" } } }' ``` Call `POST /brainstore/automation/reset-cursors` with the `id` from the previous step to register the export with the data plane and set its starting point. Without this call, the export won't run. Note that this endpoint is not under `/v1/`. ```bash curl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} curl --request POST \ --url https://api.braintrust.dev/brainstore/automation/reset-cursors \ --header "Authorization: Bearer $BRAINTRUST_API_KEY" \ --header "Content-Type: application/json" \ --data '{ "automation_id": "AUTOMATION_ID", "object_id": "project_logs:PROJECT_ID", "start_xact_id": "0" }' ``` After this call, the export runs on its configured schedule. On data plane v2.0 and earlier, this call is `POST /automation/cron` instead of `POST /brainstore/automation/reset-cursors`. ### Export data types When creating an export automation, you can choose to export trace summaries, spans, or data that matches a custom SQL query. Exports one row per trace, with scores and metrics aggregated across every span under the root. Child spans are not individually represented in the output. See the SQL [`summary` shape](/docs/reference/sql/query-structure#summary). Useful for analytics, cost dashboards, or warehouse joins. Exports one row per span, with root and child spans both present. See the SQL [`spans` shape](/docs/reference/sql/query-structure#spans). Useful for data science, fine-tuning, or data archival. Exports data according to a custom SQL query. Useful for consumers that need a filtered subset of your data. Custom queries read from `project_logs('')`, where `` is the automation's own project ID. Other sources, including datasets, experiments, cross-project logs, and subqueries, aren't supported. `GROUP BY` and aggregates are also rejected. The exporter manages sort, limit, and cursor internally, so don't set those in your query. Custom queries can also return *all* spans from any trace that contains at least one *matching* span. This can be useful for retrieving full context around failures, instead of just the failing spans in isolation. To do this, use SQL's [`traces` shape](/docs/reference/sql/query-structure#traces). Example export queries: * To export every LLM span from your project logs: ```sql theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} SELECT id, span_id, root_span_id, input, output, metadata, created FROM project_logs('your-project-id', shape => 'spans') WHERE span_attributes.type = 'llm' ``` * To export every span from any trace that contains at least one low-scoring span: ```sql theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} SELECT id, span_id, root_span_id, input, output, scores, metadata, created FROM project_logs('your-project-id', shape => 'traces') WHERE scores.correctness < 0.5 ``` ### Identity and access management (IAM) Braintrust authenticates to your bucket using credentials you create: an IAM role on AWS, or a service account on GCP. In the export configuration dialog, expand **Setup instructions** and follow the guided steps for your data plane's cloud. Google Cloud Storage export is available on self-hosted GCP deployments running data plane v2.1+. Setup requires a Terraform change to your data plane module to authorize [service account impersonation](/docs/admin/self-hosting/deploy#service-account-impersonation). ## Cloud folder structure Files are organized using Hive partitioning, a standard layout supported by many warehouses and query engines (see docs for [DuckDB](https://duckdb.org/docs/current/data/partitioning/hive_partitioning#hive-partitioning) and [BigQuery](https://docs.cloud.google.com/bigquery/docs/hive-partitioned-queries)). The partition key is `date`, set to the UTC calendar date the row was logged (not the date the export ran). A single run can write into multiple folders when catching up on old data. ``` {prefix}/ date=2026-04-19/ c1f1e5a8-....jsonl.gz 9de8b4a0-....jsonl.gz date=2026-04-20/ 42d0fe3c-....jsonl.gz ``` Exported files have the following characteristics: * **Row order** - Rows within a file are sorted ascending by `_xact_id`, an internal ID assigned at write time. Rows committed together share an `_xact_id`, and newer writes get higher values. For **Logs (spans)** exports, this does not imply trace grouping or parent-before-child ordering. To get trace-grouped output, sort in the consuming application. * **File splits** - Files split when the date changes or when the file reaches an internal size target. All rows for a given `_xact_id` always land in the same file, so the exporter writes past the size target when needed to keep them together. File names are random UUIDs to avoid collisions. * **Formats** - JSON Lines files are gzipped (`.jsonl.gz`). Parquet files are ZSTD-compressed (`.parquet`). On data plane v2.0 and earlier, files use the layout `{prefix}/YYYY-MM-DD/{file}` (no `date=` prefix), and the folder date is the date the export ran rather than the row's transaction date. ## Export status and history Monitor exports through the **Export status** dialog. The export doesn't emit notifications, so check here to confirm runs succeeded and surface errors. Include the automation ID and any error text when contacting support. 1. Go to ** Settings** > **Project** > [** Data management**](https://www.braintrust.dev/app/~/configuration/data-management). 2. To open the **Export status** dialog, click the status icon next to your export. 3. Review run history, row counts, bytes written, and duration. These values are cumulative since the export was created or last rewound. Byte counts represent compressed, on-disk sizes in the storage bucket. ## Rewind to a specific date Requires data plane Rewind moves an existing export's cursor backward to a chosen moment, so subsequent runs re-emit everything from that moment forward. Use it to re-export a window after a downstream bug, or to refill your bucket after deleting its contents. To rewind an existing export, follow these steps: 1. Open the **Export status** dialog and click **Rewind...**. 2. In the rewind panel, select a time using the time range selector. Choose a preset or enter a custom date and time. 3. To align with the UTC-based `date=YYYY-MM-DD/` folder structure in cloud storage, enable **Use UTC time** in the same dropdown. 4. Click **Rewind to selected time** to rewind to that moment, or **Re-export all data** to rewind to the beginning. A few things to expect when you rewind: * **The next run starts immediately,** and then the export cadence resumes (with the next run scheduled one interval later). * **Exports are incremental.** Runs happen back-to-back while there's a backlog, with each run continuing from where the previous one stopped. The normal interval cadence resumes once caught up. * **Existing files are kept.** Rewind doesn't delete, modify, or replace files already in your bucket. Subsequent runs write fresh UUID-named files alongside the old ones with overlapping rows, so deduplicate downstream on `id` or `span_id`, not on filename. * **Deleted logs can't be re-exported.** Rewinding past your [retention cutoff](/docs/admin/data-management/retention) has no effect on logs that were already deleted. * **Totals reset.** The **All runs** values in the Export status dialog (rows, data, duration) reset to zero after a rewind and rebuild as new runs land. ## Troubleshooting **Query timeouts**: For trace exports timing out: * Ensure you're on data plane v1.1.27 or later. * In the **Export status** dialog, click **Rewind...** and choose an earlier point in time to re-process from. * If problems persist, create a new trace export automation. **AWS IAM errors**: If test automation fails with permission errors: * Verify the IAM role ARN is correct. * Check the trust policy includes the correct external ID. * Ensure the S3 policy grants the required permissions on the target bucket and prefix. * Confirm the bucket and prefix exist. **GCP IAM errors**: If test automation fails with a 403 on `generateAccessToken`: * Verify that the service account email pasted into the UI matches the one you created. * Confirm the service account is listed in `brainstore_impersonation_targets` in your Braintrust data plane Terraform module, and that Terraform was applied after adding it. See [Service account impersonation](/docs/admin/self-hosting/deploy#service-account-impersonation). * Confirm the Brainstore service account has the `Service Account Token Creator` role on your export service account. Braintrust's Terraform module sets this automatically; if you configured IAM manually, verify it under the export service account's **Principals with access** tab. * Confirm your service account has `roles/storage.objectAdmin` on the target bucket. * Allow a few minutes for IAM changes to propagate. ## Migrate exports to data plane v2.1 For self-hosted customers, the changes below take effect once you upgrade to data plane v2.1+ and enable cloud storage export in your Terraform module or Helm chart. **Some of these changes may affect existing downstream pipelines or scripts that programmatically create exports.** Changes in v2.1: * Google Cloud Storage as an export destination (self-hosted GCP only). * Specify a [start-from date](#create-an-export) when creating an export. * [Rewind](#rewind-to-a-specific-date) lets you re-export historical data. After a rewind, the same row can appear in your bucket more than once; deduplicate downstream on `id`, keeping the row with the highest `_xact_id`. * Dataset and experiment export is no longer supported. To export an experiment or dataset, use [SQL](/docs/reference/sql) or the [REST API](/docs/api-reference/query#export-data) instead. * File paths use a [Hive-partitioned layout](#cloud-folder-structure): `{prefix}/date=YYYY-MM-DD/` instead of `{prefix}/YYYY-MM-DD/`. The date is now the row's transaction date, not the date the export ran. Old-layout files stay in place; scripts matching the old shape will miss new files. * Files split by size rather than by a 100,000-row cap. Historical catch-up spreads across many `date=` partitions instead of piling into the current run's folder, so consumers watching only today's folder will miss backfill output. * API change for programmatic creation. The second registration call is now `POST /brainstore/automation/reset-cursors` instead of `POST /automation/cron`. Scripts that create exports via the API need to be updated. ## Limitations * For [self-hosted deployments](/docs/admin/self-hosting), Rewind, start-from date, and Google Cloud Storage destinations require data plane v2.1+ and must be explicitly enabled in your Terraform module or Helm chart (on AWS, set `brainstore_enable_export = true`). See [Migrate exports to data plane v2.1](#migrate-exports-to-data-plane-v2-1). * Only log data can be exported. * Braintrust-hosted customers can only export their data to AWS S3. * Self-hosted customers can only export data to the cloud associated with their data plane (AWS S3 or Google Cloud Storage). Azure Blob Storage is not supported. * Braintrust's export writer does not use VPC endpoints or Private Service Connect. Traffic uses the cloud's public object-store endpoints. On data plane v2.0 and earlier, datasets and experiments could also be exported. ## Next steps * [Configure data retention](/docs/admin/data-management/retention) to control how long data is kept and comply with regulations * [Set up alerts](/docs/observe/alerts) to monitor data quality * [View logs](/docs/observe/view-logs) to understand what gets exported * [Export data](/docs/annotate/export) via the API for one-time exports # Protect sensitive data Source: https://braintrust.dev/docs/admin/data-management/protect-sensitive-data Detect and redact PII during project-log ingestion, or remove sensitive data before upload with SDK masking functions or span customizers. Traces can contain personal information, credentials, and other sensitive data. Braintrust supports two approaches to removing that data from your logs, which you can use separately or together: * [Redact during ingestion](#redact-during-ingestion): Braintrust receives your logs and sends unredacted text from selected fields to a detection model hosted on Baseten. Braintrust replaces detected sensitive text before storing the logs. * [Redact before upload](#redact-before-upload): An SDK masking function, span customizers, or `bt` span plugins remove sensitive values within your application or on your machine. Values they remove are not sent to Braintrust. You can use both approaches together. If a value must never leave your application, remove it before upload. ## Redact during ingestion Ingestion redaction detects sensitive text, such as names, email addresses, and credentials, in incoming project logs and replaces it before storage. Choose which fields to scan and which types of sensitive information to redact in each field. Enable it for one project or your entire organization without changing your application code. in [private preview](/docs/feature-lifecycle), available to a limited set of customers. To request access, [contact Braintrust](https://braintrust.dev/contact). Requires data plane Ingestion redaction ### How redaction works Redaction runs on incoming project logs before storage: 1. **Select text:** Braintrust applies your project and organization policies to identify which fields and entity types to scan. Selected fields include strings nested in objects and arrays. 2. **Detect sensitive information:** Braintrust sends unredacted text from the selected fields to a Braintrust-managed detection model hosted on Baseten, a [subprocessor](https://www.braintrust.dev/legal/dpa). The model identifies text matching your configured entity types. 3. **Replace and store:** Braintrust replaces the text identified by the model with category markers such as `[REDACTED_EMAIL]`, then stores the records. Text outside the detected spans and the object and array structure are preserved. Only incoming project logs are scanned. Object keys, non-string values, existing records, direct writes to datasets or experiments, and image, audio, or file attachment contents are excluded. For organizations on the EU data plane, unredacted text selected for detection is sent to a Baseten-hosted model endpoint outside the EU. Using the EU data plane does not keep ingestion redaction processing within the EU. ### Enable redaction To configure redaction, you need the [**Update** permission](/docs/admin/access-control#object-permissions) on the organization or project. Braintrust recommends enabling redaction at the project level during private preview. Choose whether to redact logs for one project or every project in your organization. * **One project**: Select the project, then open ** Settings** > [** Advanced**](https://www.braintrust.dev/app/~/configuration/advanced). Find **PII redaction**. * **All projects**: Open ** Settings** > [** Logging**](https://www.braintrust.dev/app/~/configuration/org/logging). Find **Security**. This scope includes existing and new projects. Redaction settings inherited from your organization cannot be changed at the project level. Redaction permanently replaces detected sensitive text in stored logs. You cannot recover the original text from those logs, even after disabling redaction. Queries, exports, and scorers receive only the replacement text. Turn on **Redact PII in new logs**. By default, redaction scans five log fields: `input`, `output`, `expected`, `error`, and `metadata`. Each uses the same default **entity types** (categories of sensitive text). * **Choose fields**: Use ** Field** to add or remove fields. Changes save immediately. You can also click a field's trash icon and confirm with **Remove**. * **Choose entity types**: Click a field's pencil icon, search or browse the types, select at least one, and click **Save**. Each field can have a different selection. Removed fields and entity types are no longer redacted in new logs. Organization-level redaction still applies. | Category | Selected by default | Available to add | | - | - | - | | Names and contact information | None | Names, email addresses, phone numbers, full names, first names, middle names, last names | | Locations | None | Addresses, street addresses, cities, states or regions, postal codes, countries | | Government identifiers | National ID numbers, passport numbers | Government IDs, driver's license numbers, license numbers, tax IDs, tax numbers | | Financial information | Card numbers, card security codes, bank accounts | Account numbers, routing numbers, IBANs, payment cards, card expiration dates | | Accounts and credentials | Passwords, secrets | IP addresses, API keys, access tokens, usernames, account IDs, sensitive account IDs, recovery codes | | Dates | None | Dates of birth, sensitive dates, document dates, expiration dates, transaction dates | These defaults also apply when you add a field. Existing fields retain their saved selections. Send representative test logs, then go to [** Logs**](https://www.braintrust.dev/app/~/logs) and inspect the covered fields. In LLM message views, rendered Markdown text, and the data editor, category markers appear as labeled pills. For example, `[REDACTED_EMAIL]` displays as **Redacted email address**. Exports, API queries, and scorers receive the category marker text. Detection is model-based and can miss sensitive values or treat identical inputs differently. Check for missed values and incorrect redactions before relying on the results. Missed values are stored unredacted. For values you can match reliably, such as known credential formats, also [redact before upload](#redact-before-upload). ### Failure behavior Ingestion redaction fails open: * If redaction fails for a log record, Braintrust stores that record without ingestion redaction. Successfully redacted records in the same batch retain their redactions. * If policy lookup or evaluation fails, Braintrust can store the entire batch without ingestion redaction. Sensitive content can be stored when ingestion redaction fails. To prevent specific values from reaching Braintrust, [redact them before upload](#redact-before-upload). ## Redact before upload To remove sensitive values before they reach Braintrust, change data before it leaves your application or machine. Braintrust SDKs provide masking functions and span customizers for data logged by your application, and the `bt` CLI provides span plugins for coding-agent traces: * **Masking functions**: Functions that replace sensitive parts of field values, one value at a time, in a fixed set of fields: `input`, `output`, `expected`, `metadata`, `context`, `scores`, and `metrics`. The SDK runs your function on every record it logs. If it fails, the SDK uploads an error message in place of the affected field instead of the unmasked value. * **Span customizers**: Code that changes, adds, or removes fields on a span, including `tags` and `error`. The SDK runs your customizers before it uploads each span. Which spans they cover, and what happens when one fails, depends on the SDK. * **Span plugins**: JavaScript modules that change [coding-agent traces](/docs/reference/cli/trace), which don't pass through an SDK. The `bt` CLI runs your plugins on your machine before it uploads the traces. See [Transform spans with JavaScript](/docs/reference/cli/trace#transform-spans-with-javascript). Support for masking functions and span customizers varies by SDK. See the Redact sensitive data page for [TypeScript](/docs/sdks/typescript/redact-sensitive-data), [Python](/docs/sdks/python/redact-sensitive-data), [Go](/docs/sdks/go/redact-sensitive-data), or [Java](/docs/sdks/java/redact-sensitive-data). The Ruby and C# SDKs don't support masking functions or span customizers. To remove sensitive data from their traces, use [ingestion redaction](#redact-during-ingestion), or remove it before you log it. ## Next steps * [Explore advanced tracing](/docs/instrument/advanced-tracing) to customize how your application records traces. * [Configure access control](/docs/admin/access-control) to manage who can read your projects and logs. * [Set retention policies](/docs/admin/data-management/retention) to control how long records are kept. # Configure data retention Source: https://braintrust.dev/docs/admin/data-management/retention Extend your retention window to keep data longer, or automate and manually delete data to meet compliance obligations and store less. Retention controls how long Braintrust keeps your logs, experiments, and datasets before they're deleted. Your [plan](/docs/plans-and-limits) sets a base retention window for logs and playground data (14-day retention on Starter and 30-day retention on Pro) that applies across your organization. From there, you can: * [**Extend the retention window**](#extend-the-retention-window) to keep data longer for an additional storage charge (Pro and Enterprise). * [**Automate data deletion**](#automate-data-deletion) to keep less data on specific projects, or across your whole organization (Enterprise). * [**Delete data manually**](#delete-data-manually) to remove specific data on demand (any plan). Retention windows and automations don't apply to [audit logs](/docs/admin/audit-logs). ## Extend the retention window Pro and Enterprise plans can extend the base retention window to keep logs and playground data, including each row's [activity history](/docs/observe/examine-traces#view-activity-history), queryable for longer, for an additional [storage charge](/docs/admin/billing/monitor-usage#retention-storage). * **Starter**: 14-day retention window. Not extendable. * **Pro**: 30-day retention window. Extend up to 180-day retention in ** Settings** > [** Billing**](https://www.braintrust.dev/app/~/configuration/org/billing), in 30-day increments. * **Enterprise**: Custom retention window, up to 365-day retention, configured with your account team. Extending applies org-wide. To keep less data on specific projects, [automate data deletion](#automate-data-deletion). Retention storage is billed on the highest amount retained during the billing month, so extending mid-month isn't prorated. See [Retention storage](/docs/admin/billing/monitor-usage#retention-storage) for how it's measured. Increasing your retention window takes effect immediately. Decreasing it, or removing the extension, is scheduled for the end of the current billing period and can be canceled before then. ## Automate data deletion only available on the [Enterprise plan](/docs/plans-and-limits#plans). A retention automation permanently deletes one object type once its data is older than a period you set. Use it to keep less data than your plan-level window, for example to meet a compliance requirement. Create a separate automation for each object type (logs, experiments, or datasets) you want to limit. You can scope an automation to a single project or to the whole organization. Retention automations only shorten retention. When more than one automation applies, the shortest retention period wins, so an automation can't keep data longer than your organization's retention window. To keep data longer, [extend the retention window](#extend-the-retention-window) instead. ### Set retention for one project 1. Go to ** Settings** > **Project** > [** Data management**](https://www.braintrust.dev/app/~/configuration/data-management). 2. Click **Create automation**. If another automation already exists, click **Automation**. 3. Set **Type** to **Data retention**. 4. Configure settings: * **Retention target**: Logs, Experiments, or Datasets. * **Retention period (days)**: Number of days to keep data before deletion. 5. Click **Create automation**. ### Set retention for all projects An organization-level retention automation applies a retention period for one object type to every project in the organization, including projects created later. Use it to enforce consistent retention without configuring each project. 1. Go to ** Settings** > **Organization** > [** Data management**](https://www.braintrust.dev/app/~/configuration/org/data-management). 2. Click **Create automation**. If another automation already exists, click **Automation**. 3. Select a **Retention target** and enter the **Retention period (days)**. 4. Click **Create**. A project-level automation can shorten retention further for a specific object type, but it can't extend retention beyond the organization-level automation. On self-hosted deployments, organization-level retention requires data plane v2.9.0 or later. Retention automations permanently delete data. Deleted logs, experiments, and dataset rows cannot be recovered. ### What gets deleted * **Logs**: Individual logs older than the retention period are deleted. * **Experiments**: The entire experiment (metadata and all rows) is deleted once its creation timestamp passes the retention period. * **Datasets**: Individual dataset rows older than the retention period are deleted. The dataset itself remains and can accept new rows. ### Example automations | Scenario | Retention period | Rationale | | - | - | - | | Production logs | 90-day retention | Keep recent logs for debugging, delete older data you no longer need. | | Development logs | 7-day retention | Keep short-term history for active development, clean up test data quickly. | | Experiments | 180-day retention | Retain completed evaluations for half a year, then archive or delete. | | Compliance | 30-day retention | Automatically delete user data after the 30-day retention period to meet regulatory requirements. | ## Delete data manually To remove specific data before it ages out of your retention window, delete it directly. Manual deletion is available on all plans. * **Traces and logs**: Select traces in the log table and delete them. See [Delete traces](/docs/observe/view-logs#delete-traces). * **Dataset records**: Select records in a dataset and delete them, or delete them programmatically. See [Delete records](/docs/annotate/datasets/manage#delete-records). * **Experiment rows**: Select rows in an experiment table and delete them, or delete them through the API. ## Self-hosted deployments ### Soft deletion For [self-hosted deployments](/docs/admin/self-hosting) (data plane v1.1.21+), data is soft-deleted by marking it unused. A background process purges unused files within 24 hours, providing a grace period to restore accidentally deleted data. Configure a service token for your data plane to enable retention. See [Data plane manager](/docs/admin/self-hosting/configure/telemetry#data-retention) for details. ### Cloud provider policies For [self-hosted deployments](/docs/admin/self-hosting), review your cloud provider's bucket retention policies at the project and organization level before configuring retention in Braintrust. Org-wide policies are common and can conflict with Braintrust retention settings or delete data independently of Braintrust. All Braintrust deletes are recorded in a delete log. If data goes missing, contact [Braintrust support](mailto:support@braintrust.dev) and Braintrust will audit these logs to determine whether the deletion was caused by a Braintrust retention automation or another process. Braintrust recommends enabling audit logging on your storage bucket to help trace the root cause of unexpected deletes. ## Next steps * [Export to cloud storage](/docs/admin/data-management/export) for long-term archival * [Set up alerts](/docs/observe/alerts) to monitor data quality * [Self-hosting configuration](/docs/admin/self-hosting/configure) for data plane configuration # Bring your own cloud (BYOC) Source: https://braintrust.dev/docs/admin/deployment/byoc Have Braintrust operate the data plane inside your own cloud account, keeping data in your boundary without owning day-2 operations. BYOC is a Braintrust-operated data plane deployed into a cloud account or project that you own. You keep the cloud boundary, billing relationship, audit policy, and governance controls. Braintrust manages the deployment, upgrades, service health, and operational response for the Braintrust data-plane services. BYOC suits teams that need sensitive AI data to stay inside their own cloud, but don't want to take on the day-2 operations of running the data plane themselves. For a comparison with SaaS and self-hosted, see [Deployment options](/docs/admin/deployment). BYOC requires an Enterprise plan. See [Manage billing](/docs/admin/billing) for plan details. ## Responsibility boundary * **Braintrust** provisions the data plane in your account, and operates upgrades, scaling, monitoring, and incident response for the Braintrust data-plane services. * **Your team** owns the cloud account, billing, governance, and audit policy, and provides the access path and cloud inputs Braintrust needs to operate the deployment. ## What you provide With BYOC, you typically provide: * A dedicated cloud account, project, or subscription with billing enabled and an agreed deployment region. * Completion of a reviewed bootstrap or preparation step that creates the required access path for Braintrust automation and approved support events. * Confirmation that required quotas, regional capacity, and cloud service limits are ready for the selected deployment region. * Any optional customer-specific networking, DNS, TLS, ingress, private-connectivity, or egress requirements that Braintrust should account for during deployment planning. By default, Braintrust uses its standard managed networking and endpoint pattern for the selected cloud. You also remain responsible for reviewing and managing the cloud governance controls in that environment, including organization policies, IAM constraints, allowed regions, audit-log export, SIEM routing, and billing alerts. Braintrust may ask you to confirm whether any of these controls affect the agreed deployment path. ## BYOC or self-hosted BYOC and self-hosted both keep your data in your own cloud. They differ in who operates the data plane: * Choose **BYOC** when you want Braintrust to operate the data plane on your behalf. * Choose [**self-hosted**](/docs/admin/self-hosting) when your team operates the data plane and can consume Braintrust's published Terraform modules and Helm charts. If your organization prohibits external operator access to your cloud environment, BYOC is usually not the right fit, because Braintrust needs an access path to operate the deployment. Self-hosted is the cleaner starting point. ## Next steps * [Deployment options](/docs/admin/deployment) — Compare SaaS, BYOC, and self-hosted. * [Architecture](/docs/admin/self-hosting/architecture) — Understand the control plane, data plane, and where data is stored. * [Contact Braintrust](mailto:support@braintrust.dev) — Start a BYOC onboarding conversation. # Deployment options Source: https://braintrust.dev/docs/admin/deployment/index Compare SaaS, BYOC, and self-hosted deployments and choose the right data residency and operational model for your team. Braintrust supports several deployment options for teams with different requirements around data residency, compliance, operational ownership, and cloud control. This page summarizes the supported options and the responsibility boundaries for each. ## Options at a glance Braintrust separates the **control plane** from the **data plane**. The control plane is always operated by Braintrust and provides the product UI, authentication, organization metadata, and platform management. The data plane stores and processes your sensitive AI data. The deployment options differ in where the data plane runs and who operates it. * **SaaS**: Braintrust operates both the control plane and the data plane, and your data is stored in Braintrust-managed infrastructure. Best for teams that want the fastest path with no cloud operations. * **BYOC**: Braintrust operates the control plane and the data plane, but the data plane runs inside your dedicated cloud account or project. Best for teams that need customer-owned data residency without owning day-2 operations. * **Self-hosted**: Braintrust operates the control plane, and you deploy and operate the data plane in your own cloud using Braintrust's published deployment artifacts. Best for teams with platform capacity and a requirement to operate the environment themselves. * **Non-standard self-hosted**: Braintrust operates the control plane, and you operate the data plane, but your deployment cannot use the standard artifacts or requires material deviations from the reference architecture. This path needs review because it increases upgrade and support risk. SaaS is available on Starter, Pro, and Enterprise plans. BYOC and self-hosted deployments require an Enterprise plan. Non-standard self-hosted deployments also require review to confirm they are viable. See [Manage billing](/docs/admin/billing) for plan details. ## Comparison | Dimension | SaaS | BYOC | Self-hosted | Non-standard self-hosted | | :- | :- | :- | :- | :- | | Plan availability | Starter, Pro, Enterprise | Enterprise only | Enterprise only | Enterprise only, review required | | Data location | Braintrust-managed infrastructure | Your cloud | Your cloud | Your cloud | | Data-plane operator | Braintrust | Braintrust | You | You | | Infrastructure provisioning | Braintrust | Braintrust, in your account or project | You, using Braintrust artifacts | You, with custom mapping | | Upgrade owner | Braintrust | Braintrust | You | You, with more manual work | | Operational burden | Lowest | Low | High | Highest | | Customization model | Standard service | Standardized managed deployment | Published module or chart configuration | Review required | | Best fit | Fastest path, no cloud ops | Customer-owned data boundary without day-2 ops | Customer-operated platform | Standard path blocked by constraints | ## How to choose * Choose **SaaS** when you don't need a customer-owned data boundary and want the lowest operational burden. * Choose **BYOC** when sensitive AI data must remain in your cloud account or project, but you want Braintrust to operate the data plane. * Choose **self-hosted** when you must operate the environment yourself and can use Braintrust's published deployment artifacts. * Review **non-standard self-hosted** when you cannot use the standard artifacts or require material architecture changes. If your organization prohibits external operator access to your cloud environment, BYOC is usually not the right fit. Self-hosted is the cleaner starting point. ## Next steps * [Bring your own cloud (BYOC)](/docs/admin/deployment/byoc) — Have Braintrust operate the data plane in your cloud account. * [Self-hosting Braintrust](/docs/admin/self-hosting) — Deploy and operate the data plane yourself. * [Architecture](/docs/admin/self-hosting/architecture) — Understand the control plane, data plane, and where data is stored. # Administration Source: https://braintrust.dev/docs/admin/index Manage organizations, access control, and infrastructure Administer your Braintrust organization, control access to resources, configure integrations, and manage self-hosted deployments. ## Manage organizations Organizations represent your team or business. Configure organization-wide settings including: * **Members**: Invite users and assign permission groups * **API keys**: Create and manage authentication credentials * **Service tokens**: Set up system integrations with service accounts * **AI providers**: Configure API keys for OpenAI, Anthropic, Google, and others * **Environment variables**: Set secrets for functions across your organization Go to **Settings** to manage these settings. See [Manage organizations](/docs/admin/organizations) for details. ## Control access Braintrust provides flexible access control at multiple levels: * **Organization level**: Assign users to the Owners, Engineers, Viewers, and All AI Provider Access built-in groups * **Project level**: Grant specific permissions to custom permission groups * **Object level**: Control access to individual experiments, datasets, or prompts Create custom permission groups to match your team's needs. Service accounts enable secure system integrations with granular permissions. See [Control access](/docs/admin/access-control) for details. ## Manage projects Projects organize AI features in your application. Each project contains logs, experiments, datasets, and functions. Configure project settings including: * **Tags**: Organize and filter logs across your project * **Human review scores**: Define manual review criteria * **Aggregate scores**: Combine multiple metrics into single values * **Online scoring**: Automatically evaluate production logs * **Comparison keys**: Customize experiment comparisons See [Manage projects](/docs/admin/projects) for details. ## Separate production and staging There are several ways to handle production vs. staging data: **Use separate projects (recommended)**: Split production and staging into different projects so they're isolated and code changes to staging cannot affect production. This also allows you to enforce [access controls](/docs/admin/access-control) at the project level. **Use tags within one project**: If it's easier to keep everything in one project (e.g., to triage issues in one place), use [tags](/docs/observe/view-logs#organize-with-tags) to separate environments. Filter by tags to view production or staging data independently. **Use separate organizations**: For physical isolation, create separate organizations for production and staging, each mapping to a different deployment. Experiments, prompts, and playgrounds can use data across projects. For example, you can reference a prompt from your production project in your staging logs, or evaluate using a dataset from staging in a different project. ## Set up automations Automate routine tasks and stay informed about production issues: * **Alerts**: Get notified when metrics exceed thresholds or errors spike * **Data management**: Configure retention policies and archiving rules Automations run in the background, keeping your data clean and your team informed. See [data management](/docs/admin/data-management/export) and [alerts](/docs/observe/alerts) for details. ## Choose a deployment option Braintrust supports several deployment options that differ in where your data plane runs and who operates it: * **SaaS**: Braintrust operates both the control plane and the data plane in Braintrust-managed infrastructure. * **BYOC**: Braintrust operates the data plane inside your own cloud account or project. * **Self-hosted**: Your team deploys and operates the data plane in your cloud using Braintrust's published artifacts. Each option keeps the same managed control plane. See [Deployment options](/docs/admin/deployment) to compare them and choose the right fit. ## Manage billing Control your Braintrust subscription and costs: * **Change plans**: Upgrade or downgrade between Starter, Pro, and Enterprise * **Monitor usage**: Track trace spans, processed data, and scores * **Cost alerts**: Get notified when spending exceeds thresholds * **Payment methods**: Update billing information and access invoices See [Manage billing](/docs/admin/billing) for details. ## Configure authentication Braintrust supports multiple authentication methods: * **Email/password**: Standard authentication for individuals * **SSO**: Integrate with your identity provider (Okta, Google Workspace, Azure AD) * **API keys**: Authenticate SDK and API requests * **Service tokens**: Authenticate service accounts for system integrations See [Configure authentication](/docs/admin/authentication) for details. ## Review audit logs Audit logs show who changed projects, access controls, API keys, etc. Members of the [**Owners** permission group](/docs/admin/access-control#built-in-permission-groups) can read audit logs by default. To let anyone else read them, grant the **Read audit logs** permission to a permission group. See [Audit logs](/docs/admin/audit-logs) for the event reference. # Manage organizations Source: https://braintrust.dev/docs/admin/organizations Configure organization settings and integrations Organizations represent your team or business in Braintrust. Each organization contains projects, users, and organization-wide settings. You can create multiple organizations to organize projects differently, and users can be members of multiple organizations. Configure your organization by going to **Settings**. You can also customize organization settings using the [API](/docs/api-reference). ## Create an organization To create a new organization: 1. Click your organization name in the upper-left corner and select **+ Create organization**. 2. Enter a name for your organization. 3. Select a **[data plane region](#data-plane-region)** (US or EU) from the dropdown. The region is pre-selected based on your location, but you can change it before continuing. After creating an organization, you cannot change its data plane region. 4. Click **Continue**. After creating an organization, you'll see onboarding information you can follow to instrument your code and start logging traces to Braintrust. ### Data plane region Braintrust's architecture has two main components: * The **data plane** stores all sensitive data, including experiment records, logs, traces, spans, datasets, and prompt completions. It consists of the Braintrust API, a PostgreSQL database, Redis cache, object storage, and Brainstore (a high-performance query engine for real-time trace ingestion). * The **control plane** provides the web UI, authentication, user management, and metadata storage (project names, experiment names, organization settings). The control plane does not store or process your sensitive data. | Data | Location | | - | - | | Experiment records (input, output, expected, scores, metadata, traces, spans) | Data plane | | Log records (input, output, expected, scores, metadata, traces, spans) | Data plane | | Dataset records (input, output, metadata) | Data plane | | Prompt playground prompts | Data plane | | Prompt playground completions | Data plane | | Human review scores | Data plane | | Project-level LLM provider secrets (encrypted) | Data plane | | Org-level LLM provider secrets (encrypted) | Control plane | | API keys (hashed) | Control plane | | Experiment and dataset names | Control plane | | Project names | Control plane | | Project settings | Control plane | | Git metadata about experiments | Control plane | | Organization info (name, settings) | Control plane | | Login info (name, email, avatar URL) | Control plane | | Auth credentials | [Clerk](https://clerk.com/) | Each data plane region has its own endpoint: | Region | API URL | MCP endpoint | | - | - | - | | US | `https://api.braintrust.dev` | `https://api.braintrust.dev/mcp` | | EU | `https://api-eu.braintrust.dev` | `https://api-eu.braintrust.dev/mcp` | `https://gateway.braintrust.dev` is the global hosted Gateway endpoint. It uses DNS latency routing and health checks across Braintrust-hosted Gateway regions. Gateway routing is separate from your organization's [data plane region](/docs/admin/organizations#data-plane-region). When [logging is enabled](/docs/deploy/gateway#enable-logging), Gateway logs are written to your organization's configured data plane. After creating an organization, you can view its data plane region and API URL at ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). If your organization requires storing data in your own cloud, see [Deployment options](/docs/admin/deployment) to compare BYOC and self-hosting. ## Find your organization ID Your organization ID is a UUID used in API calls, [self-hosting configuration](/docs/admin/self-hosting/configure/security#configure-organization-authorization), and audit log queries. To copy your organization ID or name to the clipboard: 1. Click your organization name in the upper-left corner. 2. Hover over the current organization (marked with a checkmark) to open its submenu. 3. Select **Copy organization ID** or **Copy organization name**. ## Manage members ### Invite members To add members to your organization: 1. Go to ** Settings** > [** Members**](https://www.braintrust.dev/app/~/configuration/org/team). 2. Click **Invite**. 3. Enter email addresses (one per line for multiple invites). 4. Select one or more permission groups. 5. Click **Send invites**. You can also invite members directly from the [projects page](/docs/admin/projects#view-projects) by selecting the **Members** > **Invite members**. Invited members receive an email with a link to join your organization. They must be assigned to at least one permission group. When [SCIM provisioning](/docs/admin/scim) manages your organization's membership, **Invite** is disabled and the API rejects membership changes. Add the user to your identity provider's mapped group instead. Service accounts are exempt and can still be added by hand. ### Remove members To remove members from your organization: 1. Go to ** Settings** > [** Members**](https://www.braintrust.dev/app/~/configuration/org/team). 2. Find the member you want to remove. 3. Click the icon next to their name. Organizations must have at least one member. You cannot remove yourself from an organization. To transfer ownership or leave an organization, ensure another member is present and can manage the organization after you leave. When [SCIM provisioning](/docs/admin/scim) manages your organization's membership, the remove control is disabled and the API rejects membership changes. Remove the user from your identity provider's mapped group instead. Service accounts are exempt and can still be removed by hand. ## Configure permission groups Permission groups are the core of Braintrust's access control system. They are collections of users that can be granted specific permissions to projects, experiments, and datasets. 1. Go to ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups). 2. View existing groups or create new ones. 3. Assign users to groups when inviting members. For detailed information about creating and managing permission groups, see [Control access](/docs/admin/access-control). ## Manage API keys Create API keys for authentication: 1. Go to ** Settings** > [** API keys**](https://www.braintrust.dev/app/~/configuration/org/api-keys). 2. Click **+ API key**. 3. Enter a name to identify the key. 4. Choose an expiration. It defaults to one year and cannot be changed after creation. 5. Click **Create**. 6. Copy the key immediately (it won't be shown again). API keys inherit permissions from the user who created them. Organization owners can view and manage all API keys in the organization. API keys can be created only in the Braintrust UI, not through the API. Store API keys securely. Anyone with an API key can access Braintrust with the permissions of the key's creator. ### Disable API key creation To prevent members from creating new API keys: 1. Go to ** Settings** > [** API keys**](https://www.braintrust.dev/app/~/configuration/org/api-keys). 2. Toggle **Disable API key creation** on. When API key creation is disabled, members cannot create new keys. Existing keys continue to work. Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can change this setting. ## Create service tokens Service tokens enable system integrations without tying credentials to individual users: 1. Go to ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens). 2. Click **+ Service token**. 3. Enter a name for the service account. 4. Assign permission groups or grant specific permissions. 5. Enter a name for the service token. 6. Choose an expiration. It defaults to one year and cannot be changed after creation. 7. Click **Create**. 8. Copy the token immediately (it won't be shown again). Service tokens use the `bt-st-` prefix. Pass a service token as the bearer token in API requests or as the API key in SDK calls. Requests use the service account's permissions. Service tokens also support [user impersonation](/docs/api-reference#impersonate-users) when the service account has the `Owner` role in every organization the target user belongs to. The target user must belong to at least one organization. These requirements apply to the service account, not the person who created the token. Only organization owners can create service tokens, either at ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens) in the Braintrust UI or by calling [`POST /v1/service_token`](/docs/api-reference/servicetokens/create-service_token) and authenticating with a service token that has organization-owner permissions. User API keys cannot be used to create service tokens. ## Configure AI providers Configure organization-wide AI provider credentials so playgrounds, experiments, and the gateway can call models without users needing individual API keys. See [Configure AI providers](/docs/admin/ai-providers). ## Configure Loop ### Choose models for Loop [Loop](/docs/loop) runs on Braintrust's [built-in models](/docs/admin/ai-providers#manage-built-in-models) or on your organization's own OpenAI-compatible [AI providers](/docs/admin/ai-providers). ### Control Loop logging By default, Braintrust logs frontend Loop interactions for product improvement and support. To opt out: 1. Go to ** Settings** > [** Loop**](https://www.braintrust.dev/app/~/configuration/org/loop). 2. Toggle **Loop logging** off. When Loop logging is off, Loop continues to work normally, but Braintrust does not log frontend interactions. Self-hosted organizations have Loop logging disabled by default. ## Enable Slack integration Connect Slack to send alerts and notifications to your channels. Only members of the **Owners** [permission group](/docs/admin/access-control), or a custom permission group with the **Manage settings** organization permission, can manage the Slack workspace connection. 1. Go to ** Settings** > [** Integrations**](https://www.braintrust.dev/app/~/configuration/org/integrations). 2. Click **Enable Slack**. 3. Authorize Braintrust to access your Slack workspace. 4. Select which channels Braintrust can access. Once enabled, you can configure alerts to send notifications to specific Slack channels. Braintrust also sends Slack direct messages to assignees when spans or dataset rows are assigned to them for review, if they have linked their Slack accounts. See [Set up alerts](/docs/observe/alerts), [Set up spend alerts](/docs/admin/billing/monitor-usage#set-up-spend-alerts), and [Assign rows for review](/docs/annotate/human-review/manage-review-work#assign-rows-for-review) for details. If the Braintrust Slack app is updated with new permissions, a warning banner and **Refresh permissions** button will appear next to your connected workspace. Click **Refresh permissions** to re-authorize through the OAuth flow and enable the latest features. ### Braintrust link previews in Slack When the Slack integration is enabled, Braintrust automatically expands links to logs, traces, and experiments that you paste into Slack. What each preview shows depends on whether you have linked your Slack account. Supported link types: * Log URLs. * Trace URLs. * Experiment URLs. Braintrust shows a public preview with the resource title and description for all Slack users, including those who have not linked their account. Private resource details are only shown to Slack users who have linked their Slack account to a Braintrust account with access to the resource. To link your accounts, run `/braintrust link` in any Slack channel, or click **Link Braintrust account** in the prompt Braintrust sends when you open the Slack App Home or first interact with a link preview. When you link from the App Home prompt and the Braintrust account you are signed in to uses a different email address than your Slack account, the authorization page displays a warning with a **Switch account** option. To unlink later, run `/braintrust unlink` in Slack. ## Set environment variables Define secrets available to all functions in your organization: 1. Go to ** Settings** > [** Env variables**](https://www.braintrust.dev/app/~/configuration/org/env-vars). 2. Click **Add variable**. 3. Enter a key and value. 4. Click **Save**. Environment variables are accessible from prompts, scorers, and tools. Use them for API keys, database credentials, or configuration values. Each variable lists a redacted preview of its value and a **Last updated** timestamp that tracks when the value itself was last changed, along with the user who made the change. Renaming a variable or editing other metadata does not bump this timestamp. ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} // Access in TypeScript functions const apiKey = process.env.MY_API_KEY; ``` ```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} # Access in Python functions import os api_key = os.environ["MY_API_KEY"] ``` ## Configure API URLs (self-hosted) only available on the [Enterprise plan](/docs/plans-and-limits#plans). If your organization leaves the Enterprise plan, existing data plane URLs stay in effect, but you can no longer change them. This applies to both this settings page and the API. For [self-hosted deployments](/docs/admin/self-hosting), set custom API URLs: 1. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). 2. Enter your URLs: * **API URL**: Main API endpoint. * **Gateway URL**: Gateway endpoint. * **Realtime URL**: Realtime API endpoint. 3. If your data plane is deployed behind a VPN or on a private network (not accessible from the public internet), enable **Data plane is on a private network**. This checkbox enables detection of Chrome Local Network Access permission issues. When enabled, if fetch requests fail due to blocked LNA permissions, Braintrust will display instructions on how to grant the necessary permissions in Chrome settings. 4. Click **Save**. Test connectivity using the provided test commands. When you configure these API URLs, Braintrust automatically provisions a service token for your data plane. This enables features like [data retention](/docs/admin/data-management/retention) without requiring manual service token setup. You can view and manage this token in the [Service tokens](#create-service-tokens) section. Braintrust-hosted organizations (US or EU) can view their region and API URL by navigating to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). ## Set git metadata logging Control which git metadata fields are logged: 1. Go to ** Settings** > [** Logging**](https://www.braintrust.dev/app/~/configuration/org/logging). 2. Enable **Collect git metadata**. 3. Select the fields to log. Your changes save automatically. Git metadata helps track which code version generated specific logs or experiment results. Diff content is opt-in. By default, Braintrust collects standard git metadata but not the **Diff** field. To capture diffs, enable **Diff** here. The server strips `git_diff` from logged repo info unless your org has enabled it, regardless of what the SDK sends (data plane v2.2.0 or later). Until you save a policy here, the Python SDK (v0.23.0+) collects no git metadata by default. Pass `git_metadata_settings` in code to override. ## Redact sensitive log data [Ingestion redaction](/docs/admin/data-management/protect-sensitive-data#redact-during-ingestion) lets you apply a shared policy to detect sensitive text in incoming project logs. Configure it for one project or all projects in your organization. See [Enable redaction](/docs/admin/data-management/protect-sensitive-data#enable-redaction) for setup and permissions. ## Control image rendering External images in logs can send sensitive data to untrusted servers when your browser loads them. Configure image rendering to control this risk: 1. Go to ** Settings** > [** Logging**](https://www.braintrust.dev/app/~/configuration/org/logging). 2. Under **Security**, select an **Image rendering** mode: * **Auto-load images**: Images are loaded automatically (default). * **Click to load**: Show placeholder, click to load each image. * **Block all images**: Never load external images. Some image URLs are never loaded automatically, whichever mode you select. This covers URLs that name a host on your local or private network, such as `localhost`, a `.local` name, or a private or link-local IP address, and URLs using a scheme that a browser should never load an image from, such as `javascript:` or `file:`. In **Auto-load images** and **Click to load**, these images show a **Load image** button, so loading one is always a deliberate choice. In **Block all images**, they are blocked along with every other external image. Inline `data:` and `blob:` images are unaffected, because they never leave the browser. ## Configure environments Create environments to version prompts and functions: 1. Go to ** Settings** > [** Environments**](https://www.braintrust.dev/app/~/configuration/org/environments). 2. Click **+ Environment**. 3. Enter environment name (e.g., "production", "staging", "dev"). 4. Click **Create environment**. Assign prompt and function versions to environments to separate development from production. See [Manage environments](/docs/deploy/environments) for details. ## Monitor infrastructure (self-hosted) Organization owners and admins can view infrastructure metrics for self-hosted deployments directly from the Braintrust UI. See [Monitor your infrastructure](/docs/admin/self-hosting#monitoring) for details. ## Delete an organization To delete an organization, [contact Braintrust](https://braintrust.dev/contact). ## Next steps * [Control access](/docs/admin/access-control) with permission groups * [Manage projects](/docs/admin/projects) within your organization * [Use the Braintrust gateway](/docs/deploy/gateway) for centralized provider access * Set up [alerts](/docs/observe/alerts) and [data management](/docs/admin/data-management/export) # Personal settings Source: https://braintrust.dev/docs/admin/personal-settings Configure your individual user preferences Personal settings control your individual user experience across all projects in your organization. Access personal settings by going to ** Settings** > [** Profile**](https://www.braintrust.dev/app/~/configuration/personal/user). ## Profile Update your name and avatar to help team members identify you in comments, reviews, and activity logs. ## Appearance Choose between light, dark, or system appearance modes. ## Default data display format Set your preferred format for viewing data in span fields across all traces: * **Pretty** - Parses objects deeply and renders values as Markdown (optimized for readability) * **JSON** - JSON highlighting and folding * **YAML** - YAML highlighting and folding * **Tree** - Hierarchical tree view for nested data structures When viewing individual span fields, additional format-specific views become available for certain data types, including LLM (formatted AI messages), LLM Raw (unformatted AI messages), and HTML (rendered HTML content). You can override this default for individual span fields when viewing traces. These overrides are remembered per field type. To clear all overrides and restore your default preferences across all fields, use the reset button. ## Next steps * [View your logs](/docs/observe/view-logs) to see how data views work in practice * [Manage organizations](/docs/admin/organizations) to configure organization-wide settings # Manage projects Source: https://braintrust.dev/docs/admin/projects Configure project settings and features Projects organize AI features in your application. Each project contains logs, experiments, datasets, prompts, and other functions. Configure project-specific settings to customize behavior for your use case. ## View projects Select your organization from the top-left menu to see a list of all projects in your organization, including key project details such as: * Name * Creator * Counts of playgrounds, experiments, datasets, logs, and spans * Log size estimates for different time periods * Token usage and duration metrics You can use the Braintrust API to change the name, description, or creator of an existing project. See the API reference for [updating projects](/docs/api-reference/projects/partially-update-project). From the [`bt` CLI](/docs/reference/cli/quickstart), list all projects or view details for a specific one: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} bt projects list bt projects view my-project ``` `bt projects view` opens the project in your browser, unlike `bt functions view` which displays details in the terminal. Use the projects list to organize and find projects: * **Search**: Use the search box to find projects by name * **Filter**: Use the **Filter** menu to add custom filtering. * **Sort**: Choose sorting options in column headers. * **Configure columns**: Customize which columns are visible in the table * **Save views**: Save your table configuration as a custom table view. See [Create custom table views](#create-custom-table-views). ### Create custom table views To create or update a custom table view: 1. Apply the filters and display settings you want. 2. Open the menu and select **Save view\...** or **Save view as...**. Custom table views are visible to all project members. Creating or editing a table view requires the **Update** project permission. ### Set default table views You can set default views at three levels: * **Organization default**: Visible to all members when they open the page. This applies per page. For example, you can set separate organization defaults for Logs, Experiments, and Review. To set an organization default, you need the **Manage settings** organization permission (included by default in the **Owner** role). See [Access control](/docs/admin/access-control) for details. * **Project default**: Overrides the organization default for everyone viewing this project. To set a project default, you need the project-level **Update** permission. Project admins can set project defaults even without organization-level permissions. See [Access control](/docs/admin/access-control) for details. * **Personal default**: Overrides the project and organization defaults for you only. Personal defaults are stored in your browser, so they do not carry over across devices or browsers. To set a default view: 1. Switch to the view you want by selecting it from the menu. 2. Open the menu again and hover over the currently selected view to reveal its submenu. 3. Choose **Set as personal default view**, **Set as project default view**, or **Set as organization default view**. To clear a default view: 1. Open the menu and hover over the currently selected view to reveal its submenu. 2. Choose **Clear personal default view**, **Clear project default view**, or **Clear organization default view**. Default view settings are mutually exclusive on a given view. Setting one type of default on a view automatically clears any other default that was previously set on the same view. When you open a page, Braintrust loads the first match in this order: personal default, project default, organization default, then the standard "All ..." view (for example, **All traces view**). ## Analyze projects with Loop Use Loop to analyze a project in depth. Open [** Loop**](https://www.braintrust.dev/app/~/loop) from the left sidebar for a full-page thread, or start a chat from a project's ** Overview** page. See [Analyze projects](/docs/loop/capabilities#analyze-projects) for example prompts, and [What Loop can do](/docs/loop/capabilities) for the rest of Loop's capabilities. ## Create a project 1. Navigate to your organization's project list 2. Click **+ Project** 3. Enter a project name 4. Optionally add a description 5. Click **Create** If a project already exists, `projects.create()` returns a handle. There is no separate `.get()` method. ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} import * as braintrust from "braintrust"; // Get a handle to the project (creates if it doesn't exist) const project = braintrust.projects.create({ name: "my-project" }); // Use the project to create functions project.prompts.create({...}); project.tools.create({...}); ``` ```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} import braintrust # Get a handle to the project (creates if it doesn't exist) project = braintrust.projects.create(name="my-project") # Use the project to create functions project.prompts.create(...) project.tools.create(...) ``` To create the project inside an existing [project group](/docs/admin/access-control/manage-permissions#manage-project-groups) when it doesn't exist yet, pass `projectGroupName` (TypeScript SDK v3.36.0 or later) or `project_group_name` (Python SDK v0.44.0 or later). You need permission to create projects in that group. [`bt functions push`](/docs/reference/cli/functions#bt-functions-push) (v0.22.2 or later) also uses this project group when it creates the project. ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} const project = braintrust.projects.create({ name: "my-project", projectGroupName: "my-group", }); ``` ```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} project = braintrust.projects.create(name="my-project", project_group_name="my-group") ``` Projects are automatically created when initializing experiments or loggers: ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} import * as braintrust from "braintrust"; // Creates "my-project" if it doesn't exist const experiment = braintrust.init("my-project", { experiment: "my-experiment" }); ``` ```python theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} import braintrust # Creates "my-project" if it doesn't exist experiment = braintrust.init( project="my-project", experiment="my-experiment" ) ``` For more details, see the SDK reference for [Python](/docs/sdks/python/api-reference) or [TypeScript](/docs/sdks/typescript/api-reference). Create a project with the [`bt` CLI](/docs/reference/cli/quickstart): ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} bt projects create my-project ``` ## Configure AI providers Configure project-scoped AI provider credentials to override organization-level defaults for a specific project. See [Configure AI providers](/docs/admin/ai-providers#project-providers). ## Add tags Tags help organize and filter logs, datasets, and experiments: 1. Go to ** Settings** > [** Tags**](https://www.braintrust.dev/app/~/configuration/tags). 2. Click **Add tag**. 3. Enter tag details: * **Name**: Tag identifier. * **Color**: Visual indicator. * **Description**: Optional explanation. 4. Click **Save**. Use tags to track data by user type, feature, environment, or any custom category. Tags are configured at the project level and shared across all objects — logs, experiments, dataset records, and entire datasets. For more information about using tags, see [Tag traces](/docs/annotate/labels#apply-tags) and [Tag and star datasets](/docs/annotate/datasets/manage#tag-and-star-datasets). ## Configure human review Review scores appear in all logs and experiments in a project. Use them for quality control, data labeling, or [feedback collection](/docs/annotate/human-review/score-traces). only available on [Pro and Enterprise plans](/docs/plans-and-limits#plans). 1. Go to ** Settings** > [** Human review**](https://www.braintrust.dev/app/~/configuration/review). 2. Click **+ Human review score**. 3. Enter a name and description for your score. Descriptions support Markdown. 4. Select a score type: * **Categorical score**: Predefined options with assigned scores. Each option gets a unique percentage value between 0% and 100% (stored as 0 to 1). Use for classification tasks like sentiment or correctness categories. Also supports writing to the `expected` field instead of creating a score. * **Continuous score**: Numeric values between 0% and 100% with a slider input control. Use for subjective quality assessments like helpfulness or tone. * **Free-form input**: String values written to `metadata` or `expected` at a specified path. Use for explanations, corrections, or structured feedback. Enable **Write to expected field instead of metadata** when configuring a free-form score to store reviewer text at `expected.`, including nested paths such as `expected.rubric.notes`. 5. (Optional) Expand **Score visibility** to configure who sees this score during review: * Select members or permission groups to limit visibility to specific reviewers. If you don't select anyone, the score is visible to everyone. * Click **+ Condition** to show the score only when a filter condition is true, such as when another score exceeds a threshold. See [Show scores conditionally](/docs/annotate/human-review#show-scores-conditionally) for details. 6. Click **Save**. Score visibility controls which reviewers see a score in the review modal. It declutters the review experience for large teams. It is not an access control or security boundary: any reviewer with hidden scores can reveal them with the **Show all scores** toggle. You can also create human review scores as you review traces. In the trace view, click **+ Human review score** and define the score as described above. For more information, see [Add human feedback](/docs/annotate/human-review). ## Create aggregate scores Combine multiple scores into a single metric: 1. Go to ** Settings** > [** Aggregate scores**](https://www.braintrust.dev/app/~/configuration/aggregate-scores). 2. Click **Add aggregate score**. 3. Define the aggregation: * **Name**: Score identifier. * **Type**: Weighted average, minimum, or maximum. * **Selected scores**: Scores to aggregate. * **Weights**: For weighted averages, set score weights. * **Description**: Optional explanation. 4. Click **Save**. Aggregate scores appear in experiment summaries and comparisons. Use them to create composite quality metrics or overall performance indicators. ## Set up online scoring Define project-level scoring rules that automatically evaluate production logs as they arrive. These rules can be created here or when creating and editing scorers and classifiers. 1. Go to ** Settings** > [** Automations**](https://www.braintrust.dev/app/~/configuration/automations). 2. Click **+ Automation**. 3. Configure the rule: * **Name**: Rule identifier. * **Scorers**: Select which scorers and classifiers to run. * **Sampling rate**: Percentage of logs to evaluate (1-100%). * **Filter**: Optional SQL query to select specific logs. * **Span type**: Apply to root spans or all spans. 4. Click **Save**. Online scoring rules run asynchronously in the background. View results in the logs page alongside other scores. Rules can also be created and managed when working with individual scorers. For more information, see [Create scoring rules](/docs/evaluate/score-online#create-scoring-automation-rules). ## Configure span iframes Render custom iframe URLs as fields in trace spans: 1. Go to ** Settings** > [** Span iframes**](https://www.braintrust.dev/app/~/configuration/span-iframes). 2. Click **Create span iframe**. If the project already has a span iframe, click **Span iframe** instead. 3. Configure the iframe: * **Field name**: The name of the field, displayed at the top of the section in the span. Must be unique within the project. * **Description**: Optional description of the iframe. * **URL**: The URL to load, with optional mustache `{{parameters}}` referencing span data paths, such as `https://example.com/{{input.id}}?q={{metadata.foo.bar}}`. Expand **Preview (without parameters)** to render the URL before you save. * **Post span data**: Post a message with the span data to the iframe on load. Enable this when the data is too large to fit in the URL, or when the iframe writes data back to the span. 4. Click **Create span iframe**. Use span iframes to render HTML, charts, or custom visualizations directly in trace views. For more information, see [Customize span rendering](/docs/instrument/advanced-tracing#customize-span-rendering). ## Set comparison key Customize how experiments match test cases: 1. Go to ** Settings** > [** Advanced**](https://www.braintrust.dev/app/~/configuration/advanced). 2. Enter a SQL expression (default: `input`). 3. Click **Save**. Examples: * `input.question` - Match by question field only. * `input.user_id` - Match by user. * `[input.query, metadata.category]` - Match by multiple fields. The comparison key determines which test cases are considered the same across experiments. For more information, see [Compare experiments](/docs/evaluate/compare-experiments#set-a-comparison-key). ## Set a default preprocessor Override the preprocessor used by functions that take one, such as [custom facets](/docs/observe/topics/custom-facets): 1. Go to ** Settings** > [** Advanced**](https://www.braintrust.dev/app/~/configuration/advanced). 2. Under **Default preprocessor**, select an option: * **thread**: The built-in preprocessor that formats traces as conversation threads. * A saved preprocessor function defined in this project. Functions that use preprocessors, such as the built-in [Topics facets](/docs/observe/topics), apply this preprocessor instead of their built-in default, unless explicitly overridden at invocation time. ## Configure Monitor chart time dimension By default, Monitor charts use the `created` field (the ingestion timestamp) as the time dimension for grouping spans into time-series buckets. If you batch-ingest historical spans, their `created` timestamps all reflect the ingestion time rather than when the events occurred, which compresses the chart into a narrow window. To use the span's actual start time (`metrics.start`) instead: 1. Go to ** Settings** > [** Advanced**](https://www.braintrust.dev/app/~/configuration/advanced). 2. Turn on **Use metrics.start for monitor charts time bucketing**. 3. Click **Save**. When enabled, time series charts on the [** Dashboards**](https://www.braintrust.dev/app/~/dashboards) page use `metrics.start` for bucketing. Topic summary charts always use `created` regardless of this setting. For query performance, Monitor charts still apply a `created` (ingestion time) filter with a one-day buffer, even when bucketing by `metrics.start`. The time range you select on the Monitor page must include the spans' ingestion time, not only their event time. This setting works best when events are ingested close to when they occurred. If event time and ingestion time are far apart, such as a backfill of spans whose events occurred months before ingestion, select a range wide enough to cover both, or the charts may appear empty or incomplete. ## Set a default baseline Set a project-wide default baseline to automatically compare all experiments against a specific reference run: 1. Go to [** Experiments**](https://www.braintrust.dev/app/~/experiments). 2. Select the checkbox next to the experiment you want to use as the baseline 3. Click **Set as default baseline** in the toolbar To clear the default baseline, select the same experiment and click **Clear default baseline**. The default baseline applies to all experiments in the project unless an experiment has its own baseline configured. For more information, see [Set a baseline](/docs/evaluate/compare-experiments#set-a-baseline). ## Speed up log filtering If you frequently filter on the same custom fields, you can index them to reduce query latency. Braintrust offers a full-text index for broad search, subfield indexes for specific fields you filter on most, and a metric columnstore for trace-level metrics. 1. Go to ** Settings** > [** Advanced**](https://www.braintrust.dev/app/~/configuration/advanced). 2. Turn on **Enable log search optimization** to build a full-text index that speeds up text-based filter queries. 3. Turn on **Enable shingled search optimization** to also index multi-word phrases, which speeds up phrase and multi-word `search()` queries. 4. Turn on **Enable metric columnstore** to write trace-level metrics in a columnar format, which speeds up the trace queries behind the ** Logs** page on high-volume projects. 5. Under **Subfield indexing**, click **+ Add subfield index** for each field you filter on frequently. Braintrust auto-discovers candidate fields from your data (e.g., `metadata.session_id`). If a field doesn't appear, you can type it in directly. Subfield paths must start with `input`, `output`, `expected`, `metadata`, or `span_attributes`. 6. Click **Save and index**. 7. Enter how many days back to backfill (default: 3) and click **Save and backfill**. The **Index status** section shows backfill progress as indexing runs in the background. Use [`search()`](/docs/reference/sql/query-structure#full-text-search) in SQL filters to query all text fields at once. It gets automatic bloom filter acceleration when log search optimization is on. ## Edit project details Update project name and description: 1. Go to ** Settings** > [** General**](https://www.braintrust.dev/app/~/configuration/general). 2. Modify name and description. The description field supports Markdown formatting. 3. Click **Save**. ## Delete a project Deleting a project permanently removes all logs, experiments, datasets, and functions. This cannot be undone. 1. Go to ** Settings** > [** General**](https://www.braintrust.dev/app/~/configuration/general). 2. Click **Delete project**. 3. Confirm by typing the project name. 4. Click **Delete**. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} bt projects delete my-project ``` ## Next steps * [Control access](/docs/admin/access-control) to projects with permission groups * [Set up automations](/docs/observe/alerts) for project-specific alerts * [View logs](/docs/observe/view-logs) filtered by tags and metadata * [Run evaluations](/docs/evaluate/run-evaluations) on project datasets * [Filter and search logs](/docs/observe/filter) with optimized indexes # SCIM provisioning Source: https://braintrust.dev/docs/admin/scim Provision organization members and permission groups from your identity provider, and keep Braintrust access in sync with it. SCIM (System for Cross-domain Identity Management) provisioning lets your identity provider (IdP) control who belongs to your Braintrust organization and, optionally, which permission groups they belong to. Instead of inviting and removing members by hand, you manage groups in your identity provider and Braintrust follows. in [private preview](/docs/feature-lifecycle), available to a limited set of customers. To request access, [contact Braintrust](https://braintrust.dev/contact). only available on the [Enterprise plan](/docs/plans-and-limits#plans). ## How SCIM provisioning works Braintrust derives access from the SCIM groups a user belongs to: * **Organization membership**: You map one SCIM group to your Braintrust organization. Users in that group are members of the organization, and users who leave it are removed. * **Permission groups**: You can optionally map each Braintrust permission group to its own SCIM group. A user is added to a permission group when they are in the mapped SCIM group and are a member of the organization. Braintrust only changes organizations and permission groups that you explicitly map. See [Unmapped organizations and permission groups](#unmapped-organizations-and-permission-groups). SCIM provisioning replaces [domain mappings](/docs/admin/authentication#domain-mappings) for an organization rather than working alongside them. Once SCIM manages an organization's membership, domain mappings no longer add users to it, so a user who signs in with a matching email domain joins only if your identity provider puts them in the mapped SCIM group. Only members of the **Owners** [permission group](/docs/admin/access-control) can view or change SCIM settings. ## Connect your identity provider Braintrust connects your identity provider to your organization for you. Contact [support@braintrust.dev](mailto:support@braintrust.dev) with the name or ID of each Braintrust organization you want to provision. A Braintrust support engineer sends you two values through a secure, short-lived link: * A **SCIM endpoint URL**. * A **bearer token**. The bearer token covers every Braintrust organization your identity provider manages. If you manage more than one Braintrust SSO application in your IdP, create a separate application dedicated to SCIM rather than enabling SCIM on one of them. ** Settings** > [** SCIM**](https://www.braintrust.dev/app/~/configuration/org/scim) appears under **Organization** for every Enterprise organization, and invites you to contact support until your identity provider is connected. ## Set up provisioning in your identity provider These steps enable SCIM on your existing Braintrust SSO application. If you created a dedicated SCIM application instead, apply the same provisioning settings to that application. ### Okta From the Okta admin dashboard, open your Braintrust SSO application and select the **General** tab. Click **Edit**, then enable **SCIM** under **Provisioning**. Open the **Provisioning** tab that now appears and enable the SCIM integration. Set **SCIM connector base URL** to the endpoint URL Braintrust sent you, set **Unique identifier field for users** to `userName`, and enter the bearer token in the **API Token** field. Enable **Push New Users**, **Push Profile Updates**, and **Push Groups**. Click **Test API Credentials** to verify the connection, then click **Save**. Under **To App**, enable **Create Users**, **Update User Attributes**, and **Deactivate Users**, then click **Save**. ### Microsoft Entra ID In the Microsoft Entra admin center, go to **Enterprise applications** and select your existing **Braintrust** application. Select **Provisioning**, then set **Provisioning Mode** to **Automatic**. Under **Admin Credentials**, set **Tenant URL** to the endpoint URL Braintrust sent you and **Secret Token** to the bearer token. Select **Test Connection**, then **Save**. Create a group containing a single test user and assign it to the Braintrust application under **Users and groups**. In the provisioning settings, set **Scope** to **Sync only assigned users and groups**, confirm under **Mappings** that both user and group provisioning are enabled, and confirm that Entra maps your chosen user identifier to the SCIM `userName` attribute. Use **Provision on demand** to test the user, then the test group. Once both succeed, select **Start provisioning** to begin ongoing synchronization. SCIM is now enabled, but no users reach Braintrust until you provision a group from your IdP and map it in your Braintrust SCIM settings. ### Provision a group to Braintrust SCIM groups appear in the Braintrust group pickers only after your IdP provisions them. For each group you plan to map: 1. Create or select a group in your IdP and add every user who should be covered by that mapping. 2. Make sure every member of the group is assigned to your Braintrust application in your IdP. 3. Provision the group. In Okta, add it under **Push Groups** and select **Push now**. In Microsoft Entra ID, use **Provision on demand**, or start provisioning and wait for the group to sync. Reload the Braintrust SCIM settings page and the group appears in the picker. Create these groups and populate them with the right members before you configure SCIM in Braintrust, so the whole setup can be completed in one sitting. In Okta, keep the group you assign to the application separate from the group you push. ## Configure Braintrust Go to ** Settings** > [** SCIM**](https://www.braintrust.dev/app/~/configuration/org/scim). Until you turn on **Enable SCIM updates**, the page guides you through three numbered steps: **Organization membership**, **Permission groups**, and **SCIM updates**. Braintrust ignores incoming events from your IdP and changes no memberships until you enable SCIM updates in the last step, so you can map groups and review the results first. **Permission groups** and **SCIM updates** appear once you select an organization membership group. Each step that changes who SCIM manages runs a [membership check](#check-members) against your identity provider before you can save it. ### Map organization membership Create one SCIM group per Braintrust organization you administer, even if you administer only one. To map it: 1. Under **Organization membership**, select the SCIM group whose members should belong to this Braintrust organization. Braintrust checks every current organization member against it. 2. Review the [check](#check-members), then click **Save**. The button is enabled once the check passes or you acknowledge the members who need attention. Changing the organization membership group makes the new group authoritative for this organization. Once **Enable SCIM updates** is on, anyone who isn't in the new SCIM group is removed from the organization and their API keys for that organization are deleted. Review everyone the check lists under **Needs attention** before you click **Change group**. To change the group: 1. Click **Change**. 2. In the **Change organization membership group** dialog, select the new SCIM group. Braintrust checks every current organization member against it. 3. Review the [check](#check-members), then click **Change group**. The button is enabled once the check passes or you acknowledge the results. To stop managing organization membership through SCIM: 1. Click **Change**. 2. Clear the group so that the picker shows **No group (stop managing membership)**. 3. Click **Change group**. This also turns off **Enable SCIM updates**, and no one is removed from the organization. If the organization membership SCIM group is deleted in your identity provider, the group field highlights in red and a warning banner at the top of the page lists **Organization membership**. Click **Change** to choose its replacement. SCIM groups are matched by ID, not name, so a new group with the same name doesn't replace the deleted one automatically. #### Check members A membership check compares Braintrust members against your identity provider and a SCIM group. A member is **Ready** only when they're active in your identity provider application, linked to the correct sign-in identity, in the identity provider application connected to this organization, and in the SCIM group. Every other member is listed under **Needs attention** with the reason, such as: * Deactivated in your identity provider application * Found in a different identity provider application * Not found in your identity provider application * Not in the SCIM group * Linked to a different sign-in identity or Braintrust account * Has API keys but hasn't signed in yet, so SCIM can't manage the account Braintrust runs the check before you save an organization membership group, add or replace a permission group mapping, or turn on SCIM updates. That step's save button is enabled when every member is **Ready**. To resolve members who need attention: 1. Review the reason listed for each member under **Needs attention**. 2. Resolve each issue in your identity provider. 3. Click **Refresh** to run the check again. To continue without resolving every issue, select the checkbox confirming that you've reviewed the members who need attention. Refreshing the check clears the checkbox. Membership checks cover up to 1,000 members. With more, the check reports "Too many members to check (limit 1,000)" and can't verify anyone, and **Check members...** doesn't appear next to **Change**. If a check can't verify members, for this or any other reason, you can continue only by selecting the checkbox confirming that some members could not be verified. You can also run the check at any time. Once a membership group is set, **Check members...** appears next to **Change** and checks every current organization member against the saved group. Each permission group mapping has its own check, described under [Map permission groups](#map-permission-groups). ### Map permission groups The **Permission groups** step is optional. Choose how members get their permission groups: add every newly provisioned member to a default group, or turn on **Manage permission groups via SCIM** and map each permission group to a SCIM group. If you map permission groups before you enable SCIM updates, the check that runs when you turn updates on includes them. #### Set a default permission group If you don't plan to manage permission groups through SCIM, use **Default group** to choose the permission group that newly provisioned members are added to. This setting is available when **Manage permission groups via SCIM** is off. #### Manage permission groups via SCIM To manage permission group membership from your identity provider: 1. Turn on **Manage permission groups via SCIM**. This toggle is available once an organization membership group is set. 2. Click **Add mapping**. 3. Select a Braintrust permission group and the SCIM group to pair it with. Braintrust checks the permission group's current members against the SCIM group. 4. Review the [check](#check-members), then click **Add mapping**. The button is enabled once the check passes or you acknowledge the results. Each permission group maps to one SCIM group. A user must also be in the organization membership group to be added to a mapped permission group. Adding or replacing a mapping makes the SCIM group authoritative for that permission group. Once **Enable SCIM updates** is on, anyone in the Braintrust group who isn't in the mapped SCIM group is removed from it. Review everyone the check lists under **Needs attention** before you save the mapping. If **Enable SCIM updates** is already on and you have saved mappings, turning on **Manage permission groups via SCIM** opens a confirmation dialog, because your saved mappings start applying immediately. Braintrust checks every saved mapping, and **Turn on** is enabled once the checks pass or you acknowledge the results. Each mapping row has these controls: * **Check members...** runs the same checks as the [organization check](#check-members), and also requires each permission group member to be in both the organization membership SCIM group and the mapped SCIM group. Members who pass show as **Ready**, and everyone else is listed under **Needs attention** with the reason. It can't check a permission group with more than 1,000 members. * **Remove** stops syncing that permission group after you confirm. Braintrust doesn't change who is in the group, and current members keep their access. #### Replace a deleted SCIM group If a mapped SCIM group is deleted in your identity provider, its row highlights in red, is marked **Deleted in your identity provider**, and shows the deleted group's ID. The warning banner at the top of the page also lists the mapping. To replace it: 1. Click **Replace** on the row. 2. Select the replacement SCIM group. If exactly one unmapped SCIM group has the same name as the deleted group, it's pre-selected. Confirm it's the replacement you created in your identity provider. 3. Review the [check](#check-members), then click **Replace mapping**. The button is enabled once the check passes or you acknowledge the results. ### Enable SCIM updates Review the membership of every mapped SCIM group and confirm that each one grants the access you intend. Then turn on **Enable SCIM updates**. This toggle is available once an organization membership group is set. Turning on the toggle opens the **Enable SCIM updates** dialog instead of saving immediately. Braintrust checks your organization members against the organization membership group and, if **Manage permission groups via SCIM** is on, the members of each mapped permission group against its SCIM group. Each check appears as its own row. Click **Details** on a row to see the members who need attention. **Enable SCIM updates** is enabled once every check passes or you acknowledge the results. Once SCIM updates are on, changes from your identity provider add and remove members of this organization and its mapped permission groups. Members removed from the organization lose their API keys for it. While the toggle is off, Braintrust ignores incoming events from your IdP. Turning it off takes effect immediately. ## What SCIM events change With **Enable SCIM updates** on, these actions in your IdP update Braintrust: * **A user is added to the organization membership group.** The user is provisioned into the Braintrust organization and added to any mapped permission groups they qualify for. * **A user is removed from the organization membership group.** The user is removed from the organization, and their API keys for that organization are deleted. * **A user is added to or removed from a mapped permission group's SCIM group.** The user is added to or removed from the corresponding Braintrust permission group. * **A user is deactivated in the IdP.** The user is removed from every Braintrust organization that IdP manages, and their API keys for those organizations are deleted. To force a full resynchronization between your IdP and Braintrust, provision the groups again. In Okta, select **Push now** under **Push Groups**. In Microsoft Entra ID, use **Provision on demand**, or let the scheduled provisioning cycle run. ## Manual membership changes Once **Enable SCIM updates** is on and an organization membership group is set, your identity provider becomes the source of truth for membership. Braintrust disables the controls that would let the two drift apart: * On ** Settings** > [** Members**](https://www.braintrust.dev/app/~/configuration/org/team), **Invite** and the remove control are disabled. * On ** Settings** > [** Permission groups**](https://www.braintrust.dev/app/~/configuration/org/groups), adding and removing members of a mapped group is disabled. * In the **Edit permission groups** dialog on a member's row, mapped groups can't be added or removed. Hovering a disabled control explains why: "Membership is managed by your identity provider (SSO)." Owners also see how to lift the lock. The lock applies to the API as well as the UI. `PATCH /v1/organization/members` is rejected with a 403 for a managed organization, and so is a `PATCH /v1/group/{group_id}` that changes a managed group's members through `add_member_users`, `remove_member_users`, `add_member_groups`, or `remove_member_groups`. Other fields on that endpoint, such as the group's name and description, are unaffected. Two exceptions keep working: * **Unmapped permission groups.** A permission group locks only when **Manage permission groups via SCIM** is on and that group is mapped to a SCIM group. Everything else stays manually editable. * **Service accounts.** Because they don't exist in your identity provider, you can add and remove service accounts by hand even while SCIM manages the organization. To make a manual change, turn off **Enable SCIM updates**, make the change, then turn it back on. Turning it back on runs the [membership check](#check-members) again, so a member you added by hand who isn't in the matching SCIM group appears under **Needs attention**. Braintrust reconciles each user against your identity provider when it next receives an event for them, so a manual change that conflicts with your SCIM groups doesn't persist. ## Unmapped organizations and permission groups Braintrust only adds and removes access for organizations and permission groups that are explicitly mapped on the SCIM settings page. Anything you haven't mapped is left alone: * If a user belongs to a Braintrust organization that has no organization membership group configured, deactivating that user in your IdP does not remove them from that organization. * If your organization contains permission groups with no SCIM mapping, SCIM events do not add or remove members from those groups. Braintrust also never removes access based on missing group data. If your IdP sends an event without group information, Braintrust makes no changes rather than treating the absent data as an empty group. ## Next steps * Set up [SSO](/docs/admin/authentication#single-sign-on-sso) so provisioned members can sign in with your identity provider. * Review [access control](/docs/admin/access-control) to decide which permission groups to map. * Learn how [API keys](/docs/admin/authentication#api-authentication) inherit their user's permissions. # Architecture Source: https://braintrust.dev/docs/admin/self-hosting/architecture Braintrust separates the **control plane** from the **data plane**. The control plane (UI, authentication, and platform management) runs as Braintrust-managed SaaS. The data plane, where your AI data actually lives, runs in Braintrust-managed infrastructure or in a cloud account you own, and is operated either by Braintrust or by your team, depending on your [deployment option](/docs/admin/deployment). This page describes the data plane as deployed for self-hosting, where it runs in your cloud account or region and your team operates it. Control plane and data plane architecture ## Control plane The control plane manages the web UI, authentication, user management, and metadata storage (project names, experiment names, organization settings). It lives in Braintrust's managed service and is delivered as SaaS. The control plane does not store or process your sensitive AI data — it communicates with your data plane only for authentication and metadata synchronization. ## Data plane The data plane stores all sensitive AI data — experiment logs, traces, spans, datasets, and prompt completions. When you use Braintrust's SDKs, they send data directly to your data plane, never touching Braintrust's servers. When you use the web UI, your browser communicates directly with your data plane via CORS. For hardware sizing requirements for all data plane components, see [Hardware requirements](/docs/admin/self-hosting#hardware-requirements). ### API service The API service is the entry point for all SDK and browser requests to the data plane, built in TypeScript/Node.js. On AWS, it runs on ECS ([Terraform module v6.0](/docs/admin/self-hosting/upgrade/v6) or later), split into three services by workload (`braintrust-api` for general traffic, `braintrust-api-ingest` for ingestion, and `braintrust-api-background` for evals, function invocation, and proxy traffic). Versions prior to v6.0 run on Lambda. On GCP and Azure, the API runs as Kubernetes containers that you manage and scale. ### PostgreSQL PostgreSQL stores metadata required to operate the platform, including pointers to raw data in object storage and aggregate statistics about the data. It is not the primary store for your AI data. Traces, spans, and logs live in Brainstore and object storage. ### Redis Redis provides caching and coordination for session management, rate limiting, and Brainstore transaction ID assignment, which ensures consistent ordering of concurrent writes. ### Object storage Object storage (S3 on AWS, GCS on GCP, Azure Blob on Azure) is the durable, long-term home for all AI data. Brainstore writes every ingested span to object storage as a write-ahead log (WAL) entry and compacts those entries into indexed segments that also live on object storage. Because object storage is the source of truth, Brainstore nodes are stateless and can be replaced without data loss. See [Trace data durability](/docs/instrument#trace-data-durability) for ingestion guarantees and retry behavior. ### Brainstore Brainstore is Braintrust's high-performance database for ingesting and querying AI data. It uses object storage and a streaming Rust engine to load spans in real time, cutting down on latency and enabling fast full-text search over large volumes of trace data. Each customer's data lives in its own partition, and Brainstore treats semi-structured fields as a first-class citizen rather than shredding them into columns — which breaks down at the payload sizes AI traces produce. Brainstore runs as three distinct node types: * **Writers** ingest incoming spans and traces and write them to object storage. * **Readers** serve ad-hoc queries, including those from the API and user-defined BTQL queries. * **Fast readers** serve predictable UI queries — paginated viewers, span and trace lookups — in isolation from standard reader nodes, keeping the UI responsive while resource-intensive queries run on readers. #### Write path Every write appends to a WAL on object storage, allowing high throughput without coordination bottlenecks. In the background, two steps convert the WAL into a queryable index: **Processing** assigns records to time-ordered segments (ensuring all spans for a trace land together), and **compacting** converts those segments into efficient indexed formats including inverted indexes, row stores, column stores, and bloom filters. Indexing is asynchronous and continuous — the system is always improving its read-optimized representation of the data. #### Read path Brainstore runs an explicit query pipeline: Parsing SQL into an AST, binding names to schemas and fields, optimizing by pushing down filters and pruning unneeded fields early, then executing against the index to stream results. Full-text search across prompts and responses is a first-class query path. Reads are not forced to wait for compaction — when you query, Brainstore merges data from the WAL, processed-but-not-yet-compacted entries, and fully indexed segments, keeping queries real-time as background compaction proceeds. ### Braintrust Gateway Optionally, the data plane can run the [Braintrust Gateway](/docs/deploy/gateway), which gives your applications a single OpenAI-compatible API to reach any LLM provider. Running it in the data plane keeps LLM requests and cached completions on your own infrastructure. Requests egress from your VPC directly to providers, and completions are cached in your Redis. Provider credentials follow the same split as other data: project-level provider secrets are stored in the data plane, and organization-level provider secrets are stored in the control plane. It is disabled by default. See [Deploy the Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) for setup. ## Where data is stored | Data | Location | | - | - | | Experiment records (input, output, expected, scores, metadata, traces, spans) | Data plane | | Log records (input, output, expected, scores, metadata, traces, spans) | Data plane | | Dataset records (input, output, metadata) | Data plane | | Prompt playground prompts | Data plane | | Prompt playground completions | Data plane | | Human review scores | Data plane | | Project-level LLM provider secrets (encrypted) | Data plane | | Org-level LLM provider secrets (encrypted) | Control plane | | API keys (hashed) | Control plane | | Experiment and dataset names | Control plane | | Project names | Control plane | | Project settings | Control plane | | Git metadata about experiments | Control plane | | Organization info (name, settings) | Control plane | | Login info (name, email, avatar URL) | Control plane | | Auth credentials | [Clerk](https://clerk.com/) | ## Next steps * [Deploy your data plane](/docs/admin/self-hosting/deploy) — Set up Braintrust in your cloud environment. * [Hardware requirements](/docs/admin/self-hosting#hardware-requirements) — Size your data plane components for production. * [Self-hosting overview](/docs/admin/self-hosting) — Understand shared responsibility, monitoring, and upgrades. # Configure your deployment Source: https://braintrust.dev/docs/admin/self-hosting/configure/index Tune a self-hosted Braintrust data plane after deployment with security, networking, scaling, and telemetry settings. After you [deploy](/docs/admin/self-hosting/deploy) a self-hosted data plane, use these settings to tune it for your environment. Most are optional, with defaults suitable for typical deployments, and are grouped by concern: * **[Security and access](/docs/admin/self-hosting/configure/security)**: Control which organizations can authenticate, secure customer data, validate outbound URLs, harden pods, and enable audit headers. * **[Networking and connectivity](/docs/admin/self-hosting/configure/networking)**: Configure client URLs, outbound request security, outbound and inbound request rate limits, load balancer timeouts, and the Braintrust Gateway. * **[Scaling and storage](/docs/admin/self-hosting/configure/scaling)**: Size the API and Brainstore compute for your workload. * **[Telemetry and data retention](/docs/admin/self-hosting/configure/telemetry)**: Control the telemetry your data plane sends to Braintrust and enable data retention. You can also turn on these optional services, both disabled by default: * **[Loop runtime](/docs/admin/self-hosting/configure/loop-runtime)**: Run [Loop](/docs/loop) threads in your own data plane on AWS, in isolated MicroVM sandboxes. * **[Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway)**: Serve [Gateway](/docs/deploy/gateway) requests from your own infrastructure instead of `gateway.braintrust.dev`. To change the infrastructure the module provisions instead, such as networking, encryption, resource tags, or cloud identity, see [Customize the infrastructure](/docs/admin/self-hosting/deploy#customize-the-infrastructure) on the Deploy page. Most settings are applied through Terraform (AWS) or Helm (GCP and Azure), and each notes whether it applies to AWS, GCP and Azure, or all clouds, along with the minimum module or chart version required. A few are client-side or API-level settings rather than deployment configuration, and are labeled as such. # Loop runtime Source: https://braintrust.dev/docs/admin/self-hosting/configure/loop-runtime Run Loop threads inside a self-hosted AWS data plane with an opt-in ECS service that executes each thread in an isolated MicroVM sandbox. The Loop runtime is an opt-in ECS Fargate service that runs [** Loop**](/docs/loop) threads in your own data plane, using isolated AWS Lambda MicroVM sandboxes for command execution and investigation tasks. It is disabled by default, so Loop, the Debugger, and Patterns are unavailable until you enable it. The Loop runtime requires AWS Terraform module v6.5.2 or later and a [supported AWS region](#supported-aws-regions). It is not available on GCP or Azure. ## What the runtime provides Loop does its investigation work in a sandbox, an isolated environment where it reads, searches, and lists trace files directly and runs commands against your data. That is what lets it work through a trace too large to read in one pass, and keep going on a long investigation instead of being bound to a browser session. Threads persist, so users can leave and resume them. The Loop runtime and its isolated sandboxes also support the [Debugger](/docs/observe/debug-traces), which investigates individual traces, and [Loop automations](/docs/loop/automations), which perform recurring work on a schedule. [** Patterns**](/docs/observe/patterns) runs as a scheduled Loop automation, so it requires the runtime too. See What Loop can do for the full range. On Braintrust-hosted deployments, Braintrust provides the sandbox. Self-hosted deployments provide it themselves by enabling this service. ## Supported AWS regions The Loop runtime is not available in every [region supported by the Braintrust data plane](/docs/admin/self-hosting/deploy#1-configure-the-terraform-module). For both self-hosted and [BYOC](/docs/admin/deployment/byoc) deployments on AWS, it requires regional support for the `AWS::Lambda::MicrovmImage` CloudFormation resource. The data plane regions that lack this support are `ca-central-1`, `eu-west-2`, `eu-west-3`, and `sa-east-1`. Enabling the runtime in one of these regions can fail with `Template format error: Unrecognized resource types: [AWS::Lambda::MicrovmImage]`. Leave `enable_loop_runtime` set to `false` there and [contact Braintrust](mailto:support@braintrust.dev) for deployment guidance. AWS [continues to add MicroVM regions](https://aws.amazon.com/about-aws/whats-new/2026/08/lambda-microvms-5-additional-regions/), so confirm coverage for your region with Braintrust. ## Enable the runtime Set two variables in your data plane module and apply. One deploys the ECS service and its sandbox, and the other decides whether those sandboxes can reach the network. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" enable_loop_runtime = true loop_runtime_sandbox_egress_mode = "restricted" # Recommended. No outbound network access from sandboxes. # ... other configuration ... } ``` **`enable_loop_runtime`** deploys the ECS service and its MicroVM sandbox. It is `false` by default. It doesn't depend on `enable_ecs_api`, which controls only whether CloudFront routes API traffic to ECS. The runtime works either way, because its URL reaches both the ECS API tasks and the Lambda API handlers. **`loop_runtime_sandbox_egress_mode`** controls whether sandbox MicroVMs can reach the network. Braintrust recommends `"restricted"`, which is the module default and how Braintrust runs its own deployment. * **`"restricted"`**: The module creates a dedicated VPC for sandboxes, with a security group that has no egress rules and a Route 53 Resolver DNS firewall that blocks every domain. Sandboxes have no outbound network access and can't resolve names. Two constraints come with it: the VPC uses the `10.255.0.0/16` CIDR block, which you can't configure, and you can't supply an existing VPC for it, including your quarantine VPC, which has no DNS firewall. * **`"internet"`**: Sandboxes use an AWS-managed internet egress connector and can reach the public internet. Any value other than exactly `"internet"` is treated as restricted. On module v6.7.0 and later, the default is `"restricted"`. Modules v6.6.0 and earlier default to `"internet"`, so set the variable explicitly before upgrading if your sandboxes need outbound internet access. On module v6.5.2 and later, the runtime sends Loop's LLM calls through your deployment's own AI proxy. Earlier versions routed them to `gateway.braintrust.dev` once `enable_ecs_api` was set. Upgrade to v6.5.2 or later before enabling the runtime if your deployment must keep inference traffic inside your network. See [Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway). The Loop runtime image isn't pinned to the module version. The module tracks the latest 2.x release, so a new runtime image can roll out to your deployment without a Terraform change. Set `loop_runtime_version_override` to pin an exact tag. ## Enable Patterns Applying the runtime makes Loop and the Debugger available across your organization, but it doesn't start Patterns. Each project turns Patterns on separately, and the person who does it needs permission to create project automations. See [Enable Patterns](/docs/observe/patterns/enable). ## Telemetry On module v6.7.0 and later, the Loop runtime sends `metrics` and `traces` for its own service to Braintrust's control plane, in addition to whatever [telemetry](/docs/admin/self-hosting/configure/telemetry) types your deployment configures. Braintrust uses them to diagnose runtime problems while the service stabilizes, and you can't turn them off while the runtime is enabled. These traces describe the runtime's operation and can include model and tool identifiers, timing information, and errors. If your deployment can't send this telemetry, leave the Loop runtime disabled and [contact Braintrust](mailto:support@braintrust.dev). ## Configuration reference | Variable | Default | Description | | - | - | - | | `enable_loop_runtime` | `false` | Deploy the Loop runtime ECS service and its MicroVM sandbox. | | `loop_runtime_version_override` | `null` | Pin the Loop runtime image and MicroVM guest artifact to a specific version tag. When unset, the module tracks the latest 2.x release. | | `loop_runtime_task_cpu` | `2048` | CPU units for each Loop runtime task. 1024 CPU units equal 1 vCPU. | | `loop_runtime_task_memory` | `8192` | Memory (MiB) for each Loop runtime task. | | `loop_runtime_ephemeral_storage_gib` | `null` | Task ephemeral storage in GiB (21 to 200). `null` uses the Fargate default of 20. | | `loop_runtime_min_capacity` | `1` | Minimum number of running Loop runtime tasks. | | `loop_runtime_max_capacity` | `4` | Maximum number of running Loop runtime tasks. | | `loop_runtime_target_cpu_utilization` | `40` | Target average CPU utilization percentage for autoscaling. | | `loop_runtime_target_memory_utilization` | `50` | Target average memory utilization percentage for autoscaling. | | `loop_runtime_log_retention_days` | `14` | CloudWatch log retention in days. Must be a valid CloudWatch Logs retention value. | | `loop_runtime_enable_execute_command` | `false` | Enable ECS Exec on the Loop runtime service. | | `loop_runtime_alb_deregistration_delay` | `900` | Deregistration delay in seconds for the Loop runtime ALB target group. | | `loop_runtime_extra_env_vars` | `{}` | Extra environment variables merged into the Loop runtime container. | | `loop_runtime_sandbox_egress_mode` | `"restricted"` | Outbound network mode for sandbox MicroVMs. `"restricted"` blocks all outbound network access. Exactly `"internet"` uses AWS-managed internet egress. Any value other than `"internet"` is treated as restricted. | | `loop_runtime_microvm_minimum_memory_mib` | `2048` | Minimum memory (MiB) provisioned for each sandbox MicroVM. | | `loop_runtime_microvm_max_idle_duration_seconds` | `900` | Seconds without traffic before a sandbox MicroVM auto-suspends. | | `loop_runtime_microvm_suspended_duration_seconds` | `28800` | Seconds a suspended MicroVM remains resumable before termination. | | `loop_runtime_microvm_maximum_duration_seconds` | `28800` | Maximum MicroVM lifetime across running and suspended states. | | `loop_runtime_microvm_auth_token_expiration_minutes` | `30` | Endpoint auth token lifetime in minutes for MicroVM invocations. | | `enable_loop_runtime_microvm_runtime_logs` | `false` | Export MicroVM stdout and stderr to CloudWatch. Output can include sandbox contents. | On module v6.8.1 and later, the runtime serves the organization named by `braintrust_org_name`. Earlier versions used a separate `loop_runtime_org_name` variable. Remove it from your configuration before you upgrade, or the plan fails with `An argument named "loop_runtime_org_name" is not expected here`. ## Next steps * Learn [what Loop can do](/docs/loop/capabilities) once the runtime is running. * Review [thread and sandbox limits](/docs/loop/manage#sandboxes) that apply to your users. * Configure [networking and connectivity](/docs/admin/self-hosting/configure/networking) for the rest of your deployment. # Networking and connectivity Source: https://braintrust.dev/docs/admin/self-hosting/configure/networking Configure URLs, rate limits, load balancer settings, network access paths, and the Gateway for a self-hosted Braintrust data plane. Use these settings to change where SDKs connect, tune incoming request behavior, control which external and private services the data plane can reach, and optionally run the Braintrust Gateway in your own infrastructure. ## Client routing ### Customize the webapp URL The SDKs guide users to `https://www.braintrust.dev` (or the `BRAINTRUST_APP_URL` variable) to view their experiments. In some advanced configurations, you can reverse proxy traffic to the `BRAINTRUST_APP_URL` from the SDKs while pointing users to a different URL. To do this, you can set the `BRAINTRUST_APP_PUBLIC_URL` environment variable to the URL of your webapp. By default, this variable is set to the value of `BRAINTRUST_APP_URL`, but you can customize it as you wish. This variable is *only* used to display information, so even its destination does not need to be accessible from the SDK. Set it through the `braintrust_api_extra_env_vars` passthrough map (Terraform module v6.0.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_extra_env_vars = { BRAINTRUST_APP_PUBLIC_URL = "https://braintrust.example.com" } ``` Deployments still served by the API Lambda (before module v6.0.0, or v6 with `enable_ecs_api = false`) set the same variable through `service_extra_env_vars.APIHandler` instead. Add it to `api.extraEnvVars` in your `values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: BRAINTRUST_APP_PUBLIC_URL value: "https://braintrust.example.com" ``` ### Constrain SDKs to the data plane If you're self-hosting the data plane, you can also constrain the SDKs to only communicate with your data plane. Normally, they communicate with the control plane to: * Get your data plane's URL * Register and retrieve metadata (e.g. about experiments) * Print URLs to the webapp The data plane can proxy the endpoints that the SDKs use to communicate with the control plane, allowing your SDKs to only communicate with the data plane directly. Set the `BRAINTRUST_APP_URL` environment variable to the URL of your data plane and `BRAINTRUST_APP_PUBLIC_URL` to `https://www.braintrust.dev` (or the URL of your webapp). ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_extra_env_vars = { BRAINTRUST_APP_URL = "https://dataplane.example.com" BRAINTRUST_APP_PUBLIC_URL = "https://www.braintrust.dev" } ``` ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: BRAINTRUST_APP_URL value: "https://dataplane.example.com" - name: BRAINTRUST_APP_PUBLIC_URL value: "https://www.braintrust.dev" ``` ## Inbound traffic ### Set HTTPS on the API load balancer On AWS with the ECS API, an internal Application Load Balancer (ALB) fronts the API services. CloudFront terminates TLS for inbound traffic and reaches this ALB over a private VPC origin, so the ALB does not need its own certificate in a standard deployment. By default, the ALB serves plain HTTP on port 80 using its AWS-assigned DNS name. To serve HTTPS on a custom domain instead, set both `braintrust_api_alb_certificate_arn` and `braintrust_api_alb_custom_domain` (available in Terraform module v6.0.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_alb_certificate_arn = "arn:aws:acm:us-east-1:123456789012:certificate/abc123" braintrust_api_alb_custom_domain = "braintrust.internal.example.com" ``` When both are set, the ALB serves HTTPS on port 443, plain HTTP is disabled, and all API URLs use `https://`. The certificate must cover the custom domain, and the domain must resolve to the ALB. These two variables must both be set or both be null. Setting only one fails at plan time. Enabling HTTPS after the ECS API exists will cause disruption: AWS will not change the protocol of a CloudFront VPC origin that is attached to a distribution, and rolling back `enable_ecs_api` does not recover it. If you enable it, you must do so during the initial apply of the [v6 upgrade](/docs/admin/self-hosting/upgrade/v6). [Contact Braintrust](mailto:support@braintrust.dev) before enabling HTTPS on the ALB in an existing deployment. ### Set the HTTP keep-alive timeout When the API server runs behind a load balancer, you may need to configure the HTTP keep-alive timeout to prevent connection resets. Load balancers typically have an idle timeout for connections, and if the API server's keep-alive timeout is shorter than the load balancer's timeout, the API server closes the connection while the load balancer still considers it open. When the load balancer tries to reuse that backend connection, it encounters a closed socket, resulting in connection reset errors and 502 responses. The API server exposes the following environment variable to configure the keep-alive timeout: * `TS_API_KEEP_ALIVE_TIMEOUT_SECONDS`: The HTTP keep-alive timeout in seconds. Default: `65` The default value of 65 seconds is designed to work with most load balancers, including AWS Application Load Balancer (which has a default idle timeout of 60 seconds). However, if your load balancer has a longer idle timeout, you should set this value to match or exceed your load balancer's timeout. For example, to match an AWS ALB configured with a 300-second idle timeout: Set it through the `braintrust_api_extra_env_vars` passthrough map (Terraform module v6.0.0 or later). This applies to the ECS API services: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_extra_env_vars = { TS_API_KEEP_ALIVE_TIMEOUT_SECONDS = "300" } ``` Add it to `api.extraEnvVars` in your `values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: TS_API_KEEP_ALIVE_TIMEOUT_SECONDS value: "300" ``` ### Set the CloudFront origin timeout On AWS, requests are served through CloudFront, which closes a connection and returns `504 Gateway Timeout` if the origin takes too long to respond. Long-running scorers or tools invoked through `/function/invoke` can exceed the default 60-second origin read timeout. Raise it with the `cloudfront_origin_read_timeout` Terraform variable (available in Terraform module v5.3.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} cloudfront_origin_read_timeout = 120 ``` The value must be between 1 and 180 seconds. CloudFront caps the origin read timeout at 180 seconds, and values above 60 seconds can require an AWS Support request to raise the account quota. ### Set the CloudFront minimum TLS version On AWS deployments that supply a custom certificate through `custom_certificate_arn`, CloudFront negotiates a minimum TLS version of `TLSv1.3_2025` with viewers. To support clients that cannot negotiate TLS 1.3, lower the floor with the `cloudfront_minimum_protocol_version` Terraform variable (available in Terraform module v6.5.2 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} cloudfront_minimum_protocol_version = "TLSv1.2_2021" ``` Accepted values are `TLSv1`, `TLSv1_2016`, `TLSv1.1_2016`, `TLSv1.2_2018`, `TLSv1.2_2019`, `TLSv1.2_2021`, and `TLSv1.3_2025`. The setting applies to the whole distribution. Leave it unset to keep the `TLSv1.3_2025` default. Set `cloudfront_minimum_protocol_version` only alongside `custom_certificate_arn`. Terraform rejects the configuration otherwise, because CloudFront cannot set a minimum protocol version on the default certificate. ### Set inbound request rate limits The API server can rate-limit log ingestion, SQL queries, and function invocation. Configure each surface separately with its own environment variables for limits, window length, and enforcement. All three surfaces behave the same way in these respects: * **Windows**: Each limit uses a fixed window that starts when the first matching request is counted, not on a clock boundary. The counter resets after the configured number of seconds, and the next matching request starts a new window. Rejected requests still count toward the limit. * **Enforcement**: With enforcement disabled, requests over a limit are allowed and the API server logs a warning. With enforcement enabled, they fail with HTTP 429 and a `Retry-After` header, and the response body reports the configured limit, the window length, and the requests consumed. A limit of `0` is a real zero-request limit, so with enforcement enabled every matching request is rejected. * **Replicas**: Each limit applies across all API server replicas combined. * **Restarts**: Rate limit configuration is read once at process start. Restart or redeploy the API services after changing any of these variables. None of these variables has a dedicated Terraform variable or Helm value, so pass them through your deployment's environment variable map, as shown in the examples below. Variables that take `=` pairs accept a comma-separated list of pairs. #### Limit log ingestion Log ingestion limits apply per organization and per project, and both are disabled by default. A project limit replaces the organization limit rather than adding to it, so a project with its own entry ignores the organization limit entirely. With no limit configured, ingestion is uncapped. **Window and enforcement** * `RATELIMIT_API_LOGS_ORG_WINDOW_SECS`: Window length in seconds. Default `60`. Despite the name, this also sets the window for project-based limits. * `RATELIMIT_API_LOGS_ORG_ENFORCE`: Return HTTP 429 when a limit is exceeded. Default `false` (log a warning and allow the request). Despite the name, this also governs the enforcement of project-based limits. **Organization limits** * `RATELIMIT_API_LOGS_ORG`: Per-organization limits, as `=` pairs. Find an organization's ID in the [organization switcher](/docs/admin/organizations#find-your-organization-id). **Project limits** Project-scoped limits require data plane v2.2.1 or later. * `RATELIMIT_API_LOGS_PROJECT`: Per-project limits, as `=` pairs. Find a project's ID under ** Settings** > [** General**](https://www.braintrust.dev/app/~/configuration/general). * `RATELIMIT_API_LOGS_PROJECT_DEFAULT`: Limit for every project without an entry in `RATELIMIT_API_LOGS_PROJECT`. Set it only if you want every project capped. Set these variables through the `braintrust_api_extra_env_vars` passthrough map (Terraform module v6.0.0 or later). For example: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_extra_env_vars = { RATELIMIT_API_LOGS_ORG = "=1000" RATELIMIT_API_LOGS_PROJECT = "=500" RATELIMIT_API_LOGS_ORG_WINDOW_SECS = "60" RATELIMIT_API_LOGS_ORG_ENFORCE = "true" } ``` Deployments still served by the API Lambda (before module v6.0.0, or v6 with `enable_ecs_api = false`) set the same variables through `service_extra_env_vars.APIHandler` instead. Add these variables to `api.extraEnvVars` in the `values.yaml` you pass to the Helm chart. For example: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: RATELIMIT_API_LOGS_ORG value: "=1000" - name: RATELIMIT_API_LOGS_PROJECT value: "=500" - name: RATELIMIT_API_LOGS_ORG_WINDOW_SECS value: "60" - name: RATELIMIT_API_LOGS_ORG_ENFORCE value: "true" ``` Setting a log ingestion limit to `0` with enforcement on rejects every matching logging request and stores nothing. Any spans the client does not retry are lost. A request counts as one unit against whichever limit applies, regardless of how many spans it carries. If any organization or project referenced in the request exceeds its limit, the entire payload is rejected. #### Limit SQL queries SQL query limits apply per organization and per project, and both are disabled by default. A project limit adds to the organization limit rather than replacing it, so it can only tighten the effective limit. Organizations without their own entry fall back to `RATELIMIT_BTQL_DEFAULT`, and with no default configured, queries are uncapped. **Window and enforcement** * `RATELIMIT_BTQL_WINDOW_SECS`: Window length in seconds. Default `60`. * `RATELIMIT_BTQL_ENFORCE`: Return HTTP 429 when a limit is exceeded. Default `false` (log a warning and allow the query). **Organization limits** * `RATELIMIT_BTQL_ORG`: Per-organization limits, as `=` pairs. Find an organization's ID in the [organization switcher](/docs/admin/organizations#find-your-organization-id). * `RATELIMIT_BTQL_DEFAULT`: Limit for every organization without an entry in `RATELIMIT_BTQL_ORG`. * `RATELIMIT_BTQL_DEFAULT_FUNCTIONS`: Separate default for queries against prompts and functions. Defaults to 20 times `RATELIMIT_BTQL_DEFAULT`, including when set to `-1`. **Project limits** Project-scoped limits require data plane v2.2.1 or later. * `RATELIMIT_BTQL_PROJECT`: Per-project limits, as `=` pairs. Setting one above the organization limit has no effect. With no organization limit configured, the project limit is the only one that applies. Find a project's ID under ** Settings** > [** General**](https://www.braintrust.dev/app/~/configuration/general). Set these variables through the `braintrust_api_extra_env_vars` passthrough map (Terraform module v6.0.0 or later). For example: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_extra_env_vars = { RATELIMIT_BTQL_ORG = "=1000" RATELIMIT_BTQL_PROJECT = "=100" RATELIMIT_BTQL_WINDOW_SECS = "60" RATELIMIT_BTQL_ENFORCE = "true" } ``` Deployments still served by the API Lambda (before module v6.0.0, or v6 with `enable_ecs_api = false`) set the same variables through `service_extra_env_vars.APIHandler` instead. Add these variables to `api.extraEnvVars` in the `values.yaml` you pass to the Helm chart. For example: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: RATELIMIT_BTQL_ORG value: "=1000" - name: RATELIMIT_BTQL_PROJECT value: "=100" - name: RATELIMIT_BTQL_WINDOW_SECS value: "60" - name: RATELIMIT_BTQL_ENFORCE value: "true" ``` A query counts once against the organization limit and once against each project it reads, so a query spanning several projects must stay under every limit it touches. These limits apply to queries from the API, the SDKs, and the MCP server. Queries issued from the Braintrust UI are exempt. #### Limit function invocation Invocation limits cover a project's prompts, scorers, tools, and other custom code functions. They are project-scoped, and there is no organization-scoped limit. Two independent limits apply: a per-project limit that is disabled by default, and a per-function cap that is on by default with a fixed 10-second window and always returns HTTP 429. **Window and enforcement** * `RATELIMIT_INVOKE_WINDOW_SECS`: Window length in seconds for the project limits. Default `10`. It does not affect the per-function cap. * `RATELIMIT_INVOKE_ENFORCE`: Return HTTP 429 when a project limit is exceeded. Default `false` (log a warning and allow the invocation). It does not affect the per-function cap. **Project limits** Project-scoped limits require data plane v2.2.1 or later. * `RATELIMIT_INVOKE_PROJECT`: Per-project limits, as `=` pairs. The count covers every function in the project, across all API keys. Find a project's ID under ** Settings** > [** General**](https://www.braintrust.dev/app/~/configuration/general). * `RATELIMIT_INVOKE_PROJECT_DEFAULT`: Limit for every project without an entry in `RATELIMIT_INVOKE_PROJECT`. Set it only if you want every project capped. **Per-function limit** * `INVOKE_RATE_LIMIT_PER_10S`: Maximum invocations per function, per API key, in a 10-second window. Default `10000`. The count is per function, not per project. * `ENABLE_INVOKE_RATE_LIMIT`: Whether invocation rate limiting runs at all. Default `true`. Setting it to `false` turns off the per-function cap and the project limits. Set these variables through the `braintrust_api_extra_env_vars` passthrough map (Terraform module v6.0.0 or later). For example: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_extra_env_vars = { RATELIMIT_INVOKE_PROJECT = "=100" RATELIMIT_INVOKE_WINDOW_SECS = "10" RATELIMIT_INVOKE_ENFORCE = "true" INVOKE_RATE_LIMIT_PER_10S = "5000" } ``` Deployments still served by the API Lambda (before module v6.0.0, or v6 with `enable_ecs_api = false`) set the same variables through `service_extra_env_vars.APIHandler` instead. Add these variables to `api.extraEnvVars` in the `values.yaml` you pass to the Helm chart. For example: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: RATELIMIT_INVOKE_PROJECT value: "=100" - name: RATELIMIT_INVOKE_WINDOW_SECS value: "10" - name: RATELIMIT_INVOKE_ENFORCE value: "true" - name: INVOKE_RATE_LIMIT_PER_10S value: "5000" ``` Preprocessors are exempt from invocation rate limits. ## Outbound traffic ### Secure outbound requests The data plane makes outbound requests both to Braintrust and to URLs you or your users supply, such as webhooks, remote scorers, and integrations. **Allow traffic to Braintrust through your firewall** If you restrict outbound network traffic, allow the data plane to reach Braintrust at: ``` www.braintrust.dev braintrust.dev gateway.braintrust.dev ``` `gateway.braintrust.dev` is the Braintrust-hosted Gateway. Two kinds of request go there: * **Built-in model requests.** Braintrust serves [built-in models](/docs/admin/ai-providers#manage-built-in-models) itself, so requests for them always go to the hosted Gateway. Self-hosted organizations have built-in models off by default and opt in explicitly, and [Topics](/docs/observe/topics) requires them. * **LLM calls from user-authored code**, such as custom scorers and tools, on AWS deployments that run the API on [ECS](/docs/admin/self-hosting/upgrade/v6) (`enable_ecs_api = true`) with Terraform module v6.4.0 or earlier. Module v6.5.0 and later sends these to the deployment's own AI proxy instead, so that traffic stays in your AWS account. This is firewall guidance. The data plane does not enforce a destination allowlist itself. **Block requests to internal addresses** To stop user-supplied URLs from reaching private or reserved IP addresses (server-side request forgery), configure URL validation. See [Configure URL security](/docs/admin/self-hosting/configure/security#configure-url-security). **Trust a private certificate authority** If the internal services your custom scorers and tools call present certificates signed by a private or enterprise certificate authority, supply the CA bundle so those requests validate. See [Configure a custom CA bundle](/docs/admin/self-hosting/configure/security#configure-a-custom-ca-bundle). ### Set outbound request rate limits The Braintrust API server can rate-limit the outbound requests it makes to external domains, such as `BRAINTRUST_APP_URL`. Rate limiting prevents unintentionally overloading an external domain, which might otherwise block the API server's IP in response. It is disabled by default. When enabled, requests are counted per API auth token per destination domain within a rolling window. * `OUTBOUND_RATE_LIMIT_MAX_REQUESTS`: The maximum number of requests per window. Default `0`, which disables rate limiting. Set a value greater than `0` to enable it. * `OUTBOUND_RATE_LIMIT_WINDOW_MINUTES`: The window length in minutes before the count resets. Default `1`. Use the dedicated variables (Terraform module v1.0.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} outbound_rate_limit_max_requests = 100 outbound_rate_limit_window_minutes = 1 ``` Set the environment variables through `api.extraEnvVars`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: extraEnvVars: - name: OUTBOUND_RATE_LIMIT_MAX_REQUESTS value: "100" - name: OUTBOUND_RATE_LIMIT_WINDOW_MINUTES value: "1" ``` ### Connect to internal resources over VPC On AWS, to connect Braintrust's VPC to other internal resources (like an LLM Gateway), use one of the following approaches: * Create a VPC Endpoint Service for your internal resource, then create a VPC Interface Endpoint inside the Braintrust "Quarantine" VPC. * Set up VPC peering with the Braintrust "Quarantine" VPC. ### AWS service VPC endpoints The module creates AWS service VPC endpoints in VPCs it manages. Interface endpoints span all three private subnets and share a security group that allows HTTPS (TCP 443) from the VPC CIDR. | Service | Type | Created when | | - | - | - | | S3 | Gateway | Always, in both main and quarantine VPCs | | SSM | Interface | Main VPC with `enable_brainstore_ec2_ssm = true` (default `false`) | | Secrets Manager | Interface | Main VPC with `create_secrets_manager_vpc_endpoint = true` (default `true`) | The Secrets Manager endpoint (available in Terraform module v6.8.1 or later) is enabled by default when the module manages the main VPC (`create_vpc = true`). Private DNS redirects the standard regional Secrets Manager hostname through the endpoint, so existing SDK and CLI calls use it without URL changes. Upgrading an existing module-managed VPC adds this endpoint. Interface endpoint charges apply. To opt out: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} create_secrets_manager_vpc_endpoint = false ``` The variable has no effect when `create_vpc = false`. The module does not create a Secrets Manager endpoint in the quarantine VPC or in VPCs you supply. ## Braintrust Gateway The [Braintrust Gateway](/docs/deploy/gateway) gives your applications a single OpenAI-compatible API for reaching any supported model provider (OpenAI, Anthropic, Google, AWS, and others), using provider keys you manage centrally in Braintrust. At request time it adds completions caching, provider failover, streaming, rate limiting, usage tracking, and trace logging. Braintrust hosts a Gateway at `gateway.braintrust.dev`. Self-hosted deployments can optionally run their own instance in the data plane, alongside the API and Brainstore, so Gateway requests are served from your own infrastructure. in [public preview](/docs/feature-lifecycle) and can change before reaching general availability. ### How the Gateway works Running the Gateway yourself keeps LLM traffic on your own infrastructure: requests egress directly from your VPC to providers, completions are cached (encrypted) in the data plane's Redis, and internal features like [Topics](/docs/observe/topics) route through it too. When the Gateway is disabled, the data plane does not fall back to the Braintrust-hosted Gateway. The API service calls providers directly. On AWS, user-authored code such as custom scorers and tools runs in an isolated quarantine environment that cannot reach your main VPC directly. Terraform module v6.5.0 and later sends its LLM calls to the deployment's own AI proxy, which forwards them to the Gateway once `enable_ai_gateway` is set, so those calls stay in your AWS account. The [Loop runtime](/docs/admin/self-hosting/configure/loop-runtime) does the same on module v6.5.2 and later. Even with the Gateway enabled, two paths still leave your data plane: * **Built-in models** are served by Braintrust, so requests for them reach the Braintrust-hosted Gateway at `gateway.braintrust.dev` rather than your own. Self-hosted organizations have [built-in models](/docs/admin/ai-providers#manage-built-in-models) disabled by default and opt in explicitly. * **Operational telemetry** (status, metrics, usage, and optional logs and traces) can still be sent to Braintrust's control plane, the same as the rest of the data plane. See [Telemetry and data retention](/docs/admin/self-hosting/configure/telemetry). This is monitoring traffic, not LLM traffic. ### Enable the Gateway The Gateway is disabled by default. It reuses the Redis instance and Brainstore license key the data plane already provisions, so no additional secrets are required. For Terraform module versions before v6.7.2 and Helm chart versions before v6.17.0, before you deploy the Gateway, [contact Braintrust](mailto:support@braintrust.dev) to enable it for your organization. Newer module versions do not need this enablement. Requires Terraform module v6.5.0 or later, which pins the Gateway image to the data plane version the module ships and keeps quarantine LLM calls in your account. The Gateway runs as an ECS Fargate service behind an internal load balancer. Three variables control the Gateway, so you can provision the infrastructure before routing traffic to it: * `create_ai_gateway` creates the private Gateway infrastructure (an internal ALB and the Gateway ECS service). * `enable_ai_gateway` sets `GATEWAY_URL` on the APIHandler, AI Proxy, and API ECS service so internal data plane traffic routes through the Gateway. It requires `create_ai_gateway`. * `use_private_ai_gateway_origin` points the public CloudFront `/v1/proxy` endpoint at the private Gateway (a VPC origin on its internal load balancer). It requires `create_ai_gateway`. **Existing deployment** On an existing deployment, roll the variables out in separate applies so each change is verifiable before the next: 1. Provision the infrastructure without changing traffic and apply: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} create_ai_gateway = true enable_ai_gateway = false ``` 2. Route internal traffic through the Gateway. Once the Gateway is healthy, set `enable_ai_gateway = true` and apply: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} enable_ai_gateway = true ``` Public `/v1/proxy` traffic still uses its existing origin. 3. Cut over the public endpoint and apply: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} use_private_ai_gateway_origin = true ``` CloudFront then routes `/v1/proxy` and `/v1/proxy/*` to the private Gateway. **New deployment** On a new deployment, you can set `create_ai_gateway` and `enable_ai_gateway` together in the first apply, then cut over the public endpoint with `use_private_ai_gateway_origin` once the Gateway is healthy. Configure task sizing and autoscaling with `ai_gateway_cpu`, `ai_gateway_memory`, `ai_gateway_min_capacity`, and `ai_gateway_max_capacity`. To authorize additional security groups to reach the internal ALB, set `ai_gateway_authorized_security_groups`. The API and Brainstore security groups are authorized automatically. The module pins the Gateway image, the same way it manages the API and Brainstore images, so upgrading the module is how you move to a newer Gateway. Leave `ai_gateway_version_override` unset unless Braintrust instructs you to pin a specific tag. **Route quarantine LLM calls through the private Gateway** By default, LLM calls from user-authored code in the quarantine environment reach the Gateway through the AI proxy Lambda function URL. On Terraform module v6.7.0 and later, you can instead route them to the private Gateway load balancer over AWS PrivateLink, so the traffic stays inside your VPCs: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} use_private_gateway_quarantine_proxy = true ``` The module creates a network load balancer in front of the Gateway load balancer, publishes it as a VPC endpoint service, and creates the matching interface endpoint in the quarantine VPC. It then sets the quarantine proxy URL to that endpoint. This requires `create_ai_gateway`, and it requires that the module manage both VPCs (`create_vpc` is `true` and the quarantine VPC is module-created). If the module can't create the PrivateLink path and you haven't set `quarantine_proxy_url`, the apply fails rather than silently falling back. If you supply either VPC, the module creates no PrivateLink resources and the `quarantine_gateway_privatelink_service_name` output is null. Stand up the network load balancer, endpoint service, and quarantine interface endpoint yourself, then point the deployment at your endpoint with `quarantine_proxy_url`: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} quarantine_proxy_url = "http://vpce-0abc123.vpce-svc-0def456.us-east-1.vpce.amazonaws.com/v1/proxy" ``` `quarantine_proxy_url` always takes precedence over the automatically selected URL. The `quarantine_proxy_url` output reports the value the deployment ended up using. When `use_global_ai_gateway_origin` is `true`, `use_private_gateway_quarantine_proxy` has no effect and no PrivateLink resources are created. This option creates a network load balancer, which is billed separately. On a deployment where the module manages both VPCs, setting `use_private_gateway_quarantine_proxy` to `true` creates that load balancer and the rest of the PrivateLink path even when `quarantine_proxy_url` overrides the URL the deployment uses, leaving you paying for resources nothing routes through. Set one or the other, not both. When `use_private_gateway_quarantine_proxy` is `false`, quarantine LLM calls reach the Gateway through the AI proxy Lambda function URL. To send them somewhere else, such as the Braintrust-hosted Gateway, set `quarantine_proxy_url`. Requires Helm chart 6.6.0 or later. The chart pins the Gateway image to the data plane version it ships. 1. The Gateway is disabled by default. To deploy it, add the following to your `values.yaml` and apply with `helm upgrade`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} aiGateway: enabled: true ``` 2. Verify the Gateway pods are running before routing traffic to them. Enabling routing before the Gateway is ready causes API requests to fail. 3. Once the Gateway pods are running and ready, set `aiGateway.useGateway: true` to route API traffic through the Gateway, then apply again with `helm upgrade`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} aiGateway: enabled: true useGateway: true ``` Configuration options: | Field | Default | Description | | - | - | - | | `aiGateway.enabled` | `false` | Deploy the Gateway service. | | `aiGateway.useGateway` | `false` | Route API requests through the Gateway. Requires `enabled: true`. | | `aiGateway.replicas` | `2` | Number of Gateway pod replicas. | | `aiGateway.region` | `""` | Optional cloud region identifier (for example, `us-east-1`). Sets `GATEWAY_REGION` when provided. | | `aiGateway.braintrustApiUrl` | `""` | Override the API URL the Gateway uses. Defaults to the in-cluster API service URL. | | `aiGateway.resources` | `cpu: 2`, `memory: 4Gi` | Gateway pod resource requests and limits. | # Scaling and storage Source: https://braintrust.dev/docs/admin/self-hosting/configure/scaling Size the API and Brainstore compute in a self-hosted data plane with Terraform variables and Helm resource settings. These settings size the API and Brainstore compute for your workload. The defaults are suitable for most deployments. Tune them for high ingestion volume, heavy eval workloads, or cluster capacity constraints. ## Configure AWS API ECS services On AWS, Terraform module v6.0 introduces ECS as the target runtime for the API, split into three services by workload type. Each service scales automatically through Application Auto Scaling, adjusting its running task count between a minimum and maximum based on multiple metrics, including CPU target tracking and event loop timing. The module exposes optional variables to tune per-task size and task counts. The defaults are suitable for most deployments. See [Upgrade to Terraform module v6](/docs/admin/self-hosting/upgrade/v6) for the migration. 1024 CPU units equal 1 vCPU. **`braintrust-api`** handles general API traffic. | Variable | Default | | - | - | | `braintrust_api_cpu` | 1024 | | `braintrust_api_memory` | 8192 | | `braintrust_api_min_count` | 3 | | `braintrust_api_max_count` | 50 | **`braintrust-api-ingest`** handles ingestion paths (`/logs3`, `/logs3/overflow`, `/otel/v1/*`). | Variable | Default | | - | - | | `braintrust_api_ingest_cpu` | 1024 | | `braintrust_api_ingest_memory` | 8192 | | `braintrust_api_ingest_min_count` | 6 | | `braintrust_api_ingest_max_count` | 200 | **`braintrust-api-background`** handles background paths (evals, function invocation, proxy). | Variable | Default | | - | - | | `braintrust_api_background_cpu` | 1024 | | `braintrust_api_background_memory` | 8192 | | `braintrust_api_background_min_count` | 3 | | `braintrust_api_background_max_count` | 50 | For example, to raise the ingestion service's floor for a high-volume deployment: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} braintrust_api_ingest_cpu = 2048 braintrust_api_ingest_min_count = 12 ``` Raise `braintrust_api_ingest_min_count` for deployments with spiky ingestion volume or if you expect a launch that will suddenly increase log traffic, and `braintrust_api_background_min_count` for heavy eval workloads. ## Configure Helm API workload isolation This feature is available to Kubernetes deployments using Helm chart 6.13.0+. AWS ECS deployments use the Terraform-managed routing described above. Setting `api.workloadIsolation.enabled: true` creates dedicated `braintrust-api-ingest` and `braintrust-api-background` Deployments and Services alongside the default `braintrust-api` pool. The pools share the same image and base configuration while allowing independent replica counts, resources, probes, rollout settings, and disruption budgets, so classified ingestion or background load does not consume the default API pool's capacity. Your ingress or gateway must keep `braintrust-api` as its default backend and route the following paths to the specialized pools: | Pool | Paths | | - | - | | `braintrust-api-ingest` | `POST /logs3`, `POST /otel/v1/traces`, `POST /attachment`, `POST /attachment/status` | | `braintrust-api-background` | `POST /v1/eval` (exact) and `POST /v1/eval/` (prefix), `POST /function/eval`, `POST /function/sandbox`, `POST /function/use`, `POST /function/invoke-async-batch`, `POST /function/insert-functions`, `POST /automation/logs/trigger`, and all methods for `/v1/proxy/chat/completions` and `/v1/proxy/responses` | | `braintrust-api` (default) | All requests not matched by an ingest or background route | If your ingress cannot match requests by HTTP method (for example, GKE Ingress), route the listed paths for all methods instead. Pools use fixed replica counts by default. Configure them with `api.replicas`, `api.workloadIsolation.ingest.replicas`, and `api.workloadIsolation.background.replicas`. On GKE, you can scale each pool automatically instead. See [Configure GKE API autoscaling](#configure-gke-api-autoscaling). ### Roll out workload isolation for an existing deployment For an existing deployment, stage the rollout so you can verify each pool before routing traffic to it: 1. Create the pools without routing traffic to them or switching Brainstore's internal AI proxy. Set `api.workloadIsolation.brainstoreAiProxyToBackground: false` and apply: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: workloadIsolation: enabled: true brainstoreAiProxyToBackground: false ``` 2. Verify the ingest and background pools are ready. Then update your ingress or gateway to route the classified paths to the new Services, and set `brainstoreAiProxyToBackground: true` to switch Brainstore's internal AI proxy to the background pool: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: workloadIsolation: enabled: true brainstoreAiProxyToBackground: true ``` To roll back, first return the classified paths and `brainstoreAiProxyToBackground` to the default API Service and verify it is serving them, then disable workload isolation in the chart. The chart cannot update an external ingress or gateway on its own, so disabling isolation before rerouting would leave the classified paths pointed at Services that no longer exist. ### Integrate with the Istio VirtualService When using the chart-managed Istio VirtualService, set `virtualService.workloadIsolation.enabled: true` after the isolated pools are healthy. The chart then renders the ingest and background route contract before your existing `virtualService.http` rules. The classified routes take precedence, so do not use `virtualService.http` to override a classified path. Enable this only after the ingest and background pools are running and ready. This option requires both `virtualService.enabled: true` and `api.workloadIsolation.enabled: true`. The chart fails to render if either is missing. ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} virtualService: workloadIsolation: enabled: true ``` ### Configuration reference | Key | Default | Description | | - | - | - | | `api.workloadIsolation.enabled` | `false` | Create dedicated ingest and background API pools. | | `api.workloadIsolation.brainstoreAiProxyToBackground` | `true` | Route Brainstore's internal AI proxy to the background pool. Set to `false` while staging pools for an existing deployment. | | `api.workloadIsolation.ingest.name` | `braintrust-api-ingest` | Name of the ingest Deployment. | | `api.workloadIsolation.ingest.replicas` | `3` | Number of ingest pod replicas. | | `api.workloadIsolation.ingest.service.name` | `braintrust-api-ingest` | Name of the ingest Service. | | `api.workloadIsolation.ingest.podDisruptionBudget.enabled` | `true` | Enable a PodDisruptionBudget for the ingest pool. | | `api.workloadIsolation.ingest.podDisruptionBudget.minAvailable` | `1` | Minimum available ingest pods during voluntary disruptions. | | `api.workloadIsolation.ingest.topologySpread.enabled` | `true` | Enable topology spread constraints for the ingest pool. | | `api.workloadIsolation.background.name` | `braintrust-api-background` | Name of the background Deployment. | | `api.workloadIsolation.background.replicas` | `3` | Number of background pod replicas. | | `api.workloadIsolation.background.service.name` | `braintrust-api-background` | Name of the background Service. | | `api.workloadIsolation.background.podDisruptionBudget.enabled` | `true` | Enable a PodDisruptionBudget for the background pool. | | `api.workloadIsolation.background.podDisruptionBudget.minAvailable` | `1` | Minimum available background pods during voluntary disruptions. | | `api.workloadIsolation.background.topologySpread.enabled` | `true` | Enable topology spread constraints for the background pool. | | `virtualService.workloadIsolation.enabled` | `false` | Render the ingest and background route contract in the chart-managed VirtualService. Requires `virtualService.enabled: true` and `api.workloadIsolation.enabled: true`. | ### Default pool availability settings Helm chart 6.13.0 or later also exposes availability settings for the default `braintrust-api` pool. The defaults preserve the chart's previous behavior, so no configuration changes are required when upgrading. The ingest and background pools inherit these values as their base configuration. | Key | Default | Description | | - | - | - | | `api.strategy.type` | `RollingUpdate` | Rollout strategy for the API Deployment. | | `api.strategy.rollingUpdate.maxSurge` | `100%` | Maximum number of extra pods created during a rollout. | | `api.strategy.rollingUpdate.maxUnavailable` | `0` | Maximum number of pods that can be unavailable during a rollout. | | `api.podDisruptionBudget.enabled` | `false` | Enable a PodDisruptionBudget for the default API pool. | | `api.podDisruptionBudget.minAvailable` | `1` | Minimum available default-pool pods during voluntary disruptions. | | `api.topologySpread.enabled` | `false` | Enable topology spread constraints for the default API pool. | | `api.topologySpread.maxSkew` | `1` | Maximum pod-count skew across topology domains. | | `api.topologySpread.topologyKey` | `topology.kubernetes.io/zone` | Node label that defines the topology domain. | | `api.topologySpread.whenUnsatisfiable` | `ScheduleAnyway` | Whether the scheduler places pods when the constraint cannot be met. | ## Configure GKE API autoscaling This feature is available to GKE deployments using Helm chart 6.15.0 or later. It is built on GKE's `AutoscalingMetric` resource, which Google classifies as Preview (Pre-GA), so Google may change or discontinue it. Setting `api.autoscaling.enabled: true` renders a HorizontalPodAutoscaler and an `AutoscalingMetric` for each API pool. Each pool then scales on three signals: * CPU utilization scoped to the `api` container, so sidecars and `extraContainers` do not skew the measurement. * Node.js event-loop utilization, as a ratio from 0 to 1. * Mean Node.js event-loop delay, in seconds. These are the same three signals AWS uses through Application Auto Scaling. Default CPU (50%) and event-loop utilization (40%) match. Event-loop delay is the exception: AWS uses step scaling, which Kubernetes HPA does not support, so GKE uses target tracking at 50 ms instead. See [Configure AWS API ECS services](#configure-aws-api-ecs-services). ### Prerequisites * Data plane v2.9.0 or later, which serves the Prometheus metrics endpoint on the API health server. * GKE 1.35.1-gke.1396000 or later. * The Performance HPA profile enabled on the cluster. * The Autoscaling API enabled on the cluster. * `roles/autoscaling.metricsWriter` granted to every node service account. * The Autoscaling API included in your service perimeter, if you use VPC Service Controls. `helm install` and `helm upgrade` fail fast if the chart detects a missing prerequisite, rather than rendering a partial configuration: * Setting `cloud` to anything other than `google` fails, because autoscaling is supported only on GKE. * A cluster without the `autoscaling.gke.io/v1beta1` API fails. Verify with `kubectl api-resources | grep autoscalingmetric`. ### Enable autoscaling ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: autoscaling: enabled: true minReplicas: 4 maxReplicas: 50 ``` When autoscaling is enabled for a pool, the chart omits `replicas` from that pool's Deployment and the HorizontalPodAutoscaler controls the replica count instead. The chart also sets `ENABLE_PROMETHEUS_METRICS` to `true` and exposes `api.healthServer.port` (`8001` by default) on the `api` container so the `AutoscalingMetric` can scrape it. Account for that port in any network policy that restricts traffic to the API pods. ### Override autoscaling per pool With `api.workloadIsolation.enabled: true`, the ingest and background pools inherit every value under `api.autoscaling` and can override any of them under `api.workloadIsolation..autoscaling`, including the metric targets and the scaling behavior: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: autoscaling: enabled: true minReplicas: 3 maxReplicas: 40 workloadIsolation: enabled: true ingest: autoscaling: minReplicas: 3 maxReplicas: 60 eventLoopUtilization: targetAverageValue: "0.5" behavior: scaleUp: stabilizationWindowSeconds: 30 background: autoscaling: minReplicas: 3 maxReplicas: 50 ``` This example turns on workload isolation at the same time as autoscaling, which suits a new deployment. On an existing deployment, stage the pools first. The example leaves `api.workloadIsolation.brainstoreAiProxyToBackground` at its default `true`, which switches Brainstore's internal AI proxy to the background pool as soon as you apply it. See [Roll out workload isolation for an existing deployment](#roll-out-workload-isolation-for-an-existing-deployment). Setting `api.workloadIsolation..autoscaling.enabled: false` keeps that pool at a fixed replica count while its siblings continue to scale. That pool's `replicas` value is honored again. The Helm chart ships [`examples/google-api-isolation-autoscaling/values.yaml`](https://github.com/braintrustdata/helm/blob/main/braintrust/examples/google-api-isolation-autoscaling/values.yaml) as a starting point for this pattern. Combine it with an Autopilot or Standard values file, which supply the rest of the configuration. For background on the GKE resource behind this feature, see Google's [Expose custom metrics for autoscaling](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/expose-custom-metrics-autoscaling). ### Configuration reference | Key | Default | Description | | - | - | - | | `api.autoscaling.enabled` | `false` | Autoscale every API pool. Requires `cloud: google`. | | `api.autoscaling.minReplicas` | `4` | Minimum replicas for the default pool. Other pools inherit this value unless overridden. | | `api.autoscaling.maxReplicas` | `50` | Maximum replicas for the default pool. Other pools inherit this value unless overridden. | | `api.autoscaling.metricsPath` | `/metrics` | Prometheus metrics path on the API health server. | | `api.autoscaling.cpu.targetAverageUtilization` | `50` | Target CPU utilization percentage for the `api` container. | | `api.autoscaling.eventLoopUtilization.targetAverageValue` | `"0.4"` | Target event-loop utilization, as a ratio from 0 to 1. `"0.4"` is 40%. | | `api.autoscaling.eventLoopDelayMean.targetAverageValue` | `"0.05"` | Target mean event-loop delay in seconds. `"0.05"` is 50 ms. | | `api.autoscaling.behavior.scaleDown.stabilizationWindowSeconds` | `300` | Stabilization window before scaling down. Passed through to the HorizontalPodAutoscaler. | | `api.autoscaling.behavior.scaleUp.stabilizationWindowSeconds` | `60` | Stabilization window before scaling up. Passed through to the HorizontalPodAutoscaler. | | `api.annotations.hpa` | `{}` | Annotations added to each generated HorizontalPodAutoscaler. | | `api.annotations.autoscalingMetric` | `{}` | Annotations added to each generated `AutoscalingMetric`. | | `api.workloadIsolation.ingest.autoscaling.minReplicas` | `3` | Minimum replicas for the ingest pool. | | `api.workloadIsolation.ingest.autoscaling.maxReplicas` | `200` | Maximum replicas for the ingest pool. | | `api.workloadIsolation.background.autoscaling.minReplicas` | `3` | Minimum replicas for the background pool. | | `api.workloadIsolation.background.autoscaling.maxReplicas` | `50` | Maximum replicas for the background pool. | ## Configure Brainstore fast readers Fast readers are isolated Brainstore nodes dedicated to serving predictable UI queries (paginated viewers, span and trace lookups), preventing resource-intensive ad-hoc queries from making the UI unresponsive. * **GCP and Azure**: Fast readers are enabled by default starting in Helm chart v5.0.0. See the configuration reference below. * **AWS**: Fast readers are enabled by default (2 nodes) starting in Terraform module v5.5.0. On earlier module versions they are disabled by default. Set `brainstore_fast_reader_instance_count` in your Terraform configuration to control the node count, or set it to `0` to opt out (recommended for sandbox or non-production deployments). Upgrading to Helm chart v5.0.0 from an earlier version automatically creates fast reader nodes. By default, 2 fast reader nodes are created with the same resource profile as standard reader nodes (CPU: 16, memory: 32Gi). Verify that your cluster has capacity for these additional nodes before upgrading. If you have custom `brainstore.readinessProbe` overrides pointing to `/status`, remove them before upgrading to Helm chart v5.0.0+. The `/status` readiness endpoint has a bug where it never recovers after a failure, which can permanently mark Brainstore nodes as not ready. Remove any `brainstore.readinessProbe` or `brainstore.fastreader.readinessProbe` customizations and rely on the chart defaults. ### Configuration reference Fast readers are configured under the `brainstore.fastreader` key in your `values.yaml`. If you have customized `brainstore.reader` settings, mirror those customizations to `brainstore.fastreader`. | Key | Default | Description | | - | - | - | | `brainstore.fastreader.name` | `brainstore-fastreader` | Name of the Deployment, Service, and ConfigMap | | `brainstore.fastreader.replicas` | `2` | Number of fast reader pod replicas | | `brainstore.fastreader.service.port` | `4000` | Service port | | `brainstore.fastreader.service.type` | `ClusterIP` | Kubernetes service type | | `brainstore.fastreader.resources.requests.cpu` | `16` | CPU request | | `brainstore.fastreader.resources.requests.memory` | `32Gi` | Memory request | | `brainstore.fastreader.resources.limits.cpu` | `16` | CPU limit | | `brainstore.fastreader.resources.limits.memory` | `32Gi` | Memory limit | | `brainstore.fastreader.objectStoreCacheMemoryLimit` | `1Gi` | Object store memory cache limit | | `brainstore.fastreader.objectStoreCacheFileSize` | `1000Gi` | Object store file cache size | | `brainstore.fastreader.cacheDir` | `/mnt/tmp/brainstore` | Local cache mount path | | `brainstore.fastreader.volume.size` | `""` | Ephemeral storage size; required for Azure ACS | | `brainstore.fastreader.extraEnvVars` | `[]` | Additional environment variables | | `brainstore.fastreader.nodeSelector` | `{}` | Node selector for scheduling | | `brainstore.fastreader.tolerations` | `[]` | Pod tolerations | | `brainstore.fastreader.affinity` | `{}` | Pod affinity rules | Azure users must explicitly set `brainstore.fastreader.volume.size` when using Azure Container Storage (`enableAzureContainerStorageDriver: true`): ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstore: fastreader: volume: size: "100Gi" ``` On GKE Autopilot, set `brainstore.fastreader.volume.size` to configure ephemeral storage requests. Match the resource profile of your standard reader nodes: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstore: fastreader: resources: requests: cpu: "16" memory: "32Gi" limits: cpu: "16" memory: "32Gi" objectStoreCacheFileSize: "900Gi" volume: size: "1000Gi" ``` ## Brainstore rollout controls This section applies to GCP and Azure deployments using the Helm chart (6.17.1+). AWS deployments manage Brainstore rollout through Terraform and do not expose these settings. Brainstore readers, fast readers, and writers each have independently configurable rollout strategy and readiness dwell settings. These controls let you limit how many replacement pods start together and require replacements to stay ready for a minimum duration before a rollout continues. This can reduce simultaneous cache warm-up, object-storage, scheduling, and compaction pressure on cache-heavy or high-throughput deployments. The defaults preserve the rollout behavior from earlier chart versions and apply to all three Brainstore roles: | Key | Default | Description | | - | - | - | | `brainstore..strategy.type` | `RollingUpdate` | Rollout strategy for this Brainstore role's Deployment. | | `brainstore..strategy.rollingUpdate.maxSurge` | `100%` | Maximum number of extra pods created during a rollout. | | `brainstore..strategy.rollingUpdate.maxUnavailable` | `0` | Maximum pods that can be unavailable during a rollout. | | `brainstore..minReadySeconds` | `0` | Minimum seconds a replacement pod must stay ready before the rollout proceeds. | | `brainstore..progressDeadlineSeconds` | `600` | Maximum seconds the Deployment controller waits for progress. Must be greater than `minReadySeconds`. | Replace `` with `reader`, `fastreader`, or `writer`. Upgrading without overriding these values does not change rollout pacing. These settings do not change the pod template and do not start a rollout on their own. Kubernetes uses updated settings for an active rollout and future rollouts. For cache-heavy or high-throughput deployments, use a longer readiness dwell: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstore: writer: minReadySeconds: 300 progressDeadlineSeconds: 900 strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 maxUnavailable: 0 reader: minReadySeconds: 120 progressDeadlineSeconds: 600 strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 maxUnavailable: 0 fastreader: minReadySeconds: 120 progressDeadlineSeconds: 600 strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 maxUnavailable: 0 ``` `progressDeadlineSeconds` must be greater than `minReadySeconds`. The chart rejects an invalid pairing. Keep enough deadline margin for scheduling, image pulls, startup, and the readiness dwell. Longer readiness dwell reduces rollout pressure but does not confirm that a pod's local cache is fully warm. Monitor workload health until the rollout converges. `maxUnavailable: 0` requires enough capacity to schedule the configured surge. On managed clusters, even `maxSurge: 1` can require an additional node and local storage. Before upgrading, verify regional compute quota and the availability of the required node type and storage, and account for the temporary increase in compute and storage cost. Setting `strategy.type: Recreate` stops all pods in a Brainstore role before creating replacements. This causes a complete role outage, and for the writer (which defaults to a single replica) pauses background processing until the replacement becomes ready. Slow rollouts extend the period during which old and new Brainstore versions run together. Do not change `brainstoreWalFooterVersion` in the same upgrade as an image version bump, except where the documented [data plane 2.0 upgrade sequence](/docs/admin/self-hosting/upgrade/v2) explicitly permits it. The chart does not create PodDisruptionBudgets for Brainstore, so these rollout controls do not limit voluntary disruptions such as node drains or protect against involuntary pod or node failures. ## Brainstore resource configuration This section applies to GCP and Azure deployments using the Helm chart (v5.0.1+). AWS deployments manage Brainstore resources automatically. Starting in Helm chart v5.0.1, the `resources` block for each Brainstore component (`brainstore.reader`, `brainstore.writer`, `brainstore.fastreader`) is passed through as-is to the Kubernetes pod spec. You can omit `limits` entirely, set them to `{}`, or supply any valid Kubernetes resource spec. Omitting `limits` sets the pod QoS class to Burstable, which prevents CPU throttling and allows pods to use available node capacity. This can improve query performance on nodes with spare capacity, but increases the risk of resource contention if multiple pods compete for the same node. ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} # With limits (Guaranteed QoS — default) brainstore: reader: resources: requests: cpu: "16" memory: "32Gi" limits: cpu: "16" memory: "32Gi" # Without limits (Burstable QoS — allows burst CPU/memory) brainstore: reader: resources: requests: cpu: "16" memory: "32Gi" writer: resources: requests: cpu: "4" memory: "16Gi" fastreader: resources: requests: cpu: "16" memory: "32Gi" ``` ### Auto-derived Brainstore environment variables As of Helm chart v5.1.0, `BRAINSTORE_RESPONSE_CACHE_URI` and `BRAINSTORE_CODE_BUNDLE_URI` are automatically populated from your `objectStorage` configuration and do not need to be set manually. The chart derives these values as follows: | Cloud | Source key | `BRAINSTORE_RESPONSE_CACHE_URI` | `BRAINSTORE_CODE_BUNDLE_URI` | | - | - | - | - | | AWS | `objectStorage.aws.responseBucket` / `objectStorage.aws.codeBundleBucket` | `s3:///brainstore-cache` | `s3://` | | Azure | `objectStorage.azure.responseContainer` / `objectStorage.azure.codeBundleContainer` | `az:///brainstore-cache` | `az://` | | GCP | `objectStorage.google.apiBucket` | `gs:///brainstore-cache` | `gs://` | If you previously configured these via `extraEnvVars`, remove those overrides after upgrading to v5.1.0 to avoid conflicts. # Security and access Source: https://braintrust.dev/docs/admin/self-hosting/configure/security Restrict which organizations can authenticate, secure customer data, validate outbound URLs, audit data access, and harden pods. These settings control who and what can reach your self-hosted data plane, and what they can touch. Most are optional and default to permissive behavior suitable for trusted, single-organization deployments. * [Data isolation and access](#data-isolation-and-access): Protect customer data and restrict which organizations can use the deployment. * [AWS resource permissions](#aws-resource-permissions): Limit access to S3 buckets and IAM roles. * [Network security and TLS](#network-security-and-tls): Validate outbound URLs, incoming headers, and server certificates. * [Auditing and network logs](#auditing-and-network-logs): Record data access and network traffic. * [Kubernetes hardening](#pod-security-contexts-and-storage-limits): Configure pod permissions and writable storage. ## Data isolation and access ### Secure sensitive customer data Braintrust's servers and employees *do not* require access to your data plane for it to operate successfully. That means that you can protect it behind a firewall or VPN and physically isolate it from access. When you use the Braintrust web application, it communicates directly with the data plane (via CORS), and the data does not flow through any intermediate systems (the control plane, or otherwise) before reaching your browser. While the data plane does send metrics and status telemetry to the control plane, it does not send logs or customer data. Deployments that enable the [Loop runtime](/docs/admin/self-hosting/configure/loop-runtime) also send that service's own metrics and traces, which cover the runtime's operation and don't include LLM calls, tool calls, or the contents of your traces. Because of this architecture, our self-hosted customers do not generally list us as a subprocessor. Like any third-party software, it is important that you establish the appropriate controls to ensure that your deployment is secure, and Braintrust can help you do so. Ultimately, the goal of the control plane and data plane split is to provide you with the highest levels of security and compliance. ### Grant browser permissions This is a client-side step performed by end-users in their browser, not a Terraform or Helm setting. The only deployment-side control is the **Data plane is on a private network** toggle in organization settings. If your data plane is deployed behind a VPN or on a private network (not accessible from the public internet), users will need to grant browser permissions to access it. When you enable the **Data plane is on a private network** setting in [organization settings](/docs/admin/organizations#configure-api-urls-self-hosted), the Braintrust UI will check for Chrome's Local Network Access permission. When users access your Braintrust org for the first time and a private network data plane has been configured, Chrome will display a permission prompt. Users must click **Allow** to grant permission for the Braintrust UI to communicate with your data plane. If permissions are blocked, we will display a modal with the steps to correct the issue. Local Network Access is a Chrome security feature that protects users from malicious websites accessing resources on their private networks. For more information, see [Chrome's Private Network Access documentation](https://developer.chrome.com/blog/private-network-access-update). If you encounter connectivity issues, ensure that your browser is up to date, any corporate proxies or network policies allow browser access to the data plane, and CORS is properly configured on your data plane deployment (automatically handled by Braintrust Terraform modules). ### Restrict access to specific organizations By default, any Braintrust organization can authenticate against your self-hosted data plane. Use the following settings to restrict access to specific organizations. You can find an organization's ID in the Braintrust UI. See [Find your organization ID](/docs/admin/organizations#find-your-organization-id). Set `allowed_org_ids` to a comma-separated list of Braintrust organization UUIDs (available in Terraform module v5.8.1 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" allowed_org_ids = "00000000-0000-4000-8000-000000000001,00000000-0000-4000-8000-000000000002" # ... other configuration ... } ``` Only organizations whose IDs appear in the list can use the deployment. Leave the variable empty (the default) to impose no restriction. Organization IDs must be UUIDs, not organization names. If `braintrust_org_name` is set to a specific organization name, include that organization's ID in `allowed_org_ids` for forward compatibility. These settings apply to deployments using the Helm chart (6.7.1+). Two `global` values control which organizations can use the deployment. `global.primaryOrgName` identifies the primary organization for self-hosted service-token management, which enables features such as [data retention](/docs/admin/self-hosting/configure/telemetry#data-retention). It is required when `global.orgName` is empty or `"*"` (a wildcard that allows all organizations). For standard single-organization deployments with a specific `orgName`, this value is not needed. `global.allowedOrgIds` is an optional comma-separated list of Braintrust organization IDs (UUIDs, not org names) authorized to use this deployment. When `global.orgName` is a specific name, that org is implicitly included. Include its Braintrust org ID here for forward compatibility. These two values serve different purposes. `global.primaryOrgName` enables service-token management for a wildcard or empty `orgName`. It does not restrict which organizations can authenticate. To scope access on a wildcard deployment, also set `global.allowedOrgIds`. Without it, any organization can authenticate. ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} global: orgName: "*" primaryOrgName: "your-primary-org" # Required when orgName is "*" or empty allowedOrgIds: "00000000-0000-4000-8000-000000000001" # Optional org ID allowlist ``` Deploying with `global.orgName: "*"` or an empty value without setting `global.primaryOrgName` (or providing `PRIMARY_ORG_NAME` via `api.extraEnvVars`) causes a Helm template validation error. ## AWS resource permissions ### Enable S3 attribute-based access control On AWS, IAM policies for the S3 buckets the module manages grant access by bucket ARN. To base access decisions on bucket tags instead, set the `enable_s3_bucket_abac` Terraform variable to `true` (available in Terraform module v5.3.0 or later). This supports tag-based governance or service control policies that gate access by resource tags. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} enable_s3_bucket_abac = true ``` When enabled, you can reference bucket tags in IAM authorization policies, and managing those tags requires the `s3:TagResource`, `s3:UntagResource`, and `s3:ListTagsForResource` permissions. This variable is disabled by default. ### Restrict S3 VPC endpoint access On AWS, the S3 gateway VPC endpoints the module creates allow traffic to any S3 bucket by default. To limit which buckets are reachable through them, set `s3_vpc_endpoint_resource_org_ids`, `s3_vpc_endpoint_resource_account_ids`, or both (available in Terraform module v6.7.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" s3_vpc_endpoint_resource_org_ids = ["o-abc123def4"] s3_vpc_endpoint_resource_account_ids = ["123456789012"] # ... other configuration ... } ``` The two lists combine. Organization IDs are matched with the `aws:ResourceOrgID` condition key, account IDs with `aws:ResourceAccount`, and a bucket satisfying either one is allowed. Use the account list for destinations outside your organization, such as a cross-org [S3 export](/docs/admin/data-management/export) target. When both lists are empty (the default), the endpoint policy allows all S3 traffic. When either is non-empty, the account running the deployment is always allowed, so the module's own buckets keep working, and the policy also permits `s3:GetObject` on two Amazon-owned buckets the deployment depends on: the regional ECR layer bucket used for private image pulls, and the `amazoncloudwatch-agent` bucket that Brainstore instances install the CloudWatch agent from. The restriction applies to every S3 request that routes through the endpoint, including Brainstore, code bundles, Lambda responses, and export. If you later add an S3 destination in another account, add it to the allowlist first, or those requests fail. These variables apply only to endpoints the module manages: the main VPC endpoint when `create_vpc` is `true`, and the quarantine VPC endpoint when `enable_quarantine_vpc` is `true` and `existing_quarantine_vpc_id` is unset. Endpoints on VPCs you supply through the `existing_*` variables are yours to manage, and the module does not modify their policies. ### Restrict S3 export AssumeRole By default, the API handler and Brainstore roles can assume any IAM role when performing [S3 export](/docs/admin/data-management/export) operations (`sts:AssumeRole` with ExternalId `bt:*`), which supports arbitrary export destinations. To lock down which roles they can assume, set `s3_export_assume_role_arns` to an explicit allowlist (available in Terraform module v6.4.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" s3_export_assume_role_arns = [ "arn:aws:iam::123456789012:role/customer-export-role", "arn:aws:iam::*:role/braintrust-export-*", ] # ... other configuration ... } ``` When the list is empty (the default), AssumeRole remains unrestricted for backward compatibility. When set, only matching role ARNs can be assumed. Entries accept exact ARNs and IAM ARN patterns, including a wildcard account ID (`*`). ### Restrict Gateway Bedrock AssumeRole If you run the [Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) in your data plane, its task role can by default assume any IAM role for Amazon Bedrock authentication (`sts:AssumeRole` with ExternalId `bt:*`). To restrict which roles it can assume, set `ai_gateway_bedrock_assume_role_arns` (available in Terraform module v6.4.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" ai_gateway_bedrock_assume_role_arns = [ "arn:aws:iam::123456789012:role/braintrust-bedrock-role", ] # ... other configuration ... } ``` This variable applies only when `create_ai_gateway` is `true`. It follows the same rules as `s3_export_assume_role_arns`: empty leaves AssumeRole unrestricted, and entries accept exact ARNs and IAM ARN patterns. ## Network security and TLS ### Configure URL security Braintrust backends validate user-supplied URLs (such as webhook targets and external endpoints invoked by scorers or tools) before making outbound requests, to prevent server-side request forgery (SSRF). By default, requests that resolve to private or reserved IP ranges are allowed and logged with a warning rather than blocked. To block them, set the URL request mode to `reject`. The following Terraform variables configure URL security (available in module v5.8.1 or later): | Variable | Default | Description | | - | - | - | | `unsafe_url_request_mode` | `""` (application default: `warn`) | How the backend handles requests to URLs that fail security checks. `reject` blocks, `warn` allows with a log warning, `proxy` routes through the configured outbound proxy, or `off` allows. | | `url_security_dns_servers` | `""` | Comma-separated DNS resolver IPs used to validate user-supplied URLs. Set this to route validation through your VPC or corporate DNS. | | `url_security_allow_cidrs` | `""` | Comma-separated CIDR ranges to allow even if private or reserved. This allowlist cannot override metadata endpoints, link-local, multicast, unspecified, or future-use ranges, which stay blocked in `reject` mode. | To block all requests to private or reserved addresses: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" unsafe_url_request_mode = "reject" # ... other configuration ... } ``` To allow a specific internal CIDR while warning on other private ranges: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" unsafe_url_request_mode = "warn" url_security_allow_cidrs = "10.0.1.0/24" # ... other configuration ... } ``` These settings apply to deployments using the Helm chart (6.7.1+). Set the following `api` values in your Helm `values.yaml` file: | Value | Default | Description | | - | - | - | | `api.unsafeUrlRequestMode` | `""` | How to handle requests to URLs that resolve to private or reserved addresses. Leave empty to use the application default (`warn`). Set to `reject` to block, `warn` to allow with a log warning, or `off` to allow. The API and runtime fetch paths also support `proxy`, which routes the request through the configured outbound proxy. The AI Gateway supports only `reject`, `warn`, and `off`. | | `api.urlSecurityDnsServers` | `""` | Comma-separated DNS resolver IP addresses used when validating user-supplied URLs. Set to your VNet or corporate DNS servers to enforce resolution through a trusted nameserver before falling back to the host resolver. | | `api.urlSecurityAllowCidrs` | `""` | Comma-separated CIDR ranges that URL validation allows even if they are private or reserved. This allowlist cannot override metadata endpoints, link-local, multicast, unspecified, or future-use ranges, which stay blocked in `reject` mode. | To reject all requests to private or reserved addresses: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: unsafeUrlRequestMode: "reject" ``` To allow specific internal CIDRs while rejecting others: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: unsafeUrlRequestMode: "reject" urlSecurityAllowCidrs: "10.0.0.0/8" ``` To resolve user-supplied URLs through a trusted DNS server: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: urlSecurityDnsServers: "10.0.0.53" ``` Setting the URL request mode to `off` disables all URL security checks and allows outbound requests to private and reserved IP ranges. Avoid `off` for internet-facing deployments. ### Configure ALB header validation To have the API and gateway Application Load Balancers (ALBs) drop HTTP requests with invalid header names before routing them, set the following Terraform variables (available in Terraform module v6.1.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" braintrust_api_alb_drop_invalid_header_fields = true ai_gateway_alb_drop_invalid_header_fields = true # ... other configuration ... } ``` Both default to `false`. Enabling them hardens the ALBs against malformed requests, though some HTTP clients or intermediary proxies that send non-standard header names may be affected. ### Configure a custom CA bundle If your internal services present certificates signed by a private or enterprise certificate authority (CA), user-authored code such as [custom scorers](/docs/evaluate/write-scorers) and tools cannot validate them, and outbound HTTPS requests fail with a certificate verification error. Available with data plane v2.9.0 or later, the API service forwards the standard CA environment variables from its own environment into the runtimes that execute user code, for both Node.js and Python. The executing runtime must also be able to read the bundle at the configured path. On earlier versions, the variables stay on the API container and don't reach user code runtimes. Trust is established by pointing these variables at a PEM bundle that the API container and user-code runtime can read at the configured path. Each one covers a different client library, so set all six to cover both Node.js and Python user code: * `NODE_EXTRA_CA_CERTS`. * `REQUESTS_CA_BUNDLE`. * `SSL_CERT_FILE`. * `CURL_CA_BUNDLE`. * `AWS_CA_BUNDLE`. * `PIP_CERT`. Include the public root certificates your workloads need alongside your own root and intermediate certificates. `SSL_CERT_FILE`, `REQUESTS_CA_BUNDLE`, and `CURL_CA_BUNDLE` replace the default trust store for OpenSSL, Python, and curl clients, so a bundle containing only your private CA chain breaks outbound HTTPS to public endpoints for that code. `NODE_EXTRA_CA_CERTS` appends to the Node.js defaults instead of replacing them, so behavior differs by runtime even though all the variables point at the same file. AWS does not yet have a packaged equivalent to Helm's `api.customCA` configuration. In particular, Lambda quarantine runtimes do not share the API container filesystem, so baking a bundle into a custom API image and setting `braintrust_api_extra_env_vars` alone does not make that bundle available to those runtimes. For Python scorers running in Lambda quarantine, [package the CA bundle with the scorer and configure `REQUESTS_CA_BUNDLE`](/docs/kb/python-scorer-custom-ca-bundle-aws-kb). That guide is specific to Python scorers. For other AWS runtime configurations, make the bundle available in the executing runtime before configuring a CA environment-variable path. Helm chart 6.11.0 and later provide the `api.customCA` values block, which mounts the bundle and sets all six variables for you. Create a Kubernetes Secret holding your PEM-encoded CA bundle in the same namespace as the release, before you install or upgrade the chart: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} apiVersion: v1 kind: Secret metadata: name: braintrust-runtime-ca type: Opaque stringData: ca-bundle.pem: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` Kubernetes limits each Secret to 1 MiB. A combined bundle is normally well under that, but check the size of a large enterprise bundle before you roll it out. Reference the Secret from `api.customCA` and apply with `helm upgrade`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: customCA: enabled: true secretName: braintrust-runtime-ca secretKey: ca-bundle.pem mountPath: /etc/braintrust/runtime-ca filename: ca-bundle.pem ``` | Field | Default | Description | | - | - | - | | `api.customCA.enabled` | `false` | Mount the CA bundle and set the CA environment variables. | | `api.customCA.secretName` | `"braintrust-runtime-ca"` | Name of the Secret holding the bundle. | | `api.customCA.secretKey` | `"ca-bundle.pem"` | Key within the Secret whose value is the PEM bundle. Only this key is mounted, not the whole Secret. | | `api.customCA.mountPath` | `"/etc/braintrust/runtime-ca"` | Directory the bundle is mounted into. | | `api.customCA.filename` | `"ca-bundle.pem"` | Filename for the mounted bundle. The full path is `mountPath/filename`. | Every field except `enabled` is required when `enabled` is `true`. If you leave one empty, the chart fails to render rather than deploying an API pod without trust configured. The chart mounts the bundle read-only and sets the variables to `mountPath/filename`, so you don't need to add them to `api.extraEnvVars`. If code runs outside the API pod, make the same bundle available at that path in its runtime as well. This approach doesn't require root and doesn't modify the container's system certificate store. ## Auditing and network logs Use read auditing to record which data was queried, VPC Flow Logs to investigate network traffic, and legacy audit headers to obtain request details in API responses. ### Enable read auditing Braintrust logs administrative actions automatically. It can also record row-level `query.read` [audit logs](/docs/admin/audit-logs#audit-data-reads-and-modifications) that capture both SQL queries run manually and ones the Braintrust UI runs implicitly when users browse logs, experiments, and traces. These events are high volume, so they are disabled by default and require an Enterprise plan. Enable them per organization by setting one of two mutually exclusive variables to a list of Braintrust [organization UUIDs](/docs/admin/organizations#find-your-organization-id). Both default to empty, and setting both is a configuration error. * **Strict mode** writes each audit row before returning query results, guaranteeing the read is recorded before results reach the caller. * **Best-effort mode** writes audit rows asynchronously and logs any failures without blocking the query. Set one of these Terraform variables (available in module v6.0.0 or later): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} btql_audit_logs_strict_org_ids = ["00000000-0000-0000-0000-000000000000"] ``` Or for best-effort mode: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} btql_audit_logs_best_effort_org_ids = ["00000000-0000-0000-0000-000000000000"] ``` Set one of these Helm values (available in chart 6.9.2 or later): ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} btqlAuditLogsStrictOrgIds: ["00000000-0000-0000-0000-000000000000"] ``` Or for best-effort mode: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} btqlAuditLogsBestEffortOrgIds: ["00000000-0000-0000-0000-000000000000"] ``` Once enabled, read these events with the audit log. See [Query the audit log](/docs/admin/audit-logs#query-the-audit-log) for the SQL data source and example queries. ### Enable VPC Flow Logs [AWS VPC Flow Logs](https://docs.aws.amazon.com/vpc/latest/userguide/flow-logs.html) record metadata about network traffic, such as source and destination IP addresses, ports, and whether traffic was accepted or rejected. Use them to investigate connectivity problems and audit network access. They do not capture request or response contents. Terraform module v6.8.0 or later supports flow logs for the main VPC, where your data plane services run, and the quarantine VPC, where user-defined functions run in network isolation. Flow logs are disabled by default and apply only to VPCs the module creates. The main VPC requires `create_vpc = true`. The quarantine VPC requires `enable_quarantine_vpc = true` with `existing_quarantine_vpc_id` unset. For VPCs you supply, configure flow logs outside this module. #### Choose a log destination Add the settings for your destination to your existing Terraform module block, then review and apply the Terraform plan. Configure the main and quarantine VPCs independently using `main_vpc_flow_log` and `quarantine_vpc_flow_log`. To create a dedicated S3 bucket for each VPC, enable flow logs without specifying a destination: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} main_vpc_flow_log = { enabled = true } quarantine_vpc_flow_log = { enabled = true } ``` Each bucket uses the data plane's KMS key for encryption, and logs expire after 365 days. To enable flow logs for only one VPC, omit the other setting. Set `destination_arn` to your bucket ARN. To configure the quarantine VPC, use `quarantine_vpc_flow_log` instead: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} main_vpc_flow_log = { enabled = true destination_arn = "arn:aws:s3:::your-flow-logs-bucket" } ``` Before applying, attach a bucket policy granting `delivery.logs.amazonaws.com` the `s3:PutObject` and `s3:GetBucketAcl` permissions. Flow log creation can succeed even when delivery is denied, leaving the bucket empty. You manage this bucket's encryption and retention. Set `destination_type` to create a log group. To configure the quarantine VPC, use `quarantine_vpc_flow_log` instead: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} main_vpc_flow_log = { enabled = true destination_type = "cloud-watch-logs" } ``` To use an existing log group, also set `destination_arn` to its ARN without the trailing `:*`. The module creates the IAM delivery role in either case and applies `permissions_boundary_arn` if configured. For module-managed destinations, set `retention_in_days` to change the 365-day retention period, or use `0` to retain logs indefinitely. For destinations you supply, manage retention yourself. Module-managed S3 buckets are not emptied automatically. After logs are written, disabling logging, changing destinations, or destroying the stack fails with `BucketNotEmpty`. Empty the bucket to delete the flow logs, or remove the bucket and its companion resources from Terraform state to keep them. To retain logs when destroying the stack, preserve the encrypting KMS key as well as the bucket. A module-managed key becomes unusable when scheduled for deletion and is permanently deleted after seven days. Retained logs encrypted with that key become unreadable. See the module's [log retention and teardown guidance](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane/tree/v6.8.0#vpc-flow-logs). #### Configuration reference Both flow log objects accept the same fields: | Field | Type | Default | Description | | - | - | - | - | | `enabled` | Boolean | `false` | Enable flow logs for this VPC. | | `traffic_type` | String | `"ALL"` | Capture `ALL`, `ACCEPT`, or `REJECT` traffic. | | `destination_type` | String | `"s3"` | Send logs to `s3` or `cloud-watch-logs`. | | `destination_arn` | String | `null` | Existing destination ARN. Leave unset to create a destination. | | `max_aggregation_interval` | Number | `600` | Aggregation interval in seconds: `60` or `600`. | | `log_format` | String | `null` | Custom format. Leave unset for the AWS default. | | `retention_in_days` | Number | `365` | Retention for a module-managed destination. Use a nonnegative value for S3 or a supported CloudWatch Logs retention value. `0` disables expiration. | | `kms_key_arn` | String | `null` | Override the data plane KMS key for a module-managed destination. | Module-managed destinations use the data plane's KMS key by default. If you set `kms_key_arn` to a different customer-managed key, its policy must allow `delivery.logs.amazonaws.com` for S3 or `logs..amazonaws.com` for CloudWatch Logs. Do not pass the module's own `kms_key_arn` output back into its flow log configuration, which creates a Terraform dependency cycle. After deployment, use these Terraform outputs to identify the log destinations the module created. Each output contains the destination's ARN, or `null` if that resource wasn't created: * `main_vpc_flow_log_s3_bucket_arn` * `main_vpc_flow_log_cloudwatch_log_group_arn` * `quarantine_vpc_flow_log_s3_bucket_arn` * `quarantine_vpc_flow_log_cloudwatch_log_group_arn` ### Enable audit headers Audit headers are a legacy feature, retained for existing integrations. They predate Braintrust's [audit logging](/docs/admin/audit-logs) feature, which is the recommended way to track administrative actions and data access. Audit headers are enabled per request by the API client rather than through Terraform or Helm. When actions must be attributed to specific users or tracked for compliance, enable audit headers. These headers add metadata about the request and the resources it touched to the API response. To enable audit headers, include the `x-bt-enable-audit: true` header in your API request. When this header is present, the API response will include the following additional headers: * `x-bt-audit-user-id`: The ID of the user who made the request (based on the provided API key or impersonation). * `x-bt-audit-user-email`: The email of the user who made the request. * `x-bt-audit-normalized-url`: A normalized representation of the API endpoint path that was called. Path parameters like object IDs are replaced with placeholders (for example, `/v1/project/[id]`). * `x-bt-audit-resources`: A JSON-encoded, gzipped, and base64-encoded string containing a list of Braintrust resources (like projects, experiments, datasets, etc.) that were accessed or modified by the request. Each resource object includes its `type`, `id`, and `name`. The `x-bt-audit-resources` header requires specific parsing due to its encoding. Here's an example of how to parse it using the Python SDK: ```py theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} import os import braintrust import requests API_URL = "https://api.braintrust.dev/v1" # Ensure BRAINTRUST_API_KEY is set in your environment. headers = { "Authorization": "Bearer " + os.environ["BRAINTRUST_API_KEY"], "x-bt-enable-audit": "true", # Enable audit headers } # Example: Create a project. response = requests.post(f"{API_URL}/project", headers=headers, json={"name": "audit-test-project"}) response.raise_for_status() project_data = response.json() print(f"Project created: {project_data['name']} (ID: {project_data['id']})") # Access and parse audit headers. user_id = response.headers.get("x-bt-audit-user-id") user_email = response.headers.get("x-bt-audit-user-email") normalized_url = response.headers.get("x-bt-audit-normalized-url") resources_header = response.headers.get("x-bt-audit-resources") print(f"Audit User ID: {user_id}") print(f"Audit User Email: {user_email}") print(f"Normalized URL: {normalized_url}") if resources_header: try: # Use the provided utility to parse the resources header. resources = braintrust.parse_audit_resources(resources_header) print("Accessed/Modified Resources:") for resource in resources: print(f" - Type: {resource['type']}, ID: {resource['id']}, Name: {resource['name']}") except Exception as e: print(f"Error parsing resources header: {e}") else: print("No resources header found.") ``` This feature is useful for building audit logs or understanding resource usage patterns within your applications that interact with the Braintrust API. ## Kubernetes hardening This section applies to GCP and Azure deployments using the Helm chart (6.2.2+). Starting in Helm chart 6.2.2, you can harden Braintrust pods by configuring pod- and container-level security contexts, explicit ephemeral storage requests and limits, `emptyDir` size limits, and writable `/tmp` volumes for the API and all Brainstore roles (`brainstore.reader`, `brainstore.fastreader`, `brainstore.writer`). By default, the Braintrust manifests don't set these fields, so configure them only if you want to enforce a stricter security posture. Some hardened clusters require these settings. They enforce admission policies, often written as Common Expression Language (CEL) constraints, that reject any workload unless its pods declare a non-root, read-only security context and bounded storage. GKE Autopilot is one such environment, so a deployment that doesn't set these fields is blocked until you add them. The settings in this section apply to any GCP or Azure Helm deployment. For a complete reference, see the [`examples/google-autopilot-cel/values.yaml`](https://github.com/braintrustdata/helm/blob/main/braintrust/examples/google-autopilot-cel/values.yaml) example in the Helm chart, which configures them for GKE Autopilot. ### Configure security contexts Set a container-level `securityContext` on the API and each Brainstore role to enforce a read-only root filesystem, disallow privilege escalation, and drop all Linux capabilities: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: securityContext: readOnlyRootFilesystem: true allowPrivilegeEscalation: false capabilities: drop: - ALL brainstore: reader: securityContext: readOnlyRootFilesystem: true allowPrivilegeEscalation: false capabilities: drop: - ALL fastreader: securityContext: readOnlyRootFilesystem: true allowPrivilegeEscalation: false capabilities: drop: - ALL writer: securityContext: readOnlyRootFilesystem: true allowPrivilegeEscalation: false capabilities: drop: - ALL ``` You can also set an optional pod-level `podSecurityContext` (for example, `api.podSecurityContext` or `brainstore.reader.podSecurityContext`) to apply settings such as `runAsNonRoot` or `fsGroup` to the entire pod. Both contexts are omitted from the rendered manifests unless you set them. ### Provide writable temporary storage When you enable `readOnlyRootFilesystem: true`, the container can no longer write to its root filesystem. Enable a `tmpVolume` for each affected component to mount a writable `emptyDir` at `/tmp`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: tmpVolume: enabled: true sizeLimit: "1Gi" brainstore: reader: tmpVolume: enabled: true sizeLimit: "1Gi" fastreader: tmpVolume: enabled: true sizeLimit: "1Gi" writer: tmpVolume: enabled: true sizeLimit: "1Gi" ``` Any component with `readOnlyRootFilesystem: true` must also set `tmpVolume.enabled: true`. Without a writable `/tmp` volume, processes that write temporary files fail at runtime. ### Set ephemeral storage and cache limits Each Brainstore role accepts an `ephemeralStorage` budget and a `volume.sizeLimit` for its cache `emptyDir`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstore: reader: ephemeralStorage: request: "1000Gi" limit: "1000Gi" volume: size: "1000Gi" sizeLimit: "900Gi" objectStoreCacheFileSize: "900Gi" ``` * `ephemeralStorage.request` sets the pod-local storage budget that GKE Autopilot reserves. `ephemeralStorage.limit` sets the corresponding cap. * `volume.sizeLimit` caps the cache `emptyDir`. Match it to `objectStoreCacheFileSize` so the cache cannot exceed its allotted space. How the ephemeral storage request is applied depends on your cloud and mode: | Deployment | Behavior | | - | - | | GKE Autopilot (`cloud: google`, `google.mode: autopilot`) | `ephemeralStorage.request` is applied. If unset, `volume.size` is used as the fallback request. | | AWS | `ephemeralStorage.request` and `ephemeralStorage.limit` are applied when set. | | GKE Standard and Azure | Ephemeral storage requests are not injected. | # Telemetry and data retention Source: https://braintrust.dev/docs/admin/self-hosting/configure/telemetry Control the telemetry a self-hosted data plane sends to Braintrust's control plane and enable data retention. These settings control the telemetry your self-hosted data plane sends to Braintrust's control plane, and the service token that enables data retention. ## Enable or disable telemetry Braintrust can send the following types of telemetry from your self-hosted data plane to Braintrust's control plane: | Type | Description | | - | - | | `status` | Health check information (enabled by default) | | `metrics` | System metrics (CPU/memory) and Braintrust-specific metrics like indexing lag (enabled by default) | | `usage` | Billing usage telemetry for aggregate usage metrics (enabled by default) | | `memprof` | Memory profiling statistics and heap usage patterns | | `logs` | Application logs | | `traces` | Distributed tracing data | By default, `status`, `metrics`, and `usage` are enabled. You can change the defaults as follows: The [Loop runtime](/docs/admin/self-hosting/configure/loop-runtime) always sends `metrics` and `traces` for its own service, in addition to the types you configure here. You can't turn this off while the runtime is enabled. These traces cover the runtime's operation. They don't include LLM calls, tool calls, or the contents of your traces. Add the `monitoring_telemetry` variable to your `variables.tf` file, and include the types of telemetry you want to send in the validation condition as a comma-separated list: ```bash {19} theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} variable "monitoring_telemetry" { description = <<-EOT The telemetry to send to Braintrust's control plane to monitor your deployment. Should be in the form of comma-separated values. Available options: - status: Health check information (default) - metrics: System metrics (CPU/memory) and Braintrust-specific metrics like indexing lag (default) - usage: Billing usage telemetry for aggregate usage metrics - memprof: Memory profiling statistics and heap usage patterns - logs: Application logs - traces: Distributed tracing data EOT type = string default = "status,metrics,usage" validation { condition = var.monitoring_telemetry == "" || alltrue([ for item in split(",", var.monitoring_telemetry) : contains(["metrics", "logs", "traces", "status", "memprof", "usage"], trimspace(item)) ]) error_message = "The monitoring_telemetry value must be a comma-separated list containing only: metrics, logs, traces, status, memprof, usage." } } ``` Update the `controlPlaneTelemetry` setting in your Helm `values.yaml` file to include the types of telemetry you want to send: ```bash {10} theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} # Global configs global: orgName: "" # When createNamespace is true, the namespace will be created and resources will be in global.namespace # When createNamespace is false, resources will use .Release.Namespace (the namespace specified during helm install/upgrade) createNamespace: false namespace: "braintrust" namespaceAnnotations: {} labels: {} controlPlaneTelemetry: "status,metrics,usage,logs,traces,memprof" ``` Braintrust also has access to endpoints reporting metrics about the backfill and compaction status of Brainstore segments. This is metadata only, no customer data. To disable these endpoints, set the `DISABLE_SYSADMIN_TELEMETRY` environment variable to `true`. If you disable telemetry, Braintrust's ability to proactively monitor your deployment and diagnose issues will be significantly limited. Before disabling, consider the impact on support response times. ## Data retention Data retention requires a [service token](/docs/admin/access-control/manage-permissions#use-service-accounts) so the data plane can query object metadata and look up retention policies configured in your organization. Braintrust automatically provisions this token when you configure your self-hosted data plane URL in [organization settings](/docs/admin/organizations#configure-api-urls-self-hosted). It is created with read-only permissions on projects and stored securely in your data plane. To verify or refresh the token, go to ** Settings** > [** Service tokens**](https://www.braintrust.dev/app/~/configuration/org/service-tokens). If the token doesn't exist, click **Create**. To rotate it, click **Refresh** — the data plane will start using the new token automatically. Before configuring retention, review your cloud provider's bucket retention policies. See [Cloud provider retention policies](/docs/admin/data-management/retention#cloud-provider-policies) for details. # Deploy Braintrust Source: https://braintrust.dev/docs/admin/self-hosting/deploy Deploy a self-hosted Braintrust data plane on AWS, GCP, or Azure Deploy the Braintrust data plane in your AWS account using the Braintrust [Terraform module](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane). This is the recommended way to self-host Braintrust on AWS. Braintrust recommends deploying in a dedicated AWS account. AWS enforces account-level [Lambda concurrency limits](https://docs.aws.amazon.com/lambda/latest/dg/lambda-concurrency.html), and Braintrust runs Lambda functions for user-defined functions (scorers and tools) and for the AI proxy that serves their LLM calls. Sharing an account with other workloads can lead to throttling and service disruptions. A dedicated account also aligns with AWS best practices for workload isolation and security. To test infrastructure provisioning before committing to production-sized resources, use the [sandbox example](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane/tree/main/examples/braintrust-data-plane-sandbox). It uses minimal instance sizes and has deletion protection disabled for easy teardown. It is not suitable for performance or load testing. ## 1. Configure the Terraform module The Braintrust [Terraform module](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane) contains all the necessary resources for a self-hosted Braintrust data plane. 1. Copy the entire contents of the [`examples/braintrust-data-plane`](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane/tree/main/examples/braintrust-data-plane) directory from the [terraform-aws-braintrust-data-plane](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane) repository into your own repository. 2. In `provider.tf`, configure your AWS account and region. Supported regions: * `ap-northeast-1`, `ap-south-1`, `ap-southeast-2` * `ca-central-1` * `eu-central-1`, `eu-west-1`, `eu-west-2`, `eu-west-3` * `sa-east-1` * `us-east-1`, `us-east-2`, `us-west-2` If you require support for a different region, [contact Braintrust](mailto:support@braintrust.dev). 3. In `terraform.tf`, set up your remote backend (typically S3 and DynamoDB). 4. In `main.tf`, append `?ref=` with a version tag to the module `source` to pin the module version. Use v6.5.0 or later, the minimum recommended version for new deployments: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v6.7.0" ``` 5. In `main.tf`, set `enable_ecs_api = true` to serve API traffic from ECS. The example configuration does not set this variable, which leaves API traffic on the Lambda path. ECS is the recommended runtime and costs about 90% less than Lambda for high-traffic data planes. To size the ECS services, see [Scaling and storage](/docs/admin/self-hosting/configure/scaling#configure-aws-api-ecs-services). 6. In `main.tf`, customize the remaining Braintrust deployment settings. The defaults are suitable for a large production-sized deployment. Adjust them based on your needs, but keep in mind the [hardware requirements](/docs/admin/self-hosting/index#hardware-requirements). Two settings fail the apply if they are wrong: * `deployment_name` must be unique within the same AWS account (max 18 characters). The default is `"braintrust"`, so change it if you have multiple deployments. Resource names (IAM roles, RDS instances, S3 buckets) are prefixed with this value and collide if duplicated. * Brainstore instance types must have **local NVMe storage** for caching (for example, the `c8gd`, `c5d`, `m5d`, `i3`, and `i4i` families). Generic instance types without local storage (`t3`, `m5`, `c5`) are not supported and fail at plan time. To keep LLM requests and cached completions on your own infrastructure, you can also run the Braintrust Gateway in your data plane. See [Deploy the Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) for the prerequisites and values. ## 2. Initialize AWS account If you're using a new AWS account, run the [`create-service-linked-roles.sh`](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane/blob/main/scripts/create-service-linked-roles.sh) script to create all necessary IAM service-linked roles for the deployment: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} ./scripts/create-service-linked-roles.sh ``` ## 3. Configure Brainstore license Your deployment includes Brainstore, a high-performance query engine for real-time trace ingestion. Brainstore requires a license key. 1. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. If you don't see your data plane configuration, [contact Braintrust](mailto:support@braintrust.dev) to enable self-hosting. 2. Copy your Brainstore license. 3. Pass the key to Terraform. The recommended approach is to store the license key in AWS Secrets Manager and reference it using a Terraform data source: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} data "aws_secretsmanager_secret_version" "brainstore_license" { secret_id = "braintrust/brainstore-license-key" } ``` Then pass `data.aws_secretsmanager_secret_version.brainstore_license.secret_string` as the `brainstore_license_key` value in the module. Alternatively, you can pass the key without storing it in Secrets Manager: * Set `TF_VAR_brainstore_license_key=your-key` in your environment. * Pass it via command line: `terraform apply -var 'brainstore_license_key=your-key'`. * Add it to an uncommitted `terraform.tfvars` or `.auto.tfvars` file. Do not commit the license key to your git repository. ## 4. Deploy the module Initialize and apply the Terraform configuration: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform init terraform apply ``` The first `terraform apply` may fail with transient errors such as ASG health check timeouts (while instances are still booting) or Lambda rate limits. Re-running `terraform apply` resolves these. This will create all necessary AWS resources including: * Two isolated VPCs: * **Main VPC**: Hosts Braintrust services (API, database, Redis, Brainstore) * **Quarantine VPC**: Runs user-defined functions (scorers, tools) in network isolation. This creates \~30 Lambda functions across multiple runtimes. This is required for most production use cases. * ECS services and Lambda functions for the Braintrust API * Public CloudFront endpoint, API Gateway, and an internal Application Load Balancer * EC2 Auto-scaling group for Brainstore * PostgreSQL database, Redis cache, and S3 buckets * KMS key for encryption With `enable_ecs_api = true`, CloudFront routes API traffic to the internal load balancer in front of the ECS services. The API Gateway and API Lambda functions are still provisioned and kept warm for rollback. A future module release removes them. ## 5. Get your API URL After the deployment completes, get your API URL from the Terraform outputs: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform output ``` You should see output similar to: ``` api_url = "https://dx6atff6gocr6.cloudfront.net" ``` Save this URL. You'll need it to configure your Braintrust organization. ## 6. Configure your organization Connect your Braintrust organization to your newly deployed data plane. Changing your live organization's API URL can disrupt access for existing users. If you are testing, create a new Braintrust organization for your data plane instead of updating your live environment. 1. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. 2. In **API URL** area, select **Edit**. 3. Enter the API URL from the last step. 4. Leave the other fields blank. 5. If your deployment is accessed through a VPN or is otherwise on a private network (not accessible from the public internet), enable **Data plane is on a private network**. This enables Chrome's Local Network Access permission handling, which is required for browser access to private network resources. When enabled, Chrome will prompt users to grant permission for the Braintrust UI to access your self-hosted data plane. See [Grant browser permissions](/docs/admin/self-hosting/configure/security#grant-browser-permissions) for details. 6. Select **Save**. The UI will automatically test the connection to your new data plane. Verify that the ping to each endpoint is successful. ## Debug issues If you encounter issues, you can use the [`dump-logs.sh`](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane/blob/main/scripts/dump-logs.sh) script to collect logs: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} ./scripts/dump-logs.sh [--minutes N] [--service ] ``` For example, to dump 60 minutes of logs for the `bt-sandbox` deployment, run: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} ./scripts/dump-logs.sh bt-sandbox ``` This will save logs for all services to a `logs-` directory, which you can share with the Braintrust team for debugging. ## Customize the infrastructure These options control what the Terraform module provisions in your AWS account, such as deploying into an existing VPC, KMS encryption keys, and resource tags. To tune how a running deployment behaves, including access control, connectivity, scaling, and telemetry, see [Configure your deployment](/docs/admin/self-hosting/configure). ### Use an existing VPC To deploy into an existing VPC instead of creating a new one, set `create_vpc = false` and provide your VPC and subnet IDs: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" create_vpc = false existing_vpc_id = "vpc-xxxxxxxxx" existing_private_subnet_1_id = "subnet-xxxxxxxxx" existing_private_subnet_2_id = "subnet-xxxxxxxxx" existing_private_subnet_3_id = "subnet-xxxxxxxxx" existing_public_subnet_1_id = "subnet-xxxxxxxxx" # ... other configuration ... } ``` Your existing VPC must have: * At least 3 private subnets across different availability zones * At least 1 public subnet * Internet and NAT gateways with properly configured route tables The module manages its own security groups. To also use an existing quarantine VPC, set `existing_quarantine_vpc_id` and the corresponding `existing_quarantine_private_subnet_*_id` variables. ### Use custom tags To apply custom tags to all resources, pass the `custom_tags` parameter to the Braintrust module: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" custom_tags = { Environment = "production" Team = "ml-platform" CostCenter = "engineering" } # ... other configuration ... } ``` These tags will be applied to all resources including Brainstore EC2 instances, volumes, and ENIs. The deployment name variable automatically prefixes resource names and applies a `BraintrustDeploymentName` tag across all resources. Use the `custom_tags` parameter instead of the AWS provider's `default_tags` configuration. Due to a Terraform limitation, `default_tags` are not applied to resources that use launch templates, such as Brainstore instances. Terraform module v6.2.0 and later reserve the `aws-apn-id` tag key, setting it to Braintrust's AWS Partner Network identifier on every resource and overriding any `aws-apn-id` value in `custom_tags`. If your account restricts tag keys with an SCP or IAM policy, allow `aws-apn-id`, or resource creation and modification will fail. ### Redis instance sizing **Important for AWS**: Avoid using burstable Redis instances (t-family instances like `cache.t4g.micro`) in production. These instances use CPU credits that can be exhausted during high-load periods, leading to performance throttling. Instead, use non-burstable instances like `cache.r7g.large`, `cache.r6g.medium`, or `cache.r5.large` for predictable performance. Even if these instances seem oversized initially, they provide consistent performance without the risk of CPU credit exhaustion. Changes to the Redis node type or engine version normally wait for the next ElastiCache maintenance window. Set `redis_apply_immediately` to `true` (Terraform module v6.5.2 or later) to apply them as soon as you run `terraform apply` instead. Applying immediately can briefly interrupt connections to Redis, so prefer the maintenance window for production deployments. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" redis_apply_immediately = true # default false # ... other configuration ... } ``` ### Lambda memory limits The API Handler and AI Proxy Lambda functions default to 10240 MB (the Lambda maximum). You can reduce these to lower costs in environments with tighter memory quotas, though Braintrust recommends keeping the defaults for production workloads. Which variable matters depends on where API traffic runs: * `api_handler_memory_limit` applies only while API traffic runs on the Lambda path. With `enable_ecs_api = true`, size the ECS services instead. See [Scaling and storage](/docs/admin/self-hosting/configure/scaling#configure-aws-api-ecs-services). * `ai_proxy_memory_limit` applies even with API traffic on ECS. As of Terraform module v6.5.0, LLM calls from user-authored scorers and tools running in the quarantine environment route through the AI Proxy Lambda. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" api_handler_memory_limit = 10240 # default, valid range 1–10240 MB ai_proxy_memory_limit = 10240 # default, valid range 1–10240 MB # ... other configuration ... } ``` ### RDS backup and maintenance windows `postgres_backup_window` and `postgres_maintenance_window`, added in Terraform module v6.5.1, shift the RDS backup and maintenance windows to times that avoid peak traffic for your deployment. They default to `"00:00-00:30"` for daily backups and `"Mon:08:00-Mon:11:00"` for weekly maintenance, both in UTC. AWS requires that the two windows not overlap. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" postgres_backup_window = "02:00-02:30" # Format: hh24:mi-hh24:mi (UTC) postgres_maintenance_window = "Tue:08:00-Tue:11:00" # Format: ddd:hh24:mi-ddd:hh24:mi (UTC) # ... other configuration ... } ``` ### S3 access logging To capture S3 server access logs for the Brainstore, code bundle, and Lambda responses buckets, set `s3_server_access_logging` (Terraform module v6.5.2 or later). Logs are delivered to a destination bucket you own, which is commonly required for audit and compliance. The variable defaults to `null`, which disables logging. Amazon S3 delivers server access logs on a [best-effort basis](https://docs.aws.amazon.com/AmazonS3/latest/userguide/ServerLogs.html). A record can arrive long after the request it describes, never arrive at all, or be delivered more than once. Use these logs to understand access patterns rather than as a complete accounting of every request. Access logging also adds storage cost. Brainstore is request-heavy, so its bucket can produce a large volume of log records in a busy deployment. Log objects are billed at standard S3 rates, so set a lifecycle policy on the destination bucket to expire or transition them. Enable logging in this order. The module configures only the source buckets, so it cannot create or depend on the destination bucket policy for you. 1. Create the destination bucket in the same AWS account and region as the data plane. It must not have Object Lock or Requester Pays enabled, and its default encryption must be SSE-S3 (AES256). SSE-KMS prevents Amazon S3 from delivering logs you can decrypt. 2. Deploy the data plane, or use an existing deployment, so that the source bucket name outputs are available. 3. Attach a bucket policy on the destination bucket that grants `s3:PutObject` to `logging.s3.amazonaws.com`. 4. Set `s3_server_access_logging` and apply again. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane" s3_server_access_logging = { bucket = "your-audit-logs-bucket" # Replace with your destination bucket prefix = "braintrust/" } # ... other configuration ... } ``` | Field | Required | Description | | - | - | - | | `bucket` | yes | Name of the destination bucket to write access logs to. Terraform does not create it. Must be non-empty. | | `prefix` | no | Prefix applied to log object keys. Must be empty or end with `/`. Defaults to `/`. | The `brainstore`, `code-bundle`, and `lambda-responses` suffixes are appended under the prefix, so a single policy statement covering the prefix is enough. Logs use a date-partitioned key format: ``` //////
/... ``` For example, `braintrust/brainstore////2026/08/12/...`. First delivery can take a few hours after you enable logging. Attach the destination bucket policy before you set `s3_server_access_logging`. The `Resource` path in the policy must match the prefix you configure, and any `Deny` statements on the destination bucket must not block log delivery. Access logs can include object keys and requester information — restrict read access to the destination bucket accordingly. Use the `brainstore_s3_bucket_name`, `code_bundle_s3_bucket_name`, and `lambda_responses_s3_bucket_name` outputs to scope the policy to your source buckets: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} data "aws_caller_identity" "current" {} data "aws_iam_policy_document" "s3_server_access_logs" { statement { sid = "S3ServerAccessLogsPolicy" effect = "Allow" principals { type = "Service" identifiers = ["logging.s3.amazonaws.com"] } actions = ["s3:PutObject"] resources = ["arn:aws:s3:::your-audit-logs-bucket/braintrust/*"] # Replace "braintrust/" with your configured prefix condition { test = "ArnLike" variable = "aws:SourceArn" values = [ "arn:aws:s3:::${module.braintrust-data-plane.brainstore_s3_bucket_name}", "arn:aws:s3:::${module.braintrust-data-plane.code_bundle_s3_bucket_name}", "arn:aws:s3:::${module.braintrust-data-plane.lambda_responses_s3_bucket_name}", ] } condition { test = "StringEquals" variable = "aws:SourceAccount" values = [data.aws_caller_identity.current.account_id] } } } resource "aws_s3_bucket_policy" "s3_server_access_logs" { bucket = "your-audit-logs-bucket" # Replace with your destination bucket policy = data.aws_iam_policy_document.s3_server_access_logs.json } ``` ### WAL footer version The `brainstore_wal_footer_version` variable controls the WAL footer format written by Brainstore. It defaults to `""` (unset) and should not be changed outside of a planned upgrade sequence. Do not set `brainstore_wal_footer_version` without following the [upgrade guide](/docs/admin/self-hosting/upgrade/v2). Setting it at the same time as a version bump can cause Brainstore nodes still rolling out to fail to read the new WAL format. See [Enable efficient WAL format](/docs/admin/self-hosting/upgrade/v2#enable-efficient-wal-format-aws) in the v2.0 upgrade guide for the correct migration steps. ### KMS encryption When `kms_key_arn` is configured, all managed S3 buckets (Brainstore, code-bundle, and Lambda responses) enforce `blocked_encryption_types = ["NONE"]`, preventing unencrypted object uploads. This policy is applied automatically as of v4.5.0 — upgrading from an earlier version will include this change in your `terraform plan`. ### AI Proxy CORS headers As of v4.5.0, the `x-bt-use-gateway` header is included in the AI Proxy Lambda function URL CORS allowed headers. Browser clients can send this header without triggering a CORS preflight rejection. When the [Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) is enabled, requests route through it by default, and `x-bt-use-gateway: false` sends an individual request straight to the provider instead. As of v6.7.0, [`x-bt-endpoint-name`](/docs/deploy/ai-proxy#advanced-configuration) is allowed as well, so browser clients can select a specific configured endpoint without a preflight rejection. Deploy the Braintrust data plane in your GCP project using the Braintrust [Terraform module](https://github.com/braintrustdata/terraform-google-braintrust-data-plane) and [Helm chart](https://github.com/braintrustdata/helm). This is the recommended way to self-host Braintrust on GCP. ## 1. Configure the Terraform module The Braintrust [Terraform module](https://github.com/braintrustdata/terraform-google-braintrust-data-plane) contains all the necessary resources for a self-hosted Braintrust data plane. A dedicated Google Cloud project for your Braintrust deployment is recommended but not required. 1. Copy the entire contents of the [`examples/braintrust-data-plane`](https://github.com/braintrustdata/terraform-google-braintrust-data-plane/tree/main/examples/braintrust-data-plane) directory from the [terraform-google-braintrust-data-plane](https://github.com/braintrustdata/terraform-google-braintrust-data-plane) repository into your own repository. 2. In `provider.tf`, configure your Google Cloud project and region. 3. In `backend.tf`, set up your remote backend (typically a GCS bucket). 4. In `main.tf`, customize the Braintrust deployment settings. The defaults are suitable for a large production-sized deployment. Adjust them based on your needs, but keep in mind the [hardware requirements](/docs/admin/self-hosting/index#hardware-requirements). ## 2. Enable Google Cloud APIs Before deploying, enable the required Google Cloud services. Run the following in Cloud Shell: 1. In a Cloud Shell, set the project to deploy Braintrust into: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} gcloud config set project "your-project-id" ``` 2. Enable the required services: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} gcloud services enable storage-api.googleapis.com \ storage-component.googleapis.com \ storage.googleapis.com \ redis.googleapis.com \ secretmanager.googleapis.com \ servicenetworking.googleapis.com \ logging.googleapis.com \ monitoring.googleapis.com \ oslogin.googleapis.com \ dns.googleapis.com \ cloudresourcemanager.googleapis.com \ compute.googleapis.com \ cloudkms.googleapis.com \ autoscaling.googleapis.com \ iam.googleapis.com \ iamcredentials.googleapis.com \ vpcaccess.googleapis.com \ sts.googleapis.com \ container.googleapis.com \ sqladmin.googleapis.com \ artifactregistry.googleapis.com ``` Allow approximately 5 minutes for services to activate. ## 3. Deploy the Terraform module Initialize and apply the Terraform configuration: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform init terraform apply ``` This will create all necessary GCP resources including: * GKE cluster for running Braintrust services * Cloud SQL PostgreSQL database * Cloud Memorystore Redis cache * Cloud Storage buckets * VPC network and subnets * Cloud KMS key for encryption ## 4. Set up Kubernetes After the Terraform deployment completes, connect to your GKE cluster: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} gcloud auth login gcloud config set project "" gcloud container clusters get-credentials -gke-autopilot --region ``` Verify cluster connectivity: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl cluster-info ``` Create the namespace for Braintrust: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl create namespace braintrust ``` ## 5. Create Kubernetes secrets Create the required Kubernetes secrets for your deployment. The secrets needed depend on your GCS authentication method: * **API**: Can use either native GCS authentication (recommended) or S3 compatibility mode with HMAC keys (legacy) * **Brainstore**: Always uses native GCS authentication via Workload Identity See [GCS authentication options](#gcs-authentication-options) below for the specific commands. ## 6. Deploy Helm chart Create a `helm-values.yaml` file for your deployment. Refer to the [Helm chart documentation](https://github.com/braintrustdata/helm) for configuration options. Deploy the Braintrust Helm chart to your cluster: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm install braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version \ --values helm-values.yaml ``` See all Helm chart releases: [GitHub Releases](https://github.com/braintrustdata/helm/releases) To keep LLM requests and cached completions on your own infrastructure, you can also run the Braintrust Gateway in your data plane. See [Deploy the Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) for the prerequisites and values. ## 7. Configure Ingress (HTTPS) The data plane requires a publicly reachable HTTPS endpoint with a valid TLS certificate. The Helm chart deploys the API as a ClusterIP service. You are expected to provide your own ingress solution that terminates TLS and routes traffic to the braintrust-api service on port 8000. Common approaches: * GCP Application Load Balancer with a Google-managed certificate (requires a custom domain) * GKE Gateway API with cert-manager and Let's Encrypt * Cloud Run NGINX proxy with a VPC Connector for SSL termination (no custom domain required) * Istio/ASM Gateway - the Helm chart includes native VirtualService support (see virtualService in values.yaml) * Any reverse proxy or load balancer that terminates TLS and forwards HTTP to the API service After configuring your ingress, save the resulting HTTPS URL. You'll need it to configure your Braintrust organization. If your ingress uses a private or self-signed certificate, pass the CA bundle to the `bt` CLI using the `--ca-cert ` flag or the `BRAINTRUST_CA_CERT` environment variable so that bt commands can connect to your data plane. ## 8. Configure your organization Connect your Braintrust organization to your newly deployed data plane. Changing your live organization's API URL can disrupt access for existing users. If you are testing, create a new Braintrust organization for your data plane instead of updating your live environment. 1. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. 2. In **API URL** area, select **Edit**. 3. Enter the API URL from the last step. 4. Leave the other fields blank. 5. If your deployment is accessed through a VPN or is otherwise on a private network (not accessible from the public internet), enable **Data plane is on a private network**. This enables Chrome's Local Network Access permission handling, which is required for browser access to private network resources. When enabled, Chrome will prompt users to grant permission for the Braintrust UI to access your self-hosted data plane. See [Grant browser permissions](/docs/admin/self-hosting/configure/security#grant-browser-permissions) for details. 6. Select **Save**. The UI will automatically test the connection to your new data plane. Verify that the ping to each endpoint is successful. ## GCS authentication options Braintrust services use different authentication methods for Google Cloud Storage: * **API**: Can use either native GCS authentication (recommended) or S3 compatibility mode with HMAC keys (legacy) * **Brainstore**: Always uses native GCS authentication via Workload Identity ### Workload Identity setup The Terraform module automatically configures Workload Identity for your GKE cluster and creates two service accounts with the following IAM grants. The module creates two GCS buckets: * **Brainstore bucket** (`-brainstore-*`): Brainstore data storage. * **API bucket** (`-api-*`): Contains two storage paths — `code-bundle/` (API layer writes) and `brainstore-cache/` (ephemeral Brainstore cache, automatically deleted after 1 day by a GCS lifecycle rule). The `brainstore-cache/` objects are ephemeral and managed automatically. Operators do not need to manage or back up this data. | Service account | Role | Resource | Purpose | | - | - | - | - | | Braintrust API | `roles/storage.objectAdmin` | Brainstore bucket, API bucket | Read/write objects | | Braintrust API | `roles/storage.legacyBucketReader` | Brainstore bucket, API bucket | List bucket contents | | Braintrust API | `roles/iam.serviceAccountTokenCreator` | Itself | Generate short-lived tokens and signed GCS URLs | | Brainstore | `roles/storage.objectAdmin` | Brainstore bucket, API bucket | Read/write objects | | Brainstore | `roles/storage.legacyBucketReader` | Brainstore bucket, API bucket | List bucket contents | ### API authentication configuration Choose one of the following authentication methods for the API service: Native GCS authentication uses the `@google-cloud/storage` SDK and Workload Identity. This is the recommended approach for enhanced security as it eliminates the need to manage service account keys. **Requirements:** * Helm chart version 3.1.0 or later * Workload Identity configured (automatic with Terraform module) **Kubernetes secrets:** For native GCS authentication, create secrets without GCS credentials: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl create secret generic braintrust-secrets \ --from-literal=REDIS_URL="" \ --from-literal=PG_URL="" \ --from-literal=FUNCTION_SECRET_KEY="" \ --from-literal=BRAINSTORE_LICENSE_KEY="" \ --namespace=braintrust ``` Refer to the Terraform outputs for the connection strings. The Brainstore license key can be found at ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. **Helm configuration:** In your Helm values file, enable native GCS authentication and configure the Google service account: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: serviceAccount: googleServiceAccount: "prod-braintrust@your-project-id.iam.gserviceaccount.com" enableGcsAuth: true ``` The Terraform module outputs the service account email as `braintrust_service_account`. Run `terraform output braintrust_service_account` to get the full email address. The `enableGcsAuth` setting defaults to `false` for backwards compatibility. Contact Braintrust if you want to enable native GCS authentication by default. S3 compatibility mode uses HMAC keys to access GCS through the S3-compatible API. This is the legacy authentication method. **Kubernetes secrets:** For S3 compatibility mode, include HMAC credentials: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl create secret generic braintrust-secrets \ --from-literal=REDIS_URL="" \ --from-literal=PG_URL="" \ --from-literal=GCS_ACCESS_KEY_ID="" \ --from-literal=GCS_SECRET_ACCESS_KEY="" \ --from-literal=FUNCTION_SECRET_KEY="" \ --from-literal=BRAINSTORE_LICENSE_KEY="" \ --namespace=braintrust ``` Refer to the Terraform outputs for the connection strings. The Brainstore license key can be found at ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. **Helm configuration:** In your Helm values file, ensure `enableGcsAuth` is not set or set to `false`. Do not configure `googleServiceAccount` when using HMAC keys: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: enableGcsAuth: false # or omit this line ``` Helm chart v5.0.1+ automatically sets `AWS_REQUEST_CHECKSUM_CALCULATION` and `AWS_RESPONSE_CHECKSUM_VALIDATION` to `WHEN_REQUIRED` when `enableGcsAuth` is disabled. This ensures AWS SDK compatibility with the GCS S3-compatible endpoint. No additional configuration is required. ### Tune GCS retry behavior (optional) When `api.enableGcsAuth: true`, the API service uses the `@google-cloud/storage` SDK and applies the SDK's default retry policy for transient GCS errors. To override these defaults, set any of the following environment variables on the API container. If none are set, the SDK defaults apply. See [Cloud Storage retry strategy](https://docs.cloud.google.com/storage/docs/retry-strategy) for the upstream defaults and the Node.js client behavior. | Environment variable | Type | Description | | - | - | - | | `GCS_RETRY_OPTIONS_AUTO_RETRY` | boolean | Enable or disable automatic retries. | | `GCS_RETRY_OPTIONS_MAX_RETRIES` | integer | Maximum number of retry attempts per request. | | `GCS_RETRY_OPTIONS_RETRY_DELAY_MULTIPLIER` | float | Exponential backoff multiplier between retries. | | `GCS_RETRY_OPTIONS_TOTAL_TIMEOUT` | integer (seconds) | Total time allowed across all retries for a request. | | `GCS_RETRY_OPTIONS_MAX_RETRY_DELAY` | integer (seconds) | Maximum delay between consecutive retries. | | `GCS_RETRY_OPTIONS_IDEMPOTENCY_STRATEGY` | string | One of `retry-always`, `retry-conditional`, or `retry-never`. | Set these through `api.extraEnvVars` in your Helm values file: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: enableGcsAuth: true extraEnvVars: - name: GCS_RETRY_OPTIONS_MAX_RETRIES value: "5" - name: GCS_RETRY_OPTIONS_TOTAL_TIMEOUT value: "120" - name: GCS_RETRY_OPTIONS_IDEMPOTENCY_STRATEGY value: "retry-conditional" ``` Requires data plane v2.0.0 or later. These environment variables are read only when `api.enableGcsAuth: true` and have no effect in S3 compatibility mode. They apply to the API service only. Brainstore manages its GCS client independently. ### Brainstore configuration Brainstore always uses native GCS authentication. In your Helm values file, configure the Google service account for Workload Identity: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstore: serviceAccount: googleServiceAccount: "prod-brainstore@your-project-id.iam.gserviceaccount.com" ``` The Terraform module outputs the service account email as `brainstore_service_account`. Run `terraform output brainstore_service_account` to get the full email address. ## Customize the infrastructure These options control what the Terraform module and Helm chart provision, such as cloud identity, resource labels, and cluster IP ranges. To tune how a running deployment behaves, including access control, connectivity, scaling, and telemetry, see [Configure your deployment](/docs/admin/self-hosting/configure). ### Service account impersonation If Brainstore needs to access GCS buckets or other GCP resources in another project that are restricted to a specific service account identity, use `brainstore_impersonation_targets` to grant the Brainstore Kubernetes service account the ability to impersonate one or more Google Cloud service accounts. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" # ...other variables... brainstore_impersonation_targets = [ "projects/my-other-project/serviceAccounts/data-access@my-other-project.iam.gserviceaccount.com" ] } ``` This grants `roles/iam.serviceAccountTokenCreator` on each target service account to the Brainstore Kubernetes service account, enabling Brainstore to generate short-lived tokens and act as those accounts. Values must use the full resource name format `projects/{project_id}/serviceAccounts/{service_account_email}`, not bare email addresses. The default is `[]` (no impersonation). This variable is only needed if you are not separately granting the Brainstore service account IAM access to the target accounts. The Terraform executor must have `roles/iam.serviceAccountAdmin` or `roles/resourcemanager.projectIamAdmin` on each target service account (or equivalent project-level permissions). Target service accounts must also exist before running `terraform apply`. ### Custom labels To apply user-defined GCP labels to all resources created by the module, use the `custom_labels` variable: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" custom_labels = { environment = "production" team = "platform" cost-center = "ai-infra" } # ... other configuration ... } ``` Labels are applied to Cloud SQL, Memorystore Redis, Cloud Storage buckets, the GKE cluster, and the KMS key. The module's built-in `braintrustdeploymentname` label is always preserved when merged with your custom labels. Keys must start with a lowercase letter and contain only lowercase letters, numbers, underscores, or dashes (max 63 characters). Values must contain only lowercase letters, numbers, underscores, or dashes, do not require a leading letter, and may be empty (max 63 characters). Maximum 63 custom labels. ### Private Service Access range When the module creates the VPC (`create_vpc = true`), Cloud SQL and Memorystore connect over a Private Service Access range. By default, the module allocates a `/16` range and lets Google pick the starting address. To avoid overlap with existing VPCs, peering connections, or corporate networks, use `private_service_access_prefix_length` and `private_service_access_address`: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" private_service_access_prefix_length = 16 private_service_access_address = "10.10.0.0" # ... other configuration ... } ``` * `private_service_access_prefix_length` accepts a value from 8 through 24 (default `16`). Lower prefix lengths create larger ranges. Smaller ranges work, but reduce future expansion headroom, so choose the size with your Braintrust architecture team. * `private_service_access_address` sets the starting address of the range. If unset, Google selects an available range. Choose these values before your first deployment. Changing the Private Service Access range later can require rebuilding Cloud SQL, Memorystore, and other dependent resources. ### GKE Pod and Service IP ranges To control the secondary IP ranges that the GKE cluster uses for Pods and Services, set either a CIDR block or the name of a pre-existing secondary range. By default, GKE allocates these ranges automatically. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" # Option A: let the module create the ranges from CIDR blocks gke_pods_ipv4_cidr_block = "10.20.0.0/20" gke_services_ipv4_cidr_block = "10.30.0.0/22" # Option B: reference secondary ranges that already exist on the subnet # gke_pods_secondary_range_name = "braintrust-pods" # gke_services_secondary_range_name = "braintrust-services" # ... other configuration ... } ``` * The CIDR variables accept a full CIDR (`"10.20.0.0/20"`) or a netmask size only (`"/20"`). * The secondary range name variables reference ranges that already exist on the subnet. They do not create the ranges. * For each range type (Pods and Services), set the CIDR block or the secondary range name, but not both. The module rejects setting both forms for the same range type. * Most deployments should leave the Service range unset unless they intentionally need a custom Service CIDR. Choose these values before your first deployment. Changing GKE secondary ranges later requires recreating the cluster. ### VPC subnet flow logs When the module creates the VPC (`create_vpc = true`), you can enable VPC flow logs on the subnet by setting `subnet_flow_logs_config`. Each field is optional and takes the default listed below when omitted. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" subnet_flow_logs_config = { aggregation_interval = "INTERVAL_1_MIN" flow_sampling = 0.1 metadata = "INCLUDE_ALL_METADATA" } # ... other configuration ... } ``` | Field | Type | Accepted values | Description | | - | - | - | - | | `aggregation_interval` | optional `string` | `INTERVAL_5_SEC`, `INTERVAL_30_SEC`, `INTERVAL_1_MIN`, `INTERVAL_5_MIN`, `INTERVAL_10_MIN`, `INTERVAL_15_MIN` | How long to aggregate flows before writing a log entry. Longer intervals reduce log volume for long-lived connections. Defaults to `INTERVAL_5_SEC`. | | `flow_sampling` | optional `number` | 0 to 1 inclusive | Sampling rate for collected flow logs. `1.0` reports all of them and `0.0` reports none. Defaults to `0.5`, which reports half. | | `metadata` | optional `string` | `INCLUDE_ALL_METADATA`, `EXCLUDE_ALL_METADATA`, `CUSTOM_METADATA` | Which metadata fields to add to reported logs. Defaults to `INCLUDE_ALL_METADATA`. | | `metadata_fields` | optional `list(string)` | — | Specific metadata fields to include. Set this only when `metadata` is `CUSTOM_METADATA`, where it is required and must contain at least one field. | | `filter_expr` | optional `string` | — | [CEL (Common Expression Language)](https://cel.dev) expression defining which flow logs to export. Defaults to `true`, which includes everything. | `subnet_flow_logs_config` can only be set when `create_vpc = true`. To enable flow logs on a subnet you supply yourself, configure them directly on that subnet. ### GCS access logging To capture access logs for the Brainstore or API GCS buckets, set `gcs_brainstore_logging_config`, `gcs_api_logging_config`, or both. Both default to `null` (disabled). #### Create a log bucket Requires Terraform GCP module v1.5.7 or later. Set a bucket's logging configuration to `{}` to enable logging and create a destination bucket. To enable logging for both buckets: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" gcs_brainstore_logging_config = {} gcs_api_logging_config = {} # ... other configuration ... } ``` * **Destination:** Both configurations share one log bucket. The module grants `cloud-storage-analytics@google.com` the `roles/storage.objectCreator` role on it automatically. * **Prefixes:** Log object names use the fixed prefixes `brainstore` and `api`, without a trailing slash. The module ignores `log_object_prefix` when `log_bucket` is omitted or `null`. * **Output:** `access_log_bucket_name` returns the created bucket's name, or `null` when the module does not manage one. Omitting `log_bucket` or setting it to `null` within a logging configuration has the same effect as `{}`. #### Use an existing log bucket Set `log_bucket` to a non-empty bucket name. Existing configurations continue to work unchanged. Before applying, confirm the destination meets Cloud Storage's requirements, or logging fails: * The log bucket already exists. Terraform does not create it. * The log bucket and the bucket it receives logs for are in the same location (region, dual-region, or multi-region). * Both are in the same organization, or the same project if there is no organization. * Both are in the same VPC Service Controls perimeter, if you use them. * The `cloud-storage-analytics@google.com` group has the `roles/storage.objectCreator` role on the log bucket, which lets Cloud Storage write the logs. For example, to send both buckets' logs to an existing destination: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust_data_plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane" gcs_brainstore_logging_config = { log_bucket = "my-access-logs-bucket" log_object_prefix = "brainstore/" } gcs_api_logging_config = { log_bucket = "my-access-logs-bucket" log_object_prefix = "api/" } # ... other configuration ... } ``` Use `log_object_prefix` to customize log object names. If omitted, it defaults to the name of the bucket being logged. See [Set up log delivery](https://docs.cloud.google.com/storage/docs/access-logs#set-up-log-delivery) in the Cloud Storage documentation. Each bucket is configured independently: you can mix destination types or leave one logging configuration as `null` to keep logging disabled for that bucket. Deploy the Braintrust data plane in your Azure subscription using the Braintrust [Terraform module](https://github.com/braintrustdata/terraform-azure-braintrust-data-plane) and [Helm chart](https://github.com/braintrustdata/helm). This is the recommended way to self-host Braintrust on Azure. **Requirements**: Terraform >= 1.10.0 and the azurerm provider \~> 4.0 are required. If you have an existing deployment using azurerm 3.x, run `terraform init -upgrade` before applying and review the [azurerm v4 upgrade guide](https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs/guides/4.0-upgrade-guide). ## 1. Configure the Terraform module The Braintrust [Terraform module](https://github.com/braintrustdata/terraform-azure-braintrust-data-plane) contains all the necessary resources for a self-hosted Braintrust data plane. A dedicated Azure subscription for your Braintrust deployment is recommended but not required. 1. Copy the entire contents of the [`examples/default`](https://github.com/braintrustdata/terraform-azure-braintrust-data-plane/tree/main/examples/default) directory from the [terraform-azure-braintrust-data-plane](https://github.com/braintrustdata/terraform-azure-braintrust-data-plane) repository into your own repository. 2. In `provider.tf`, configure your Azure subscription and tenant details. 3. In `terraform.tf`, set up your remote backend (typically Azure Blob Storage). 4. In `main.tf`, customize the Braintrust deployment settings. The defaults are suitable for a large production-sized deployment. Adjust them based on your needs, but keep in mind the [hardware requirements](/docs/admin/self-hosting/index#hardware-requirements). The module provisions two AKS node pools: * `brainstore` pool (`aks_brainstore_pool_vm_size`): runs Brainstore pods. Must be a VM SKU with local NVMe SSD (e.g. `Standard_D32ds_v6`). The Azure Container Storage extension is automatically installed to configure RAID0 across the local disks. * `services` pool (`aks_services_pool_vm_size`): runs API and other application pods. Does not require local SSD (e.g. `Standard_D16s_v6`). 5. Initially set `enable_front_door = false` in `main.tf`. You'll enable this later after configuring the load balancer. ## 2. Configure Brainstore license Your deployment requires a Brainstore license key. 1. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. If you don't see your data plane configuration, [contact Braintrust](mailto:support@braintrust.dev) to enable self-hosting. 2. Copy your Brainstore license. 3. Pass the key to Terraform. Do not commit it to your git repository. Recommended options: * Set `TF_VAR_brainstore_license_key=your-key` in your environment before running `terraform apply`. * Pass it on the command line: `terraform apply -var 'brainstore_license_key=your-key'`. * Add it to an uncommitted `terraform.tfvars` or `.auto.tfvars` file. ## 3. Deploy the base infrastructure Initialize and apply the Terraform configuration: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform init terraform apply ``` This will create all necessary Azure resources including: * AKS cluster for running Braintrust services * Azure Database for PostgreSQL * Azure Cache for Redis * Azure Storage Account * Virtual Network * Azure Key Vault for encryption and secrets This deployment typically takes 15-20 minutes. ## 4. Restart PostgreSQL The Terraform module configures PostgreSQL extensions (`pg_cron`, `pg_partman`) and sets `cron.database_name` to the Braintrust database. These are static parameters that require a server restart to take effect. The Terraform provider is configured to not restart the server automatically, since automatic restarts on configuration changes could cause unintended downtime in production. After your first `terraform apply`, restart the PostgreSQL server before proceeding: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} az postgres flexible-server restart \ --resource-group \ --name ``` You can find the resource group and server name in your Terraform outputs. This step is only required on the initial deployment. Subsequent `terraform apply` runs do not require a restart unless you modify the PostgreSQL extension configuration. ## 5. Connect to AKS cluster After the Terraform deployment completes, connect to your AKS cluster: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} az aks get-credentials --resource-group braintrust --name braintrust-aks ``` Verify cluster connectivity: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl cluster-info ``` ## 6. Deploy Helm chart Create a `helm-values.yaml` file for your deployment. Populate it with values from your Terraform outputs: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform output workload_identity_client_id terraform output azure_tenant_id terraform output key_vault_name terraform output storage_account_name ``` Use those values to fill in your `helm-values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} cloud: "azure" global: orgName: "" azure: tenantId: "" enableAzureContainerStorageDriver: true enableAzureKeyVaultDriver: true keyVaultCSIclientID: "" keyVaultName: "" objectStorage: azure: storageAccountName: "" api: annotations: service: service.beta.kubernetes.io/azure-load-balancer-internal: "true" service: type: LoadBalancer ``` Refer to the [Helm chart documentation](https://github.com/braintrustdata/helm) for the full list of configuration options. Deploy the Braintrust [Helm chart](https://github.com/braintrustdata/helm) to your cluster: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm install braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --create-namespace \ --version \ --values helm-values.yaml ``` See all Helm chart releases: [GitHub Releases](https://github.com/braintrustdata/helm/releases) To keep LLM requests and cached completions on your own infrastructure, you can also run the Braintrust Gateway in your data plane. See [Deploy the Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) for the prerequisites and values. ## 7. Enable Front Door 1. Retrieve the load balancer IP address and frontend configuration: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} lb_ip_address=$(kubectl get service braintrust-api -n braintrust \ -o jsonpath='{.status.loadBalancer.ingress[0].ip}') az network lb list \ --query "[?frontendIPConfigurations[?privateIPAddress=='$lb_ip_address']].{Name:name, ResourceGroup:resourceGroup, FrontendIPConfig:frontendIPConfigurations[0].name, Id:frontendIPConfigurations[0].id}" \ -o table ``` 2. Update `main.tf` with the values from the previous step and apply: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} enable_front_door = true front_door_api_backend_address = "" front_door_load_balancer_frontend_ip_config_id = "" ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` 3. In the Azure Portal, find the private link service named `-aks-api-pls` and manually approve it. This manual approval step is an Azure platform requirement — Front Door cannot automatically approve private link connections to resources in a different subscription or tenant. The deployment will appear to succeed but Front Door traffic will not flow until the connection is approved. Front Door deployment takes up to 45 minutes after the Terraform apply completes. Wait for the deployment to finish before proceeding. ## 8. Get your API URL After the Front Door deployment completes, get your API URL from the Terraform outputs: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform output ``` Test the endpoint: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} curl https:/// ``` You should receive a 200 OK response. Save this URL. You'll need it to configure your Braintrust organization. ## 9. Configure your organization Connect your Braintrust organization to your newly deployed data plane. Changing your live organization's API URL can disrupt access for existing users. If you are testing, create a new Braintrust organization for your data plane instead of updating your live environment. 1. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Only [organization owners](/docs/admin/access-control#built-in-permission-groups) can access this page. 2. In **API URL** area, select **Edit**. 3. Enter the API URL from the last step. 4. Leave the other fields blank. 5. If your deployment is accessed through a VPN or is otherwise on a private network (not accessible from the public internet), enable **Data plane is on a private network**. This enables Chrome's Local Network Access permission handling, which is required for browser access to private network resources. When enabled, Chrome will prompt users to grant permission for the Braintrust UI to access your self-hosted data plane. See [Grant browser permissions](/docs/admin/self-hosting/configure/security#grant-browser-permissions) for details. 6. Select **Save**. The UI will automatically test the connection to your new data plane. Verify that the ping to each endpoint is successful. ## Next steps * [Upgrade your deployment](/docs/admin/self-hosting/upgrade) — learn how to keep your data plane up to date * [Configuration](/docs/admin/self-hosting/configure) — configure telemetry, network and URL security, rate limiting, and other options # Self-hosting Braintrust Source: https://braintrust.dev/docs/admin/self-hosting/index only available on the [Enterprise plan](/docs/plans-and-limits#plans). Braintrust offers a self-hosted deployment option that separates data storage from platform management. You deploy and control the infrastructure that stores your sensitive AI data, while Braintrust provides the managed UI, authentication, and platform updates. This gives you full control over your data without the operational overhead of running the entire platform. If you want data residency in your own cloud without operating the data plane yourself, consider [BYOC](/docs/admin/deployment/byoc), where Braintrust operates the deployment. For a comparison of all deployment options, see [Deployment options](/docs/admin/deployment). ## Use cases Self-hosting is designed for organizations with specific requirements: * **Data residency and compliance**: Meet regulatory or contractual obligations by keeping all customer data (experiment logs, traces, datasets, and prompts) within your own cloud account and region. * **Security posture and isolation**: Deploy the data plane behind your firewall or VPN, using your own IAM policies, KMS encryption keys, and audit trails. This ensures sensitive data never traverses external networks. * **Access to private resources**: Connect to internal LLM models, proprietary tools, or private APIs that are not accessible from the public internet. The data plane runs within your network and can access resources in your VPC or private network. ## How it works Braintrust's architecture has two main components: * The **data plane** stores all sensitive data, including experiment records, logs, traces, spans, datasets, and prompt completions. It consists of the Braintrust API, a PostgreSQL database, Redis cache, object storage, and Brainstore (a high-performance query engine for real-time trace ingestion). * The **control plane** provides the web UI, authentication, user management, and metadata storage (project names, experiment names, organization settings). The control plane does not store or process your sensitive data. | Data | Location | | - | - | | Experiment records (input, output, expected, scores, metadata, traces, spans) | Data plane | | Log records (input, output, expected, scores, metadata, traces, spans) | Data plane | | Dataset records (input, output, metadata) | Data plane | | Prompt playground prompts | Data plane | | Prompt playground completions | Data plane | | Human review scores | Data plane | | Project-level LLM provider secrets (encrypted) | Data plane | | Org-level LLM provider secrets (encrypted) | Control plane | | API keys (hashed) | Control plane | | Experiment and dataset names | Control plane | | Project names | Control plane | | Project settings | Control plane | | Git metadata about experiments | Control plane | | Organization info (name, settings) | Control plane | | Login info (name, email, avatar URL) | Control plane | | Auth credentials | [Clerk](https://clerk.com/) | When you self-host Braintrust, you deploy the data plane in your own infrastructure using Terraform. On AWS, this uses ECS (or Lambda in versions prior to v6.0) and EC2 instances. On GCP and Azure, this uses Kubernetes containers. Braintrust continues to host the control plane. When you use Braintrust's SDKs, they send data directly to your data plane. When you use the web UI, your browser communicates directly with your data plane via CORS. The control plane and data plane communicate only for authentication and metadata synchronization. Braintrust's servers and employees do not require access to your data plane for it to operate. When you configure your self-hosted data plane URL in [organization settings](/docs/admin/organizations#configure-api-urls-self-hosted), Braintrust automatically provisions a service token with the necessary permissions. This streamlines setup and enables features like [data retention](/docs/admin/data-management/retention) without manual configuration steps. ## Cloud providers Braintrust provides official Terraform modules for self-hosting on AWS, Google Cloud Platform (GCP), and Azure: * [**AWS**](/docs/admin/self-hosting/deploy): Terraform with ECS and EC2 * [**GCP**](/docs/admin/self-hosting/deploy#gcp): Terraform with Kubernetes and Helm * [**Azure**](/docs/admin/self-hosting/deploy#azure): Terraform with Kubernetes and Helm Braintrust strongly recommends using these Terraform modules because they are kept up-to-date with best practices, mirror the fully hosted offering (proven at scale), minimize configuration issues, and ensure Braintrust can efficiently troubleshoot performance and operational issues. If the module conflicts with your organization's infrastructure standards, you can deploy Braintrust in a dedicated cloud account or project to address these concerns. If you still can't use the standard artifacts as published, see [Non-standard self-hosted](#non-standard-self-hosted). **Legacy customers**: If you previously deployed using AWS CloudFormation, the [CloudFormation guide](/docs/admin/self-hosting/aws-cloudformation) remains available. This deployment method is not supported for new customers. ## Non-standard self-hosted Non-standard self-hosted deployments are customer-operated deployments that cannot use Braintrust's standard Terraform module and Helm chart pattern as intended, or that require material deviations from the reference architecture. Review it with Braintrust before adopting it. Examples include: * Forking or patching Braintrust Terraform modules or Helm chart templates. * Replacing standard deployment components with customer-specific infrastructure equivalents. * Using internal platform tooling that prevents the standard upgrade path from being followed. * Running with configuration choices that prevent straightforward adoption of future Braintrust releases. Choosing this path means more manual upgrade planning, more validation work, and higher support complexity. Braintrust support may need additional environment context before diagnosing issues, because the deployment differs from the standard reference architecture. If you anticipate any of these deviations, [contact Braintrust](mailto:support@braintrust.dev) first to confirm the approach is viable. ## Shared responsibility When you self-host, uptime becomes a shared responsibility between your team and Braintrust: * **Braintrust** is responsible for responding quickly when you have issues, collaboratively resolving them with you, and fixing bugs to improve quality. * **Your team** is responsible for following the documentation, assigning infrastructure resources on your team, and ensuring that in the event of an incident, you have staff who are familiar with Braintrust and can work with the Braintrust team to share context and resolve issues. ## Monitoring Braintrust monitors your self-hosted deployment through automatic telemetry and an in-app infra dashboard. ### Telemetry By default, your self-hosted data plane automatically sends the following telemetry back to the Braintrust-managed control plane: * Health check information * System metrics (CPU/memory) and Braintrust-specific metrics like indexing lag * Billing usage telemetry for aggregate usage metrics This allows Braintrust to monitor key health indicators and quickly identify issues before they cause downtime. In some cases, Braintrust may ask you to enable additional telemetry to help with troubleshooting, including logs and traces. For more details, see [Enable or disable telemetry](/docs/admin/self-hosting/configure/telemetry#enable-or-disable-telemetry). If you disable telemetry, Braintrust's ability to proactively monitor your deployment and diagnose issues will be significantly limited. Before disabling, consider the impact on support response times. ### Infra dashboard Only [organization owners](/docs/admin/access-control#built-in-permission-groups) and members with the **Manage settings** permission can access this dashboard. Go to ** Settings** > [** Infra dashboard**](https://www.braintrust.dev/app/~/configuration/org/infra) to view: * Processing throughput (bytes processed, compaction) * CPU and memory usage by reader and writer nodes * Object storage latency and operations * Realtime lag * Status checks * Query patterns for UI and API queries, grouped by object type, filter fields, source, and predicate types like `ILIKE`, `match()`, and inequalities The **Infra dashboard** option is available once Braintrust has enabled infrastructure monitoring for your organization. Contact [Braintrust support](mailto:support@braintrust.dev) to get started. ## Upgrades Braintrust ships new data plane versions 1-2 times per month. You can find the details of each release on the [Self-hosting releases](/docs/data-plane-changelog). **Braintrust recommends upgrading each time a new version is published.** New features often depend on data plane changes, and when they do, Braintrust will automatically gate those features until you upgrade. | Data plane age | Status | | - | - | | Up to date | Fully supported | | 1-3 months out of date | Supported with caveats — you may encounter functionality issues or bugs. Given the pace of the AI space, Braintrust prioritizes shipping new features while doing their best to maintain compatibility. If you hit a bug, [contact support](mailto:support@braintrust.dev) and Braintrust will prioritize a fix or workaround. | | More than 3 months out of date | Unsupported — upgrade immediately. If you contact support, the first thing Braintrust will ask you to do is upgrade. | To check which data plane version you're currently running, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). For upgrade instructions, see [Upgrade your deployment](/docs/admin/self-hosting/upgrade/routine). ## Remote access There are occasionally issues that require ad-hoc debugging or running manual commands against containers, the Postgres database, or storage buckets to repair the state of the system. Customers who provide Braintrust with remote access (as needed) have experienced much faster resolutions when such issues occur, because the Braintrust team can connect directly and resolve issues. If this is not possible, factor this into your uptime calculations. If uptime of Braintrust is a key metric for you, strongly consider making remote access available to the Braintrust team as needed. If you cannot set up remote access, ensure that you can swiftly access: * Containers directly (to update them, view logs, restart them, and view host metrics like CPU, network, memory, and disk utilization) * Postgres to run SQL queries * Redis to run commands * Storage buckets to run read, write, and list commands Your on-call staff should have basic familiarity with Braintrust and the ability to perform all of these operations. ## Hardware requirements When deploying Braintrust in production, consider these hardware requirements for reliable performance and uptime. These requirements assume typical production usage patterns. For high-utilization deployments, you may need to scale these resources up significantly. Monitor your resource utilization and adjust accordingly. ### API service The API service handles all SDK and browser requests to the data plane. This section primarily applies to GCP and Azure with Kubernetes. AWS deployments on ECS are pre-sized appropriately and scale automatically through Application Auto Scaling, but expose optional Terraform variables to tune per-task size and task counts. The defaults are suitable for most deployments. See [Configure AWS API ECS services](/docs/admin/self-hosting/configure/scaling#configure-aws-api-ecs-services). | Resource | Testing/Staging | Production | | - | - | - | | CPU | 1 vCPU | 2+ vCPUs per instance | | Memory | 2GB RAM | 8GB+ RAM | | Instance count | 1 | 4+ | **Environment variables**: * `NODE_MEMORY_PERCENT`: Set to `80`-`90` if the API is running on a dedicated instance or container orchestrator with cgroup memory limits (e.g. Kubernetes, ECS). * `TS_API_KEEP_ALIVE_TIMEOUT_SECONDS`: Configure the HTTP keep-alive timeout when running behind a load balancer. See [Set the HTTP keep-alive timeout](/docs/admin/self-hosting/configure/networking#set-the-http-keep-alive-timeout) for details. ### PostgreSQL PostgreSQL stores metadata required to operate the platform, including pointers to raw data in object storage and aggregate statistics about the data. It is not the primary store for your AI data — traces, spans, and logs live in Brainstore and object storage. | Resource | Testing/Staging | Production | | - | - | - | | CPU | 2 vCPUs | 8+ vCPUs | | Memory | 8GB RAM | 64GB+ RAM | | Storage size | 100GB | 1000GB+ (monitor for growth) | | Storage IOPS | 3,000 | 15,000+ | | Version | 15+ | 17+ | ### Redis cache Redis provides caching and coordination for session management, rate limiting, and Brainstore write ordering. | Resource | Testing/Staging | Production | | - | - | - | | CPU | 1 vCPU | 2 vCPUs | | Memory | 1GB RAM | 4GB+ RAM | | Version | 7+ | 7+ | **Important for AWS**: Avoid using burstable Redis instances (t-family instances like `cache.t4g.micro`) in production. These instances use CPU credits that can be exhausted during high-load periods, leading to performance throttling. Instead, use non-burstable instances like `cache.r7g.large`, `cache.r6g.medium`, or `cache.r5.large` for predictable performance. Even if these instances seem oversized initially, they provide consistent performance without the risk of CPU credit exhaustion. ### Brainstore Brainstore is Braintrust's high-performance database for ingesting and querying AI data. It uses object storage and a streaming Rust engine to load spans in real time, cutting down on latency and enabling fast full-text search over large volumes of trace data. Brainstore runs as separate reader and writer node types, each with distinct resource requirements. **Important** * Brainstore requires high-performance storage with at least 150,000 IOPS for both reads and writes. Use NVMe-based ephemeral storage (the storage does not need to be persistent). Do not use EBS volumes or other slower storage options like Azure's standard local disks, as these will significantly degrade performance. * For Kubernetes deployments (GCP and Azure), each Brainstore pod must run on its own dedicated node to ensure optimal performance and resource isolation. #### Readers Readers serve ad-hoc queries, including those from the API and user-defined BTQL queries. Plan for a minimum of 2 reader nodes in production to ensure high availability. A specialized reader variant — **fast readers** — serves predictable UI queries (paginated viewers, span and trace lookups) in isolation from standard reader nodes, keeping the UI responsive while resource-intensive queries run on readers. On **GCP and Azure**, fast readers are enabled by default with 2 replicas starting in Helm chart v5.0.0. On **AWS**, fast readers are enabled by default with 2 nodes starting in Terraform module v5.5.0; on earlier module versions they are disabled by default. Set `brainstore_fast_reader_instance_count` to `0` to opt out. When planning cluster capacity, account for these additional nodes. See [Configure Brainstore fast readers](/docs/admin/self-hosting/configure/scaling#configure-brainstore-fast-readers) for configuration details. | Resource | Testing/Staging | Prod: readers | Prod: fast readers | | - | - | - | - | | CPU | 4 vCPUs | 16 vCPUs | 16 vCPUs | | Memory | 8GB RAM | 32GB RAM | 32GB RAM | | Storage size | 128GB | 1024GB+ | 1024GB+ | | Storage type | SSD | NVMe (ephemeral) | NVMe (ephemeral) | | Storage IOPS | — | 150,000+ read/write | 150,000+ read/write | | Instance count | 1 | 2+ | 2+ | #### Writers Writers ingest incoming spans and traces and write them to object storage. Writers don't serve interactive requests, so a single writer node is sufficient for production. | Resource | Testing/Staging | Production | | - | - | - | | CPU | 4 vCPUs | 32 vCPUs | | Memory | 8GB RAM | 64GB RAM | | Storage size | 128GB | 1024GB+ | | Storage type | SSD | NVMe (ephemeral) | | Storage IOPS | — | 150,000+ read/write | | Instance count | 1 | 1+ | # Upgrade your deployment Source: https://braintrust.dev/docs/admin/self-hosting/upgrade/routine Upgrade a self-hosted Braintrust data plane on AWS, GCP, or Azure Upgrading to data plane v2.0? It requires a multi-step migration with irreversible infrastructure changes. Follow the [Upgrade to v2.0](/docs/admin/self-hosting/upgrade/v2) guide instead. This guide shows the routine process for upgrading the Braintrust data plane in your AWS account. Upgrading from AWS Terraform module v5.x to v6? It requires a two-step apply to move the API from Lambda to ECS. Follow the [Upgrade to v6](/docs/admin/self-hosting/upgrade/v6) guide instead. ## How it works On AWS, the Braintrust data plane runs on ECS and EC2 (Terraform module versions before v6.0 run the API on Lambda). The data plane version is **bundled into the Terraform module** — each module release pins specific versions of the Braintrust API and Brainstore in `VERSIONS.json`. When you update your module source to a newer version and run `terraform apply`, Terraform automatically deploys the corresponding data plane. To check which data plane version you're currently running, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) in the Braintrust UI. ## 1. Update the module version Check the [Self-hosting releases](/docs/data-plane-changelog) page for your target data plane version to find the minimum Terraform module version required. Then update the `?ref=` in your module source: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=vX.Y.Z" # ... other configuration ... } ``` See all module releases: [GitHub Releases](https://github.com/braintrustdata/terraform-aws-braintrust-data-plane/releases) AWS Terraform module v6.2.0 and later add a mandatory `aws-apn-id` tag to every resource. If your account restricts tag keys with an SCP or IAM policy, allow `aws-apn-id` before you apply, or the upgrade will fail. See [Use custom tags](/docs/admin/self-hosting/deploy#use-custom-tags). If your deployment runs the API on [ECS](/docs/admin/self-hosting/upgrade/v6) (`enable_ecs_api = true`), upgrade to AWS Terraform module v6.5.0 or later. On v6.4.0 and earlier, running the API on ECS also sends LLM calls from user-authored scorers and tools to the Braintrust-hosted Gateway, outside your AWS account. See [Upgrade to Terraform module v6](/docs/admin/self-hosting/upgrade/v6). AWS Terraform module v6.7.0 removes the `enable_brainstore` variable. Brainstore is always enabled, and there is no opt-out. If your configuration sets `enable_brainstore`, delete the line before you apply, or the plan fails with an unsupported argument error. After applying AWS Terraform module v6.8.0, downgrading to v6.7.x or earlier destroys and recreates API ECS listener rules, which can disrupt traffic. Review the Terraform plan before any rollback. See the [v6.8.0 release notes](/docs/data-plane-changelog#terraform-aws-module-releases). ## 2. Review and apply As of v5.6.0, the `modules/services` sub-module requires the `hashicorp/http` provider `~> 3.3`. If your `.terraform.lock.hcl` pins an older 3.x version, run `terraform init -upgrade` before `terraform plan` or `terraform apply`. Review the planned changes before applying: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform plan ``` Carefully review the output of `terraform plan` before applying. If you see something unexpected, like deletion of a database or S3 bucket, [contact Braintrust](mailto:support@braintrust.dev) for help. Apply the changes: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` This updates both the infrastructure and the data plane components in a single apply. This guide shows the routine process for upgrading the Braintrust data plane in your GCP account. ## How it works GCP deployments have two independently upgradeable components: * **Terraform module** (`terraform-google-braintrust-data-plane`): provisions your cloud infrastructure — the GKE cluster, Cloud SQL database, Redis, and storage buckets. You upgrade this by changing the `?ref=` version in your module source and running `terraform apply`. * **Helm chart**: deploys the Braintrust application (API and Brainstore containers) onto your GKE cluster. You upgrade this by running `helm upgrade` with a new `--version`. The Helm chart `--version` is the chart version, not the data plane version. The data plane version is set separately via image tags in your `values.yaml`. Upgrading the Helm chart alone does not change which data plane version is running — you must also update the image tags. To check which data plane version you're currently running, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) in the Braintrust UI. ## 1. Decide what to upgrade Before upgrading, check the [Self-hosting releases](/docs/data-plane-changelog) page for your target data plane version. Each entry specifies its requirements: * **"Requires: Helm X.Y.Z+"** — you must upgrade the Helm chart to at least that version before or alongside updating your image tags. Update Terraform first if infrastructure changes are also listed. * **No requirements listed** — you can upgrade by updating image tags in `values.yaml` and running `helm upgrade`. No Terraform or chart version change needed. * **Helm-only release ("No data plane version change")** — update only the Helm chart version. No image tag change needed. ## 2. Update infrastructure (Terraform) Run this step when a release lists infrastructure requirements, or when you want to apply infrastructure changes independently. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` Carefully review the output of `terraform plan` before applying any changes to your deployment. If you see something unexpected, like deletion of a database or storage bucket, [contact Braintrust](mailto:support@braintrust.dev) for help. To pin to a specific Terraform module version, update the `?ref=` in your module source: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane?ref=vX.Y.Z" # ... other configuration ... } ``` See all module releases: [GitHub Releases](https://github.com/braintrustdata/terraform-google-braintrust-data-plane/releases) ## 3. Update services (Helm) Run this step to deploy a new data plane version or apply a Helm chart update. Set the data plane version by updating the image tags for the API and Brainstore in your `values.yaml`. Refer to the [Helm chart values reference](https://github.com/braintrustdata/helm/blob/main/braintrust/values.yaml) for the exact field names. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version \ --values helm-values.yaml ``` Replace `` with the Helm chart version required by the release (or your current chart version if no chart upgrade is needed). When a release requires both Terraform and Helm changes, apply Terraform first, then run the Helm upgrade. See all Helm chart releases: [GitHub Releases](https://github.com/braintrustdata/helm/releases) This guide shows the routine process for upgrading the Braintrust data plane in your Azure account. ## How it works Azure deployments have two independently upgradeable components: * **Terraform module** (`terraform-azure-braintrust-data-plane`): provisions your cloud infrastructure — the AKS cluster, Azure Database for PostgreSQL, Redis, and storage. You upgrade this by changing the `?ref=` version in your module source and running `terraform apply`. * **Helm chart**: deploys the Braintrust application (API and Brainstore containers) onto your AKS cluster. You upgrade this by running `helm upgrade` with a new `--version`. The Helm chart `--version` is the chart version, not the data plane version. The data plane version is set separately via image tags in your `values.yaml`. Upgrading the Helm chart alone does not change which data plane version is running — you must also update the image tags. To check which data plane version you're currently running, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) in the Braintrust UI. ## 1. Decide what to upgrade Before upgrading, check the [Self-hosting releases](/docs/data-plane-changelog) page for your target data plane version. Each entry specifies its requirements: * **"Requires: Helm X.Y.Z+"** — you must upgrade the Helm chart to at least that version before or alongside updating your image tags. Update Terraform first if infrastructure changes are also listed. * **No requirements listed** — you can upgrade by updating image tags in `values.yaml` and running `helm upgrade`. No Terraform or chart version change needed. * **Helm-only release ("No data plane version change")** — update only the Helm chart version. No image tag change needed. ## 2. Update infrastructure (Terraform) Run this step when a release lists infrastructure requirements, or when you want to apply infrastructure changes independently. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` Carefully review the output of `terraform plan` before applying any changes to your deployment. If you see something unexpected, like deletion of a database or storage account, [contact Braintrust](mailto:support@braintrust.dev) for help. To pin to a specific Terraform module version, update the `?ref=` in your module source: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-azure-braintrust-data-plane?ref=vX.Y.Z" # ... other configuration ... } ``` See all module releases: [GitHub Releases](https://github.com/braintrustdata/terraform-azure-braintrust-data-plane/releases) Some Terraform Azure module upgrades include one-time breaking changes, such as renamed node pool variables, an azurerm provider upgrade, or a Terraform state migration. Before bumping the module across major versions, review the **Terraform Azure module releases** entries on the [Self-hosting releases](/docs/data-plane-changelog) page for the versions you're crossing. ## 3. Update services (Helm) Run this step to deploy a new data plane version or apply a Helm chart update. Set the data plane version by updating the image tags for the API and Brainstore in your `values.yaml`. Refer to the [Helm chart values reference](https://github.com/braintrustdata/helm/blob/main/braintrust/values.yaml) for the exact field names. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version \ --values helm-values.yaml ``` Replace `` with the Helm chart version required by the release (or your current chart version if no chart upgrade is needed). When a release requires both Terraform and Helm changes, apply Terraform first, then run the Helm upgrade. See all Helm chart releases: [GitHub Releases](https://github.com/braintrustdata/helm/releases) ## Next steps * [Self-hosting releases](/docs/data-plane-changelog) — review release notes and infrastructure requirements for each data plane version * [Configuration](/docs/admin/self-hosting/configure) — configure telemetry, network access, rate limiting, and other options # Upgrade to data plane v2.x Source: https://braintrust.dev/docs/admin/self-hosting/upgrade/v2 Upgrade your self-hosted Braintrust deployment to the data plane v2 series Data plane v2.0 introduced many new features, including [trace-level scorers](/docs/changelog#trace-level-scorers), [sandboxes for agent evals](/docs/changelog#sandboxes-for-agent-evals), and [log indexing and full-text search](/docs/changelog#log-indexing-and-full-text-search). The current recommended target for the v2.x cutover is data plane v2.1.1. The module and chart versions below deploy v2.1.1 images unless noted otherwise. This guide assumes a standard Braintrust deployment using the Braintrust Terraform modules or Helm chart to manage storage, identity, networking, and generated configuration. If your deployment customizes those layers, confirm the equivalent resources, permissions, network paths, and runtime environment are in place before applying the version upgrade. Braintrust components must be able to list, read, write, and delete the generated object-storage paths, including versioned deletes where object versioning is enabled. Avoid bundling unrelated configuration or workload changes into the upgrade window. Validate the version upgrade first, then separately ramp new automations, scoring traffic, credential rotations, or storage and network policy changes. This guide shows the process for upgrading the Braintrust data plane to v2.x in your AWS account. ## 1. Prepare for v2.x Before upgrading to data plane v2.x, you need to reach Terraform module v4.5.0 (data plane v1.1.32) and enable the efficient WAL format. v1.1.32 is the required foundation for v2.x. It ensures all Brainstore configuration is in place and verified before the upgrade. ### Upgrade Terraform module The goal of this step is to reach v4.5.0. Check your current module version in `main.tf`: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v..." # your current version } ``` Then follow the path for your starting point: | Starting version | Sub-steps to follow | | - | - | | TF module \< v3.0.0 | 1 → 2 → 3 | | TF module v3.x | 2 → 3 | | TF module v4.x (not v4.5.0) | 3 only | | TF module v4.5.0 | Skip to [Enable efficient WAL format](#enable-efficient-wal-format-aws) | | TF module v5.0.0, v5.1.0, or v5.2.0 | Skip to [Upgrade to v2.x](#upgrade-to-v2-x-aws) | | TF module v5.2.1+ | Skip to [Enable no-PostgreSQL mode](#enable-no-postgresql-mode-aws) | This upgrades the data plane from v1.1.25 to v1.1.27. There are no breaking changes. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v3.1.3" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` This step upgrades the AWS provider from v5 to v6. **The Terraform state file changes made in this step are irreversible.** Apply at this version explicitly before proceeding, so that if anything goes wrong, it is clear whether the provider upgrade or the subsequent data plane bump caused it. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v4.0.0" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` The data plane version does not change in this step. For straightforward deployments, you can upgrade directly from v3.1.3 to v4.5.0, skipping the v4.0.0 stop. The intermediate stop is recommended for complex deployments to isolate the AWS provider migration. If something goes wrong in a single jump, it's harder to identify the cause. This upgrades the data plane to v1.1.32. v4.5.0 also introduces the `skip_pg_for_brainstore_objects` variable used in the Topics enablement step. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v4.5.0" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` ### Enable efficient WAL format This must be a separate `terraform apply` from the v4.5.0 upgrade. If you set `brainstore_wal_footer_version` in the same apply, Brainstore nodes that are still rolling out won't be able to read the new WAL format. Wait until all nodes are confirmed running v1.1.32 before applying this step. Add `brainstore_wal_footer_version = "v1"` to your module and apply: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v4.5.0" brainstore_wal_footer_version = "v1" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` After applying, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and check the WAL footer version (`BRAINSTORE_WAL_FOOTER_VERSION`) row. On v1.1.32, it will show "Unable to verify setting. Please upgrade to configure." The setting is active; Brainstore v1.1.32 does not report it back to the control plane. It will show correctly after upgrading to v2.x. This step is strongly recommended but not mechanically enforced. Deployments running mixed WAL formats become increasingly difficult to tune at scale. ### Run the preflight check Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Verify the following expected statuses before proceeding: | Setting | Expected status | | - | - | | License | Configured | | Telemetry | Configured | | Data plane service token | Configured | | Brainstore default mode (`BRAINSTORE_DEFAULT`) | Configured | | Insert logs (`INSERT_LOGS2`) | Configured | | Insert row references (`BRAINSTORE_INSERT_ROW_REFS`) | Configured | | Realtime WAL (`BRAINSTORE_REALTIME_WAL_BUCKET`) | Configured | | Ephemeral WAL (`BRAINSTORE_XACT_MANAGER_URI`) | Configured | | Proxy URL (`BRAINSTORE_AI_PROXY_URL`) | Configured | | Redis (`BRAINSTORE_REDIS_URI`) | Configured | | Brainstore direct writes | "Upgrade to configure Brainstore direct writes" | On v1.1.32, items like WAL footer version, response cache URI, and code bundle URI are not displayed. They appear only after upgrading to v2.x. The HTTP/2 check may show a warning on v1.1.32. This resolves after upgrading. If any of the first ten rows are not green, resolve them before proceeding. Contact [Braintrust support](mailto:support@braintrust.dev) if you need help. ## 2. Upgrade to v2.x With preflight complete, upgrade to Terraform module v5.2.1, which ships the recommended data plane v2.1.1 images. The WAL footer also bumps to v3 in the same apply. If you are already on Terraform module v5.2.1 with WAL footer version set to v3, skip to [Enable no-PostgreSQL mode](#enable-no-postgresql-mode-aws). ### Apply the upgrade Update your module to v5.2.1 and bump the WAL footer to v3. You can do both in the same apply. As long as WAL footer is already on v1 from the previous step, all v2.x nodes will understand the v3 format on startup. Because the latest module defaults no-PostgreSQL mode on, explicitly disable it for the base upgrade. Enable it only in the next step. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v5.2.1" brainstore_wal_footer_version = "v3" skip_pg_for_brainstore_objects = "" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` ### Verify settings Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and verify the following are all green: | Setting | Expected status | | - | - | | License | Configured | | Telemetry | Configured | | Data plane service token | Configured | | Brainstore default mode (`BRAINSTORE_DEFAULT`) | Configured | | Insert logs (`INSERT_LOGS2`) | Configured | | Insert row references (`BRAINSTORE_INSERT_ROW_REFS`) | Configured | | Realtime WAL (`BRAINSTORE_REALTIME_WAL_BUCKET`) | Configured | | Ephemeral WAL (`BRAINSTORE_XACT_MANAGER_URI`) | Configured | | Proxy URL (`BRAINSTORE_AI_PROXY_URL`) | Configured | | Redis (`BRAINSTORE_REDIS_URI`) | Configured | | Response cache URI (`BRAINSTORE_RESPONSE_CACHE_URI`) | Configured | | Code bundle URI (`BRAINSTORE_CODE_BUNDLE_URI`) | Configured | | HTTP/2 | Data plane is using HTTP/2 or higher | **Brainstore direct writes** will show "Upgrade to configure Brainstore direct writes." This is enabled in the next step. ### Validate the base upgrade Before enabling no-PostgreSQL mode, validate the workloads your team relies on. For example, trace one production-shaped request, run an eval suite you use regularly, and invoke or iterate on a prompt. If you use online scorers or automations, trigger one through its normal online path. For trace-level scorers, choose one that uses full trace context. If these workflows round-trip cleanly through the data plane and look right in the UI, continue to no-PostgreSQL mode. ### Enable no-PostgreSQL mode No-PostgreSQL mode removes PostgreSQL from the Brainstore write path, delivering higher ingestion throughput. This step is a prerequisite for [Topics](/docs/observe/topics) but does not enable them. The switch to no-PostgreSQL mode is one-way: Once Brainstore is writing data directly, rolling back requires downtime to drain and re-ingest. Enable it only when you are ready to commit. Keep your module on v5.2.1 and set `skip_pg_for_brainstore_objects = "all"`: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v5.2.1" brainstore_wal_footer_version = "v3" skip_pg_for_brainstore_objects = "all" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` Repeat the same workload validation after enabling no-PostgreSQL mode. If the same representative workflows round-trip cleanly through the data plane and look right in the UI, proceed to the next environment. ### Enable Topics [Topics](/docs/observe/topics) for self-hosted deployments has specific deployment, inference, billing, and maintenance requirements, and Braintrust must grant access to your organization. Completing the full v2.x upgrade and no-PostgreSQL mode migration is a prerequisite, and the following conditions must also be met before Topics can be enabled: * **Standard deployment required.** Topics is supported on standard Terraform/Helm deployments only. Custom Kubernetes clusters using the Helm chart on non-standard infrastructure are not supported. * **Braintrust-hosted inference required.** Topics runs LLM calls on models hosted by Braintrust with Zero Data Retention (ZDR). Bring-your-own-model is not supported. Currently US-only. * **Built-in models must be enabled.** Self-hosted organizations have [built-in models](/docs/admin/ai-providers#manage-built-in-models) disabled by default, so no trace data leaves your network boundary unless you opt in. Because Topics depends on built-in models, an admin must explicitly enable them before Topics can run. * **Pricing is not included by default.** Topics usage is not covered by standard contracts. Discuss pricing with your account team before proceeding. * **Ongoing maintenance required.** Self-hosted Topics deployments run on a frequent-release build track separate from the tagged data plane versions and will require regular upgrades to receive the latest fixes and improvements. If you meet all of these requirements and want to proceed, contact your account team or [Braintrust support](mailto:support@braintrust.dev). Until Braintrust enables Topics access for your organization, the Topics page and Topics automations remain blocked, even after the upgrade and no-PostgreSQL migration are complete. ## Configuration reference | Variable | Source | Since | | - | - | - | | `SERVICE_TOKEN_SECRET_KEY` | Auto-generated 32-character key in Secrets Manager | v3.0.0 | | `BRAINSTORE_REDIS_URI` | ElastiCache endpoint and port | v3.1.0 | | `BRAINSTORE_DEFAULT` | Hardcoded `"force"` | v3.0.0 | | `INSERT_LOGS2` | Hardcoded `"true"` | v3.0.0 | | `BRAINSTORE_INSERT_ROW_REFS` | Hardcoded `"true"` | v3.0.0 | | `BRAINSTORE_REALTIME_WAL_BUCKET` | Auto-created S3 bucket | v3.0.0 | | `BRAINSTORE_AI_PROXY_URL` | Fetched from SSM at boot | v4.0.3 | | `BRAINSTORE_XACT_MANAGER_URI` | Same ElastiCache endpoint as Redis | v4.4.0 | | `BRAINSTORE_RESPONSE_CACHE_URI` | Module-created S3 bucket (at `/brainstore-cache` prefix) | v4.4.0 | | `BRAINSTORE_CODE_BUNDLE_URI` | Module-created S3 bucket | v4.4.0 | | Variable | How to set | Since | | - | - | - | | `BRAINSTORE_WAL_FOOTER_VERSION` | Set `brainstore_wal_footer_version = "v1"` in module | v4.5.0 | Allowed values: `""` (disabled, default), `"v1"`, `"v2"`, `"v3"`. Set to `"v1"` as part of the pre-upgrade steps. Only set to `"v3"` after successfully deploying v2.x. | Variable | How to set | Since | | - | - | - | | `SKIP_PG_FOR_BRAINSTORE_OBJECTS` | Set `skip_pg_for_brainstore_objects` in module | v4.5.0 | | `BRAINSTORE_ASYNC_SCORING_OBJECTS` | Auto-derived from `skip_pg_for_brainstore_objects` | v4.5.0 | | `BRAINSTORE_LOG_AUTOMATIONS_OBJECTS` | Auto-derived from `skip_pg_for_brainstore_objects` | v4.5.0 | | `BRAINSTORE_WAL_USE_EFFICIENT_FORMAT` | Auto-enabled when either `brainstore_wal_footer_version` or `skip_pg_for_brainstore_objects` is set | v4.5.0 | Setting `skip_pg_for_brainstore_objects = "all"` is a prerequisite for enabling Topics. The module automatically fans this value out to `BRAINSTORE_ASYNC_SCORING_OBJECTS` and `BRAINSTORE_LOG_AUTOMATIONS_OBJECTS`. This guide shows the process for upgrading the Braintrust data plane to v2.x in your GCP account. ## 1. Prepare for v2.x Before upgrading to data plane v2.x, you need to reach Helm chart v5.0.1 (data plane v1.1.32) and enable the efficient WAL format. v1.1.32 is the required foundation for v2.x. It ensures all Brainstore configuration is in place and verified before the upgrade. ### Upgrade Helm chart Check your current Helm chart version: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm list -n braintrust ``` Then follow the path for your starting point: | Starting version | Sub-steps to follow | | - | - | | Helm chart \< v3.1.1 | 1 → 2 → 3 | | Helm chart v3.x | 2 → 3 | | Helm chart v4.x | 3 only | | Helm chart v5.0.1 | Skip to [Enable efficient WAL format](#enable-efficient-wal-format-gcp) | | Helm chart v6.0.0, v6.1.0, or v6.2.0 | Skip to [Upgrade to v2.x](#upgrade-to-v2-x-gcp) | | Helm chart v6.2.1+ | Skip to [Enable no-PostgreSQL mode](#enable-no-postgresql-mode-gcp) | This upgrades the data plane from v1.1.27 to v1.1.31. There are no breaking changes. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 3.1.1 \ --values helm-values.yaml ``` After applying, restart pods: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl rollout restart deployment -n braintrust ``` This step upgrades the data plane to v1.1.32 and introduces bidirectional communication between Brainstore and API pods. Brainstore now needs to reach the API pods. The chart auto-sets `BRAINSTORE_AI_PROXY_URL` to `http://braintrust-api:8000`. If your cluster has network policies or security rules restricting traffic between pods, allow Brainstore pods to reach the API service on port 8000 before upgrading. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 4.0.0 \ --values helm-values.yaml ``` After applying, restart pods: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl rollout restart deployment -n braintrust ``` This step enables fast readers, isolated Brainstore nodes that handle common UI queries. Fast readers are enabled by default with 2 replicas and use the same node spec as standard readers, **increasing your cluster node count by 2**. You can upgrade directly to v5.0.1 from any v4.x version. Before upgrading, account for 2 additional fast reader nodes in your cluster capacity. If you have customized `brainstore.reader` settings in `values.yaml`, mirror those customizations to the new `brainstore.fastreader` block. See [Configure Brainstore fast readers](/docs/admin/self-hosting/configure/scaling#configure-brainstore-fast-readers) for details. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 5.0.1 \ --values helm-values.yaml ``` After applying, restart pods: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl rollout restart deployment -n braintrust ``` ### Enable efficient WAL format This must be a separate `helm upgrade` after confirming all pods are running v1.1.32. Setting this during the same upgrade as the version bump risks Brainstore nodes still rolling out being unable to read the new WAL format. Set `brainstoreWalFooterVersion` in your `helm-values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstoreWalFooterVersion: "v1" ``` Then apply: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 5.1.0 \ --values helm-values.yaml ``` On chart v5.1.0+, pods will automatically restart after applying the upgrade. Once the pods have cycled, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and check the WAL footer version (`BRAINSTORE_WAL_FOOTER_VERSION`) row. On v1.1.32, it will show "Unable to verify setting. Please upgrade to configure." The setting is active; Brainstore v1.1.32 does not report it back to the control plane. It will show correctly after upgrading to v2.x. This step is strongly recommended but not mechanically enforced. Deployments running mixed WAL formats become increasingly difficult to tune at scale. ### Run the preflight check Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Verify the following expected statuses before proceeding: | Setting | Expected status | | - | - | | License | Configured | | Telemetry | Configured | | Data plane service token | Configured | | Brainstore default mode (`BRAINSTORE_DEFAULT`) | Configured | | Insert logs (`INSERT_LOGS2`) | Configured | | Insert row references (`BRAINSTORE_INSERT_ROW_REFS`) | Configured | | Realtime WAL (`BRAINSTORE_REALTIME_WAL_BUCKET`) | Configured | | Ephemeral WAL (`BRAINSTORE_XACT_MANAGER_URI`) | Configured | | Proxy URL (`BRAINSTORE_AI_PROXY_URL`) | Configured | | Redis (`BRAINSTORE_REDIS_URI`) | Configured | | Brainstore direct writes | "Upgrade to configure Brainstore direct writes" | On v1.1.32, items like WAL footer version, response cache URI, and code bundle URI are not displayed. They appear only after upgrading to v2.x. The HTTP/2 check may show a warning on v1.1.32. This resolves after upgrading. If any of the first ten rows are not green, resolve them before proceeding. Contact [Braintrust support](mailto:support@braintrust.dev) if you need help. ### Update Workload Identity bindings Data plane v2.0 introduces API pod endpoints that access GCS directly. Prior versions only accessed GCS from Brainstore pods. The GCP IAM binding must be updated to grant the `braintrust-api` Kubernetes service account Workload Identity access before upgrading. * **If using the Braintrust GCP Terraform module**, update to [v1.5.2](https://github.com/braintrustdata/terraform-google-braintrust-data-plane/releases/tag/v1.5.2) and run `terraform apply` before upgrading Helm to v6.2.1. The updated module includes the binding for both the `brainstore` and `braintrust-api` service accounts. ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-google-braintrust-data-plane?ref=v1.5.2" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` * **If managing IAM manually**, add the binding: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} gcloud iam service-accounts add-iam-policy-binding \ @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[/braintrust-api]" ``` **Apply these IAM changes before upgrading to v2.x.** Without this binding, the API pod will return errors on endpoints that access GCS directly. ## 2. Upgrade to v2.x With preflight complete, upgrade to Helm chart v6.2.1, which ships the recommended data plane v2.1.1 images. The WAL footer also bumps to v3 in the same apply. If you are already on Helm chart v6.2.1 with WAL footer version set to v3, skip to [Enable no-PostgreSQL mode](#enable-no-postgresql-mode-gcp). ### Apply the upgrade Update your `helm-values.yaml` to bump the WAL footer to v3. Because the latest chart defaults no-PostgreSQL mode on, explicitly disable it for the base upgrade. Enable it only in the next step. ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstoreWalFooterVersion: "v3" skipPgForBrainstoreObjects: "" ``` **Breaking change:** Chart v6.2.1 disables code function execution by default. To preserve code-backed scorers, tools, or functions on a single-org deployment, add: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: allowCodeFunctionExecution: true ``` If your deployment uses `ORG_NAME=*` or hosts multiple Braintrust organizations, [contact Braintrust](mailto:support@braintrust.dev) before enabling code function execution. Apply Helm chart v6.2.1: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 6.2.1 \ --values helm-values.yaml ``` Pods will automatically restart after applying the upgrade. ### Verify settings After the pods have cycled, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and verify the following are all green: | Setting | Expected status | | - | - | | License | Configured | | Telemetry | Configured | | Data plane service token | Configured | | Brainstore default mode (`BRAINSTORE_DEFAULT`) | Configured | | Insert logs (`INSERT_LOGS2`) | Configured | | Insert row references (`BRAINSTORE_INSERT_ROW_REFS`) | Configured | | Realtime WAL (`BRAINSTORE_REALTIME_WAL_BUCKET`) | Configured | | Ephemeral WAL (`BRAINSTORE_XACT_MANAGER_URI`) | Configured | | Proxy URL (`BRAINSTORE_AI_PROXY_URL`) | Configured | | Redis (`BRAINSTORE_REDIS_URI`) | Configured | | Response cache URI (`BRAINSTORE_RESPONSE_CACHE_URI`) | Configured | | Code bundle URI (`BRAINSTORE_CODE_BUNDLE_URI`) | Configured | | HTTP/2 | Data plane is using HTTP/2 or higher | **Brainstore direct writes** will show "Upgrade to configure Brainstore direct writes." This is enabled in the next step. ### Validate the base upgrade Before enabling no-PostgreSQL mode, validate the workloads your team relies on. For example, trace one production-shaped request, run an eval suite you use regularly, and invoke or iterate on a prompt. If you use online scorers or automations, trigger one through its normal online path. For trace-level scorers, choose one that uses full trace context. If these workflows round-trip cleanly through the data plane and look right in the UI, continue to no-PostgreSQL mode. ### Enable no-PostgreSQL mode No-PostgreSQL mode removes PostgreSQL from the Brainstore write path, delivering higher ingestion throughput. This step is a prerequisite for [Topics](/docs/observe/topics) but does not enable them. The switch to no-PostgreSQL mode is one-way: Once Brainstore is writing data directly, rolling back requires downtime to drain and re-ingest. Enable it only when you are ready to commit. Set `skipPgForBrainstoreObjects` in your `helm-values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} skipPgForBrainstoreObjects: "all" ``` Keep Helm chart v6.2.1 and apply the updated values: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 6.2.1 \ --values helm-values.yaml ``` Then repeat the same workload validation after enabling no-PostgreSQL mode. If the same representative workflows round-trip cleanly through the data plane and look right in the UI, proceed to the next environment. ### Enable Topics [Topics](/docs/observe/topics) for self-hosted deployments has specific deployment, inference, billing, and maintenance requirements, and Braintrust must grant access to your organization. Completing the full v2.x upgrade and no-PostgreSQL mode migration is a prerequisite, and the following conditions must also be met before Topics can be enabled: * **Standard deployment required.** Topics is supported on standard Terraform/Helm deployments only. Custom Kubernetes clusters using the Helm chart on non-standard infrastructure are not supported. * **Braintrust-hosted inference required.** Topics runs LLM calls on models hosted by Braintrust with Zero Data Retention (ZDR). Bring-your-own-model is not supported. Currently US-only. * **Built-in models must be enabled.** Self-hosted organizations have [built-in models](/docs/admin/ai-providers#manage-built-in-models) disabled by default, so no trace data leaves your network boundary unless you opt in. Because Topics depends on built-in models, an admin must explicitly enable them before Topics can run. * **Pricing is not included by default.** Topics usage is not covered by standard contracts. Discuss pricing with your account team before proceeding. * **Ongoing maintenance required.** Self-hosted Topics deployments run on a frequent-release build track separate from the tagged data plane versions and will require regular upgrades to receive the latest fixes and improvements. If you meet all of these requirements and want to proceed, contact your account team or [Braintrust support](mailto:support@braintrust.dev). Until Braintrust enables Topics access for your organization, the Topics page and Topics automations remain blocked, even after the upgrade and no-PostgreSQL migration are complete. ## Configuration reference | Variable | Source | Pods | | - | - | - | | `BRAINSTORE_REDIS_URI` | From `REDIS_URL` secret | reader, writer, fastreader | | `BRAINSTORE_DEFAULT` | Hardcoded `"force"` | API | | `INSERT_LOGS2` | Hardcoded `"true"` | API | | `BRAINSTORE_INSERT_ROW_REFS` | Hardcoded `"true"` | API | | `BRAINSTORE_REALTIME_WAL_BUCKET` | From `objectStorage` values | API | | `BRAINSTORE_AI_PROXY_URL` | Built from `api.name` and `api.service.port` | reader, writer, fastreader | | `BRAINSTORE_XACT_MANAGER_URI` | From `REDIS_URL` secret | reader, writer, fastreader | | `BRAINSTORE_RESPONSE_CACHE_URI` | Auto-derived from `objectStorage` config | reader, writer, fastreader | | `BRAINSTORE_CODE_BUNDLE_URI` | Auto-derived from `objectStorage` config | reader, writer, fastreader | | Variable | How to set | Since | Pods | | - | - | - | - | | `BRAINSTORE_WAL_FOOTER_VERSION` | Set `brainstoreWalFooterVersion` in `helm-values.yaml` | v5.1.0 | API | | `BRAINSTORE_WAL_USE_EFFICIENT_FORMAT` | Auto-enabled when either `brainstoreWalFooterVersion` or `skipPgForBrainstoreObjects` is set | v5.1.0 | API | Set `brainstoreWalFooterVersion: "v1"` as part of the pre-upgrade steps. Only set to `"v3"` after successfully deploying v2.x. | Variable | How to set | Since | Pods | | - | - | - | - | | `SKIP_PG_FOR_BRAINSTORE_OBJECTS` | Set `skipPgForBrainstoreObjects` in `helm-values.yaml` | v5.1.0 | API | | `BRAINSTORE_ASYNC_SCORING_OBJECTS` | Auto-derived from `skipPgForBrainstoreObjects` | v5.1.0 | reader, writer, fastreader | | `BRAINSTORE_LOG_AUTOMATIONS_OBJECTS` | Auto-derived from `skipPgForBrainstoreObjects` | v5.1.0 | reader, writer, fastreader | Setting `skipPgForBrainstoreObjects: "all"` is a prerequisite for enabling Topics. The chart automatically fans this value out to `BRAINSTORE_ASYNC_SCORING_OBJECTS` and `BRAINSTORE_LOG_AUTOMATIONS_OBJECTS`. This guide shows the process for upgrading the Braintrust data plane to v2.x in your Azure account. ## 1. Prepare for v2.x Before upgrading to data plane v2.x, you need to reach Helm chart v5.0.1 (data plane v1.1.32) and enable the efficient WAL format. v1.1.32 is the required foundation for v2.x. It ensures all Brainstore configuration is in place and verified before the upgrade. ### Upgrade Helm chart Check your current Helm chart version: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm list -n braintrust ``` Then follow the path for your starting point: | Starting version | Sub-steps to follow | | - | - | | Helm chart \< v3.1.1 | 1 → 2 → 3 | | Helm chart v3.x | 2 → 3 | | Helm chart v4.x | 3 only | | Helm chart v5.0.1 | Skip to [Enable efficient WAL format](#enable-efficient-wal-format-azure) | | Helm chart v6.0.0, v6.1.0, or v6.2.0 | Skip to [Upgrade to v2.x](#upgrade-to-v2-x-azure) | | Helm chart v6.2.1+ | Skip to [Enable no-PostgreSQL mode](#enable-no-postgresql-mode-azure) | This upgrades the data plane from v1.1.27 to v1.1.31. There are no breaking changes. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 3.1.1 \ --values helm-values.yaml ``` After applying, restart pods: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl rollout restart deployment -n braintrust ``` This step upgrades the data plane to v1.1.32 and introduces bidirectional communication between Brainstore and API pods. Brainstore now needs to reach the API pods. The chart auto-sets `BRAINSTORE_AI_PROXY_URL` to `http://braintrust-api:8000`. If your cluster has network policies or security rules restricting traffic between pods, allow Brainstore pods to reach the API service on port 8000 before upgrading. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 4.0.0 \ --values helm-values.yaml ``` After applying, restart pods: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl rollout restart deployment -n braintrust ``` This step enables fast readers, isolated Brainstore nodes that handle common UI queries. Fast readers are enabled by default with 2 replicas and use the same node spec as standard readers, **increasing your cluster node count by 2**. You can upgrade directly to v5.0.1 from any v4.x version. Before upgrading, account for 2 additional fast reader nodes in your cluster capacity. If you have customized `brainstore.reader` settings in `values.yaml`, mirror those customizations to the new `brainstore.fastreader` block. See [Configure Brainstore fast readers](/docs/admin/self-hosting/configure/scaling#configure-brainstore-fast-readers) for details. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 5.0.1 \ --values helm-values.yaml ``` After applying, restart pods: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} kubectl rollout restart deployment -n braintrust ``` ### Enable efficient WAL format This must be a separate `helm upgrade` after confirming all pods are running v1.1.32. Setting this during the same upgrade as the version bump risks Brainstore nodes still rolling out being unable to read the new WAL format. Set `brainstoreWalFooterVersion` in your `helm-values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstoreWalFooterVersion: "v1" ``` Then apply: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 5.1.0 \ --values helm-values.yaml ``` On chart v5.1.0+, pods will automatically restart after applying the upgrade. Once the pods have cycled, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and check the WAL footer version (`BRAINSTORE_WAL_FOOTER_VERSION`) row. On v1.1.32, it will show "Unable to verify setting. Please upgrade to configure." The setting is active; Brainstore v1.1.32 does not report it back to the control plane. It will show correctly after upgrading to v2.x. This step is strongly recommended but not mechanically enforced. Deployments running mixed WAL formats become increasingly difficult to tune at scale. ### Run the preflight check Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url). Verify the following expected statuses before proceeding: | Setting | Expected status | | - | - | | License | Configured | | Telemetry | Configured | | Data plane service token | Configured | | Brainstore default mode (`BRAINSTORE_DEFAULT`) | Configured | | Insert logs (`INSERT_LOGS2`) | Configured | | Insert row references (`BRAINSTORE_INSERT_ROW_REFS`) | Configured | | Realtime WAL (`BRAINSTORE_REALTIME_WAL_BUCKET`) | Configured | | Ephemeral WAL (`BRAINSTORE_XACT_MANAGER_URI`) | Configured | | Proxy URL (`BRAINSTORE_AI_PROXY_URL`) | Configured | | Redis (`BRAINSTORE_REDIS_URI`) | Configured | | Brainstore direct writes | "Upgrade to configure Brainstore direct writes" | On v1.1.32, items like WAL footer version, response cache URI, and code bundle URI are not displayed. They appear only after upgrading to v2.x. The HTTP/2 check may show a warning on v1.1.32. This resolves after upgrading. If any of the first ten rows are not green, resolve them before proceeding. Contact [Braintrust support](mailto:support@braintrust.dev) if you need help. ### Update Terraform module Update the Braintrust Azure Terraform module to [v1.1.0](https://github.com/braintrustdata/terraform-azure-braintrust-data-plane/releases/tag/v1.1.0) and apply before upgrading Helm to v6.2.1: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-azure-braintrust-data-plane?ref=v1.1.0" # ... other configuration } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` ## 2. Upgrade to v2.x With preflight complete, upgrade to Helm chart v6.2.1, which ships the recommended data plane v2.1.1 images. The WAL footer also bumps to v3 in the same apply. If you are already on Helm chart v6.2.1 with WAL footer version set to v3, skip to [Enable no-PostgreSQL mode](#enable-no-postgresql-mode-azure). ### Apply the upgrade Update your `helm-values.yaml` to bump the WAL footer to v3. Because the latest chart defaults no-PostgreSQL mode on, explicitly disable it for the base upgrade. Enable it only in the next step. ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} brainstoreWalFooterVersion: "v3" skipPgForBrainstoreObjects: "" ``` **Breaking change:** Chart v6.2.1 disables code function execution by default. To preserve code-backed scorers, tools, or functions on a single-org deployment, add: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} api: allowCodeFunctionExecution: true ``` If your deployment uses `ORG_NAME=*` or hosts multiple Braintrust organizations, [contact Braintrust](mailto:support@braintrust.dev) before enabling code function execution. Then apply Helm chart v6.2.1: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 6.2.1 \ --values helm-values.yaml ``` Pods will automatically restart after applying the upgrade. ### Verify settings After the pods have cycled, go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and verify the following are all green: | Setting | Expected status | | - | - | | License | Configured | | Telemetry | Configured | | Data plane service token | Configured | | Brainstore default mode (`BRAINSTORE_DEFAULT`) | Configured | | Insert logs (`INSERT_LOGS2`) | Configured | | Insert row references (`BRAINSTORE_INSERT_ROW_REFS`) | Configured | | Realtime WAL (`BRAINSTORE_REALTIME_WAL_BUCKET`) | Configured | | Ephemeral WAL (`BRAINSTORE_XACT_MANAGER_URI`) | Configured | | Proxy URL (`BRAINSTORE_AI_PROXY_URL`) | Configured | | Redis (`BRAINSTORE_REDIS_URI`) | Configured | | Response cache URI (`BRAINSTORE_RESPONSE_CACHE_URI`) | Configured | | Code bundle URI (`BRAINSTORE_CODE_BUNDLE_URI`) | Configured | | HTTP/2 | Data plane is using HTTP/2 or higher | **Brainstore direct writes** will show "Upgrade to configure Brainstore direct writes." This is enabled in the next step. ### Validate the base upgrade Before enabling no-PostgreSQL mode, validate the workloads your team relies on. For example, trace one production-shaped request, run an eval suite you use regularly, and invoke or iterate on a prompt. If you use online scorers or automations, trigger one through its normal online path. For trace-level scorers, choose one that uses full trace context. If these workflows round-trip cleanly through the data plane and look right in the UI, continue to no-PostgreSQL mode. ### Enable no-PostgreSQL mode No-PostgreSQL mode removes PostgreSQL from the Brainstore write path, delivering higher ingestion throughput. This step is a prerequisite for [Topics](/docs/observe/topics) but does not enable them. The switch to no-PostgreSQL mode is one-way: Once Brainstore is writing data directly, rolling back requires downtime to drain and re-ingest. Enable it only when you are ready to commit. Set `skipPgForBrainstoreObjects` in your `helm-values.yaml`: ```yaml theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} skipPgForBrainstoreObjects: "all" ``` Keep Helm chart v6.2.1 and apply the updated values: ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} helm upgrade braintrust \ oci://public.ecr.aws/braintrust/helm/braintrust \ --namespace braintrust \ --version 6.2.1 \ --values helm-values.yaml ``` Then repeat the same workload validation after enabling no-PostgreSQL mode. If the same representative workflows round-trip cleanly through the data plane and look right in the UI, proceed to the next environment. ### Enable Topics [Topics](/docs/observe/topics) for self-hosted deployments has specific deployment, inference, billing, and maintenance requirements, and Braintrust must grant access to your organization. Completing the full v2.x upgrade and no-PostgreSQL mode migration is a prerequisite, and the following conditions must also be met before Topics can be enabled: * **Standard deployment required.** Topics is supported on standard Terraform/Helm deployments only. Custom Kubernetes clusters using the Helm chart on non-standard infrastructure are not supported. * **Braintrust-hosted inference required.** Topics runs LLM calls on models hosted by Braintrust with Zero Data Retention (ZDR). Bring-your-own-model is not supported. Currently US-only. * **Built-in models must be enabled.** Self-hosted organizations have [built-in models](/docs/admin/ai-providers#manage-built-in-models) disabled by default, so no trace data leaves your network boundary unless you opt in. Because Topics depends on built-in models, an admin must explicitly enable them before Topics can run. * **Pricing is not included by default.** Topics usage is not covered by standard contracts. Discuss pricing with your account team before proceeding. * **Ongoing maintenance required.** Self-hosted Topics deployments run on a frequent-release build track separate from the tagged data plane versions and will require regular upgrades to receive the latest fixes and improvements. If you meet all of these requirements and want to proceed, contact your account team or [Braintrust support](mailto:support@braintrust.dev). Until Braintrust enables Topics access for your organization, the Topics page and Topics automations remain blocked, even after the upgrade and no-PostgreSQL migration are complete. ## Configuration reference | Variable | Source | Pods | | - | - | - | | `BRAINSTORE_REDIS_URI` | From `REDIS_URL` secret | reader, writer, fastreader | | `BRAINSTORE_DEFAULT` | Hardcoded `"force"` | API | | `INSERT_LOGS2` | Hardcoded `"true"` | API | | `BRAINSTORE_INSERT_ROW_REFS` | Hardcoded `"true"` | API | | `BRAINSTORE_REALTIME_WAL_BUCKET` | From `objectStorage` values | API | | `BRAINSTORE_AI_PROXY_URL` | Built from `api.name` and `api.service.port` | reader, writer, fastreader | | `BRAINSTORE_XACT_MANAGER_URI` | From `REDIS_URL` secret | reader, writer, fastreader | | `BRAINSTORE_RESPONSE_CACHE_URI` | Auto-derived from `objectStorage` config | reader, writer, fastreader | | `BRAINSTORE_CODE_BUNDLE_URI` | Auto-derived from `objectStorage` config | reader, writer, fastreader | | Variable | How to set | Since | Pods | | - | - | - | - | | `BRAINSTORE_WAL_FOOTER_VERSION` | Set `brainstoreWalFooterVersion` in `helm-values.yaml` | v5.1.0 | API | | `BRAINSTORE_WAL_USE_EFFICIENT_FORMAT` | Auto-enabled when either `brainstoreWalFooterVersion` or `skipPgForBrainstoreObjects` is set | v5.1.0 | API | Set `brainstoreWalFooterVersion: "v1"` as part of the pre-upgrade steps. Only set to `"v3"` after successfully deploying v2.x. | Variable | How to set | Since | Pods | | - | - | - | - | | `SKIP_PG_FOR_BRAINSTORE_OBJECTS` | Set `skipPgForBrainstoreObjects` in `helm-values.yaml` | v5.1.0 | API | | `BRAINSTORE_ASYNC_SCORING_OBJECTS` | Auto-derived from `skipPgForBrainstoreObjects` | v5.1.0 | reader, writer, fastreader | | `BRAINSTORE_LOG_AUTOMATIONS_OBJECTS` | Auto-derived from `skipPgForBrainstoreObjects` | v5.1.0 | reader, writer, fastreader | Setting `skipPgForBrainstoreObjects: "all"` is a prerequisite for enabling Topics. The chart automatically fans this value out to `BRAINSTORE_ASYNC_SCORING_OBJECTS` and `BRAINSTORE_LOG_AUTOMATIONS_OBJECTS`. ## Next steps * [Self-hosting releases](/docs/data-plane-changelog): release notes and infrastructure requirements for each data plane version * [Configuration](/docs/admin/self-hosting/configure): telemetry, network access, rate limiting, and other options # Upgrade to Terraform module v6 Source: https://braintrust.dev/docs/admin/self-hosting/upgrade/v6 Upgrade your self-hosted Braintrust AWS deployment from Terraform module v5.x to v6 and cut API traffic over from Lambda to ECS. Terraform module v6 moves the Braintrust API and AI Proxy workloads from Lambda to ECS. For high-traffic data planes, ECS costs about 90% less than Lambda. This is an AWS-only infrastructure change and the move itself does not affect the data plane version. Each v6 release pins its own data plane images, so check [Self-hosting releases](/docs/data-plane-changelog) for the version your target release ships. Upgrade to v6.5.0 or later, not to v6.0.0. Go directly from v5.x to the latest v6 release. On module v6.4.0 and earlier, cutting API traffic over to ECS (`enable_ecs_api = true`) also routes LLM calls from user-authored scorers and tools running in the quarantine environment to the Braintrust-hosted Gateway at `gateway.braintrust.dev`, sending that traffic outside your AWS account. Module v6.5.0 and later routes those calls through the deployment's own AI proxy, or through your own [Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) if you run one. If you already cut over on an earlier v6 release, upgrade to v6.5.0 or later to keep that traffic in your account. Always upgrade one major version at a time. If you are on v4.x, upgrade to v5.x first before upgrading to v6. See the [routine upgrade guide](/docs/admin/self-hosting/upgrade/routine) for the standard process. ## What changed APIHandler and AIProxy now run as ECS services alongside the existing Lambdas. The ECS API is split into three services for different workload types: * `braintrust-api`: general API traffic * `braintrust-api-ingest`: ingestion paths * `braintrust-api-background`: background paths (evals, function invoke, proxy) The module routes each path to the right service for you. For the current path assignments and the variables that size each service, see [Scaling and storage](/docs/admin/self-hosting/configure/scaling#configure-aws-api-ecs-services). During the transition, the API Lambdas remain deployed and are kept warm. A future module release removes them. After cutover, none of the primary Braintrust services handling traffic run on Lambda. Braintrust continues to use Lambda only for small one-off maintenance tasks, such as database migrations and other automations. ## Upgrade steps Update your module source to a v6 release (v6.5.0 or later), leaving `enable_ecs_api` at its default (`false`): ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v6.7.0" # ... other configuration ... } ``` The ECS services use new variables. If you customized the Lambda equivalents on an earlier module version, carry those values into the ECS variables before you apply, or the ECS services will start without them: * Values in `service_extra_env_vars.APIHandler` or `service_extra_env_vars.AIProxy` must be duplicated into `braintrust_api_extra_env_vars`. * A version pinned with `lambda_version_tag_override` must be copied into `braintrust_api_version_override`. CloudFront terminates TLS for the data plane, so the API load balancer does not need its own certificate. If you do configure HTTPS on it with `braintrust_api_alb_certificate_arn` and `braintrust_api_alb_custom_domain`, set them in this first apply. Enabling them after the ECS API infrastructure exists will cause disruption. See [Set HTTPS on the API load balancer](/docs/admin/self-hosting/configure/networking#set-https-on-the-api-load-balancer). ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform init -upgrade terraform plan ``` Carefully review the output of `terraform plan` before applying. If you see something unexpected, like deletion of a database or S3 bucket, [contact Braintrust](mailto:support@braintrust.dev) for help. ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` This creates the ECS services, ALB, and related infrastructure. CloudFront continues to send traffic to Lambda. ECS and Lambda run side by side while ECS warms up. After the apply completes, exercise the data plane with a few calls in the Braintrust UI to confirm traffic is still flowing correctly through Lambda. Go to ** Settings** > [** Data plane**](https://www.braintrust.dev/app/~/configuration/org/api-url) and confirm all settings show green status. Set `enable_ecs_api = true` in your module configuration and apply again: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v6.7.0" enable_ecs_api = true # ... other configuration ... } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` CloudFront will route API traffic to the ECS ALB instead of API Gateway and Lambda. Exercise the data plane again. Run API requests, ingest traces, and run an eval to confirm traffic is flowing correctly through ECS. ## Rollback If you need to revert to Lambda after cutting over, set `enable_ecs_api = false` and apply: ```hcl theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} module "braintrust-data-plane" { source = "github.com/braintrustdata/terraform-aws-braintrust-data-plane?ref=v6.7.0" enable_ecs_api = false # ... other configuration ... } ``` ```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} terraform apply ``` CloudFront will revert to the Lambda path. This rollback is available as long as the Lambdas remain deployed. The custom CloudFront origin request policy (`-all-viewer-with-forwarded-proto`) is left in place after rollback, because CloudFront cannot delete a policy in the same apply that detaches it from the distribution. Terraform module v6.5.2 and later always creates the policy and attaches it only when API traffic routes to the ECS origin. On v6.5.1 and earlier, the rollback apply fails while trying to delete the policy, so upgrade to v6.5.2 or later before you roll back. ## Next steps * [Self-hosting releases](/docs/data-plane-changelog) — review release notes and infrastructure requirements for each data plane version. * [Configuration](/docs/admin/self-hosting/configure) — configure telemetry, network access, rate limiting, and other options. * [Deploy the Braintrust Gateway](/docs/admin/self-hosting/configure/networking#braintrust-gateway) — run the Gateway in your data plane (module v6.5.0 or later). Enabling it is independent of this cutover, so you can do either first. # Service health Source: https://braintrust.dev/docs/admin/service-health Monitor Braintrust service availability and subscribe to incident notifications The [Braintrust status page](https://status.braintrust.dev/) provides real-time health information for all centrally-hosted Braintrust services, including historical uptime and incident history. ## Monitored services | Service | Description | | - | - | | **Web UI / Control Plane** | The braintrust.dev UI and Control Plane API | | **Centrally-hosted data plane (US)** | Hosted data plane for reads and writes in the US region | | **Centrally-hosted data plane (EU)** | Hosted data plane for reads and writes in the EU region | | **AI Gateway** | Centrally-hosted AI Gateway for accessing LLM providers | If you use a [self-hosted data plane](/docs/admin/self-hosting), the status page reflects the health of Braintrust's control plane and the centrally-hosted components your deployment depends on. Your data plane's availability is managed within your own infrastructure. ## Subscribe to updates To receive notifications when incidents are created or updated: 1. Go to [status.braintrust.dev](https://status.braintrust.dev/). 2. Click **Subscribe to updates**. 3. Choose **email** or **RSS** and follow the prompts. ## Report a problem If you observe an issue not reflected on the status page, contact [info@braintrust.dev](mailto:info@braintrust.dev). # Create custom views Source: https://braintrust.dev/docs/annotate/custom-views Help your team understand and act on traces and dataset rows faster. Describe the interface you need in natural language, and Loop generates a customizable React component you can embed anywhere. Custom views transform complex traces and dataset rows into interfaces anyone on your team can use. Describe what you want in natural language and [** Loop**](/docs/loop) generates an interactive React component you can customize or embed anywhere. ## Common use cases Build custom annotation interfaces for large-scale human review tasks, surfacing only relevant information for annotators and subject matter experts. Replace JSON with intuitive UI components like carousels, playlists, or structured summaries to make traces accessible to PMs, legal reviewers, and domain experts. Create views that mirror your product experience: * Playlist-style views for music applications * Interactive source-and-answer layouts * Custom dashboards for internal evaluations Aggregate and display data across conversation turns to analyze dialogue flow and long-running interactions. ## Create views To create a custom view using [** Loop**](/docs/loop): 1. Select a trace or open a dataset row. 2. Select **Views**. 3. Describe how you want to view your data. After [** Loop**](/docs/loop) generates your view, refine it by describing additional changes or [edit the React component code](#edit-view-react-code) directly. To start without a prompt, select **Create default view** to load a pre-built template that renders the trace or dataset fields as formatted JSON. Example prompts: * "Create a view that renders a list of all tools available in this trace and their outputs" * "Build an interface to review each trace one by one with easy switching between traces" * "Create a conversation-style view that highlights user messages and assistant responses" * "Render the video url from the trace's metadata field and show simple thumbs up/down buttons" * "Create a side-by-side view of the input and expected output" * "Build an interface to review each row with easy switching between rows" * "Highlight differences between the input and expected fields" Self-hosted deployments: If you restrict outbound access, allowlist `https://www.braintrustsandbox.dev` to enable custom views. This domain hosts the sandboxed iframe that securely renders custom view code. ## Share views Custom views are stored differently depending on whether they've been saved: * **Unsaved views** exist only in the browser where you created them. If you close the browser, clear browser data, or switch to a different device, the view will not be available. * **Saved views** are stored in Braintrust, versioned, and available to all team members in the project. To save your view and share it with your team: 1. Select **Save** in the view editor. 2. Choose **Save as new view version**. 3. Select **Update** to make it available project-wide. All team members can then use the shared view when reviewing traces or dataset rows. Custom views integrate with Braintrust workflows. Use them during [human review](/docs/annotate/human-review), write annotations that flow into [datasets](/docs/annotate/datasets), and combine with [Loop](/docs/loop) for analysis. Selecting a saved view records it in the page URL as a `tv` parameter for traces or a `dv` parameter for datasets. Copy the URL to share a link that opens that exact view, even if the project has multiple saved views. Unsaved views aren't written to the URL. ## Edit view React code Custom views are React components that run inside Braintrust. You can edit the component code directly to customize behavior beyond what Loop generates. To edit the React code: 1. Go to the custom view. 2. Select in the lower left of the view. 3. Select **Edit**. Your React component receives props based on the view type: | Prop | Type | Description | | - | - | - | | `trace` | object | Contains all spans and methods for the trace. Attachment references in span data are automatically signed for rendering. | | `span` | object | The currently selected span with full data | | `update` | function | Legacy helper for selected-span metadata. Use `trace.update` instead: `update('field', value)` | | `selectSpan` | function | Navigate to a different span: `selectSpan(spanId)` | The `trace` object includes: * `rootSpanId`, `selectedSpanId` - Current span context * `spanOrder` - All span IDs in execution order * `spans` - Map of span\_id → span (IDs/relationships only) * `fetchSpanFields` - Fetch full data for multiple spans (see [Access data from multiple spans](#access-data-from-multiple-spans)) * `update` - Write supported fields back to a trace span When the view is open on a session grouped from multiple traces (using [**Group by**](/docs/observe/view-logs#group-related-traces) in the row type selector), the `trace` object contains traces from the session, not just the selected trace. `spanOrder` and `spans` include every span across those traces in session order, `fetchSpanFields` fetches across all of them, and `selectSpan` and `trace.update` operate on spans in any trace of the session. A session includes at most 1,000 traces. Spans from traces beyond that limit are not included. | Prop | Type | Description | | - | - | - | | `input` | object or string | The row's input field. | | `expected` | object or string | The row's expected field. | | `metadata` | object | The row's metadata field. | | `tags` | string\[] | Tags applied to the row. | | `id` | string | The row's unique identifier. | | `update` | function | Write supported fields back to the dataset row. | Attachment references in `input`, `expected`, and `metadata` are automatically signed for rendering. See [Render attachments](#render-attachments). `input` and `expected` are passed through as stored, which for most rows is a JSON string. When the stored value is a JSON object or array that contains an attachment reference, Braintrust parses it into an object so the attachment can be signed. Handle both shapes: ```javascript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} function parseField(value) { if (typeof value !== 'string') return value; try { return JSON.parse(value); } catch { return value; } } ``` The component can be copied and embedded in your own applications, enabling you to: * Reuse custom views outside of Braintrust * Integrate review interfaces into internal tools * Build standalone annotation applications * Create consistent review experiences across different contexts ### Add interactive controls Custom views support interactive elements that write data back to traces and dataset rows. Add buttons, inputs, or custom controls to collect: * Thumbs up/down feedback. * Custom metadata fields. * Custom score fields. * Dataset input corrections. * Dataset expected output corrections. * Tags. Use `trace.update` to write `metadata`, `scores`, or `tags` back to a trace span. In dataset views, use the `update` prop to write `metadata`, `input`, `expected`, or `tags` back to the current dataset row. Braintrust applies the update only when the current view is writable. Dataset fields like `input` and `expected` require the current view to support editing that field. | Field | Value | Notes | | - | - | - | | `metadata` | object | Supported on writable trace spans and dataset rows when metadata is editable. Updates the provided metadata keys. Omitted metadata keys are unchanged. | | `scores` | object | Supported on writable trace spans. Updates the provided score keys with numbers from 0 to 1 inclusive, or null. Omitted score keys are unchanged. | | `input` | any JSON-serializable value | Supported on dataset rows when the current view supports editing input. | | `expected` | any JSON-serializable value | Supported on dataset rows when the current view supports editing expected output. | | `tags` | string\[] or null | Supported on trace spans and dataset rows when tags are editable. Replaces the target row or span's tags. | When a field you write back contains an attachment that Braintrust signed for rendering, Braintrust restores the durable attachment reference before saving, so the row keeps the original attachment instead of a temporary signed URL. Inline attachments you add yourself with an `http` or `data:` URL are saved unchanged. By default, `trace.update` writes to the selected span. Use the optional `target` field to write to another span: | Target | Description | | - | - | | Omitted or `"selected"` | Update the currently selected span. | | `"root"` | Update the root span for the trace. | | `{ spanId: "..." }` | Update the span with the matching `span_id` or row `id`. | ```javascript title="Example: Add thumbs up/down buttons" theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} function FeedbackView({ trace, span }) { const handleFeedback = async (isPositive) => { await trace.update({ metadata: { user_feedback: isPositive ? 'positive' : 'negative', reviewed_at: new Date().toISOString(), }, scores: { user_feedback_score: isPositive ? 1 : 0, }, tags: ['reviewed'], }); }; return (

Review this output

        {JSON.stringify(span.data.output, null, 2)}
      
); } module.exports = FeedbackView; ``` ```javascript title="Example: Update a dataset row" theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} function DatasetReview({ update }) { const saveCorrection = async () => { await update({ metadata: { reviewed: true, reviewed_at: new Date().toISOString(), }, input: "corrected input value", expected: "corrected expected value", tags: ['reviewed'], }); }; return ; } module.exports = DatasetReview; ``` ### Access data from multiple spans By default, only the selected span has full data (input, output, expected, metadata). To access data from other spans, use `fetchSpanFields`: ```javascript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} // Fetch all fields for one span const data = await trace.fetchSpanFields(spanId); // Fetch specific fields for multiple spans const data = await trace.fetchSpanFields(trace.spanOrder, ['input', 'output']); ``` ```javascript title="Example: Display all span inputs" theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} function AllInputsView({ trace, span }) { const [spanData, setSpanData] = React.useState(null); const [loading, setLoading] = React.useState(true); const [error, setError] = React.useState(null); React.useEffect(() => { if (!trace?.fetchSpanFields) return; trace.fetchSpanFields(trace.spanOrder, ['input']) .then(setSpanData) .catch((err) => { console.error('Failed to fetch span data:', err.message); setError(err.message); }) .finally(() => setLoading(false)); }, [trace]); if (loading) return
Loading...
; if (error) return
Error: {error}
; return (
{trace.spanOrder.map((id) => (
          {JSON.stringify(spanData?.[id]?.input, null, 2)}
        
))}
); } module.exports = AllInputsView; ``` ### Render attachments [Attachments](/docs/instrument/attachments) (images, videos, audio, and other binary data) logged in your traces or dataset rows can be displayed directly in custom views. Braintrust automatically converts attachment references to `inline_attachment` objects with pre-signed URLs ready for rendering. Attachments are automatically transformed into objects with this structure: ```typescript theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} { type: "inline_attachment", src: "https://signed-url...", // Pre-signed URL ready to use content_type: "image/jpeg", // MIME type filename: "example.jpg", // Optional filename data?: string // Pre-fetched text content (JSON, text, CSV, XML, Markdown attachments) } ``` The `type` field identifies the object as an attachment, `src` contains a pre-signed URL that works directly in image, video, or audio tags, and `content_type` indicates the media type. For text-based attachment references (JSON, plain text, CSV, XML, and Markdown) that Braintrust signs on the viewer's behalf, Braintrust pre-fetches the content and populates the `data` field with the text string. Structured `inline_attachment` objects you log directly with an `http` or `data:` URL in `src` are passed through unchanged and will not have `data` populated. When `data` is present, render it directly instead of loading from `src`. Size limits apply to pre-fetching: 1MB per individual attachment and 5MB in aggregate across the entire payload being processed. In trace views, the initial custom view load signs the full `{trace, span}` object (which includes all spans in the trace), and each `fetchSpanFields` call signs all spans in that response together. In dataset views, the row's `input`, `expected`, and `metadata` fields are signed together as a single payload. Once the 5MB aggregate is reached within a single payload, remaining text attachments will have `data` omitted even if each individual attachment is under 1MB. The content is still accessible via `src` in all cases. For example, the following code creates an input/output verification view that automatically detects and renders attachments alongside regular data: ```javascript title="Example: Input/output verification" expandable theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}} function InputOutputVerification({ trace, span }) { // Helper to check if a value is an attachment const isAttachment = (value) => { return value && typeof value === 'object' && value.type === 'inline_attachment' && value.src; }; // Helper to render a value, handling attachments and regular data const renderValue = (value, label) => { if (!value && value !== 0 && value !== false) { return (
No {label.toLowerCase()}
); } // Check if it's an attachment if (isAttachment(value)) { const { src, filename, content_type, data } = value; if (content_type?.startsWith('image/')) { return (
{filename {filename &&

{filename}

}
); } if (content_type?.startsWith('video/')) { return