Eval library

All research

Model comparisons, agent setups, cost, and modalities, measured on real tasks.

Paper MCP vs Figma MCP for frontend agents

An independent eval of the Paper and Figma MCP servers as design tools for a coding agent, scored on visual quality, consistency, and cost.

20 July 2026
Evaluating the GPT-5.6 family

I mapped the GPT-5.6 family across task families and difficulty, with an Anthropic comparison, to find the cheapest model that clears your reliability bar.

10 July 2026
Evaluating speech-to-text models

I ran a controlled eval across six speech-to-text providers, 240 audio clips, and eight content domains to find where voice agents break and which STT model comes out ahead.

9 July 2026
Evaluating the USA vs Belgium World Cup matchup

Applying the best-performing Parallel configuration from the World Cup eval to inspect the USA vs Belgium matchup as a source-backed research map.

6 July 2026
From World Cup matchups to research maps: evaluating Parallel's web research agents

We used Braintrust to evaluate six Parallel web research configurations across 48 World Cup matchups

2 July 2026
Benchmarking GLM-5.2 vs Opus 4.8 for real-world long-context retrieval

A Braintrust-native eval comparing GLM-5.2, Opus 4.8, and Sonnet 5 on exact long-context retrieval, cost, and latency.

30 June 2026
GLM-5.2 vs. Opus 4.8 technical report

The full methodology, hypotheses, and findings for Braintrust's GLM-5.2, Opus 4.8, (and Sonnet 5!) long-context retrieval eval.

30 June 2026
Using OSS models to save on inference costs without cutting quality

Call GLM-5.2 directly in playgrounds, prompts, and scorers, with no configuration, through July 31.

30 June 2026
Using Braintrust to eval agentic setups from large-scale Hugging Face data

We pulled 1,781 real agent traces from Hugging Face into Braintrust, scored every run, and found that the wrapper around your model explains 7× more variation in success than the model itself.

24 June 2026
How to test agent cost-efficiency with Braintrust

How the cheapest model is not always the cheapest system, and how evals plus control logic reduce cost per resolved request.

17 June 2026
Testing if "bash is all you need"

Testing whether filesystems and bash provide the optimal abstraction for AI agents through rigorous evaluation.

22 January 2026
Claude Sonnet 4.5 analysis

Learn how aspirational evals can help you figure out when new AI models unlock new product opportunities.

29 September 2025
GPT-5 vs. Claude Opus 4.1

Which one you should ship with, and how to know for sure.

8 August 2025
Building with Grok 4

xAI recently announced Grok 4. We put it to the ultimate test.

11 July 2025
Evaluating Gemini models for vision

Faster, more efficient, and highly accurate for real-world applications.

14 November 2024

Subscribe to the Department of Evals

A newsletter for unfiltered thoughts on eval methodology, analysis, and failures

Subscribe