Independent eval research. Open methodology, published datasets, and the tools to run your own studies.

Original, open-source studies

Model comparisons, agent setups, cost, and modalities, measured on real tasks. Read the methodology and results, or rerun a study on your own data.

1 October 2026
Resend vs Postmark MCP for transactional email

Codex ran 45 transactional email tasks twice per MCP server. Seven deterministic scorers checked acceptance, content, recipients, and verification.

Emails accepted
Resend0%
Postmark0%

240 emails per providerhigher is better

Records available on first lookup
Resend0.0%
Postmark0.0%

240 emails per providerhigher is better

30 September 2026
GPT-6.1 Sol vs GPT-6 Sol for problem solving and teaching

We tested GPT-6.1 Sol on 175 MathTutorBench cases. It improved most at giving students a useful next step without revealing the answer.

Problems solved correctly
0.0%
+2.9 points
vs. GPT-6 Sol

125 paired cases · 0.6 points behind GPT-6 Astra

Tutoring guidance
0.0%
+13.5 points
vs. GPT-6 Sol

50 paired cases · useful next steps without giving away the answer

22 September 2026
Opus 5.5 vs new GPT-6 models for writing quality

Models answered 175 MathTutorBench cases covering mathematical problem solving and tutoring. Deterministic scorers checked correctness, and GLM 5.3 Flash judged writing quality.

Problems solved
Opus 5.50.00%
Fable 5.10.00%
GPT-6 Luna0.00%
GPT-6 Sol0.00%

125 objectively scored caseshigher is better

Writing quality
Opus 5.50.00%
GPT-6 Sol0.00%
GPT-6 Luna0.00%
Fable 5.10.00%

GLM 5.3 judgehigher is better

21 September 2026
When Jev holds up as a judge
9 September 2026
Moonshot vs Fireworks for Kimi K3 frontend agents
4 September 2026
Loop vs Braintrust MCP + Codex for production agent investigations
31 August 2026
You.com vs built-in web search
20 August 2026
Behavior scoring vs output scoring for coding agents
12 August 2026
Compare Kimi K3 and DeepSeek V4

Run evals like us

Each skill packages best practices from leading AI teams and researchers into a plain SKILL.md file your coding agent can read. Use them to get started running evals on your data.

Claude Code
Codex
Cursor
Gemini CLI
GitHub Copilot
opencode
Browse the repo

Point your agent at a skill

Ask it to do what the skill describes, like validating a scorer against human labels or finding hidden failure modes.

gh skill install braintrustdata/eval-library braintrust-validate-eval-scorer
skills/braintrust-validate-eval-scorer/SKILL.md
---
name: braintrust-validate-eval-scorer
description: Validate automated eval scorers and LLM judges against expert-reviewed reference data.
---
# Validate the scorer
  1. 1.Name the reference tier before computing anything.
  2. 2.Verify alignment: scorer outputs and reference labels must line up at the item and criterion level.
  3. 3.Report agreement (κ or α) with uncertainty, not raw accuracy.
  4. 4.Lead with the most decision-relevant false acceptance before any aggregate, and enumerate the dangerous cells case by case.
  5. 5.Break errors down by class and severity.
  6. 6.Test sensitivity and shortcuts: inject known regressions and improvements and confirm the scorer moves; probe whether length, confidence, or polish raise the score independent of quality; probe whether text addressed to the judge moves it; slice agreement by subgroup.

Subscribe to the Department of Evals

A newsletter for unfiltered thoughts on eval methodology, analysis, and failures

Subscribe