Model comparisons, agent setups, cost, and modalities, measured on real tasks. Read the methodology and results, or rerun a study on your own data.
Codex ran 45 transactional email tasks twice per MCP server. Seven deterministic scorers checked acceptance, content, recipients, and verification.
240 emails per providerhigher is better
240 emails per providerhigher is better
We tested GPT-6.1 Sol on 175 MathTutorBench cases. It improved most at giving students a useful next step without revealing the answer.
125 paired cases · 0.6 points behind GPT-6 Astra
50 paired cases · useful next steps without giving away the answer
Models answered 175 MathTutorBench cases covering mathematical problem solving and tutoring. Deterministic scorers checked correctness, and GLM 5.3 Flash judged writing quality.
125 objectively scored caseshigher is better
GLM 5.3 judgehigher is better
Each skill packages best practices from leading AI teams and researchers into a plain SKILL.md file your coding agent can read. Use them to get started running evals on your data.
Ask it to do what the skill describes, like validating a scorer against human labels or finding hidden failure modes.