30 September 2026

GPT-6.1 Sol vs GPT-6 Sol for problem solving and teaching

Izzy Hurley8 min
Key takeaways
+13.5 percentage points in pedagogy for GPT-6.1 Sol over GPT-6 Sol
GPT-6.1 Sol gave better guidance without revealing the answer, closing most of the teaching-quality gap to Astra.
+2.9 percentage points in accuracy for GPT-6.1 Sol over GPT-6 Sol
The modest problem-solving gain came mostly from correcting students’ mistakes.
≈1/5 of Astra’s generation cost for GPT-6.1 Sol
GPT-6.1 Sol cost about $0.0012 per response versus Astra’s $0.0059 in this workload, while reaching 94% of Astra’s pedagogy score. Grading costs are excluded.

GPT-6.1 Sol looks like a targeted update to certain behaviors, and we certainly see that show up in our teaching metric. It moved much closer to Astra on the behavior that mattered most in this eval without moving into Astra’s price range.

The 6.1 update largely preserved the behaviors we saw in GPT-6 Sol and showed major improvements in how it guided users through the interaction. If your workflow depends on a model following a specific behavioral approach, guiding someone through a problem, or following prescriptive instructions, it’s worth exploring whether this Sol update can replace your Astra workflow.


One week after releasing GPT-6 Sol, OpenAI has already released an update. GPT-6.1 Sol promises performance close to Astra on coding, computer use, and professional work, at about one-fifth of the token cost. But how does it compare with last week’s release? A paired audit of three Codex tasks documented missed context and one incomplete synchronization task with GPT-6 Sol, alongside a separate case where GPT-6 Sol found a serious defect that GPT-5.6 Sol missed. We wanted to explore what this update changes.

Last week, we compared the new GPT-6 models and Opus 5.5 using MathTutorBench, which tests problem solving alongside writing and teaching. We returned to the same dataset because it separates three things that can easily get collapsed into “model quality”: getting an answer right, explaining it clearly, and helping a student learn without simply giving them the answer.

The gains were narrowly concentrated. GPT-6.1 Sol’s problem-solving accuracy rose from 75.8% to 78.6%. Its pedagogy score climbed from 64.1% to 77.6%, much closer to Astra’s 82.2%. Writing quality stayed roughly flat.

Most of the accuracy improvement came from mistake correction, while most of the pedagogy improvement came from giving better scaffolding without revealing the answer. That behavior is relevant beyond tutoring, since agents also need to follow a requested approach instead of rushing toward a plausible result. MathTutorBench cannot tell us whether GPT-6.1 Sol is easier to steer through a long coding or browser workflow, but it gives us a specific behavior to test.

Methodology

MathTutorBench covers seven tasks. Some ask the model to solve a problem or inspect a student’s work. Others ask it to tutor: give a useful hint, ask a productive question, or follow a teaching strategy without handing over the answer. That mix is useful because a model can solve the problem and still be a bad tutor.

We ran GPT-6 Sol, GPT-6.1 Sol, and GPT-6 Astra on the same 175 sampled cases, with 25 cases from each task. Each model produced four independent responses per case, for 700 responses per model. All three used the same settings: medium reasoning and low verbosity. We averaged the four responses within each case before comparing models, so every case retained equal weight.

We scored three dimensions separately:

  • Problems solved uses the benchmark’s task-specific checks for numerical answers, yes-or-no judgments, mistake locations, corrections, and teaching questions.
  • Writing quality averages clarity, answer-first structure, responsiveness, naturalness, and verifiability across all seven tasks.
  • Teaching quality measures whether an open-ended tutoring response moves the student forward without giving away the answer.

Teaching quality is the teaching-quality score in the charts below. It is also a helpful indicator of instruction following in this benchmark, because the tutoring tasks tell the model how to help and score whether it followed that request instead of solving the problem outright.

We used GLM-5.3-Flash as an independent judge for writing and tutoring. It graded every model with the same rubric and did not evaluate its own responses.¹ We estimated the confidence intervals by resampling the cases 10,000 times. We also ran two-sided paired t-tests on the per-case averages, treating p < 0.05 as statistically distinguishable. For statistically distinguishable differences, we report paired Cohen’s d to show the size of the effect.

Results

Cost versus correctness

First, we put GPT-6.1 Sol back into the full field from the original eval. It reached 78.6% problem-solving accuracy at an observed generation cost of about $0.12 per 100 responses. That is slightly cheaper than GPT-6 Sol in this run and far cheaper than Astra, Opus 5.5, or Fable 5.1.

GPT-6.1 Sol is cheaper than GPT-6 Sol and more accurate
78.6% accuracy at $0.12 per 100 responses
GPT-6 models in color, other models in gray · Lines show 95% confidence intervals · The vertical scale is zoomed
Problems solved
70%
75%
80%
85%
$0.01
$0.10
$1.00
Observed model cost per 100 responses (log scale)

The accuracy intervals overlap widely. Because every model answered the same cases, the paired comparison can still distinguish GPT-6.1 Sol from GPT-6 Sol by measuring the difference case by case. It does not establish a new accuracy leader across the full field. GPT-6 Luna is still the cheapest model by a wide margin.

Cost versus writing quality

Writing tells a different story. GPT-6 Luna had the strongest GPT-6 writing score at 83.8%, followed by GPT-6 Sol at 83.6%. GPT-6.1 Sol scored 82.9%, slightly ahead of Astra’s 82.5%, while costing about one-fifth as much as Astra in this workload.

Writing quality stays in the same cluster
GPT-6.1 Sol scored 82.9%, within 0.7 points of GPT-6 Sol
GPT-6 models in color, other models in gray · Lines show 95% confidence intervals · The vertical scale is zoomed
Writing quality
80%
82%
84%
86%
$0.01
$0.10
$1.00
Observed model cost per 100 responses (log scale)

Opus 5.5 still has the highest writing point estimate. GPT-6 Luna remains unusually competitive for its price. GPT-6.1 Sol lands in the middle of that cluster rather than pushing the writing frontier forward.

The GPT-6 family comparison

Putting the three GPT-6 models side by side makes the tradeoff easier to see.

GPT-6.1 Sol gained most on teaching quality
Mean score on the same cases, GPT-6 Sol compared with GPT-6.1 Sol, with GPT-6 Astra for reference
Change is GPT-6.1 Sol minus GPT-6 Sol, paired by case, with its 95% confidence interval · Hover a dot for that model’s interval
Change
Problems solved125 cases
Writing quality175 cases
Teaching quality50 cases
60%
70%
80%
90%
Mean score
−0.7 pts
−1.4 pts to +0.1 pts
GPT-6 Sol
GPT-6.1 Sol
GPT-6 Astra

Problem solving barely moved

GPT-6 Sol solved 75.8% of the objectively scored cases. GPT-6.1 Sol reached 78.6%, and Astra reached 79.2%. The difference between GPT-6 Sol and GPT-6.1 Sol was 2.9 points, with a 95% confidence interval from +0.6 to +5.6 points. A paired t-test found that difference statistically distinguishable from zero, though small in magnitude. Astra led GPT-6 Sol by 3.5 points, with an interval from +1.5 to +5.8 points.

The four-trial run changes the conclusion from the first pass. GPT-6.1 Sol’s accuracy improvement is modest, but both the confidence interval and paired t-test now separate it from zero.

The task breakdown shows where that lift came from. Both Sol versions solved every direct problem-solving case and scored 92% on judging whether a student’s solution was correct. GPT-6.1 Sol improved from 90% to 97% on mistake correction, from 73% to 76% on locating mistakes, and from 23.8% to 28.2% on Socratic questioning. The largest contribution came from correcting a known mistake, rather than from a broad jump across every task.

The improvement showed up in pedagogy

GPT-6.1 Sol’s pedagogy score rose from 64.1% to 77.6%, closing most of the gap with Astra’s 82.2%. The 13.5-point gain was statistically significant (paired t-test, p = 0.0026; Cohen’s d = 0.45).

Most of that improvement came from scaffolding, where the score jumped from 40.0% to 66.3%. This is the work of helping a student get unstuck: giving a hint, asking a useful question, or helping them revisit a mistake while leaving them room to figure it out. The goal is to give enough support to move them forward without taking over the problem.

The score for following explicit teaching instructions barely changed, moving from 88.3% to 89.0%. So the clearest gain was in how the model guided the interaction when it had less detailed direction about how to teach.

What changed inside the writing score

The overall writing score barely changed, but the individual scores show a small shift in style: GPT-6.1 Sol stayed clear and responsive, while sounding slightly less natural and getting to the answer less directly.

GPT-6.1 Sol stays clear, but answers less directly
Five writing rubric dimensions across all 175 cases
Change is GPT-6.1 Sol minus GPT-6 Sol · Hover a dot for that model’s 95% confidence interval
Change
Clarity
Answer first
Responsiveness
Naturalness
Verifiability
50%
60%
70%
80%
90%
100%
Mean score
+0.6 pts
−2.4 pts
+0.7 pts
−1.6 pts
−0.8 pts
GPT-6 Sol
GPT-6.1 Sol
GPT-6 Astra

The biggest drop was in answer-first structure, from 85% to 82%. Naturalness and verifiability also dipped slightly. Astra showed a similar pattern: strong clarity and responsiveness, with less consistent performance on leading with the answer and giving readers enough detail to check it.

Why GPT-6.1 Sol cost less in this run

OpenAI prices GPT-6.1 Sol at the same standard token rates as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. In our runs, GPT-6.1 Sol cost about $0.0012 per response, compared with $0.0015 for GPT-6 Sol and $0.0059 for Astra.

The lower observed cost relative to GPT-6 Sol came from shorter outputs and fewer reasoning tokens, since the two models have the same list price. Compared with Astra, GPT-6.1 Sol delivered 94% of its pedagogy score at about one-fifth of the generation cost.

These costs include only the three models’ responses. Grading costs are excluded so the comparison reflects the models being evaluated.

What this does and does not tell us

MathTutorBench does not test the headline agentic workloads from OpenAI’s launch, but it does expose a representative interaction many people have with models every day where they are asking questions, iterating on responses, and landing at an answer.


¹ Judge configuration. zai-org/GLM-5.3-Flash, served through Baseten via the Braintrust Gateway, using Chat Completions, low reasoning effort, and max_completion_tokens=2048. The same model graded writing and pedagogy in separate calls.

Share

Read more evals

Opus 5.5 vs new GPT-6 models for writing quality
22 September 2026
When Jev holds up as a judge
21 September 2026
Moonshot vs Fireworks for Kimi K3 frontend agents
9 September 2026

Subscribe to the Department of Evals

A newsletter for unfiltered thoughts on eval methodology, analysis, and failures

Subscribe