Izzy Hurley8 minGPT-6.1 Sol looks like a targeted update to certain behaviors, and we certainly see that show up in our teaching metric. It moved much closer to Astra on the behavior that mattered most in this eval without moving into Astra’s price range.
The 6.1 update largely preserved the behaviors we saw in GPT-6 Sol and showed major improvements in how it guided users through the interaction. If your workflow depends on a model following a specific behavioral approach, guiding someone through a problem, or following prescriptive instructions, it’s worth exploring whether this Sol update can replace your Astra workflow.
One week after releasing GPT-6 Sol, OpenAI has already released an update. GPT-6.1 Sol promises performance close to Astra on coding, computer use, and professional work, at about one-fifth of the token cost. But how does it compare with last week’s release? A paired audit of three Codex tasks documented missed context and one incomplete synchronization task with GPT-6 Sol, alongside a separate case where GPT-6 Sol found a serious defect that GPT-5.6 Sol missed. We wanted to explore what this update changes.
Last week, we compared the new GPT-6 models and Opus 5.5 using MathTutorBench, which tests problem solving alongside writing and teaching. We returned to the same dataset because it separates three things that can easily get collapsed into “model quality”: getting an answer right, explaining it clearly, and helping a student learn without simply giving them the answer.
The gains were narrowly concentrated. GPT-6.1 Sol’s problem-solving accuracy rose from 75.8% to 78.6%. Its pedagogy score climbed from 64.1% to 77.6%, much closer to Astra’s 82.2%. Writing quality stayed roughly flat.
Most of the accuracy improvement came from mistake correction, while most of the pedagogy improvement came from giving better scaffolding without revealing the answer. That behavior is relevant beyond tutoring, since agents also need to follow a requested approach instead of rushing toward a plausible result. MathTutorBench cannot tell us whether GPT-6.1 Sol is easier to steer through a long coding or browser workflow, but it gives us a specific behavior to test.
MathTutorBench covers seven tasks. Some ask the model to solve a problem or inspect a student’s work. Others ask it to tutor: give a useful hint, ask a productive question, or follow a teaching strategy without handing over the answer. That mix is useful because a model can solve the problem and still be a bad tutor.
We ran GPT-6 Sol, GPT-6.1 Sol, and GPT-6 Astra on the same 175 sampled cases, with 25 cases from each task. Each model produced four independent responses per case, for 700 responses per model. All three used the same settings: medium reasoning and low verbosity. We averaged the four responses within each case before comparing models, so every case retained equal weight.
We scored three dimensions separately:
Teaching quality is the teaching-quality score in the charts below. It is also a helpful indicator of instruction following in this benchmark, because the tutoring tasks tell the model how to help and score whether it followed that request instead of solving the problem outright.
We used GLM-5.3-Flash as an independent judge for writing and tutoring. It graded every model with the same rubric and did not evaluate its own responses.¹ We estimated the confidence intervals by resampling the cases 10,000 times. We also ran two-sided paired t-tests on the per-case averages, treating p < 0.05 as statistically distinguishable. For statistically distinguishable differences, we report paired Cohen’s d to show the size of the effect.
First, we put GPT-6.1 Sol back into the full field from the original eval. It reached 78.6% problem-solving accuracy at an observed generation cost of about $0.12 per 100 responses. That is slightly cheaper than GPT-6 Sol in this run and far cheaper than Astra, Opus 5.5, or Fable 5.1.
The accuracy intervals overlap widely. Because every model answered the same cases, the paired comparison can still distinguish GPT-6.1 Sol from GPT-6 Sol by measuring the difference case by case. It does not establish a new accuracy leader across the full field. GPT-6 Luna is still the cheapest model by a wide margin.
Writing tells a different story. GPT-6 Luna had the strongest GPT-6 writing score at 83.8%, followed by GPT-6 Sol at 83.6%. GPT-6.1 Sol scored 82.9%, slightly ahead of Astra’s 82.5%, while costing about one-fifth as much as Astra in this workload.
Opus 5.5 still has the highest writing point estimate. GPT-6 Luna remains unusually competitive for its price. GPT-6.1 Sol lands in the middle of that cluster rather than pushing the writing frontier forward.
Putting the three GPT-6 models side by side makes the tradeoff easier to see.
GPT-6 Sol solved 75.8% of the objectively scored cases. GPT-6.1 Sol reached 78.6%, and Astra reached 79.2%. The difference between GPT-6 Sol and GPT-6.1 Sol was 2.9 points, with a 95% confidence interval from +0.6 to +5.6 points. A paired t-test found that difference statistically distinguishable from zero, though small in magnitude. Astra led GPT-6 Sol by 3.5 points, with an interval from +1.5 to +5.8 points.
The four-trial run changes the conclusion from the first pass. GPT-6.1 Sol’s accuracy improvement is modest, but both the confidence interval and paired t-test now separate it from zero.
The task breakdown shows where that lift came from. Both Sol versions solved every direct problem-solving case and scored 92% on judging whether a student’s solution was correct. GPT-6.1 Sol improved from 90% to 97% on mistake correction, from 73% to 76% on locating mistakes, and from 23.8% to 28.2% on Socratic questioning. The largest contribution came from correcting a known mistake, rather than from a broad jump across every task.
GPT-6.1 Sol’s pedagogy score rose from 64.1% to 77.6%, closing most of the gap with Astra’s 82.2%. The 13.5-point gain was statistically significant (paired t-test, p = 0.0026; Cohen’s d = 0.45).
Most of that improvement came from scaffolding, where the score jumped from 40.0% to 66.3%. This is the work of helping a student get unstuck: giving a hint, asking a useful question, or helping them revisit a mistake while leaving them room to figure it out. The goal is to give enough support to move them forward without taking over the problem.
The score for following explicit teaching instructions barely changed, moving from 88.3% to 89.0%. So the clearest gain was in how the model guided the interaction when it had less detailed direction about how to teach.
The overall writing score barely changed, but the individual scores show a small shift in style: GPT-6.1 Sol stayed clear and responsive, while sounding slightly less natural and getting to the answer less directly.
The biggest drop was in answer-first structure, from 85% to 82%. Naturalness and verifiability also dipped slightly. Astra showed a similar pattern: strong clarity and responsiveness, with less consistent performance on leading with the answer and giving readers enough detail to check it.
OpenAI prices GPT-6.1 Sol at the same standard token rates as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. In our runs, GPT-6.1 Sol cost about $0.0012 per response, compared with $0.0015 for GPT-6 Sol and $0.0059 for Astra.
The lower observed cost relative to GPT-6 Sol came from shorter outputs and fewer reasoning tokens, since the two models have the same list price. Compared with Astra, GPT-6.1 Sol delivered 94% of its pedagogy score at about one-fifth of the generation cost.
These costs include only the three models’ responses. Grading costs are excluded so the comparison reflects the models being evaluated.
MathTutorBench does not test the headline agentic workloads from OpenAI’s launch, but it does expose a representative interaction many people have with models every day where they are asking questions, iterating on responses, and landing at an answer.
¹ Judge configuration. zai-org/GLM-5.3-Flash, served through Baseten via the Braintrust Gateway, using Chat Completions, low reasoning effort, and max_completion_tokens=2048. The same model graded writing and pedagogy in separate calls.
A newsletter for unfiltered thoughts on eval methodology, analysis, and failures
Subscribe