GPT-6.1 Sol vs Claude Sonnet 5.5 is the most interesting model matchup of the month, because for the first time the two labs have shipped mid-tier models at exactly the same list price. Both cost $2 per million input tokens and $10 per million output tokens, and they launched one day apart.
The short answer: Claude Sonnet 5.5 is the stronger model, and GPT-6.1 Sol is the cheaper one to actually run. Those are not the same thing, and the gap between them is the whole story of this comparison.
I read OpenAI’s launch material, Anthropic’s Sonnet 5.5 numbers, and the independent results Artificial Analysis published this week. I have not run either model on my own benchmark suite yet, so everything below is sourced, and I flag which numbers are vendor-reported.
What GPT-6.1 Sol Actually Is
GPT-6.1 Sol is OpenAI’s mid-priced reasoning model, released at DevDay on September 29, 2026, as a direct upgrade to GPT-6 Sol. It replaced Sol less than two weeks after Sol shipped on September 22, which tells you how fast this tier is moving.
The specs, from OpenAI’s launch pricing and the coverage that followed it:
- Price: $2 input / $10 output per million tokens, unchanged from GPT-6 Sol
- Cached input: $0.10 per million, half of GPT-6 Sol’s $0.20
- Long prompts: above 272K input tokens, input is billed at 2x and output at 1.5x
- Context window: 1.05 million tokens, with 128K max output
- Where: the API for all developers, plus Codex and ChatGPT Work for Plus, Pro, Business, Enterprise and Edu plans
The cached input price is the change I care about most. Agentic coding sessions re-send the same files and instructions on every turn, so most of the input bill is cache reads. Halving that rate cuts real Codex costs more than any headline change would.
The catch is availability inside ChatGPT. According to Fello AI’s breakdown of the launch, GPT-6.1 Sol is in the agent surfaces only, and regular ChatGPT chat users do not have it yet. For developers that hardly matters, since Codex and the API are where the work happens.
GPT-6.1 Sol is a price-performance release, and OpenAI built it to make Astra unnecessary for most coding work.
GPT-6.1 Sol Benchmarks: What OpenAI Claims
OpenAI’s pitch is that 6.1 Sol nearly matches its flagship. These are the vendor-reported numbers as relayed by LMSpedia and Layer3 Labs from OpenAI’s launch post:
| Benchmark | GPT-6.1 Sol | Comparison |
|---|---|---|
| DeepSWE v1.1 (agentic coding) | 75.2% | GPT-6 Sol 68.8%, matches Astra |
| OSWorld 2.0 (computer use) | 71.4% | within 2.1 points of Astra |
| GDP.pdf (professional work) | 32.0% | Astra 32.2% |
| AutomationBench 1.0.6 | +2.2 pts vs Opus 5.5 | at medium effort |
| Factual error rate (low effort) | 7.7% | GPT-6 Sol 11.4% |
(Source: OpenAI launch post via LMSpedia and Layer3 Labs, September 2026. All figures vendor-reported.)
The 6.4-point DeepSWE jump over GPT-6 Sol is big for a point release. For context, Google’s own table puts Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1% on the same benchmark, which I covered in the Gemini 4 Argon breakdown. On OpenAI’s numbers, a $2 model now sits level with $4 and $10 flagships on agentic coding.
The cost figure is the one that made me read the table twice. Szymon Paluch’s DevDay pricing analysis reports the average Terminal-Bench Science task costing $5.47 on GPT-6.1 Sol against $23.80 on Astra. If those numbers hold, Astra is now very hard to justify for anything but the hardest tasks.
There is a gap in this table, and it is the one that matters for this post. OpenAI does not compare GPT-6.1 Sol to Claude Sonnet 5.5 anywhere in its launch material. It compares to Opus 5.5 on selected benchmarks and to its own models on the rest. Anthropic publishes Terminal-Bench 4.0 and FrontierCode, OpenAI publishes DeepSWE and OSWorld 2.0, and neither lab reports the other’s headline benchmark.
So the vendor tables cannot settle this matchup. For that you need someone who runs both.
GPT-6.1 Sol vs Claude Sonnet 5.5 on Independent Tests
Artificial Analysis ran both models through its Intelligence Index, and this is the closest thing to a neutral comparison available right now.
| Metric | GPT-6.1 Sol | Claude Sonnet 5.5 |
|---|---|---|
| Intelligence Index (max effort) | 52 | 56 |
| Engineering category | 54 | 58 |
| Output speed | 55 tokens/s | 139 tokens/s |
| Cost per task (low effort) | $0.13 | $0.42 |
| List price (in / out per 1M) | $2 / $10 | $2 / $10 |
| Context window | 1.05M | 1M |
(Source: Artificial Analysis model comparison and Sonnet 5.5 article, October 2026.)
Claude Sonnet 5.5 leads GPT-6.1 Sol by 4 points on the Intelligence Index and wins every category Artificial Analysis reports. Sonnet 5.5 sits at number two overall, behind only Claude Opus 5.5 at 58. GPT-6.1 Sol scores 52, one point behind GPT-6 Astra’s 53 and four points up on the original GPT-6 Sol’s 48.
That last comparison is the one OpenAI should be pleased with. A one-point gap to its own flagship at one-fifth of the price is a strong result. It just does not beat Anthropic’s mid-tier model.
Speed favours Sonnet too. At 139 tokens per second against 55, Sonnet 5.5 streams output more than twice as fast.
Then the cost row flips the whole table. At low effort, GPT-6.1 Sol costs $0.13 per task against $0.42 for Sonnet 5.5, more than three times cheaper at an identical list price.
Why the Same Price Is Not the Same Cost
The explanation is token usage. Artificial Analysis says Sonnet 5.5 used about 193,000 output tokens per task on its index, the highest it has measured, roughly 60% more than Opus 5.5 at max effort and about seven times more than GPT-6 Astra.
Sonnet 5.5 thinks out loud, a lot. That is a big part of why it scores well, and it is also why per-token pricing tells you very little about what a task will cost.
This shows up in a strange place in the Artificial Analysis data. Sonnet 5.5 generates tokens 2.5x faster, yet its end-to-end response time on their test was about 432 seconds against 292 for GPT-6.1 Sol. A faster model that writes far more tokens can still finish later.
This is the pattern with every reasoning-heavy model, and it is the part of model pricing that launch posts never mention. The bill tracks how much the model decides to think, not the price page. If you are budgeting a team rollout, measure cost per completed task on your own work for a week before you commit. The Claude API cost reduction techniques I wrote about earlier, mostly caching and effort tuning, apply directly to Sonnet 5.5.
At equal list prices, the model that uses fewer tokens wins on cost, and right now that is GPT-6.1 Sol by roughly 3x.
Claude Sonnet 5.5 vs GPT-6.1 Sol for Agentic Coding
For hard, multi-step coding work, I would still pick Sonnet 5.5.
Anthropic reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0, higher than Opus 5.5’s 66.4%, and Artificial Analysis’ own Terminal-Bench 4.0 run also has Sonnet ahead of Opus, 64% to 60%. Those are long, terminal-driven agent tasks, which is what Claude Code and Codex actually do all day. The 4-point engineering lead on the independent index points the same way.
GPT-6.1 Sol’s DeepSWE score is impressive, but it is OpenAI’s number on a benchmark Anthropic has not reported for Sonnet. I cannot line them up honestly, so I am not going to pretend the comparison is closer than the independent data says.
Where GPT-6.1 Sol makes more sense:
- High-volume pipelines, such as CI triage, test generation across hundreds of files, or bulk refactors where each task is simple and the bill is the constraint
- Factual or document-heavy work, where OpenAI reports a lower error rate than the previous Sol
- Teams already on Codex, where 6.1 Sol is now the obvious default below Astra
- Long-context jobs that stay under 272K tokens, because the long-prompt surcharge kicks in above that line
Where Sonnet 5.5 earns its higher per-task cost:
- Open-ended multi-file changes where one wrong turn wastes an hour of review
- Long agent runs in Claude Code, where the Terminal-Bench lead matters most
- Anything interactive, where 139 tokens per second feels noticeably faster in the terminal
My comparison of Claude Code, Cursor and Codex covers the tool side of this choice. The model choice now mostly follows the tool: Sonnet 5.5 if you live in Claude Code, GPT-6.1 Sol if you live in Codex.
For agentic coding quality, Sonnet 5.5 wins. For cost per task, GPT-6.1 Sol wins, and by a wider margin than Sonnet wins on quality.
Where Opus 5.5 and Astra Fit Now
Both mid-tier releases put pressure on their own flagships.
Claude Opus 5.5, at $4 / $20, still leads the Artificial Analysis index at 58 and leads on FrontierCode and CursorBench, according to Anthropic. I made the case for it in the Claude Opus 5.5 vs GPT-6 comparison, and that holds for the hardest open-ended work.
GPT-6 Astra is in a worse spot. At $10 / $50 it scores 53 on the independent index, one point above a model that costs a fifth as much. Unless a specific task measurably fails on GPT-6.1 Sol and passes on Astra, I cannot find a coding job where Astra is the rational pick this month.
The mid-tier is where the real competition is now, and both labs’ $2 models are good enough for most daily coding.
The Verdict: GPT-6.1 Sol or Claude Sonnet 5.5?
Use Claude Sonnet 5.5 if quality per task matters more than cost per task. It scores higher on every independent measure, streams faster, and leads on the terminal benchmarks that look most like real agent work.
Use GPT-6.1 Sol if you run a lot of tasks and pay per token. Three times cheaper per task at low effort is a large difference, and on OpenAI’s numbers it gives up very little to Astra.
If you run a team, the sensible setup is both. Route the hard, ambiguous work to Sonnet 5.5 and the high-volume, well-specified work to GPT-6.1 Sol. Inside Claude Code, the new Claude Code mods can even change the model for a single request, which makes per-step routing a few lines of TypeScript. The labs have priced these two models identically, but they have not built them to do the same job.
I will update this comparison when independent DeepSWE and Terminal-Bench runs include both models.