Claude Opus 5.5 vs GPT-6 is the model decision every developer using Claude Code or Codex has to make this month. Both labs shipped their new families within the same week in September 2026, and both claim to be the best at coding.

On the published numbers, Claude Opus 5.5 beats GPT-6 Astra on agentic coding at 40% of the price. That is not the whole story, though. OpenAI’s cheaper GPT-6 Sol and Luna tiers are better value than Astra for a lot of real work, and Anthropic’s own Sonnet 5.5 outscores Opus 5.5 on one major benchmark.

I went through both labs’ launch tables, worked out per-task costs, and checked which claims hold up. Here is the comparison I would want if I were choosing for a team today.

Claude Opus 5.5 vs GPT-6: The Short Answer

  • Hard, multi-file agentic coding: Claude Opus 5.5
  • Daily coding at mid-tier prices: Claude Sonnet 5.5 or GPT-6 Sol, both $2 / $10
  • High-volume, cheap tasks (classification, simple edits, CI triage): GPT-6 Luna at $0.10 / $0.50
  • Only if you need OpenAI-specific features: GPT-6 Astra, at $10 / $50 it is hard to justify on cost

For live prices across every model, see the AI model pricing tracker. The rest of this post explains where those recommendations come from, and where each one breaks down.

The Lineups: What Actually Launched in September 2026

Both labs now sell a tiered family rather than a single flagship. That makes the comparison less “Claude vs GPT” and more “which tier for which job”.

Anthropic

  • Claude Opus 5.5 (September 22): $4 input / $20 output per million tokens, cache reads $0.20. Fast mode runs up to 2.5x faster at $8 / $40. Five effort levels from low to max, with medium as the default. Available in the Claude API, AWS, Google Cloud, Azure, Claude Code, and all paid Claude plans.
  • Claude Sonnet 5.5 (September 28): $2 / $10, the same price as Sonnet 5. 1 million token context window and 128K output, according to MarkTechPost’s launch coverage.

OpenAI

  • GPT-6 Astra (early September): $10 input / $50 output, cached input $1. Context window of about 1.05 million tokens. Requests over 272K input tokens are billed at 2x input and 1.5x output for the whole request, which is easy to miss.
  • GPT-6 Sol (September 22): $2 / $10.
  • GPT-6 Luna (September 22): $0.10 / $0.50.

Sol and Luna are available in Codex for ChatGPT Plus, Pro, Business, Enterprise and Edu users, and in the API. Free and Go users can try Luna in the ChatGPT desktop app.

The most notable fact in that list is that Opus 5.5 costs 40% of Astra’s per-token price. A year ago, Anthropic was the expensive option. That has reversed.

Claude Opus 5.5 vs GPT-6 Coding Benchmarks

Here are the agentic coding results from Anthropic’s Opus 5.5 and Sonnet 5.5 launch tables. GPT-6 Sol’s FrontierCode figure comes from the Sonnet 5.5 table:

BenchmarkOpus 5.5Sonnet 5.5GPT-6 AstraGPT-6 Sol
Terminal-Bench 4.066.4%70.6%57.9%—
FrontierCode v1.154.4%52.1%53.3%49.3%
CursorBench 4.057.8%55.5%——
OSWorld 2.1 (computer use)81.8%80.1%——

DeepSWE v1.1, from OpenAI’s launch materials and Google’s Gemini 4 Argon comparison:

ModelDeepSWE v1.1
Claude Opus 5.574.2%
GPT-6 Astra74.1%
GPT-6 Sol (max effort)68.8%
GPT-6 Luna (max effort)66.6%
Terminal-Bench 4.0 score (vendor-reported)
  • Claude Sonnet 5.5 70.6%
  • Claude Opus 5.5 66.4%
  • GPT-6 Astra 57.9%

Source: Anthropic, Introducing Claude Opus 5.5 (Sep 2026)

Two things about these tables.

The Terminal-Bench 4.0 gap is the biggest one here. Opus 5.5 leads Astra by 8.5 points, and Sonnet 5.5 leads it by 12.7. Terminal-Bench measures long shell-driven workflows: running builds, reading errors, retrying. That is the closest benchmark to what a coding agent actually does all day. Earlier this year, in my Codex CLI vs Claude Code comparison, OpenAI led this benchmark by 13 points. That lead has reversed.

On DeepSWE and FrontierCode, Opus 5.5 and Astra are effectively tied. A 0.1-point or 1.1-point gap is noise, especially when the numbers come from a competitor’s harness.

The usual warning applies, and it matters more than usual here. Most of these figures come from Anthropic’s own tables, comparing its models against OpenAI’s at settings Anthropic chose. OpenAI’s launch posts make similar comparisons in the other direction, and they tend to look better for OpenAI. Treat single-digit gaps as provisional until independent labs re-run them.

The Sonnet 5.5 Surprise

The detail that surprised me most this launch cycle is that Claude Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0, 70.6% to 66.4%, at half the price.

Anthropic is careful to say that “Opus 5.5 remains clearly stronger on complex, open-ended work”, and the FrontierCode and CursorBench numbers support that. On structured terminal tasks with clear success criteria, though, the cheaper model wins.

In practice, this matches how I would split work anyway. Sonnet handles “fix this failing test”, “add this endpoint” and “update these imports”. Opus handles “redesign this module’s data flow” and “work out why this race condition only happens in production”. The benchmark result says you should lean on Sonnet more than you probably do.

For Claude Code users, the default choice for most work should now be Sonnet 5.5, with Opus 5.5 kept for the hard problems.

What a Coding Session Actually Costs

Per-token prices are misleading for agent work, because caching dominates the bill. Agents re-send the same files and instructions on every turn, and most of those input tokens are cache reads.

So here is an illustrative agentic coding session: 10 million input tokens, 90% served from cache, and 300,000 output tokens. I am ignoring cache write costs and differences in how many tokens each model uses, so treat these as rough estimates for comparison only:

ModelFresh inputCached inputOutputSession total
Claude Opus 5.5$4.00$1.80$6.00$11.80
Claude Sonnet 5.5$2.00$1.80$3.00$6.80
GPT-6 Astra$10.00$9.00$15.00$34.00
GPT-6 Sol$2.00~$1.80$3.00~$6.80
GPT-6 Luna$0.10~$0.09$0.15~$0.34

(Sol and Luna cached input assumes the “up to 90% off” discount OpenAI described at launch.)

A few conclusions from that table:

  • Astra costs nearly 3x as much as Opus 5.5 for the same session, and does not lead on any coding benchmark in the tables above. If Astra hits the 272K long-context surcharge, the gap widens further.
  • Sonnet 5.5 and Sol cost exactly the same at list price. The choice between them comes down to quality and tooling, not cost.
  • Luna is in a different price class. At about 34 cents for a session that costs $11.80 on Opus, it is the obvious choice for bulk work like lint fixes, test generation over hundreds of files, or PR summaries in CI. OpenAI says Luna scores 66.6% on DeepSWE, which is not far behind the frontier.

Token efficiency moves these numbers in practice. Anthropic says Opus 5.5 completes tasks at about 40% lower total cost than Opus 5, partly by using fewer tokens. My earlier post on cutting Claude API costs covers the caching and context habits that matter more than the model choice.

Per task, Claude Opus 5.5 is the cheapest frontier-quality option available right now, and GPT-6 Luna is by far the cheapest acceptable one.

Where GPT-6 Wins

This is not a one-sided comparison, and I would be misleading you if I made it look like one.

Luna has no Anthropic equivalent. Anthropic does not currently have anything near $0.10 / $0.50 with competitive coding scores. For teams running thousands of small agent calls a day, that settles the decision.

Codex availability is broad. GPT-6 Sol and Luna are included in ChatGPT Plus and above. If your team already pays for ChatGPT, using Codex with Sol costs nothing extra.

Context handling at the top end. Astra’s 1.05 million token context is comparable to Sonnet 5.5’s 1 million. For whole-repository analysis, both families are now in the same range. Context is no longer a deciding factor between them.

OpenAI claims a big drop in factual errors. OpenAI says Sol makes about 50% fewer factual errors than its predecessor. If that holds up, it addresses the hallucination gap I wrote about in the GPT-5.5 vs Claude Opus 4.7 comparison, which was the strongest argument against OpenAI for coding agents. I would want independent confirmation before relying on it.

Which Should You Use? My Setup Recommendation

For a professional team today, this is the setup I would recommend:

  1. Default agent model: Claude Sonnet 5.5 in Claude Code. It costs $2 / $10, leads Terminal-Bench 4.0, and handles the large majority of daily tasks.
  2. Escalation model: Claude Opus 5.5 for architecture changes, difficult debugging, and large migrations. Anthropic cites one tester finishing a 680,000-line migration in under a day.
  3. Bulk and CI work: GPT-6 Luna through the API or Codex, for anything high-volume where a small quality drop is fine.
  4. Skip GPT-6 Astra unless a specific OpenAI-only workflow needs it. Nothing in the published coding numbers justifies paying 2.5x Opus’s per-token price.

If your team is already on Codex and ChatGPT, use GPT-6 Sol as the default and Luna for bulk work. That setup is good, and changing tools has its own cost. The patterns in the agentic coding guide apply equally to both ecosystems.

The Bottom Line

Claude Opus 5.5 vs GPT-6 is closer than the marketing on either side suggests, and the tiers matter more than the brands. Anthropic currently leads on frontier coding quality per dollar. OpenAI leads clearly on the low-cost tier with Luna. GPT-6 Astra is the model that is hardest to recommend.

One more thing to watch: Google’s Gemini 4 Argon claims 77.9% on DeepSWE, ahead of both families, but it is not generally available yet. My Gemini 4 Argon breakdown covers what that launch means. When it opens up, this comparison will need a third column.