AI StrategyChoosely EditorialEarly assessment

GPT-6 Sol vs Claude Opus 5.5: OpenAI Cut the Price. Anthropic Raised the Bar.

Anthropic and OpenAI released their latest models within roughly 90 minutes of each other. Claude Opus 5.5 now leads independent testing, while GPT-6 Sol costs half as much per token — the useful story is where their cost-capability curves cross.

← Back to AI Radar
GPT-6 Sol vs Claude Opus 5.5 comparison with ChatGPT and Claude logos

Anthropic and OpenAI released their latest models within roughly 90 minutes of each other on September 22. The timing made the comparison inevitable, but the more interesting story is not which company won launch day. It is what happened to the economics of frontier AI.

Claude Opus 5.5 now reaches the highest score of the models compared here on Artificial Analysis's independent Intelligence Index. GPT-6 Sol does not reach the same measured ceiling, but OpenAI cut its standard API price sharply compared with GPT-5.6 Sol's promotional rate. GPT-6 Luna pushes that efficiency strategy even further down the cost curve.

The simple interpretation is that Anthropic gained capability while OpenAI attacked price. The numbers reveal something more useful.

Sol owns more of the cheaper end of the current cost-capability curve. Opus becomes economically interesting much sooner than its raw token price suggests.

That matters if you are deciding what to put into a product, coding workflow or AI stack.

The short version

At standard API rates, GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens. Claude Opus 5.5 costs $4 and $20. On the rate card alone, Opus is twice the price.

OpenAI says Sol's new pricing is 50% below GPT-5.6 Sol's promotional rates. Anthropic cut Opus pricing by 20% from Opus 5, reduced cache-read pricing by 60% and says typical workloads at default settings can cost around 40% less.

That Anthropic figure needs context. Opus 5.5 defaults to Medium effort, while Opus 5 defaulted to High, so the vendor's default-to-default saving is not a same-effort comparison. Anthropic attributes the reduction to lower token pricing, cheaper cache reads and fewer tokens used per task.

Independent testing also makes the headline price gap less straightforward.

Artificial Analysis currently measures GPT-6 Sol Max at an Intelligence Index score of 48 for about $1.06 per benchmark task. Claude Opus 5.5 Medium reaches 51 for about $1.34.

That means the first measured step beyond Sol's current ceiling costs around 26% more per task, rather than the 100% premium suggested by the raw token rates.

Lower down the curve, the advantage reverses. Sol High scores 43 for about $0.37 per task, while Opus Low scores 42 for about $0.55.

That is where this comparison gets interesting.

GPT-6 Sol vs Claude Opus 5.5: the specs that matter

GPT-6 SolClaude Opus 5.5GPT-6 Luna
Standard input$2 / 1M$4 / 1M$0.10 / 1M
Standard output$10 / 1M$20 / 1M$0.50 / 1M
Cached input$0.20 / 1M$0.20 / 1M$0.01 / 1M
Advertised context1.05M total1M1.05M total
Max output128K128K128K
Default effortMediumMediumMedium
AA Intelligence Index, max485837
AA cost per task, max$1.06$5.98$0.07
Long-context pricingHigher rates above 272K inputStandard listed rates across 1MHigher rates above 272K input
Primary roleSerious coding and agentsHigh-end coding and knowledge workFocused high-volume work

OpenAI's 1.05 million-token context figure needs a qualification. Once a Sol or Luna request exceeds 272K input tokens, OpenAI charges twice the standard input and cache rates and 1.5 times the output rate for the entire request.

Anthropic lists Opus 5.5's 1M context window at its standard pricing. A slightly larger advertised window therefore does not automatically make Sol or Luna cheaper for very long-context workloads.

Opus leads the benchmark, with some important context

Artificial Analysis currently gives Claude Opus 5.5 the strongest result of the models compared here.

ConfigurationAA Intelligence IndexApprox. cost per task
Claude Opus 5.5 Max58$5.98
Claude Opus 5.5 High54$1.82
Claude Opus 5.5 Medium51$1.34
GPT-6 Sol Max48$1.06
GPT-6 Sol High43$0.37
Claude Opus 5.5 Low42$0.55
GPT-6 Sol Medium40$0.25
GPT-6 Luna Max37$0.07
GPT-6 Luna Medium29$0.02

Artificial Analysis evaluates both Opus 5.5 and Fable 5.1 with Anthropic's default fallback enabled. Cost-per-task figures are benchmark-specific estimates, not universal production costs.

Those Opus results should not be read as a perfectly isolated test of Opus 5.5 alone. Artificial Analysis labels the model as tested with Anthropic's default fallback, while Anthropic says its production safeguards can route certain tasks to earlier Claude models.

When those safeguards intervened in Anthropic's own evaluations, cybersecurity tasks were completed by Opus 4.8, while biology and frontier-model-development tasks were completed by Opus 5. Anthropic says those fallbacks likely reduce the measured scores.

That does not make the results invalid. It does mean the tested object is closer to Anthropic's production system than a pure model-in-isolation experiment.

The Fable comparison also needs precision. Artificial Analysis currently scores Opus 5.5 Max at roughly 58 and Fable 5.1 Max at roughly 53, so the Max-to-Max benchmark gap is meaningful. Opus 5.5 High, however, lands around 54, roughly level with Fable 5.1 Max.

Anthropic also says that in its own use, the practical gap between Opus 5.5 and Fable 5.1 is narrower than benchmark scores might suggest.

Opus 5.5 is clearly a strong release. Small differences between particular effort settings still should not be treated as proof of a universal real-world advantage.

The more interesting number is $1.34

The most useful Opus configuration may not be Max.

At its default Medium effort, Opus 5.5 scores 51 on the Artificial Analysis Intelligence Index for about $1.34 per task. Sol requires Max effort to reach 48, at about $1.06.

That is around a 26% increase in measured cost for a move beyond Sol's current ceiling.

At lower capability levels, Sol is the better economic choice. Sol High scores 43 for $0.37 per task while Opus Low scores 42 for $0.55. Sol Medium reaches 40 for $0.25.

The cost curve therefore changes shape as the job gets harder.

There is no universal threshold where every workload should suddenly switch providers, and Artificial Analysis is one benchmark suite rather than a calculator for production traffic. But the data does show why a simple comparison of $2/$10 against $4/$20 is incomplete.

Across the middle of the curve, Sol delivers capability cheaply. Once a workload needs more than Sol Max can reliably provide, Opus can become economically competitive surprisingly quickly.

That is a much more useful buying signal than declaring one model cheaper and the other smarter.

GPT-6 Sol is more than a price cut

OpenAI's clearest message is efficiency, but the underlying Sol results moved in several directions.

Artificial Analysis describes GPT-6 Sol's overall Intelligence Index performance as effectively level with GPT-5.6 Sol. Under that flat headline score, however, some important capabilities changed.

Sol's Coding Agent Index increased by two points, using Artificial Analysis's Codex-based coding harness, with gains including Terminal-Bench and SWE-Atlas-QnA.

Its hallucination behavior changed even more.

On AA-Omniscience, Artificial Analysis measured Sol Max's hallucination rate falling from 92% to 60%. Part of that improvement appears to come from greater caution: the model attempted 83% of questions compared with 99% for GPT-5.6 Sol, while raw accuracy fell from 59% to 54%.

That suggests Sol is more willing to withhold an answer when uncertain instead of filling the gap.

For many production systems, that can be valuable. A model that declines more often may be easier to build guardrails around than one that confidently produces weak information.

But another part of the independent testing moved backward.

GPT-6 Sol dropped roughly 100 Elo on Artificial Analysis's GDPval-AA professional-work benchmark. Artificial Analysis says the regression was often associated with weaker presentation, shorter deliverables and outputs that omitted elements required by the evaluation rubric.

That distinction matters.

If GPT-5.6 Sol already produces long reports, polished client deliverables, presentations or structured professional work that you trust, the lower token price is not enough reason to switch blindly. Run representative examples through GPT-6 Sol and compare the finished work.

For coding and agentic workloads, the evidence gives a stronger reason to test the upgrade.

Opus 5.5 is cheaper, faster and more practical than Opus 5

Anthropic cut standard Opus pricing from $5/$25 to $4/$20 per million input and output tokens. Cache reads fell from $0.50 to $0.20, and Anthropic says output generation is more than 30% faster than Opus 5.

The company also says typical workloads at default settings can cost about 40% less overall.

That figure combines several effects: cheaper token pricing, cheaper cache reads and lower token use. It also compares different default effort levels, because Opus 5.5 defaults to Medium while Opus 5 defaulted to High.

Anthropic separately says Opus 5.5 at Medium can match or exceed Opus 5 at High on some coding and knowledge-work evaluations.

That is more meaningful than the benchmark headline alone. Historically, the problem with Opus-class models has not simply been whether they can perform difficult work. It has been whether their additional capability justifies routing enough real traffic to them.

Opus 5.5 makes that decision less extreme.

At Medium effort, the independent cost-per-task numbers put it close enough to Sol Max that teams dealing with complex code, expensive failures or demanding knowledge work have a stronger reason to test it as a premium default rather than reserve it only for rare escalation.

There is still an integration catch.

Anthropic's migration guide lists four breaking changes for applications moving from Opus 5. Thinking can no longer be disabled, forced tool selection behaves differently, thinking blocks have tighter model and conversation constraints, and some computer-use integrations require migration changes.

Anthropic explicitly recommends testing in a development environment before switching production traffic.

Evaluate quickly. Migrate carefully.

Astra still has a job

Opus 5.5 does not make GPT-6 Astra irrelevant.

Artificial Analysis currently puts Astra Max at roughly 53 on its Intelligence Index, placing it in the same broad frontier tier while still below Opus 5.5 Max on that benchmark.

Anthropic's own comparison also gives Astra two notable wins. Astra scores 41.4% against Opus 5.5's 40.0% on AutomationBench, and 64.6% against 58.7% on Terminal-Bench-Science.

OpenAI continues to position Astra as its highest-capability model.

That leaves a sensible OpenAI ladder. Luna is the extreme efficiency play. Sol covers a much wider range of serious work at substantially reduced cost. Astra remains the escalation path when a task justifies paying for OpenAI's highest tier.

The mistake would be routing every difficult task to Astra automatically because it sits at the top of the product line.

GPT-6 Luna pushes capable reasoning into the cents-per-task range

Luna is easy to overlook beside two premium models, but its economics may be the more important infrastructure story.

Artificial Analysis scores Luna Max at 37 for about $0.07 per task. At Medium, it scores 29 for roughly $0.02. OpenAI's standard API price is $0.10 per million input tokens and $0.50 per million output tokens.

Those economics open a different class of use cases. Classification, extraction, first-pass analysis, focused summarization and other high-volume operations can become cheap enough that Luna does not need to replace a frontier model to create value.

But Luna did not improve everywhere.

Artificial Analysis found its Coding Agent Index fell by two points compared with GPT-5.6 Luna, using the same Codex-based harness. SWE-Atlas-QnA and DeepSWE both moved lower.

That makes Luna compelling for high-volume work, but it should not be treated as a miniature Sol for every agentic task.

What normal ChatGPT and Claude users actually get

The API comparison is only part of the story because these models also arrive through subscription products.

Claude Opus 5.5 is available to Claude Pro, Max, Team and Enterprise users, alongside its API and cloud-platform availability.

OpenAI's rollout is different. At launch, GPT-6 Sol and Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Luna also reaches Free and Go users through the desktop app.

As of September 23, OpenAI says Sol and Luna are not selectable in ordinary ChatGPT conversations.

That distinction matters for readers comparing subscriptions rather than APIs. Seeing "GPT-6 Sol launched" does not mean every ChatGPT user can simply select Sol in a standard chat.

Which model fits which workload?

If your priority is...Model worth testing firstWhy
High-end coding or difficult knowledge workClaude Opus 5.5 MediumStrong independent capability without Max-level cost
Cost-sensitive serious agent workGPT-6 Sol High43 AA Index at about $0.37 per task
Maximum Sol capabilityGPT-6 Sol Max48 AA Index at about $1.06 per task
Moving beyond Sol's measured ceilingClaude Opus 5.5 Medium51 at about $1.34 per task
Very high-volume focused workGPT-6 LunaExtremely low token and measured task cost
OpenAI tasks that justify maximum capabilityGPT-6 AstraOpenAI's highest-tier model and stronger on some evaluations
Existing GPT-5.6 Sol report workflowsTest before migratingGPT-6 Sol regressed on an independent professional-work benchmark
Existing Opus 5 API integrationsTest migration before productionOpus 5.5 introduces breaking API behavior

These are routing starting points, not universal rankings. The best model for a production workload remains the one that clears your own quality bar at the lowest acceptable cost.

The bigger shift is compression of the premium tier

The most important part of these releases is not another model reaching the top of another leaderboard.

Premium capability is getting cheaper.

Anthropic has lowered Opus pricing, cache cost and token use while pushing its measured capability higher. OpenAI has cut Sol's headline API pricing sharply while maintaining roughly the same overall independent intelligence score as GPT-5.6 Sol and improving some coding and factual-reliability measures. Luna pushes useful reasoning into fractions of that cost again.

That puts pressure on the old habit of choosing the smartest model you can afford and sending everything through it.

A better question in late 2026 is:

What is the cheapest model that reliably clears the quality bar for this specific task?

Sometimes that will be Luna. Sometimes Sol. Sometimes paying the extra 26% from Sol Max to Opus Medium will make sense because the workload is sitting near the edge of Sol's measured capability.

And sometimes the cost of failure matters enough that Astra, Opus High or Opus Max remains justified.

Model selection is becoming a routing problem.

Choosely verdict

Claude Opus 5.5 currently reaches the stronger independent capability result. GPT-6 Sol currently offers stronger economics across much of the middle of the curve.

The useful discovery is where those two observations start to overlap.

At Sol High versus Opus Low, OpenAI is cheaper at roughly comparable measured capability. At Sol Max versus Opus Medium, Anthropic moves ahead on the benchmark for only about 26% more measured task cost.

That does not establish a universal crossover point for real production work. It does establish that raw token pricing is no longer enough to compare premium models.

For developers and teams, evaluate cost per accepted result. Existing GPT-5.6 Sol users should test coding and agent workflows aggressively, while checking report and presentation workloads carefully before migrating. Existing Opus 5 users have a strong reason to evaluate 5.5, but the API changes make this a genuine version migration rather than a silent swap.

For high-volume work, Luna deserves attention even if Sol and Opus attract most of the launch headlines.

The frontier is still moving. The bigger commercial shift is that more of it is becoming affordable.

FAQ

Is Claude Opus 5.5 better than GPT-6 Sol?

Artificial Analysis currently measures a higher capability ceiling for Opus 5.5. That does not mean it is better for every workload. Sol is materially cheaper across several lower-effort configurations, and individual benchmarks show different strengths.

Is GPT-6 Sol really half the price of Claude Opus 5.5?

At standard short-context API token rates, yes. Sol costs $2/$10 per million input/output tokens and Opus 5.5 costs $4/$20.

Actual workload economics can be much closer because models use different numbers of tokens and reasoning settings. Artificial Analysis currently estimates Sol Max at about $1.06 per benchmark task and Opus Medium at about $1.34.

Did GPT-6 Sol improve over GPT-5.6 Sol?

Artificial Analysis describes their overall Intelligence Index scores as effectively level, but the underlying results changed. GPT-6 Sol improved its Coding Agent Index and substantially reduced measured hallucination, while regressing on GDPval-AA professional work.

Is Claude Opus 5.5 cheaper than Claude Opus 5?

Yes. Standard input/output pricing fell 20% to $4/$20 per million tokens, and cache reads fell 60% to $0.20.

Anthropic says typical default-setting workloads can cost about 40% less overall, although that comparison includes a change in default effort from High on Opus 5 to Medium on Opus 5.5.

Can I use GPT-6 Sol in normal ChatGPT conversations?

As of September 23, 2026, OpenAI says GPT-6 Sol and Luna are available through ChatGPT Work and Codex but are not selectable in regular ChatGPT conversations. Availability can change quickly, so this should be treated as a launch-period limitation.

Should existing Opus 5 developers switch immediately?

Not without testing. Opus 5.5 changes several API behaviors, including always-on thinking and tool-use handling. Anthropic's migration guide recommends testing in development before moving production traffic.

Keep your AI stack from going stale

Model names, pricing and capability can change faster than most teams re-evaluate the tools they are paying for.

Choosely Stack Intelligence helps you compare the tools already in your workflow, find overlapping capability and see where a cheaper or stronger model may now fit.

Review your AI stack →

Sources

The Change Brief

Get the week’s AI changes in one clear read

Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.

Stay ahead of AI without following it all day. We’ll send you what matters each week.

Continue reading

Related reads