AI StrategyChoosely EditorialEarly assessment

GPT-6 Astra Explained: What the 99.9% ARC-AGI-3 Score Really Means

OpenAI's GPT-6 Astra is a major step forward in computer use, adaptive reasoning and long multi-step work. But the viral 99.9% ARC-AGI-3 result needs an important qualification: under ARC Prize's provider-neutral Standard harness, Astra scored 62.7%. With OpenAI's own context-management system wrapped around it, the score jumped to 99.9%.

← Back to AI Radar
Choosely Chimp beside the OpenAI GPT-6 Astra interface in a cinematic blue AI environment.

OpenAI's GPT-6 Astra is a major step forward in computer use, adaptive reasoning and long multi-step work. But the viral 99.9% ARC-AGI-3 result needs an important qualification: under ARC Prize's provider-neutral Standard harness, Astra scored 62.7%. With OpenAI's own context-management system wrapped around it, the score jumped to 99.9%.

Early assessment

OpenAI released GPT-6 Astra on September 3, 2026, calling it its most capable model yet for computer use, software engineering, cybersecurity, science and professional work.

This is not a small model refresh.

Astra scores 59.3% on Agents' Last Exam, 72.6% on OpenAI's OSWorld 2.0 evaluation, 57.9% on Terminal-Bench 4.0 and 41.4% on AutomationBench. It also reaches 97.6% on FrontierMath Tier 4 and 100% on ExploitBench. OpenAI's launch material contains the full comparison set.

But the result dominating the launch is 99.9% on ARC-AGI-3, a benchmark designed to test whether AI systems can enter unfamiliar interactive environments, infer hidden rules and adapt.

That number is genuine.

It is also only half the story.

ARC Prize independently tested Astra under two different harnesses, and the difference between them may be more important than the headline score itself. ARC Prize's technical write-up explains why the foundation will report the two conditions separately.

GPT-6 Astra at a glance

GPT-6 Astra
ReleasedSeptember 3, 2026
API modelgpt-6-astra
Context window1,050,000 tokens
Maximum output128,000 tokens
Standard API input$10 / 1M tokens
Cached input$1 / 1M tokens
Standard API output$50 / 1M tokens
Knowledge cutoffApril 30, 2026
Reasoning levelsLow, medium, high, xhigh, max
ChatGPT rolloutPlus, Pro, Business, Enterprise

OpenAI is initially rolling Astra out to a limited group of organizations. Broader ChatGPT access for Plus, Pro, Business and Enterprise users, plus API, Microsoft Azure and AWS Bedrock availability, is planned over the coming days. If Astra is missing from your model picker today, that is expected. OpenAI's Astra announcement is the current first-party rollout source.

What actually changed from GPT-5.6 Sol?

The biggest gains appear when the model has to complete work, rather than simply answer a difficult question.

BenchmarkGPT-6 AstraGPT-5.6 Sol
Agents' Last Exam59.3%53.6%
OSWorld 2.072.6%65.7%
ScreenSpot-Pro92.7%76.9%
AutomationBench41.4%18.1%
Terminal-Bench 4.057.9%37.3%
Terminal-Bench Science 0.164.6%22.4%
FrontierMath Tier 4 v297.6%83.0%
GPQA Diamond96.0%94.6%

These are predominantly OpenAI-reported evaluations, so they should not be mistaken for a complete independent comparison. Still, the pattern is difficult to miss. OpenAI publishes the underlying launch table here.

The GPQA gain is modest.

The AutomationBench, ScreenSpot-Pro, Terminal-Bench and computer-use gains are not.

Astra's generational shift appears to be less about simply knowing more and more about turning reasoning into completed action.

That distinction matters as frontier models move from answering prompts toward operating software, maintaining state and carrying work across multiple steps.

Computer use is becoming the real product

OpenAI's launch examples show Astra operating software rather than merely explaining how a human should operate it.

The model is demonstrated working with circuit-board design in KiCad, spreadsheets, browser workflows, frontend QA, documents and other professional environments.

On OSWorld 2.0, OpenAI reports Astra scoring 72.6% at roughly 40 minutes per task, compared with Sol at 65.7% and roughly 75 minutes.

That is about 47% less elapsed time per task. OpenAI reports both the score and task-time comparison in its Astra launch material.

For chat, shaving a few seconds off a response is pleasant.

For an autonomous workflow that might otherwise occupy a computer for more than an hour, halving execution time starts to become an operational advantage.

This is where Astra looks most consequential.

One model, two very different ARC-AGI-3 results

Here is the part the 99.9% headline tends to leave out.

ARC Prize evaluated Astra under two conditions.

The Standard harness gives every provider the same minimal interface. The model receives what it needs to solve the environment, but it is responsible for deciding what information to preserve in its visible notes.

The Provider Adapter harness lets OpenAI use context-management infrastructure designed specifically for Astra. It preserves opaque reasoning state between requests and uses compaction to manage longer conversations. ARC Prize documents the harness distinction here.

The scores are striking:

Reasoning effortStandard harnessProvider Adapter
Max62.7%98.6%
XHigh59.3%98.4%
High54.8%99.9%
Medium38.6%98.4%
Low17.5%98.0%
None35.2%96.7%

ARC Prize's full effort table is the most important evidence behind the headline number.

That last row deserves attention.

With no reasoning effort selected, Astra scores 35.2% under the Standard harness and 96.7% with the Provider Adapter.

Across the Standard harness, reasoning effort produces a wide spread in performance.

With the Provider Adapter, that spread compresses into only a few percentage points.

That does not mean the 99.9% result is fake, or that the adapter is somehow solving ARC-AGI-3 on Astra's behalf.

It means the two evaluations answer meaningfully different questions.

The Standard harness asks how well the model adapts when every provider operates through the same minimal interface.

The Provider Adapter asks how capable Astra becomes when paired with OpenAI's own context-management system.

ARC Prize is therefore reporting the two separately.

And there is another interesting efficiency result.

Across 167 game and reasoning-level pairs solved under both conditions, ARC Prize says Provider Adapter runs were about 3.66 times faster and used 49% fewer total tokens. The foundation reports those efficiency figures in the same technical analysis.

The important story is no longer simply that a model scored 99.9%.

It is that model capability and system capability are becoming harder to separate.

For agentic AI, the memory, context handling, tools and orchestration around the model may increasingly determine what the overall system can actually accomplish.

The 99.9% result is not an AGI certificate

OpenAI President Greg Brockman reportedly finished the Astra press briefing with the line:

"Welcome to the AGI era."

It is a good launch line.

It is not a measurement standard.

ARC Prize itself is much more restrained. Its benchmark consists of closed-ended interactive environments with deterministic rules. Strong performance demonstrates unusual adaptive reasoning, but it does not establish that an AI system can generalize across the messiness, uncertainty and open-ended complexity of the real world. ARC Prize explicitly cautions against treating saturation as proof of AGI.

Astra's Standard-harness result is arguably impressive enough without needing the AGI label.

GPT-5.6 Sol previously scored around 7.8% on ARC-AGI-3.

Claude Opus 5 reached roughly 30%.

Astra reaches 62.7% under the same broad provider-neutral testing philosophy.

That is a large capability movement.

There is no need to inflate it into something the benchmark itself does not claim.

Independent testing also complicates the "world's smartest model" story

Astra's ARC result is enormous.

Its performance across broader independent intelligence testing is much less dominant.

Artificial Analysis currently scores GPT-6 Astra at roughly 61 on its Intelligence Index, essentially level with GPT-5.6 Sol.

Claude Fable 5.1 sits at 66, while Claude Opus 5 scores 63. Artificial Analysis' Astra benchmarking provides the broader independent context.

That is a useful sanity check.

The same model that nearly saturates ARC-AGI-3 with its Provider Adapter does not suddenly run away from the frontier on every independent measure of intelligence.

Artificial Analysis reaches a different conclusion on coding agents, where Astra's efficiency improvements are much stronger. It scores 67 on the Coding Agent Index, roughly alongside other frontier coding systems while using substantially fewer tokens than Sol.

Again, the pattern points toward agentic efficiency and execution, not universal domination.

What GPT-6 Astra means for the Choosely AI Progress Index

Astra also arrives at an awkwardly perfect moment for Choosely.

The first canonical Choosely AI Progress Index snapshot launched at 40.2 / 100, measuring demonstrated progress toward broadly autonomous digital work across Reasoning & Adaptation, Real-World Work and Autonomy & Agency.

Its frozen Reasoning pillar includes ARC-AGI-2 and ARC-AGI-3. The current launch snapshot records the previous frontier observation from GPT-5.6 Sol before Astra appeared. The exact evidence rules are public in the AI Progress Index methodology.

ARC Prize has now verified Astra at:

  • 95.0% on ARC-AGI-2
  • 62.7% on ARC-AGI-3 under the Standard harness
  • 99.9% on ARC-AGI-3 using the separate Provider Adapter

ARC Prize's canonical Astra results page records the verified observations.

That makes Astra immediately relevant to the next AI Progress evidence review.

It does not mean Choosely simply substitutes 99.9 into the Index.

The Index methodology freezes exact evidence surfaces, system boundaries, evaluators and benchmark conditions. A new observation has to pass those eligibility rules before it can alter the canonical score.

Editorial coverage also cannot pre-announce an unpublished Index movement. An interesting benchmark result and an approved canonical snapshot are deliberately different things.

So, for now, the public score remains 40.2.

Astra has produced new evidence.

The measurement process now decides what that evidence is actually allowed to move.

That is precisely why the Index exists.

GPT-6 Astra is also OpenAI's first Critical-rated cyber model

Astra crossed a second threshold that deserves more attention than a normal model upgrade.

OpenAI says it is its first broadly deployed model to reach the Critical cybersecurity capability level under the company's Preparedness Framework. OpenAI's Astra safety overview explains the deployment implications.

OpenAI says that, with appropriate tools and access, Astra can discover previously unknown vulnerabilities and develop exploits across well-protected systems without a human guiding every step.

Without production safeguards, Astra scored 100% on ExploitBench, compared with 78.5% for GPT-5.6 Sol. OpenAI also says Astra discovered and used two previously unknown zero-day vulnerabilities during evaluation and that both were disclosed to the relevant maintainers. OpenAI reports those results in the launch material.

That is not merely an interesting benchmark statistic.

It is an example of model capability becoming strong enough to change the security controls required to deploy the model safely.

How much does GPT-6 Astra cost?

The API price is substantial:

$10 per million input tokens

$1 per million cached input tokens

$50 per million output tokens

The model supports a 1.05 million-token context window and up to 128,000 output tokens. Requests above 272,000 input tokens attract higher long-context rates. OpenAI's GPT-6 Astra API documentation is the current pricing and specification source.

GPT-5.6 Sol's current standard token prices are $4 input and $20 output per million tokens, so Astra's headline token rates are 2.5 times higher.

That sounds painful, but token price alone is not the right comparison for agentic work.

Artificial Analysis found Astra using roughly 10% fewer output tokens than Sol on its Intelligence Index, and dramatically fewer tokens in its coding-agent evaluation. OpenAI also reports major execution-time gains on computer-use tasks.

The useful commercial question is therefore not:

Is Astra cheaper per token?

It isn't.

The question is:

Does Astra require fewer attempts, less supervision and less elapsed time to finish the job?

For expensive or failure-prone workflows, the answer could matter more than the token rate.

For rewriting an email, probably not.

Is GPT-6 Astra available on ChatGPT Plus?

Not generally yet.

OpenAI says Astra is initially rolling out to a limited group of organizations, with access coming over the following days to ChatGPT Plus, Pro, Business and Enterprise, as well as the API, Azure and AWS Bedrock. The launch announcement remains the primary rollout source.

OpenAI's current ChatGPT release notes still describe Astra as not yet generally available. ChatGPT release notes should be rechecked as the rollout moves.

So if you have Plus and cannot select Astra today, nothing appears to be wrong with your account.

The rollout is simply not complete.

This is a fast-decay detail and should be checked again before changing plans or buying additional usage.

The Choosely verdict

GPT-6 Astra looks like a major frontier-model release, but not for the reason the 99.9% headline suggests.

The most interesting evidence is the pattern around it.

Astra makes a huge independently verified jump on ARC-AGI-3 even under the provider-neutral Standard harness. It improves computer use, agentic coding and professional workflow completion. It executes some computer tasks materially faster. And with OpenAI's own context-management infrastructure wrapped around it, its adaptive performance becomes dramatically stronger and more consistent.

That last point may prove to be the most important.

The next stage of AI progress may be increasingly difficult to measure by asking which model is smartest.

The better question may become which system can maintain context, use tools, recover from mistakes and keep working until the job is actually finished.

For complex agentic work, Astra deserves testing quickly.

For ordinary ChatGPT use, there is no reason to panic if it has not appeared yet, and no evidence that GPT-5.6 Sol suddenly became inadequate overnight.

And if someone tells you a 99.9% benchmark score means AGI has officially arrived, ask one annoying but useful question:

Which harness?

What Choosely verified

This is an Early assessment, not a hands-on Choosely review.

Choosely verified the launch details and model specifications against OpenAI's current announcement, API documentation, release notes and safety material. ARC-AGI claims were checked against the ARC Prize Foundation's independent results and technical write-up. Broader intelligence comparisons were checked against Artificial Analysis.

Choosely has not yet completed a controlled hands-on evaluation of GPT-6 Astra.

Availability, plan access, usage limits and pricing are fast-decay facts and will be rechecked as the rollout continues.

Frequently asked questions

What is GPT-6 Astra?

GPT-6 Astra is OpenAI's new flagship model for complex reasoning, computer use, software engineering, research and professional multi-step work. It was announced on September 3, 2026.

Did GPT-6 Astra really score 99.9% on ARC-AGI-3?

Yes. ARC Prize independently verified a 99.9% result using its Provider Adapter harness. Under the provider-neutral Standard harness, Astra's best verified result was 62.7%. The two conditions are reported separately.

Why is there such a large difference between 62.7% and 99.9%?

The Standard harness uses the same minimal interface across providers and leaves the model responsible for preserving visible notes. The Provider Adapter allows Astra to use OpenAI-designed context management, including preserved opaque reasoning state and compaction.

Is GPT-6 Astra AGI?

No benchmark result establishes that. ARC Prize explicitly cautions against treating ARC-AGI-3 saturation as proof of AGI because the benchmark uses bounded, deterministic environments rather than unrestricted real-world tasks.

Is GPT-6 Astra available on ChatGPT Plus?

OpenAI says Astra is coming to Plus users over the coming days, but broader rollout is not yet complete. If it is missing from your model picker today, that is expected.

How much does GPT-6 Astra cost?

OpenAI's standard API price is $10 per million input tokens and $50 per million output tokens, with cached input at $1 per million tokens.

Is GPT-6 Astra better than GPT-5.6 Sol?

Astra shows large gains over Sol on computer use, automation, coding-agent and adaptive reasoning evaluations. On Artificial Analysis' broader Intelligence Index, however, the two currently score almost identically. Astra also costs substantially more per token.

Your AI stack shouldn't go stale. AI models, pricing, access and capabilities now move quickly enough that a good decision can become stale within weeks. Choosely tracks those changes and helps keep the decisions behind your stack current. Build your AI stack free →

Sources

The Change Brief

Get the week’s AI changes in one clear read

Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.

Stay ahead of AI without following it all day. We’ll send you what matters each week.

Continue reading

Related reads