Grok 4.7 is a genuine capability upgrade, particularly for coding agents and longer professional work.
SpaceXAI's September 21 release uses a larger base model than Grok 4.6 and a longer reinforcement-learning run weighted toward difficult tasks that can take hours to complete. The company says the model checks its own work more carefully, handles longer context better and was trained to understand the Grok Build harness natively. Standard API pricing remains unchanged at $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens.
More capability at the same token price sounds straightforward. The economics are not.
Artificial Analysis finds that Grok 4.7 can consume substantially more tokens and cost more per evaluated task than Grok 4.6. Yet its current data also reveals something more useful than another warning about token consumption: Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index v4.3.2 at both high and xhigh reasoning. High is cheaper and is also SpaceXAI's default API setting.
That makes the buyer decision less about whether Grok 4.7 is better. The evidence says it is, especially for agentic work. The more practical question is how much Grok 4.7 reasoning the job actually deserves.
Choosely has not tested Grok 4.7 hands-on. This early assessment uses SpaceXAI's published materials and independent benchmark data available immediately after launch.
What actually improved?
SpaceXAI is aiming Grok 4.7 directly at coding, agentic tasks and professional knowledge work.
Its launch material reports gains across CursorBench 4.0, AA-Briefcase, Terminal-Bench, Harvey's Legal Agent Benchmark, EEBench and other evaluations designed around coding or longer professional tasks. On CursorBench 4.0, SpaceXAI highlights a score of 46.3% for Grok 4.7 against 40.4% for Grok 4.6.
There is an important benchmark caveat. The main launch table compares Grok 4.7 at xhigh reasoning with Grok 4.6 at high for several headline results, while at least one Grok 4.7 result uses high instead. Those figures still tell us something about the model's ceiling, but they are not all controlled generation-to-generation comparisons at the same reasoning effort.
Independent evidence broadly supports the direction of travel.
Artificial Analysis scores Grok 4.7 at 46 on Intelligence Index v4.3.2, up from 44 for Grok 4.6. The strongest gains are concentrated around agentic knowledge work and coding rather than a universal jump across every evaluation.
That distinction matters. Grok 4.7 looks like a model trained to do more on difficult, sustained work, rather than one that became dramatically better at everything.
The coding-agent improvement is real
The cleanest generation-to-generation result comes from Artificial Analysis testing both models with the same Grok Build coding harness at xhigh reasoning.
Grok 4.7 scores 56 on Coding Agent Index v1.5, up from 47 for Grok 4.6.
Every component moves upward:
- DeepSWE v1.1: 65% to 73%
- Terminal-Bench 4.0: 18% to 33%
- SWE-Atlas-QnA: 58% to 63%
That is a meaningful improvement for developers using Grok as an agent rather than as an autocomplete box.
It also comes from a much larger agent run.
| Grok Build run | Coding Agent Index | Cost per evaluated task | Avg. time | Tokens per task |
|---|---|---|---|---|
| Grok 4.6 xhigh | 47 | $3.57 | 19.5 min | 5.5M |
| Grok 4.7 xhigh | 56 | $8.82 | 39.2 min | 14.3M |
Artificial Analysis therefore measures a nine-point index improvement alongside roughly 2.6 times the total tokens, twice the elapsed time and about 2.5 times the estimated cost per evaluated task.
That does not tell us the extra computation is wasted. It tells us that same API rate and same job cost are different claims.
High versus xhigh may be the most useful result
The strongest reason not to reduce Grok 4.7 to a token-burn story comes from Artificial Analysis' own general benchmark data.
| Setting | Intelligence Index v4.3.2 | Cost per evaluated task | Output tokens per task |
|---|---|---|---|
| Grok 4.7 high | 46 | $2.73 | 66K |
| Grok 4.7 xhigh | 46 | $3.74 | 81K |
| Grok 4.6 high | 44 | $1.86 | 36K |
| Grok 4.6 xhigh | 44 | $2.32 | 38K |
The rate card is identical between Grok 4.6 and 4.7. The measured task economics are not.
More interestingly, high already captures the full two-point Intelligence Index gain that xhigh produces.
That does not mean xhigh never helps. The composite hides differences underneath it. Artificial Analysis has Grok 4.7 xhigh slightly ahead of high on AA-Briefcase, AutomationBench-AA and Terminal-Bench 4.0, while high edges xhigh on SciCode and GDP.pdf. AA-LCR is tied.
So the expensive setting is buying something in particular workloads. It simply is not buying a higher overall composite score.
For general API use, that makes SpaceXAI's default high setting a sensible starting point. Maximum reasoning should have to earn its keep.
Two coding benchmarks tell different cost stories
This is where things get particularly useful for buyers.
Artificial Analysis' Grok Build evaluation makes 4.7 look substantially more expensive than 4.6. CursorBench 4.0 tells a very different story.
On Cursor's live leaderboard, Grok 4.7 at xhigh scores 46.3% at $6.01 per evaluated task. Grok 4.6 at the same xhigh setting scores 41.4% at $6.10 per task.
In that harness, Grok 4.7 gets materially better while task cost stays essentially flat. Cursor also places Grok 4.7 xhigh close to Claude Fable 5.1 Medium, which scores 46.8% at $7.05 per task, and ahead of GPT-5.6 Sol Max at 41.7% and $8.23.
That matters because it stops the cost argument becoming too neat.
Artificial Analysis' Grok Build benchmark says the newer model performs more work through a longer and substantially more expensive agent loop. CursorBench says the newer model improves over 4.6 at virtually the same measured task cost.
Neither result needs to be discarded. They are evaluating different agent environments and workloads.
That is arguably the larger lesson from Grok 4.7: there may be no useful universal answer to what the model "costs per task." The harness and the job can change the answer dramatically.
Same token price still does not mean same workload price
Grok 4.7's standard API rates are unchanged:
- Input: $2 per million tokens
- Cached input: $0.50 per million tokens
- Output: $6 per million tokens
High is the default reasoning setting, with low, medium and xhigh also available. The model has a 500,000-token context window.
Those numbers are useful, but they are only unit prices.
Artificial Analysis estimates the weighted Intelligence Index task cost at $1.86 for Grok 4.6 high and $2.73 for Grok 4.7 high. At xhigh, the comparison is $2.32 versus $3.74.
The tariff stayed still. The amount of computation used did not.
For a production team, the more useful metric is therefore the cost of obtaining an acceptable result. That includes tokens, caching, retries, tool calls, latency and the amount of human cleanup still required.
A cheap token can still produce an expensive job.
SpaceXAI has a legitimate counterargument
There is a straightforward explanation for much of the heavier computation: SpaceXAI trained the model to do more of it.
The company says Grok 4.7 received a longer reinforcement-learning run weighted toward problems that can take many hours to complete, with more emphasis on checking its work and managing long context.
Some independent results suggest that extra work is productive.
The Coding Agent Index improves by nine points. AA-Briefcase moves substantially higher than Grok 4.6 high. Artificial Analysis also reports Grok 4.7 xhigh at a lower hallucination rate than Grok 4.6 high, although that particular comparison uses different reasoning settings and should be treated accordingly.
The improvement is uneven, which is just as important. In the same AA analysis, Grok 4.7 xhigh regresses against Grok 4.6 high on AA-LCR and AutomationBench-AA. Those are mismatched reasoning settings, so they are signals rather than clean same-effort regressions.
The fair reading is that longer traces are producing more useful agentic work in some of the areas SpaceXAI targeted. They are not producing proportional gains everywhere.
"Same speed" is not settled yet
SpaceXAI says Grok 4.7 is served at the same price and speed as Grok 4.6. The price part is easy to verify from the API tariff. Speed is messier.
Artificial Analysis' early measurements vary depending on what is being measured. Its current model data shows different output throughput between high and xhigh, while its launch analysis separately reports much higher answer-output throughput on long prompts. Some current AA measurements also show Grok 4.7 producing output more slowly than Grok 4.6.
That does not necessarily measure the same serving-speed concept SpaceXAI is describing. It does mean "same speed" should remain a first-party claim for now rather than an independently settled fact.
For buyers, wall-clock completion time matters more anyway. The Grok Build benchmark is a good example: 4.7 takes 39.2 minutes per evaluated task against 19.5 minutes for 4.6 while completing a larger agent run and scoring substantially higher.
Fast, deep and cheap remain three different knobs.
The 500K context window has a pricing cliff
Grok 4.7 supports 500,000 tokens of context, but the standard rate only applies below the long-context threshold.
Once the prompt reaches 200,000 tokens, pricing moves to:
- Input: $4 per million tokens
- Cached input: $1 per million tokens
- Output: $12 per million tokens
More importantly, SpaceXAI's pricing documentation says the long-context rates apply to all tokens in the request once the prompt reaches that threshold, not just the portion beyond 200K.
That gives context management a direct financial consequence.
A 500K window is useful capacity for difficult agentic work. It is not a free invitation to keep stuffing history into the prompt.
Caching, compaction and sensible state management can matter as much to the eventual bill as the headline model price.
So should Grok 4.6 users switch?
For API developers, Grok 4.7 deserves a proper workload test.
There is enough independent evidence to call it a real upgrade, and its standard per-token pricing has not increased.
For general workloads, start with high rather than automatically reaching for xhigh. Artificial Analysis currently gives both settings the same Intelligence Index score, while high consumes fewer output tokens and costs less per evaluated task.
For difficult agentic work, the answer is less tidy. Xhigh does edge high on some of the agentic components, and Grok 4.7's same-harness Coding Agent Index improvement over 4.6 is substantial.
The correct choice therefore depends on the job rather than on the existence of a more powerful toggle.
How does Grok 4.7 compare with the coding leaders?
Artificial Analysis says Grok 4.7 with Grok Build now ranks fourth among models in their native coding harnesses, behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5. It also moves ahead of GPT-5.6 Sol on Coding Agent Index v1.5.
Cost complicates that picture again.
On the same Coding Agent Index, GPT-6 Astra in Codex scores 62 at $7.47 per evaluated task, while Grok 4.7 in Grok Build scores 56 at $8.82. GPT-5.6 Sol in Codex scores 55 at $6.35. These are not apples-to-apples model comparisons because both model and agent harness change, but they are useful reminders that a lower API token rate does not guarantee a lower agent-task bill.
On the broader Artificial Analysis Intelligence Index v4.3.2, Claude Fable 5.1 and GPT-6 Astra currently lead at 53, compared with Grok 4.7 at 46.
Grok 4.7 has moved closer. It has not turned the frontier into a one-model race.
That is increasingly how model comparisons should be read. The leaders separate by workload, reasoning effort, harness, latency and economics rather than lining up neatly from best to worst.
Choosely verdict
Grok 4.7 is a real upgrade.
The strongest evidence appears where SpaceXAI says it put the work: coding agents, sustained professional tasks and workloads that benefit from longer reasoning and self-checking.
Its unchanged standard API tariff is attractive. But the rate card does not tell you what the finished job will cost.
Artificial Analysis finds substantially higher task cost in its Grok Build evaluation. CursorBench finds a strong improvement over Grok 4.6 at essentially the same task cost. Those results are not contradictory once the harness and workload are taken seriously.
The high-versus-xhigh result is perhaps the most useful clue for everyday buyers. Grok 4.7 scores 46 on Artificial Analysis Intelligence Index v4.3.2 at both settings, while high costs less and uses fewer tokens. Xhigh still helps in some agentic evaluations, but maximum reasoning is not automatically maximum value.
So the decision is no longer simply whether Grok 4.7 is worth using.
It is how much Grok 4.7 reasoning the job actually deserves.
Start with representative work. Start with high. Measure the outcome and the bill. Then make the model earn every step upward.
Sources
- SpaceXAI: Grok 4.7 launch announcement
- SpaceXAI: Grok 4.7 model documentation
- SpaceXAI: API pricing
- Artificial Analysis: Benchmarking Grok 4.7
- Artificial Analysis: Grok 4.7 model page
- Artificial Analysis: Coding Agents leaderboard
- Cursor: Evaluations
The Change Brief
Get the week’s AI changes in one clear read
Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.
Stay ahead of AI without following it all day. We’ll send you what matters each week.
Continue reading
Related reads
AI Strategy
GPT-6 Astra Explained: What the 99.9% ARC-AGI-3 Score Really Means
GPT-6 Astra makes a huge leap in adaptive reasoning and agentic work, but its viral 99.9% ARC-AGI-3 result depends on a provider-specific context-management harness.
AI Strategy
Claude Fable 5.1 vs Opus 5: Is Anthropic’s Best Model Worth the Cost?
Claude Fable 5.1 raises Anthropic’s capability ceiling, but its value against Opus 5 depends on whether the workload is output-heavy or cache-heavy.
