Evidence-based analysis
OpenAI has cut GPT-5.6 Luna API pricing by 80%. From 30 July 2026, Luna costs US$0.20 per million input tokens and US$1.20 per million output tokens, down from US$1 and US$6.
That is large enough to reopen model-routing decisions for high-volume work. It is not a reason to replace every model automatically. The useful question is whether Luna can now produce an accepted result at a lower total cost once retries, latency and human correction are included.
This week also brought three changes with clear deadlines or migration work: GitHub Models has shut down, Grok Voice will move its latest alias to a new model on 5 August, and GitHub is preparing a default-on model policy for Copilot Business and Enterprise.
OpenAI changes the GPT-5.6 cost equation
OpenAI's 30 July update reduced prices for two GPT-5.6 tiers:
| Model | Previous input/output price | New input/output price | Change |
|---|---|---|---|
| GPT-5.6 Luna | US$1 / US$6 | US$0.20 / US$1.20 | 80% lower |
| GPT-5.6 Terra | US$2.50 / US$15 | US$2 / US$12 | 20% lower |
| GPT-5.6 Sol | US$5 / US$30 | Unchanged | No change |
Prices are per million tokens through the API.
OpenAI says ChatGPT and Codex subscription prices and quota budgets are unchanged. Luna and Terra will, however, consume fewer credits in ChatGPT Work and Codex. That distinction matters. A cheaper model does not create a cheaper subscription, but it can allow more work inside the same usage budget.
The company also replaced Priority Processing for GPT-5.6 Sol with Fast mode. OpenAI says Fast mode can run up to 2.5 times faster at twice the standard API price. Existing requests that use the old priority setting will continue to work.
OpenAI attributes the lower prices to inference efficiency improvements. Its performance claims remain vendor-reported. Choosely has not independently tested the updated models or Fast mode.
Choosely's read: reroute selectively
Luna's new price makes it a credible candidate for routine, repeated agent work where volume matters more than extracting the last increment of capability. That could include first-pass classification, structured extraction, document triage and low-risk workflow steps.
The right comparison is cost per accepted result, not cost per token.
A lower token price can disappear quickly if a workflow needs more retries, larger prompts or extra human correction. Before changing a default, test the same representative jobs on the current model and Luna. Record:
- successful completions;
- retry and escalation rates;
- response time;
- token or credit use;
- human correction time.
Keep stronger models on tasks where planning depth, reliability or error cost matters more than throughput. The price cut makes routing more attractive. It does not remove the need for evaluation.
GitHub Models has shut down
GitHub completed the retirement of GitHub Models on 30 July. The model playground, catalog, inference API and bring-your-own-key access are no longer available, including for customers with active usage.
GitHub recommends Microsoft Foundry for direct model access and GitHub Copilot for development workflows.
Teams that used GitHub Models for prototypes, evaluations or internal tools should now identify every remaining dependency. Check API endpoints, secrets, playground links and documentation. A replacement service may have different model names, authentication, rate limits and billing.
This is a completed shutdown, not an upcoming deprecation. Any surviving dependency is already operational debt.
Grok Voice changes its `latest` alias on 5 August
xAI introduced Grok Voice Think Fast 2.0 on 29 July. The company prices it at US$0.08 per audio minute.
On 5 August, the grok-voice-latest alias will move from Think Fast 1.0 to 2.0 automatically. Teams that are not ready for the change should pin the exact 1.0 model before that date.
xAI reports improvements in instruction following, emotional expressiveness and tool use. Those claims are vendor-reported and have not been independently tested by Choosely.
The practical risk is not the announced capability. It is an unreviewed model change reaching a customer-facing voice workflow. Run a short regression set against accents, interruptions, tool calls and failure handling before accepting the new alias.
Copilot admins get a new default model policy
GitHub announced a new global default policy for generally available Copilot models on 29 July. The setting is visible now and is scheduled to take effect on 26 August for Copilot Business and Enterprise.
When the policy becomes active, models without an explicit organization-level choice will inherit the default. If the default is enabled, newly generally available models can become available without a separate administrator action.
Existing per-model choices will remain in place. GitHub also says open-weight models and models outside its data protection agreement exclusions will remain excluded from the default.
Administrators should review the setting before 26 August. Organizations that require model-by-model legal, security or procurement approval should disable the global default and keep explicit controls.
What to do before next Friday
- 1Run one high-volume workflow on Luna and compare cost per accepted result with the current default.
- 2Search for any remaining GitHub Models endpoints, keys or documentation and assign a replacement.
- 3Test Grok Voice Think Fast 2.0 or pin version 1.0 before the
latestalias changes on 5 August. - 4Review the Copilot model default before it takes effect on 26 August.
The decision
The headline this week is not simply that an API became cheaper. It is that Luna's price reduction is large enough to justify a measured routing test.
At the same time, GitHub and xAI are reminding operators that model access can disappear or change through aliases and defaults. Lower cost is useful. Controlled change is what keeps it useful.
Sources
- OpenAI: Advancing the price-performance frontier with GPT-5.6 (30 July 2026)
- OpenAI: Introducing GPT-5.6
- GitHub: GitHub Models is now retired (30 July 2026)
- xAI: Grok Voice Think Fast 2.0 (29 July 2026)
- GitHub: Default model enablement for Copilot Business and Enterprise (29 July 2026)
Sources were checked on 31 July 2026.
The Change Brief
Get the week’s AI changes in one clear read
Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.
Stay ahead of AI without following it all day. We’ll send you what matters each week.
Continue reading
Related reads
AI Strategy
GPT-5.6 vs Grok 4.5 vs Claude Fable 5: Early Comparison
GPT-5.6, Grok 4.5 and Claude Fable 5 are now available for practical testing. Here is how they compare on pricing, coding, agents and professional work.
AI Strategy
The Unlimited AI Subscription Is Dying
AI subscriptions are shifting toward credits, token metering and usage-based billing. Here is what Copilot, Codex and Claude reveal about the real cost of AI workflows.
