Product UpdateChoosely Team

Gemini 3.6 Flash Cuts Agent Costs: The AI Tool Changes That Matter This Week

Google's new Gemini Flash tiers sharpen the economics of production agents, while Kimi K3, Grok's Microsoft 365 add-ins and OpenAI Presence move capable models deeper into real work.

← Back to AI Radar
A monumental Gemini symbol above two luminous model-routing paths in a dark Choosely technology chamber

Google has launched two new Gemini Flash models aimed squarely at production AI agents, while Kimi, SpaceXAI and OpenAI all pushed capable models deeper into real work. Here is the verified Choosely read on what changed, who should care and what to test before changing your stack.

Quick take: Gemini 3.6 Flash is the headline because Google has paired a lower price with claims of stronger coding, knowledge work and computer use. The cheaper 3.5 Flash-Lite tier makes the routing decision even more interesting. Kimi K3 arrived as a 2.8-trillion-parameter flagship with a one-million-token context window, Grok moved into Microsoft 365 and OpenAI launched Presence for controlled enterprise-agent deployments. The pattern is not simply “better models.” It is a new competition over which model does each step, where the agent works and how much authority it receives.

What changed at a glance

  • Google launched Gemini 3.6 Flash and 3.5 Flash-Lite on July 21. Both are available through the Gemini API, Google AI Studio, Gemini Enterprise and the Gemini app, with additional rollout surfaces varying by model.
  • Kimi K3 became available across Kimi products and the API. Moonshot lists a one-million-token context window and API pricing of $3 input and $15 output per million tokens, with cached input at $0.30.
  • Grok 4.5 expanded into Microsoft 365 and consumer surfaces. SpaceXAI released Excel and Outlook add-ins, then made Grok 4.5 live across grok.com, X, iOS and Android.
  • OpenAI introduced Presence on July 22. It packages policies, guardrails, simulations, approved actions and human escalation around voice and chat agents, but is currently limited to eligible enterprise customers.

The headline change: Gemini Flash becomes a model-routing decision

Google introduced Gemini 3.6 Flash as its new workhorse model for coding, knowledge work, multimodal tasks and agentic workflows. It is priced at $1.50 per million input tokens and $7.50 per million output tokens.

Alongside it, Google launched Gemini 3.5 Flash-Lite for high-throughput workloads at $0.30 per million input tokens and $2.50 per million output tokens. Google describes Flash-Lite as the fastest model in the 3.5 series and says it reaches 350 output tokens per second according to Artificial Analysis.

The important shift is not that one new model replaces every old one. It is that Google now offers a clearer two-lane production stack:

ModelInput / 1M tokensOutput / 1M tokensPractical role
Gemini 3.6 Flash$1.50$7.50Higher-value coding, knowledge work, multimodal analysis and multi-step agent tasks
Gemini 3.5 Flash-Lite$0.30$2.50High-volume extraction, classification, search, document processing and simpler agent steps

Google says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, takes fewer reasoning steps and tool calls, and improves computer-use performance. It also reports gains on coding and knowledge-work evaluations. Those results mix independent measurements, Google evaluations and customer observations, so they are evidence to investigate—not a guarantee for every workload.

Computer use is now a built-in client-side tool through the Gemini API and Gemini Enterprise for the new Flash models. That matters because an agent can operate software interfaces without every action being represented by a purpose-built API. It also raises the stakes: browser or desktop actions need explicit scopes, safe defaults and review gates.

Choosely's read: Do not ask whether 3.6 Flash is “better” in the abstract. Ask which steps deserve it. A sensible production design may use Flash-Lite to classify, extract or prepare context, then route only difficult decisions and tool-heavy work to 3.6 Flash. The saving comes from total task cost—not from choosing the cheapest headline rate.

Who should care

  • Developers running high-volume API or agent workloads.
  • Teams using Gemini for document parsing, chart analysis, coding or computer use.
  • Operators paying flagship-model rates for routine substeps.
  • Buyers who need a cheaper fallback without changing provider infrastructure.

What to do next

Choose 50 to 100 representative tasks and compare your current model with both Flash tiers. Track completion rate, output tokens, tool calls, latency, retries, correction time and cost per accepted result. Route by measured task type, not by benchmark reputation.

Read Google's official Gemini Flash announcement.

Kimi K3 enters the frontier race with a one-million-token context window

Moonshot AI introduced Kimi K3 as a 2.8-trillion-parameter Mixture-of-Experts model for long-horizon coding, knowledge work, reasoning and native visual understanding. It is available through Kimi.com, Kimi Work, Kimi Code and the Kimi API.

Moonshot lists:

  • A one-million-token context window.
  • $3 per million input tokens.
  • $15 per million output tokens.
  • $0.30 per million cached input tokens.

Moonshot calls K3 an open model, but its launch post says the full weights and additional technical details are due by July 27. Until the files and final licence are actually available, “open” should not be interpreted as immediately downloadable or straightforward to self-host.

Kimi's documentation also says older Kimi K2.5 and Moonshot V1 models are no longer available to newly registered users, with a full platform sunset expected on August 31. Existing API users should audit hard-coded model names rather than assuming an alias will continue indefinitely.

Choosely's read: Kimi K3 is a credible test candidate for large-context coding, research and deliverable-heavy work, especially where Claude Fable 5 pricing is difficult to justify. But the real decision should wait for the released weights, licence, independent testing and a trial on your own data. A one-million-token window is useful only if retrieval, attention and output quality remain strong across the material you actually provide.

Read Moonshot AI's official Kimi K3 announcement and official Kimi K3 pricing.

Grok moves from a chat tab into Microsoft 365

SpaceXAI released Grok add-ins for Excel and Outlook, then expanded Grok 4.5 across grok.com, X, iOS and Android.

The Excel add-in can answer questions about selected data, cite workbook cells, write formulas, add charts and run scenarios. SpaceXAI says it can also use connectors to pull context from email, SharePoint or Google Drive. The add-in itself is free through Microsoft Marketplace.

The Outlook add-in can summarise threads, draft replies, read attachments and help organise a mailbox. SpaceXAI says nothing sends until the user presses send. It is available to paid X and SuperGrok users.

The broader Grok 4.5 rollout also makes the model available for everyday knowledge work on mobile and web, including spreadsheets, slides, diagrams, prose and long-PDF analysis.

Choosely's read: This is more material than another chatbot app launch because the model now sits beside the workbook and inbox where decisions happen. Test read-only tasks first. Before granting connector or mailbox access, verify exactly which files, messages and external sources the add-in can reach, and keep send, delete and bulk-edit actions behind human confirmation.

Read SpaceXAI's official Grok for Excel announcement, Grok for Outlook announcement and Grok 4.5 availability update.

OpenAI Presence packages the controls around production agents

OpenAI launched Presence on July 22 for enterprise voice and chat agents. It combines model reasoning with policies, standard operating procedures, guardrails, approved actions, simulations, evaluations and escalation rules.

A deployment begins with a narrowly defined job—such as a billing issue, insurance claim or employee IT request. The company controls what knowledge and systems the agent can access, which actions it may take, when approval is required and when a person should take over. OpenAI says Codex can analyse production signals and propose changes, but teams test and approve updates before controlled rollout.

Presence is available through a limited general availability program for eligible enterprise customers. Deployments are led by OpenAI Forward Deployed Engineers and selected systems integrators. It is not a self-serve product, and OpenAI has not published standard pricing.

OpenAI reports that Presence resolves 75% of inbound issues on its English-language phone support line without human assistance and reduced handoffs by 15 percentage points over 10 days. These are OpenAI's own deployment results, not independent benchmarks.

Choosely's read: Presence shows where the enterprise-agent market is moving: the valuable product is no longer only a capable model. It is the permission system, test harness, escalation logic and improvement loop wrapped around it. Smaller teams may not be eligible for Presence, but they should copy the operating principle—one bounded job, least-privilege access, explicit approvals and measurable escalation criteria.

Read OpenAI's official Presence announcement.

The pattern behind this week's changes

The announcements point to four practical shifts:

  1. 1Routing matters more than a single default model. Gemini's two Flash tiers make cost-aware task routing an operator decision.
  2. 2Long context is becoming common—but not automatically useful. Kimi K3's one-million-token window still needs testing for retrieval quality and task completion.
  3. 3AI is moving into the work surface. Grok is no longer waiting in a separate tab; it is beside the spreadsheet and inbox.
  4. 4Permissions are becoming part of the product. Presence treats guardrails, approvals and escalation as core infrastructure rather than post-launch paperwork.

The sensible response is not to migrate everything this weekend. Pick one measurable workflow, define what a successful result costs today, and compare the new option while keeping the existing path available.

What Choosely verified this week

Every material date, price and availability statement in this brief was checked against a first-party vendor announcement, documentation or pricing page on July 24, 2026.

Availability can vary by plan, geography, account eligibility, workspace controls and staged rollout. Recheck the relevant account and live pricing page before changing a production workflow.

The bottom line

Gemini 3.6 Flash is this week's clearest operator signal: better agent economics will increasingly come from routing each step to the right tier. Kimi K3 adds a serious long-context challenger, Grok is embedding itself in Microsoft 365, and OpenAI Presence shows how production agents are being wrapped in controls.

The winning AI stack will not be the one with the most new models. It will be the one that assigns each model a narrow job, measures the completed result and gives every agent only the authority it needs.

The Change Brief

Get the week’s AI changes in one clear read

Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.

Stay ahead of AI without following it all day. We’ll send you what matters each week.

Continue reading

Related reads