GuideChoosely Team

Gemini 3.6 Flash Cuts Agent Costs: The AI Tool Changes That Matter This Week

Google's new Gemini Flash tiers sharpen the economics of production agents, while Kimi K3, Grok's Microsoft 365 add-ins and OpenAI Presence move capable models deeper into real work.

A monumental Gemini symbol above two luminous model-routing paths in a dark Choosely technology chamber

Best for

  • Developers and operators running high-volume API or agent workloads.
  • Teams comparing model tiers for coding, document processing, computer use and knowledge work.
  • Buyers deciding whether new embedded agents deserve access to workbooks, mailboxes or company systems.

Not ideal for

  • Teams looking for one benchmark winner to replace every model without task-level testing.
  • Operators unwilling to define permissions, approvals and escalation rules before enabling agent actions.

Google has launched two new Gemini Flash models aimed squarely at production AI agents, while Kimi, SpaceXAI and OpenAI all pushed capable models deeper into real work. Here is the verified Choosely read on what changed, who should care and what to test before changing your stack.

Quick take: Gemini 3.6 Flash is the headline because Google has paired a lower price with claims of stronger coding, knowledge work and computer use. The cheaper 3.5 Flash-Lite tier makes the routing decision even more interesting. Kimi K3 arrived as a 2.8-trillion-parameter flagship with a one-million-token context window, Grok moved into Microsoft 365 and OpenAI launched Presence for controlled enterprise-agent deployments. The pattern is not simply “better models.” It is a new competition over which model does each step, where the agent works and how much authority it receives.

What changed at a glance

  • Google launched Gemini 3.6 Flash and 3.5 Flash-Lite on July 21. Both are available through the Gemini API, Google AI Studio, Gemini Enterprise and the Gemini app, with additional rollout surfaces varying by model.
  • Kimi K3 became available across Kimi products and the API. Moonshot lists a one-million-token context window and API pricing of $3 input and $15 output per million tokens, with cached input at $0.30.
  • Grok 4.5 expanded into Microsoft 365 and consumer surfaces. SpaceXAI released Excel and Outlook add-ins, then made Grok 4.5 live across grok.com, X, iOS and Android.
  • OpenAI introduced Presence on July 22. It packages policies, guardrails, simulations, approved actions and human escalation around voice and chat agents, but is currently limited to eligible enterprise customers.

The headline change: Gemini Flash becomes a model-routing decision

Google introduced Gemini 3.6 Flash as its new workhorse model for coding, knowledge work, multimodal tasks and agentic workflows. It is priced at $1.50 per million input tokens and $7.50 per million output tokens.

Alongside it, Google launched Gemini 3.5 Flash-Lite for high-throughput workloads at $0.30 per million input tokens and $2.50 per million output tokens. Google describes Flash-Lite as the fastest model in the 3.5 series and says it reaches 350 output tokens per second according to Artificial Analysis.

The important shift is not that one new model replaces every old one. It is that Google now offers a clearer two-lane production stack:

ModelInput / 1M tokensOutput / 1M tokensPractical role
Gemini 3.6 Flash$1.50$7.50Higher-value coding, knowledge work, multimodal analysis and multi-step agent tasks
Gemini 3.5 Flash-Lite$0.30$2.50High-volume extraction, classification, search, document processing and simpler agent steps

Google says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, takes fewer reasoning steps and tool calls, and improves computer-use performance. It also reports gains on coding and knowledge-work evaluations. Those results mix independent measurements, Google evaluations and customer observations, so they are evidence to investigate—not a guarantee for every workload.

Computer use is now a built-in client-side tool through the Gemini API and Gemini Enterprise for the new Flash models. That matters because an agent can operate software interfaces without every action being represented by a purpose-built API. It also raises the stakes: browser or desktop actions need explicit scopes, safe defaults and review gates.

Choosely's read: Do not ask whether 3.6 Flash is “better” in the abstract. Ask which steps deserve it. A sensible production design may use Flash-Lite to classify, extract or prepare context, then route only difficult decisions and tool-heavy work to 3.6 Flash. The saving comes from total task cost—not from choosing the cheapest headline rate.

Who should care

  • Developers running high-volume API or agent workloads.
  • Teams using Gemini for document parsing, chart analysis, coding or computer use.
  • Operators paying flagship-model rates for routine substeps.
  • Buyers who need a cheaper fallback without changing provider infrastructure.

What to do next

Choose 50 to 100 representative tasks and compare your current model with both Flash tiers. Track completion rate, output tokens, tool calls, latency, retries, correction time and cost per accepted result. Route by measured task type, not by benchmark reputation.

Read Google's official Gemini Flash announcement.

Kimi K3 enters the frontier race with a one-million-token context window

Moonshot AI introduced Kimi K3 as a 2.8-trillion-parameter Mixture-of-Experts model for long-horizon coding, knowledge work, reasoning and native visual understanding. It is available through Kimi.com, Kimi Work, Kimi Code and the Kimi API.

Moonshot lists:

  • A one-million-token context window.
  • $3 per million input tokens.
  • $15 per million output tokens.
  • $0.30 per million cached input tokens.

Moonshot calls K3 an open model, but its launch post says the full weights and additional technical details are due by July 27. Until the files and final licence are actually available, “open” should not be interpreted as immediately downloadable or straightforward to self-host.

Kimi's documentation also says older Kimi K2.5 and Moonshot V1 models are no longer available to newly registered users, with a full platform sunset expected on August 31. Existing API users should audit hard-coded model names rather than assuming an alias will continue indefinitely.

Choosely's read: Kimi K3 is a credible test candidate for large-context coding, research and deliverable-heavy work, especially where Claude Fable 5 pricing is difficult to justify. But the real decision should wait for the released weights, licence, independent testing and a trial on your own data. A one-million-token window is useful only if retrieval, attention and output quality remain strong across the material you actually provide.

Read Moonshot AI's official Kimi K3 announcement and official Kimi K3 pricing.

Grok moves from a chat tab into Microsoft 365

SpaceXAI released Grok add-ins for Excel and Outlook, then expanded Grok 4.5 across grok.com, X, iOS and Android.

The Excel add-in can answer questions about selected data, cite workbook cells, write formulas, add charts and run scenarios. SpaceXAI says it can also use connectors to pull context from email, SharePoint or Google Drive. The add-in itself is free through Microsoft Marketplace.

The Outlook add-in can summarise threads, draft replies, read attachments and help organise a mailbox. SpaceXAI says nothing sends until the user presses send. It is available to paid X and SuperGrok users.

The broader Grok 4.5 rollout also makes the model available for everyday knowledge work on mobile and web, including spreadsheets, slides, diagrams, prose and long-PDF analysis.

Choosely's read: This is more material than another chatbot app launch because the model now sits beside the workbook and inbox where decisions happen. Test read-only tasks first. Before granting connector or mailbox access, verify exactly which files, messages and external sources the add-in can reach, and keep send, delete and bulk-edit actions behind human confirmation.

Read SpaceXAI's official Grok for Excel announcement, Grok for Outlook announcement and Grok 4.5 availability update.

OpenAI Presence packages the controls around production agents

OpenAI launched Presence on July 22 for enterprise voice and chat agents. It combines model reasoning with policies, standard operating procedures, guardrails, approved actions, simulations, evaluations and escalation rules.

A deployment begins with a narrowly defined job—such as a billing issue, insurance claim or employee IT request. The company controls what knowledge and systems the agent can access, which actions it may take, when approval is required and when a person should take over. OpenAI says Codex can analyse production signals and propose changes, but teams test and approve updates before controlled rollout.

Presence is available through a limited general availability program for eligible enterprise customers. Deployments are led by OpenAI Forward Deployed Engineers and selected systems integrators. It is not a self-serve product, and OpenAI has not published standard pricing.

OpenAI reports that Presence resolves 75% of inbound issues on its English-language phone support line without human assistance and reduced handoffs by 15 percentage points over 10 days. These are OpenAI's own deployment results, not independent benchmarks.

Choosely's read: Presence shows where the enterprise-agent market is moving: the valuable product is no longer only a capable model. It is the permission system, test harness, escalation logic and improvement loop wrapped around it. Smaller teams may not be eligible for Presence, but they should copy the operating principle—one bounded job, least-privilege access, explicit approvals and measurable escalation criteria.

Read OpenAI's official Presence announcement.

The pattern behind this week's changes

The announcements point to four practical shifts:

  1. 1Routing matters more than a single default model. Gemini's two Flash tiers make cost-aware task routing an operator decision.
  2. 2Long context is becoming common—but not automatically useful. Kimi K3's one-million-token window still needs testing for retrieval quality and task completion.
  3. 3AI is moving into the work surface. Grok is no longer waiting in a separate tab; it is beside the spreadsheet and inbox.
  4. 4Permissions are becoming part of the product. Presence treats guardrails, approvals and escalation as core infrastructure rather than post-launch paperwork.

The sensible response is not to migrate everything this weekend. Pick one measurable workflow, define what a successful result costs today, and compare the new option while keeping the existing path available.

What Choosely verified this week

Every material date, price and availability statement in this brief was checked against a first-party vendor announcement, documentation or pricing page on July 24, 2026.

Availability can vary by plan, geography, account eligibility, workspace controls and staged rollout. Recheck the relevant account and live pricing page before changing a production workflow.

The bottom line

Gemini 3.6 Flash is this week's clearest operator signal: better agent economics will increasingly come from routing each step to the right tier. Kimi K3 adds a serious long-context challenger, Grok is embedding itself in Microsoft 365, and OpenAI Presence shows how production agents are being wrapped in controls.

The winning AI stack will not be the one with the most new models. It will be the one that assigns each model a narrow job, measures the completed result and gives every agent only the authority it needs.

Replacement guides

Compare more replacement options

Save the useful parts

Build your AI stack in Choosely

Save tools you're considering, keep workflow context attached, and use your account as the foundation for future stack updates.

What matters most

Gemini 3.6 Flash is priced at $1.50 input and $7.50 output per million tokens; 3.5 Flash-Lite costs $0.30 and $2.50.
Kimi K3 combines a one-million-token context window with $3 input and $15 output pricing, while its full weights remain scheduled for release.
Grok moved into Excel and Outlook, and OpenAI Presence packages policies, guardrails and escalation around enterprise agents.

This week's changes at a glance

OptionBest forWhy it winsTradeoff
Gemini FlashRouting production workloads between a stronger workhorse model and a cheaper high-throughput tier.The two-tier pricing creates a practical path to lower cost per completed agent task inside one provider stack.Google's efficiency and benchmark claims still need validation on each team's real workloads.
Kimi K3Large-context coding, research and deliverable-heavy work where flagship-model costs compound.It combines a one-million-token context window with pricing below premium flagship models.The full weights and final licence were not yet available at verification time, and independent evidence is still developing.
Embedded agentsWorkflows where the model needs to operate beside spreadsheets, email or enterprise systems.Grok and Presence reduce context switching and put permissions closer to the actual work.Broader access increases operational risk, so read-only trials and human approval gates matter.

What to do next

  1. 1Benchmark both Gemini Flash tiers on representative tasks and measure cost per accepted result.
  2. 2Audit model names and sunset dates before adopting Kimi K3 or changing production aliases.
  3. 3Test Grok add-ins with read-only, non-sensitive material before granting connector or mailbox permissions.
  4. 4Write explicit approval and escalation rules for any agent that can change external systems.

FAQ

How much does Gemini 3.6 Flash cost?

Google lists Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens.

How much does Gemini 3.5 Flash-Lite cost?

Google lists Gemini 3.5 Flash-Lite at $0.30 per million input tokens and $2.50 per million output tokens.

Is Kimi K3 open source?

Moonshot calls Kimi K3 an open model, but at verification time its launch page said the full weights and additional technical details would be released by July 27. Check the released files and licence before planning self-hosting.

Is Grok for Excel free?

SpaceXAI describes the Excel add-in as a free Microsoft 365 add-in. Access to underlying Grok services, connectors or other add-ins may still depend on the user's plan.

Can anyone buy OpenAI Presence?

No. OpenAI says Presence is available to eligible enterprise customers through a limited general availability program and is not a self-serve product.

AI stack brief

Get the weekly AI stack change brief

Pricing moves, tool launches, free-tier changes, and practical AI stack updates - written for people who actually use these tools.

Prefer account-based updates? Create a free account and use it as the foundation for stack updates as Choosely rolls out email digests.

Related reads

Browse more updates on the AI Radar hub. Looking for the right AI tool for a specific task? Try the Choosely tool finder For a related read, continue with What Is Kimi K3? The Powerful New AI Model Challenging Claude.