Coding & app building

Inception Mercury

By inceptionlabs.ai

Inception Mercury is a strong fit for extremely latency-sensitive or high-throughput llm api calls, with a profile optimized for advanced users who value medium ease-of-use and medium output quality.

Best for: Extremely latency-sensitive or high-throughput LLM API calls

What it is

Inception's closed-weights diffusion-LLM API — Mercury 2.5 is Inception's most capable production model, tuned for extremely low latency and high throughput on supporting agent calls, RAG substeps, context compaction and structured, low-cost generation.

In Choosely terms, this sits in the coding & app building lane and is commonly selected for extremely latency-sensitive or high-throughput llm api calls and high-volume supporting agent steps, rag substeps and context compaction.

Pricing

Usage-based token pricing for Mercury 2.5 (a production model): standard $0.20 per million input tokens and $0.75 per million output tokens. A temporary 80% launch discount ($0.04/M input, $0.15/M output) was advertised but should not be treated as permanent. Available via the Inception API, Baseten and OpenRouter; enterprise deployments support dedicated capacity, autoscaling, compliance controls and configurable data retention.

Basis: Usage BasedConfidence: VerifiedLast checked: September 2026

Why people pick it vs where it falls short

Why people pick it

  • Diffusion-LLM architecture tuned for very low latency; vendor claims up to 1,107 tokens/sec on widely available NVIDIA GPUs (vendor benchmark, not independently verified)
  • OpenAI-API-compatible drop-in with 260K context, parallel tool calls, schema-aligned JSON output and tunable reasoning
  • Low standard token pricing ($0.20/M input, $0.75/M output), available via the Inception API, Baseten and OpenRouter, with enterprise deployments offering dedicated capacity, autoscaling, compliance controls and configurable data retention

Where it falls short

  • Optimized for speed/cost on supporting steps, not frontier-quality reasoning; vendor quality and adoption claims are not independently verified
  • Closed-weights (no local/open-weights option); Mercury Voice and Mercury Router are previews
  • An advertised 80% launch discount ($0.04/M input, $0.15/M output) is temporary — plan for the standard $0.20/$0.75 pricing

When it is a strong fit

A strong match when your main priority is extremely latency-sensitive or high-throughput llm api calls and you need an advanced-friendly starting point.

Useful when your team values medium ease of use and fast execution over heavier setup.

Best when medium quality matters, but you still want a practical workflow rather than a complex implementation track.

How it compares in Choosely terms

  • Speed profile: Fast. This is best when you want momentum from prompt to usable output without heavy process overhead.
  • Ease profile: Medium for Advanced users. You can move quickly even if this is not your full-time specialty.
  • Control profile: High. Expect practical customization, but not an infinite-control architecture.
  • Pricing signal: Usage-based. Good for teams balancing capability with cost sensitivity.
Tradeoff: Optimized for speed/cost on supporting steps, not frontier-quality reasoning; vendor quality and adoption claims are not independently verified.

Best-fit use cases

Practical ways Inception Mercury fits the current Choosely catalog profile.

Extremely Low Latency Llm API For Supporting Agent Calls

Strong lane

Use Inception Mercury for extremely low-latency llm api for supporting agent calls when you want fast execution, medium ease of use, and medium output quality.

High Throughput Diffusion Llm API For Rag Substeps

Use Inception Mercury for high-throughput diffusion llm api for rag substeps when you want fast execution, medium ease of use, and medium output quality.

Fast Structured Json Generation With Parallel Tool Calls

Use Inception Mercury for fast structured json generation with parallel tool calls when you want fast execution, medium ease of use, and medium output quality.

Low Latency Model For Context Compaction And Routing Substeps

Strong lane

Use Inception Mercury for low-latency model for context compaction and routing substeps when you want fast execution, medium ease of use, and medium output quality.

Alternatives

ChatGPT

General-purpose conversational assistant for drafting, ideation, lightweight research, file-based work, coding help, and everyday task support — now led by the GPT-6 Astra model, with ChatGPT Images 2.5 for image generation and editing.

Choose ChatGPT if output quality consistency matters more than raw speed.

Claude

Conversational reasoning assistant especially strong for long-form writing, careful analysis, structured thinking, and document-heavy work.

Choose Claude if output quality consistency matters more than raw speed.

Next step

Point a latency-sensitive substep (RAG, tool-calling, or context compaction) at the Mercury 2.5 API via OpenRouter, budget for standard $0.20/$0.75 pricing rather than the launch discount, and reserve frontier models for the reasoning-heavy steps.

Related reads

FAQ

What is Inception Mercury best for?

Inception Mercury is best for extremely latency-sensitive or high-throughput llm api calls, high-volume supporting agent steps, rag substeps and context compaction, structured, low-cost json generation with parallel tool calls.

Is Inception Mercury beginner-friendly?

This catalog profile lists Inception Mercury at advanced skill level with medium ease of use.

What should I watch out for before choosing Inception Mercury?

Optimized for speed/cost on supporting steps, not frontier-quality reasoning; vendor quality and adoption claims are not independently verified