Mercury

Mercury

Inception Labs · Voice & Language · Coding

Mercury is a family of diffusion large language models from Inception Labs that generates text in parallel instead of one word at a time. The current release, Mercury 2.5, runs at 1,107 tokens per second with a 260K context window. Pricing sits at $0.20 per million input tokens and $0.75 per million output tokens. As a text generation model, it targets developers who need fast, low-cost output for search, voice, and coding workloads. This low-latency AI model is rented through an API rather than run on your own machine, so AI model pricing is pay-as-you-go.

Interface preview of Mercury

About Mercury

What Is Mercury

Mercury is the flagship model line from Inception Labs, a company founded by researchers from Stanford, UCLA, Cornell, Google DeepMind, Meta AI, Microsoft AI, and OpenAI. Its defining trait is the architecture underneath. Almost every commercial LLM generates one token at a time, left to right. Mercury uses diffusion, the same broad idea behind image generators, to produce many tokens at once and then refine them. That's where the speed comes from. The company says the approach delivers 5 to 7 times higher throughput and up to 70% lower cost than standard models of similar quality.

The trade-off worth understanding up front: Mercury is a text and reasoning model served over an API, not a downloadable app you install. You don't chat with it on a website the way you would with ChatGPT. You call it from your own software, paying per token. Inception markets it to teams that already have a product and need the model layer to feel instant.

Speed is the headline, but control is the quieter selling point. Because diffusion lets the model adjust tokens as a group, Inception says it can follow strict JSON schemas and semantic rules more reliably than autoregressive models, and it supports tunable reasoning and parallel tool calls. If your pipeline chains dozens of model calls per user request, that combination matters more than raw benchmark scores. Think about it. A search agent might plan, rewrite, rerank, and summarize before it ever answers.

Getting Started

  1. Create an account at the Inception Labs platform and claim an API key from the console.
  2. Pick a Mercury model in the dashboard and check the current per-token rates before you build.
  3. Send a first request with the chat completions endpoint, using the model name provided in your account.
  4. Test the schema-aligned JSON mode if your app needs structured output, and enable tunable reasoning only where it helps.
  5. Wire the key into your production environment and monitor latency and spend as traffic scales.

Product Information

A quick look at Mercury's pricing, supported platforms, and performance.

Free PlanYes
Paid Plans$0.20 - $0.75 per million tokens
PlatformAPI, Web
DeveloperInception Labs
CategoryVoice & Language · Coding
Release DateJan 2025
Latest UpdatedSep 2025
Website Visits102.5K
Website Global Rank350.8K
API AvailabilityYes

Best for

The users, tasks, and scenarios where this tool fits best.

Users

  • Backend and ML engineers
  • Startup teams on tight budgets
  • Voice and support builders

Tasks

  • Search and RAG pipelines
  • Code completion and refactoring
  • Structured data extraction

Scenarios

  • Adding a responsive assistant to an existing app without rebuilding your infrastructure around latency.
  • Prototyping an AI feature on a small budget, then scaling it up as usage grows.
  • Running high-volume batch jobs where cost per task matters more than squeezing out the last bit of quality.

Key features

Parallel token generation

Mercury writes tokens in parallel rather than one at a time, the core reason it's several times faster than conventional models. Inception reports a sub-300ms time to first token and 1,107 tokens per second on widely available NVIDIA GPUs. That gap sounds technical until you use it. For a user, that's the difference between an assistant that feels instant and one that makes them wait while a spinner turns, which is exactly the kind of small delay that quietly pushes people away from a product.

Long 260K context window

The Mercury 2.5 model accepts up to 260,000 tokens of context in a single request. That's enough to hold long documents, big code files, or a lengthy conversation history without chopping it up. Long context plus high speed is an unusual pairing, since most fast models keep their windows modest, and teams that work with large codebases or lengthy reports tend to feel that gap immediately. Huge deal for big files.

Tunable reasoning

You can dial how much internal reasoning the model does per request. Turn it up for hard problems that need step-by-step thinking; turn it down for simple lookups where extra reasoning just adds latency and cost. That control is rare in hosted models, where reasoning is usually on or off. Nice.

Schema-aligned JSON output

Mercury can constrain its output to match a JSON schema you define. Diffusion lets it fix up tokens as a group, which Inception says makes structured output more dependable. Fewer broken replies. Developers get less malformed output to clean up when they feed results into the next step of a pipeline, which is the kind of papercut that adds up fast once a workflow runs thousands of times a day in production.

Parallel tool calls

The model can fire off multiple tool or function calls at once instead of waiting for each to finish. In an agent that checks a calendar, queries a database, and searches the web, that parallelism trims real seconds off the response. It fits the multi-step workflows Mercury is pitched for, and anyone who has watched an agent grind through five sequential calls knows how much that wait adds up over a single conversation. Big win for agents.

Diffusion architecture

Mercury is a diffusion LLM, or dLLM. Inception describes it as among the largest diffusion language models trained to date. The architecture also opens a path to mixing text with audio, images, and video in one framework. Today's shipped product is text-first.

Pros and cons

Pros

  • Throughput of 1,107 tokens per second and sub-300ms first-token latency, fast even by frontier-model standards.
  • Pricing of $0.20 per million input and $0.75 per million output tokens, well below most comparable models.
  • 260K context window handles long documents and codebases without splitting them.
  • Schema-aligned JSON and tunable reasoning give developers finer control over output.

Cons

  • API-only: there's no desktop app or mobile app, so casual users can't just open Mercury and chat.
  • Quality is pitched as comparable to cost-optimized frontier models, not the absolute top tier, so the hardest reasoning tasks may still favor bigger models.
  • Adoption is still early, which means fewer community examples and third-party integrations than the mainstream LLM ecosystems.
  • Rates and model names change quickly, so you'll want to check current pricing before committing.

Frequently asked questions

There's a free tier to try it out through the Inception Labs platform, but the model itself is billed per token in production. At launch, Mercury 2.5 was offered at an 80% discount, working out to $0.04 per million input tokens and $0.15 per million output tokens.

Related content

Explore related tools, skills, and articles for Mercury.

Mercury Alternatives

Prosp

Prosp

Prosp · Writing · Voice & Language · Marketing

Prosp is an AI LinkedIn outreach tool built for agencies and sales teams, and it writes the message and the voice note in your own voice for each prospect so they actually reply. You connect your accounts, find leads, and let the AI draft and send personalized messages at scale, all from one inbox. It's built for people running outbound at volume. That's the whole pitch. Every touchpoint still has to feel human.

Paid / $30.99 - $79.99 per account/moView details
Wordly AI Translation

Wordly AI Translation

Wordly · Voice & Language · Productivity

Wordly AI Translation is a real-time AI translation and captioning platform built for meetings, conferences, and events. It delivers live translation, captions, transcripts, and summaries in more than 60 languages, and attendees join by scanning a QR code or opening a link instead of using dedicated headsets. The platform works with Zoom, Microsoft Teams, Google Meet, and Webex, and it's designed for organizations that want multilingual access without hiring human interpreters for every session. Simple as that.

Paid / $0 - $150/moView details
Musicful

Musicful

Musicful AI · Voice & Language · Video

Musicful is an AI music generator and AI music video maker that turns text to music in minutes. Give it a text prompt, a set of lyrics, or a hummed melody and it returns a finished track with vocals and instruments. It also doubles as an AI song generator, produces music videos from the songs you create, and offers an AI cover tool plus a developer API. The platform runs in a web browser and through an Android app, so you can start a song on desktop and pick it up on your phone.

Free / $0 - $20/moView details