Inception: Mercury Coder
This model is no longer available.Add AI to your application with Puter.js.
Explore Other ModelsModel Card
Mercury Coder is a code-specialized diffusion language model from Inception Labs, built on the same parallel token refinement architecture as Mercury. It is available in Mini and Small sizes.
On fill-in-the-middle tasks, Mercury Coder Small scored 84.8% average accuracy, exceeding Codestral 2501 (82.5%). On MultiPL-E, it reaches 82.0% in C++, 83.9% in JavaScript, and 82.6% in TypeScript. In Copilot Arena human evaluations, Mercury Coder Mini ranked second in user preference with an average latency of just 25 milliseconds.
It is the go-to choice for real-time code completion, autocomplete, and apply-edit workflows where both speed and accuracy are critical.
Context Window 128K
tokens
Max Output 32K
tokens
Input Cost $0.25
per million tokens
Output Cost $0.75
per million tokens
Release Date Mar 31, 2025
Code Example
Add AI to your app with the Puter.js AI API — no API keys or setup required.
// npm install @heyputer/puter.js
import { puter } from '@heyputer/puter.js';
puter.ai.chat("Explain quantum computing in simple terms").then(response => {
document.body.innerHTML = response.message.content;
});
<html>
<body>
<script src="https://js.puter.com/v2/"></script>
<script>
puter.ai.chat("Explain quantum computing in simple terms").then(response => {
document.body.innerHTML = response.message.content;
});
</script>
</body>
</html>
More AI Models From Inception
Mercury 2.5 Preview
Mercury 2.5 Preview is a diffusion-based reasoning language model from Inception Labs, refining tokens in parallel rather than generating them one at a time. It is the newest release in the Mercury line, following Mercury 2, and adds tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. Inception reports throughput of 1,107 tokens per second on standard GPUs and a 10+ point jump in intelligence over Mercury 2. It says quality is comparable to cost-optimized models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. With a 260K context window and 65,536 max output tokens, it targets production workloads where latency compounds, such as search agents, voice pipelines, and coding subagents. It is available through OpenRouter with an OpenAI-compatible API, priced at $0.04 per million input tokens and $0.15 per million output tokens.
ChatMercury 2
Mercury 2 is a diffusion-based reasoning language model from Inception Labs that refines all tokens in parallel rather than generating them sequentially, achieving over 1,000 tokens per second — roughly 5x faster than speed-optimized competitors like Claude Haiku and GPT-5 Mini at comparable quality. On reasoning benchmarks, Mercury 2 scores 91.1 on AIME 2025 and 73.6 on GPQA. It also placed second on the Copilot Arena leaderboard for quality while ranking first for speed overall. With a 128K context window, it is purpose-built for latency-sensitive applications — real-time assistants, high-throughput pipelines, and cost-conscious production workloads where reasoning capability matters.
Frequently Asked Questions
You can access Mercury Coder by Inception through Puter.js AI API. Include the library in your web app or Node.js project and start making calls with just a few lines of JavaScript — no backend and no configuration required. You can also use it with Python or cURL via Puter's OpenAI-compatible API.
Yes, it is free if you're using it through Puter.js. With the User-Pays Model, you can add Mercury Coder to your app at no cost — your users pay for their own AI usage directly, making it completely free for you as a developer.
| Price per 1M tokens | |
|---|---|
| Input | $0.25 |
| Output | $0.75 |
Mercury Coder was created by Inception and released on Mar 31, 2025.
Mercury Coder supports a context window of 128K tokens. For reference, that is roughly equivalent to 256 pages of text.
Mercury Coder can generate up to 32K tokens in a single response.
Yes — the Mercury Coder API works with any JavaScript framework, Node.js, or plain HTML through Puter.js. Just include the library and start building. See the documentation for more details.
Get started with Puter.js
Add AI to your application without worrying about API keys or setup.
Explore Models View Tutorials