‹ Back to Home

Mercury 2: The First Reasoning Diffusion Language Model — 5x Faster Than Claude Haiku at Half the Cost

Inception Labs launches Mercury 2, the world's first Reasoning Diffusion Language Model (dLLM), achieving 1,009 tokens per second on Nvidia Blackwell GPUs — 5x faster than Claude 4.5 Haiku and GPT-5 Mini. Priced at $0.25/M input and $0.75/M output tokens, it undercuts Gemini 3 Flash by 4x on output cost.

Keerthika 4 min read 471
Follow on Google
Updated 2 weeks ago
AI & Future Mercury 2: The First Reasoning Diffusion Language Model — 5x Faster Than Claude Haiku at Half the Cost 4 min left Follow on Google
Mercury 2: The First Reasoning Diffusion Language Model — 5x Faster Than Claude Haiku at Half the Cost

TamilTech AI summary

Inception Labs launched Mercury 2 on February 24, 2026, calling it the world’s first Reasoning Diffusion Language Model, and it generates text in a totally different way from ChatGPT or Claude. Instead of writing one token at a time like normal transformers, it starts with a rough noisy draft of the whole answer and refines many tokens in parallel through denoising steps, which is why it hits about 1,009 tokens per second—roughly 5x faster than Claude 4.5 Haiku—while costing far less per million tokens. That speed and price cut matter a lot for real-time work like agentic coding, customer support at scale, multi-step AI agents, and high-volume India use cases such as vernacular assistants, EdTech tutoring, fintech checks, and government services. You should know the company says coherence issues common to diffusion language models are largely fixed, though independent tests are still ongoing and long-form creative writing may still favor classic transformers for now. API access is live at inceptionlabs.ai, so if you’re weighing AI infrastructure costs and latency, Mercury 2 is worth a serious look.

  • What is Mercury 2 and how is it different from ChatGPT or Claude?
  • How fast is Mercury 2 compared to other AI models?
  • Can Indian developers access Mercury 2's API?
  • What are the limitations of Mercury 2's diffusion approach?

AI-assisted summary, checked by the TamilTech editorial team.

Mercury 2: The First Reasoning Diffusion Language Model — How Inception Labs Is Reinventing AI Architecture

On February 24, 2026, AI startup Inception Labs launched Mercury 2 — the world's first Reasoning Diffusion Language Model (dLLM). The benchmark numbers are staggering: 1,009 tokens per second, 5x faster than Claude 4.5 Haiku, at pricing that undercuts competitors by 2–4x. This isn't a faster version of ChatGPT or Claude. It's a fundamentally different approach to how AI generates text.

Founded by researchers from Stanford, UCLA, and Cornell who contributed to foundational diffusion model research, Inception Labs has spent years applying diffusion — the technique behind image generators like Stable Diffusion — to language. Mercury 2 is the production-ready result.

Premium Content

You've read all your free articles today. Subscribe to continue reading.

You've used 3 of 3 free articles today.

Subscribe Now

Already subscribed? Sign in

Get tomorrow’s tech news on WhatsApp

One short update a day, free. Follow the TamilTech channel.

What do you think?

people reacted

Keerthika

TamilTech editorial team · 3,344 articles

Keerthika is an editor at TamilTech, the Tamil and English technology publication founded by Praveen Kumar S. She covers AI, smartphones, gadgets, EVs, startups and cybersecurity i...

More from Keerthika

Ask TamilTech on WhatsApp

Tech doubt? Ask in Tamil or English — our WhatsApp assistant answers from TamilTech articles in seconds.

Related stories

Comments (0)

| Supports **bold**, *italic*, `code`

Be the first to comment!

Next story Who's Really Paying for Yotta and Rivals' Multi-Billion-Dollar Nvidia Orders?
Tamiltech

Tamiltech

Install app for faster access

Earn XP 🏆
WhatsApp
Notifications