On September 8, 2026, AI company Inception announced "Mercury 2.5," an AI model that generates text using a diffusion model. Unlike traditional autoregressive LLMs that determine the next token sequentially from the previous one, this model adopts a mechanism called a "Diffusion Large Language Model (dLLM)," which begins by creating the entire response in a coarse state and refines multiple tokens in parallel.
Mercury 2.5 achieves an output speed of 1,107 tokens per second on NVIDIA GPUs. Regarding performance, Inception stated that the model reaches a quality comparable to cost-optimized state-of-the-art models such as "GPT-5.6 Luna (Low)," "Gemini 3.5 Flash-Lite," and "Claude Haiku 4.5." Pricing is set at $0.20 (approximately ¥31) per million input tokens and $0.75 (approximately ¥120) per million output tokens.
Additionally, the model supports functions to adjust the strength of inference according to the use case, the ability to call multiple external tools in parallel, and response outputs that follow a specified JSON schema. Along with these, Inception also announced preview versions of "Mercury Voice" for voice AI and "Mercury Router," which distributes inputs to the most optimal model based on the content.
Source: