HeadlinesBriefing favicon HeadlinesBriefing.com

Mercury 2.5: Faster, Cheaper, Smarter AI Model

Hacker News •
×

Mercury 2.5 is the latest AI model from Inception, offering a significant quality boost over Mercury 2 while maintaining the same low-latency, low-cost profile. It's now the largest diffusion language model on the market, with a 40% increase in intelligence and speeds of 1,107 tokens per second on NVIDIA GPUs. The model supports a 260K context window and is priced at $0.20 per million input tokens and $0.75 per million output tokens, with an 80% launch discount.

Real-world deployments show dramatic improvements. Open Call, a voice AI company, reduced response latency from several minutes to just one second (P99), while Augment Code cut latency by 82% and costs by 90% for coding agents. These results highlight Mercury 2.5's ability to handle complex, multi-step tasks efficiently.

In addition to the model, Inception is previewing Mercury Voice and Mercury Router. Mercury Voice optimizes for sub-170ms time-to-first-token, ideal for voice agents. Mercury Router intelligently directs prompts to the best model based on quality, speed, and cost, ensuring optimal performance for any workload.

Mercury 2.5 is available now via the Inception API, Baseten, and Open Router. For enterprises, Inception offers dedicated deployments, autoscaling, and compliance controls. With its combination of speed, intelligence, and affordability, Mercury 2.5 is set to power the next generation of AI applications.