Hacker News
Mercury 2.5
Mercury 2.5, the newest diffusion-based LLM, delivers a 40% intelligence boost over Mercury 2, matching cost-optimized frontier models while processing 1,107 tokens per second on common NVIDIA GPUs and handling 260 K-token contexts. Priced at $0.20 per million input and $0.75 per million output (80% launch discount), it powers latency-sensitive search, voice, and coding workloads, cutting response times to sub-200 ms and reducing latency and cost by up to 90% in production use cases.