September 24, 2026
Long answers arrive faster: Mercury 2.5 delivers 780 tokens per second
In an Artificial Analysis measurement dated September 8, Mercury 2.5 delivered 780.8 tokens per second through the API. Inception gives a new account 100 million free tokens without payment details; an API key is then required.

Mercury 2 delivered 421 tokens per second in an independent Puter test. In the same test, Mercury 2.5 posted a median of 1,241 tokens per second across 10 runs.
The start is not instant. Mercury 2.5 started returning a 600-word response after about 5 seconds, while three other models started after 0.6 seconds. But Mercury completed the full request in a median 5.7 seconds, versus 9.4 seconds for Claude Haiku 4.5.
The model is called through an OpenAI-compatible API: install the inceptionai package, set INCEPTION_API_KEY, and select mercury-2.5. At launch, Inception listed pricing at $0.20 per million input tokens and $0.75 per million output tokens.
Inception already uses Mercury in Augment Code for context compression, model routing, and MCP tool discovery.
