Kilo Gateway
Inception: Mercury 2.5
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
inception/mercury-2.5Context
260K
Max output
65.5K
Input $/M
$0.20
Output $/M
$0.75
Cache read $/M
$0.02
Cache write $/M
—
Against the catalogue
Context window260K
Larger than 44% of the 7,784 models listed
Input price$0.20
Cheaper than 68% of priced models
Capabilities
Reasoningsupported
Tool callingsupported
Structured outputsupported
Attachmentsnot supported
Temperaturesupported
Open weightsnot supported
Interleaved thinkingnot supported
Modalities
Input
Text
Output
Text
Reasoning controls
effortnone, low, medium, highPricing detail
| Rate | Base $/M |
|---|---|
| Input | $0.20 |
| Output | $0.75 |
| Cache read | $0.02 |