Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Inception · Language model

Mercury 2.5

inception/mercury-2.5
  • Language
Released Sep 8, 2026

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Specifications

Context window
260,000 tokens
Max output
65,536 tokens
Input
text
Output
text
Reasoning effort
high, medium, low, none
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • max_tokens
  • reasoning
  • reasoning_effort
  • response_format
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"inception/mercury-2.5","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'