Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Inception · Language model

Mercury 2

inception/mercury-2
  • Language
Released Mar 4, 2026

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

Specifications

Context window
128,000 tokens
Max output
50,000 tokens
Input
text
Output
text
Reasoning effort
high, medium, low, none
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • max_tokens
  • reasoning
  • reasoning_effort
  • response_format
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"inception/mercury-2","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'