← All models
Inception · Language model
Mercury 2.5
inception/mercury-2.5- Language
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Specifications
- Context window
- 260,000 tokens
- Max output
- 65,536 tokens
- Input
- text
- Output
- text
- Reasoning effort
- high, medium, low, none
- Tokenizer
- Other
Supported parameters
Accepted by the model and forwarded as-is by the gateway.
- max_tokens
- reasoning
- reasoning_effort
- response_format
- stop
- structured_outputs
- temperature
- tool_choice
- tools
Use it
POST /api/v1/chat/completions · MCP chat · full reference
curl https://omnirail.org/api/v1/chat/completions \
-H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
-d '{"model":"inception/mercury-2.5","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'