Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

NVIDIA · Language model

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b
  • Language
Released Jun 4, 2026

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Specifications

Context window
262,144 tokens
Max output
182,520 tokens
Input
text
Output
text
Reasoning effort
high, medium
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • frequency_penalty
  • logit_bias
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'