Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

NVIDIA · Language model

Nemotron 3.5 Lightning

nvidia/nemotron-3.5-lightning
  • Language
Released Aug 11, 2026

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Specifications

Context window
1,000,000 tokens
Max output
131,072 tokens
Input
text
Output
text
Reasoning
supported
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • frequency_penalty
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3.5-lightning","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'