Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Meta · Language model

Llama 4 Maverick

meta-llama/llama-4-maverick
  • Language
Released Apr 5, 2025

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Specifications

Context window
1,048,576 tokens
Max output
16,384 tokens
Input
text, image
Output
text
Tokenizer
Llama4
Knowledge cutoff
2024-08-31

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • frequency_penalty
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"meta-llama/llama-4-maverick","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'