Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Z.ai · Language model

GLM 5.3 Prime

z-ai/glm-5.3-prime
  • Language
Released Sep 23, 2026

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...

Specifications

Context window
1,000,000 tokens
Max output
131,072 tokens
Input
text
Output
text
Reasoning effort
max, high, low (always on)
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • frequency_penalty
  • logprobs
  • max_tokens
  • presence_penalty
  • reasoning
  • reasoning_effort
  • response_format
  • seed
  • stop
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"z-ai/glm-5.3-prime","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'