Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Z.ai · Language model

GLM 5.3 Flash

z-ai/glm-5.3-flash
  • Language
Released Aug 26, 2026

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Specifications

Context window
1,310,720 tokens
Max output
943,717 tokens
Input
text, image, video
Output
text
Reasoning effort
max, high, low (always on)
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • frequency_penalty
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • parallel_tool_calls
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"z-ai/glm-5.3-flash","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'