Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Z.ai · Language model

GLM 5.3 FlashX

z-ai/glm-5.3-flashx
  • Language
Released Sep 18, 2026

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Specifications

Context window
1,048,576 tokens
Max output
131,072 tokens
Input
text, image, video
Output
text
Reasoning effort
max, high, low (always on)
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • max_tokens
  • reasoning
  • reasoning_effort
  • response_format
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"z-ai/glm-5.3-flashx","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'