← All models
Z.ai · Language model
GLM 5.3 FlashX
z-ai/glm-5.3-flashx- Language
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Specifications
- Context window
- 1,048,576 tokens
- Max output
- 131,072 tokens
- Input
- text, image, video
- Output
- text
- Reasoning effort
- max, high, low (always on)
- Tokenizer
- Other
Supported parameters
Accepted by the model and forwarded as-is by the gateway.
- max_tokens
- reasoning
- reasoning_effort
- response_format
- temperature
- tool_choice
- tools
- top_k
- top_p
Use it
POST /api/v1/chat/completions · MCP chat · full reference
curl https://omnirail.org/api/v1/chat/completions \
-H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
-d '{"model":"z-ai/glm-5.3-flashx","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'