← All models
Thinking Machines · Language model
Inkling Small
thinkingmachines/inkling-small- Language
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Specifications
- Context window
- 524,288 tokens
- Max output
- 262,144 tokens
- Input
- text, image, audio
- Output
- text
- Reasoning effort
- max, high, medium, low, minimal, none
- Tokenizer
- Other
Supported parameters
Accepted by the model and forwarded as-is by the gateway.
- frequency_penalty
- logit_bias
- max_tokens
- min_p
- presence_penalty
- reasoning
- reasoning_effort
- repetition_penalty
- seed
- stop
- temperature
- tool_choice
- tools
- top_k
- top_p
Use it
POST /api/v1/chat/completions · MCP chat · full reference
curl https://omnirail.org/api/v1/chat/completions \
-H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
-d '{"model":"thinkingmachines/inkling-small","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'