Skip to contentx402 is live on Robinhood Chain: agents can now pay for any model, per call, in USDG →
OmniRail
← All models

Thinking Machines · Language model

Inkling Small

thinkingmachines/inkling-small
  • Language
Released Jul 30, 2026

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Specifications

Context window
524,288 tokens
Max output
262,144 tokens
Input
text, image, audio
Output
text
Reasoning effort
max, high, medium, low, minimal, none
Tokenizer
Other

Supported parameters

Accepted by the model and forwarded as-is by the gateway.

  • frequency_penalty
  • logit_bias
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • seed
  • stop
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_p

Use it

POST /api/v1/chat/completions · MCP chat · full reference

curl https://omnirail.org/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIRAIL_KEY" -H "Content-Type: application/json" \
  -d '{"model":"thinkingmachines/inkling-small","messages":[{"role":"user","content":"Hello"}],"max_tokens":300}'