Inkling pricing & benchmarks
Released Jul 17, 2026 · served by 3 providers · ranked #27 by intelligence in our index
What Inkling is
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
- Model ID:
thinkingmachines/inkling - Modality: text+image+audio->text
- Tool calling: supported · Structured output: not supported
- Weights:
thinkingmachines/Inklingon Hugging Face
Providers and prices for Inkling
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $1.00 | $4.05 | 524K | fp8 | 96.4% | — |
| BaseTen | $1.00 | $4.05 | 1.0M | fp8 | 99.9% | — |
| Together | $1.00 | $4.05 | 524K | — | 99.2% | — |
Cheaper models in the same class
Models scoring within 4 points of Inkling on the Intelligence Index, but with a lower output price.
FAQ
How much does the Inkling API cost?
$1.00 per 1M input tokens and $4.05 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $2.21.
Is Inkling free?
No. It is a paid model starting at $1.00 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Inkling?
DeepInfra at $1.00 input / $4.05 output per 1M tokens (fp8 quantization).
Can I self-host Inkling?
Yes — weights are published as thinkingmachines/Inkling on Hugging Face, so you can run it on your own hardware.