Inkling
Thinking Machines Lab's Inkling via the Vercel AI Gateway: an open-weights multimodal MoE (975B total, 41B active) with native text, image, and audio understanding, controllable thinking effort, and function calling at a 256k context window. Available serverless on just4o.chat at the base tier. Uses 4 base requests per send before length multipliers. Supports image input. Does not support web search.
You always get the exact model you pick — we never silently route you to another.
About Inkling
Inkling is Thinking Machines Lab's first open-weights model, and on just4o.chat it runs through the Vercel AI Gateway under thinkingmachines/inkling. It is a multimodal MoE with 975B total parameters and 41B active, trained to reason over text, images, and audio in a shared space rather than bolting on separate encoders. Controllable thinking effort lets you trade latency for depth, function calling supports agentic workflows, and a 256k context window on the gateway is roomy enough for long multimodal sessions. It is a base-tier model on just4o.chat, billed at 4 base requests per send before length multipliers, at $1.00 per million input and $4.05 per million output tokens (cached input $0.17). Practical caveats: no native web search on this route, and chat attachments currently emphasize image input even though the model itself understands audio. For teams that want an open, customizable multimodal foundation with ZDR-compatible gateway routing, Inkling is a strong addition.
Best for
- Multimodal chat and analysis over text plus images
- Agentic workflows that need function calling and adjustable thinking effort
- Open-weights experimentation and fine-tuning-oriented product work
- Long-context multimodal sessions up to 256k tokens
- Privacy-sensitive gateway traffic routed through ZDR-compatible providers
Specifications
| Provider | Thinking Machines |
|---|---|
| Released | 2026-07 |
| Intelligence | High |
| Speed | Medium |
| Context window | 256,000 tokens |
| Max output | 256,000 tokens |
| Knowledge cutoff | Not officially published |
| Input price | $1.00 / 1M tokens |
| Output price | $4.05 / 1M tokens |
| Request cost | 4 base requests |
| Plan tier | Base |
| Input | Text, Image, Audio |
| Output | Text |
| Features | Cached input: $0.17 / 1M tokens, 4 base requests per send before length multipliers, Open weights (Apache 2.0), Controllable thinking effort, Function calling supported, Image input supported, Audio understanding supported by the model, Served through the Vercel AI Gateway (thinkingmachines/inkling), ZDR-compatible providers on the Vercel AI Gateway |
| Model ID | inkling |
Frequently asked questions
Through the Vercel AI Gateway's OpenAI-compatible Chat Completions endpoint under the thinkingmachines/inkling route. Gateway providers for Inkling are marked ZDR-compatible.
$1.00 per million input tokens and $4.05 per million output tokens at provider reference pricing, with cached input at $0.17 per million. On just4o.chat it is a base model that uses 4 base requests per send before length multipliers.
The model accepts text, image, and audio inputs and returns text. On just4o.chat, image attachments are supported in chat; native web search is not available on this gateway route.
Yes. Inkling is released as open weights under Apache 2.0 by Thinking Machines Lab.
Yes. Inkling exposes controllable thinking effort in the chat UI (none, low, medium, high) so you can balance speed against deeper reasoning.