Thinking MachinesBase plan

Inkling

Thinking Machines Lab's Inkling via the Vercel AI Gateway: an open-weights multimodal MoE (975B total, 41B active) with native text, image, and audio understanding, controllable thinking effort, and function calling at a 256k context window. Available serverless on just4o.chat at the base tier. Uses 4 base requests per send before length multipliers. Supports image input. Does not support web search.

You always get the exact model you pick — we never silently route you to another.

About Inkling

Inkling is Thinking Machines Lab's first open-weights model, and on just4o.chat it runs through the Vercel AI Gateway under thinkingmachines/inkling. It is a multimodal MoE with 975B total parameters and 41B active, trained to reason over text, images, and audio in a shared space rather than bolting on separate encoders. Controllable thinking effort lets you trade latency for depth, function calling supports agentic workflows, and a 256k context window on the gateway is roomy enough for long multimodal sessions. It is a base-tier model on just4o.chat, billed at 4 base requests per send before length multipliers, at $1.00 per million input and $4.05 per million output tokens (cached input $0.17). Practical caveats: no native web search on this route, and chat attachments currently emphasize image input even though the model itself understands audio. For teams that want an open, customizable multimodal foundation with ZDR-compatible gateway routing, Inkling is a strong addition.

Best for

  • Multimodal chat and analysis over text plus images
  • Agentic workflows that need function calling and adjustable thinking effort
  • Open-weights experimentation and fine-tuning-oriented product work
  • Long-context multimodal sessions up to 256k tokens
  • Privacy-sensitive gateway traffic routed through ZDR-compatible providers

Specifications

ProviderThinking Machines
Released2026-07
IntelligenceHigh
SpeedMedium
Context window256,000 tokens
Max output256,000 tokens
Knowledge cutoffNot officially published
Input price$1.00 / 1M tokens
Output price$4.05 / 1M tokens
Request cost4 base requests
Plan tierBase
InputText, Image, Audio
OutputText
FeaturesCached input: $0.17 / 1M tokens, 4 base requests per send before length multipliers, Open weights (Apache 2.0), Controllable thinking effort, Function calling supported, Image input supported, Audio understanding supported by the model, Served through the Vercel AI Gateway (thinkingmachines/inkling), ZDR-compatible providers on the Vercel AI Gateway
Model IDinkling

Frequently asked questions

Through the Vercel AI Gateway's OpenAI-compatible Chat Completions endpoint under the thinkingmachines/inkling route. Gateway providers for Inkling are marked ZDR-compatible.

$1.00 per million input tokens and $4.05 per million output tokens at provider reference pricing, with cached input at $0.17 per million. On just4o.chat it is a base model that uses 4 base requests per send before length multipliers.

The model accepts text, image, and audio inputs and returns text. On just4o.chat, image attachments are supported in chat; native web search is not available on this gateway route.

Yes. Inkling is released as open weights under Apache 2.0 by Thinking Machines Lab.

Yes. Inkling exposes controllable thinking effort in the chat UI (none, low, medium, high) so you can balance speed against deeper reasoning.