Moonshot AIPremium plan

Kimi K2.6

Moonshot AI's Kimi K2.6 via Fireworks: open-source, native multimodal agentic model for long-horizon coding, coding-driven design, autonomous execution, and task orchestration. 1T MoE, 262k context. Supports image input and function calling. Uses 2 premium requests per send before length multipliers. Does not support web search.

You always get the exact model you pick — we never silently route you to another.

About Kimi K2.6

Kimi K2.6 is Moonshot AI's open-weight coding specialist built for the kind of work that takes hours, not seconds. Its signature capability is agent swarm orchestration — coordinating up to 300 sub-agents across 4,000 execution steps — enabling autonomous refactoring sessions that developers have run for over 13 hours straight. On SWE-Bench Verified it scores 80.2%, and it edges out GPT-5.4 on SWE-Bench Pro at 58.6%, making it the strongest open-weight coding model available at its price point. Users report up to 88% cost savings on coding workloads compared to proprietary alternatives, which is the real draw for teams running code-heavy pipelines at scale. The tradeoff is speed and occasional drift: at 40.6 tokens per second — well below the category median — it is not suited to real-time use. In long-running agentic tasks, users note the model can wander into unnecessary redesigns around the three-hour mark, requiring clear, constrained prompting to keep it on track. For deep, non-interactive coding work where cost efficiency and open-weight flexibility matter more than instant responses, K2.6 occupies a position few models can match.

Best for

  • Long-horizon autonomous coding and multi-file refactoring projects
  • Multi-agent orchestration workflows requiring coordination of specialized sub-agents
  • Cost-sensitive coding deployments where price-per-result matters more than speed
  • Extended research and document analysis using the 256K context window
  • Self-hosted or fine-tuned deployments where open-weight access is required

Specifications

ProviderMoonshot AI
Released2026-04
IntelligenceHigh
SpeedSlow
Context window262,100 tokens
Max output262,144 tokens
Knowledge cutoffNot publicly disclosed
Input price$0.95 / 1M tokens
Output price$4.00 / 1M tokens
Request cost2 premium requests
Plan tierPremium
InputText, Image
OutputText
FeaturesCached input: $0.16 / 1M tokens, 2 premium requests per send before length multipliers, Function calling supported, LoRA fine-tuning supported on Fireworks
Model IDkimi-k2p6

Frequently asked questions

Input is $0.95 per million tokens and output is $4.00 per million tokens on Moonshot AI's official platform. Alternative providers like Parasail offer lower rates ($0.60 input / $2.80 output). A blended rate of roughly $0.70 per million tokens applies at a 7:2:1 input/output ratio.

256,000 tokens (256K). Maximum output is 98,304 tokens for standard tasks, or up to 262,144 tokens when using tools in long-horizon agent runs.

Autonomous coding and multi-agent orchestration. It scores 80.2% on SWE-Bench Verified and supports coordinating up to 300 sub-agents across 4,000 steps — capabilities that make it well-suited to extended, unsupervised engineering tasks.

Output speed is 40.6 tokens per second, significantly below the category median of ~70 t/s, making it unsuitable for real-time or conversational use. In long agentic sessions it can drift into unnecessary optimizations, requiring careful prompting to constrain scope. Multimodal image understanding ranks weak relative to peers.

Kimi K2.7 Code is a fine-tuned fork of K2.6 optimized specifically for code generation and tool use. If your workload is almost entirely coding, K2.7 Code is the more specialized option. K2.6 is the broader general-purpose model with stronger agent swarm and research capabilities.

Yes. It is distributed as an open-weight model on Hugging Face under a Modified MIT License, allowing self-hosting, fine-tuning, and on-premise inference.

Related models