Kimi K2.7 Code
Kimi K2.7 Code
https://huggingface.co/moonshotai/Kimi-K2.7-Code
Free to self-host; API from $0.95/$4.00 per 1M tokens
Preserve Thinking Mode — Reasoning chains persist across conversation turns; model maintains context about architectural decisions across multi-step workflows
30% Fewer Thinking Tokens — Uses ~30% fewer reasoning tokens than K2.6 to reach same conclusions; direct cost reduction on every task
MCP Tool Use Leader — 81.1% on MCP Mark Verified, beating Opus 4.8's 76.4%; reliable tool calling for coding agent workflows
Open Weights (Modified MIT) — Full model weights on Hugging Face; download, self-host, fine-tune, use commercially
Native INT4 Quantization — Ships with INT4 quantization built in; dramatically reduces memory requirements for self-hosting
OpenAI- and Anthropic-Compatible APIs — Drop into existing agents with a base URL swap; works with Claude Code, Cline, Roo Code
MoE Architecture — 1T total parameters but only 32B activate per token; frontier-level intelligence at a fraction of compute cost
256K Context Window — Large enough for substantial codebases; handles complex multi-file coding tasks
Multimodal Input — MoonViT vision encoder (400M params) handles images and video; useful for design-to-code workflows
vLLM / SGLang / KTransformers Support — Multiple serving options for self-hosting, including tensor parallelism for GPU clusters
Beats Claude Opus 4.8 on MCP tool use (81.1% vs 76.4%) — the benchmark that matters for coding agents
Open weights under Modified MIT — download, self-host, fine-tune, use commercially
~30% fewer thinking tokens than K2.6 — direct cost savings on every task
Preserve Thinking mode maintains reasoning chains across multi-turn workflows
5× cheaper per token than Opus ($0.95/$4.00 vs $5/$25)
OpenAI- and Anthropic-compatible APIs — drop into existing agents with a base URL swap
Native INT4 quantization for practical self-hosting
Multimodal input (MoonViT) for design-to-code workflows
256K context — large enough for most real-world codebases
Active open-source community building quantizations and integrations
256K context trails Claude's 1M-token window for very large codebase work
No independently verified third-party benchmarks yet (as of June 15, 2026)
Forced-thinking design can't be disabled — adds cost on trivial calls
Self-hosting requires serious multi-GPU hardware (~600GB full precision)
Modified MIT license requires legal review for large-scale commercial use
Only the coding-specialized variant exists at launch — no general-purpose K2.7 Instruct sibling yet