Key Takeaways
- Moonshot Kimi K3 is a 2.8‑trillion‑parameter sparse mixture‑of‑experts (MoE) model with a 1‑million‑token context window and native multimodal (text, image, video) support.
- K3 activates 16 of 896 experts per token, delivers ~2.5× better scaling efficiency than prior Kimi models, and cuts GPU kernel latency from 283.6 ms to 114.4 ms in internal runs.
- K3 is open‑weight: developers can download, self‑host, and fine‑tune the full weights, enabling on‑prem/sovereign‑cloud deployment, custom safety and governance, and deep compliance reviews.
- Pricing targets roughly $3 per million input tokens and $15 per million output tokens (with cached segments reducing input cost to ~$0.30/MTok), and estimated composite task cost of ~$0.94 versus $1.04 for GPT‑5.6 Sol.
What Is Moonshot Kimi K3 and Why It Matters Now
Moonshot’s Kimi K3 is a 2.8‑trillion‑parameter mixture‑of‑experts (MoE) large language model, currently the largest open‑weight AI system publicly announced. [1][6][7]
It is:
- ~75% larger than DeepSeek V4 Pro (~1.6T)
- Far bigger than Zhipu’s 744B GLM‑5 series
- At the top of the open ecosystem by scale. [4][6]
Open‑weight means developers can download, self‑host, and fine‑tune the model weights instead of using only a hosted API. [7] This differs from closed systems like OpenAI’s GPT‑5.6 or Anthropic’s Fable, where weights stay private. [3][9] For enterprises, this enables:
- On‑prem or sovereign‑cloud deployment
- Custom safety, governance, and logging
- Deep security and compliance review of the model itself.
Headline specs: a 2.8T sparse MoE, 1‑million‑token context window, and native multimodal support for text, images, and video in one system. [1][2][4] That supports:
- Full‑repo and multi‑service code analysis
- Cross‑document legal or research review
- Mixed text–image–video investigations in a single session.
Geopolitically, K3 arrives just after Anthropic’s Fable and Mythos were pulled back under US pressure, highlighting how quickly Chinese labs like Moonshot, Z.ai, and MiniMax are closing on US frontier models from OpenAI and Anthropic. [3][5][7] Benchmarks show K3:
- Competitive with Fable 5 on several tasks
- Ahead of some GPT‑5.x variants on GPU kernels and complex workflows. [3][5][8]
💡 Key takeaway: Kimi K3 is a near‑frontier, open‑weight system that raises expectations for openness and performance at trillion‑parameter scale. [1][9]
Under the Hood: Architecture, Capabilities, and Benchmarks
K3’s 2.8T parameters use a Stable Latent MoE, activating only 16 of 896 experts per token for efficient inference. [1] Kimi Delta Attention (KDA) and Attention Residuals improve long‑range information flow, yielding ~2.5× better scaling efficiency than earlier Kimi models. [1][2]
The workflow below summarizes how K3 processes large multimodal contexts and turns them into long‑horizon outputs.
flowchart TB
title Kimi K3 Architecture and Long-Context Workflow
A[Multimodal input] --> B[Stable Latent MoE]
B --> C[Delta Attention]
C --> D[Context caching]
D --> E[Structured outputs]
style A fill:#3b82f6,stroke:#0f172a
style B fill:#22c55e,stroke:#14532d
style C fill:#f59e0b,stroke:#78350f
style D fill:#3b82f6,stroke:#0f172a
style E fill:#22c55e,stroke:#14532d
Moonshot optimized K3 for long‑horizon work, not quick chat. Internal results include: [1]
- GPU training kernel latency cut from 283.6 ms to 114.4 ms over 24 hours
- MiniTriton compiler built from scratch to near‑Triton performance
- A working chip design for its nano model in 48 hours
- Reproduction of astrophysics I‑Love‑Q relations in ~2 hours (normally 1–2 weeks).
⚡ In the wild: A fintech engineer let K3 run overnight on their entire trading‑risk codebase; it traced a subtle precision bug across dozens of files and produced a patch with regression tests—something their previous LLM stack could not reliably do without heavy orchestration. [2][9]
Key technical features:
- 1‑million‑token context and automatic context caching for whole monorepos, multi‑day logs, and large legal corpora in one session. [2][3]
- Cached segments can drop input cost from about $3.00 to $0.30 per million tokens. [2][4]
- Developer‑friendly API with:
On external benchmarks, K3 ranks at or near the top on GPU kernel optimization, SWE Marathon, BrowseComp, and Automation Bench, often matching or exceeding Anthropic Fable 5 and GPT‑5.6 Sol in coding and agent tasks. [1][3][5]
Pricing: around $3/MTok input and $15/MTok output, with cheaper cached tokens—aggressive versus US frontier models but costlier than DeepSeek V4 Pro or MiniMax M3. [4][8]
📊 Data point: Artificial Analysis estimates: [8]
- K3: ~$0.94 per composite task
- GPT‑5.6 Sol: ~$1.04
- Fable 5: ~$2.75
- DeepSeek V4 Pro: ~$0.04
- MiniMax M3: ~$0.12.
Strategic Impact: What Kimi K3 Means for Open‑Weight AI
K3’s near‑frontier, open‑weight status lands amid a heated policy debate on model openness. [3][9] US CEOs warn that powerful open‑weight models raise security risks, yet many of the strongest releases now come from Chinese labs those CEOs view as core competitors. [3][9]
Economically, K3 shows:
- Open‑weight ≠ “cheapest”; it is China’s most capable and among its most expensive models
- A shift toward open systems competing on capability, control, and ROI, not just price. [8][10]
High‑value use cases include: [1][2]
- Enterprise knowledge management over millions of tokens
- Complex codebase refactoring and observability‑driven debugging
- Long‑running research and optimization agents (chip design, compilers, scientific workflows)
- Multimodal analytics across documents, diagrams, UIs, and video.
⚠️ Key point: Risks remain. Access depends on Chinese infrastructure and regulation, and future export or safety controls—after the Fable/Mythos precedent—could rapidly alter availability. [3][5] Once weights are out, they can be repurposed for automated vulnerability discovery or influence operations, which concerns security analysts. [7][9]
Conclusion: A New Line in the Sand for Open‑Weight Frontier Models
Kimi K3 is a step change: a 2.8‑trillion‑parameter open‑weight model with 1‑million‑token context, multimodal reasoning, and benchmarks close to leading closed systems from OpenAI and Anthropic. [1][3][6] It narrows the US–China capability gap while keeping full weights—at least in principle—available to researchers, startups, and sovereign clouds outside major US labs. [6][8]
For technical leaders, the next step is empirical: benchmark K3 against your stack on long‑horizon workflows, strict JSON tooling, overnight agents, and costs with caching. [2][4] Over the next 12–18 months, the July 27 weight release and successors will show how open‑weight frontier systems ultimately compare with closed APIs on capability, control, and price. [2][6][9]
Frequently Asked Questions
What technical features make Kimi K3 stand out?
How can enterprises deploy and use K3 while managing security and compliance?
What are the main risks and geopolitical implications of K3’s release?
Sources & References (10)
- 1Moonshot AI: Introducing Kimi K3 the best and largest OpenWeight model
Moonshot AI: Introducing Kimi K3 the best and largest OpenWeight model .... What does it take to convert raw compute into usable intelligence at trillion-parameter scale? Moonshot AI just answered tha...
- 2Moonshot AI just dropped Kimi K3: An open weight frontier multimodal AI model (2.8T params) MoE with a 1M context window
Moonshot AI has just launched Kimi K3. It is a 2.8 trillion-parameter Mixture of Experts (MoE) model built on a new Kimi Delta Attention (KDA) architecture. The Kimi API Platform is live right now, an...
- 3China's Moonshot unveils world's 'largest' open AI model, Kimi K3, closing in on US rivals
Chinese AI startup Moonshot on Friday unveiled Kimi K3, a 2.8 trillion-parameter model that it said is the world’s largest open-weight AI system and delivers performance approaching US giant Anthropic...
- 4Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3
Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 — a 2.8-trillion-parameter model that the company says is now the largest open-source AI ...
- 5China’s Startup Moonshot Unveils World’s Largest Open AI Model
China’s AI startup Moonshot on Friday unveiled Kimi K3, a 2.8 trillion-parameter model that it said is the world’s largest open-weight AI system and delivers performance approaching U.S. giant Anthrop...
- 6Moonshot AI unveils world’s largest open-source AI model as China narrows gap with US rivals
Ben Jiang in Beijing and Minxiao Chang in Shenzhen Published: 12:31pm, 17 Jul 2026 Updated: 2:26pm, 17 Jul 2026 Chinese start-up Moonshot AI has launched the world’s largest open-source artificial in...
- 7China's Moonshot unveils world's largest open-weight AI model, closing gap with US rivals
BEIJING, July 17 (Reuters) - Chinese AI startup Moonshot on Friday unveiled Kimi K3, a 2.8 trillion-parameter model that it said is the world's largest open-weight AI system and delivers performance a...
- 8Moonshot unveils Kimi K3, largest open-weight AI model yet
Moonshot announced a $2bn raise at a $20bn valuation in May. China is showcasing its AI prowess despite restrictive policy efforts from Washington, with Moonshot AI’s new Kimi K3 boasting a performan...
- 9China's open-weight Kimi model stuns AI world with frontier-level results
Chinese AI startup Moonshot AI stunned developers on Thursday with a massive new model that may rival the best American systems at a fraction of the cost. Why it matters: Kimi K3's early performance...
- 10Kimi K3: China’s Most Capable and Most Expensive AI Model Yet
Kimi K3 has just released, Moonshot's latest AI model. This is the first time that a Chinese labs AI model is extremely close to the benchmarks of the US frontier AI models like OpenAI's GPT 5.5 or An...
Key Entities
Generated by CoreProse in 4m 9s
What topic do you want to cover?
Get the same quality with verified sources on any subject.