[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"kb-article-moonshot-kimi-k3-inside-the-world-s-largest-open-weight-ai-model-en":3,"ArticleBody_Z5TNxlCPgKejpwtXRJlz7AlDC2OMvev99r8U3MiFg":221},{"article":4,"relatedArticles":192,"locale":66},{"id":5,"title":6,"slug":7,"content":8,"htmlContent":9,"excerpt":10,"category":11,"tags":12,"metaDescription":10,"wordCount":13,"readingTime":14,"publishedAt":15,"sources":16,"sourceCoverage":58,"transparency":60,"seo":63,"language":66,"featuredImage":67,"featuredImageCredit":68,"isFreeGeneration":72,"trendSlug":73,"trendSnapshot":74,"niche":83,"geoTakeaways":87,"geoFaq":96,"entities":106},"6a5af12b08bc9b1b28e40d5c","Moonshot Kimi K3: Inside the World’s Largest Open‑Weight AI Model","moonshot-kimi-k3-inside-the-world-s-largest-open-weight-ai-model","## What Is [Moonshot](\u002Fentities\u002F6a3e7cc4c460e8b42cde2c30-moonshot) [Kimi K3](\u002Fentities\u002F6a5af275b336bdca17d22a28-kimi-k3) and Why It Matters Now\n\nMoonshot’s Kimi K3 is a 2.8‑trillion‑parameter mixture‑of‑experts (MoE) large language model, currently the largest open‑weight AI system publicly announced. [1][6][7]  \nIt is:\n\n- ~75% larger than [DeepSeek V4 Pro](\u002Fentities\u002F6a06ff641f0b27c1f4254b62-deepseek-v4-pro) (~1.6T)  \n- Far bigger than Zhipu’s 744B GLM‑5 series  \n- At the top of the open ecosystem by scale. [4][6]\n\n**Open‑weight** means developers can download, self‑host, and fine‑tune the model weights instead of using only a hosted API. [7] This differs from closed systems like [OpenAI](\u002Fentities\u002F6939892d312dc892c4c1841a-openai)’s GPT‑5.6 or [Anthropic](\u002Fentities\u002F6939b254312dc892c4c1857e-anthropic)’s Fable, where weights stay private. [3][9] For enterprises, this enables:\n\n- On‑prem or sovereign‑cloud deployment  \n- Custom safety, governance, and logging  \n- Deep security and compliance review of the model itself.\n\nHeadline specs: a 2.8T sparse MoE, 1‑million‑token context window, and native multimodal support for text, images, and video in one system. [1][2][4] That supports:\n\n- Full‑repo and multi‑service code analysis  \n- Cross‑document legal or research review  \n- Mixed text–image–video investigations in a single session.\n\nGeopolitically, K3 arrives just after Anthropic’s Fable and [Mythos](\u002Fentities\u002F69dc1d31dc9b12943743b5f6-mythos) were pulled back under US pressure, highlighting how quickly Chinese labs like Moonshot, [Z.ai](\u002Fentities\u002F6a3ae887add847c9a8512ab7-zai), and [MiniMax](https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMinimax) are closing on US frontier models from OpenAI and Anthropic. [3][5][7] Benchmarks show K3:\n\n- Competitive with Fable 5 on several tasks  \n- Ahead of some GPT‑5.x variants on GPU kernels and complex workflows. [3][5][8]\n\n💡 **Key takeaway:** Kimi K3 is a near‑frontier, open‑weight system that raises expectations for openness and performance at trillion‑parameter scale. [1][9]\n\n## Under the Hood: Architecture, Capabilities, and Benchmarks\n\nK3’s 2.8T parameters use a Stable Latent MoE, activating only 16 of 896 experts per token for efficient inference. [1] Kimi Delta Attention (KDA) and Attention Residuals improve long‑range information flow, yielding ~2.5× better scaling efficiency than earlier Kimi models. [1][2]\n\nThe workflow below summarizes how K3 processes large multimodal contexts and turns them into long‑horizon outputs.\n\n```mermaid\nflowchart TB\n    title Kimi K3 Architecture and Long-Context Workflow\n    A[Multimodal input] --> B[Stable Latent MoE]\n    B --> C[Delta Attention]\n    C --> D[Context caching]\n    D --> E[Structured outputs]\n    style A fill:#3b82f6,stroke:#0f172a\n    style B fill:#22c55e,stroke:#14532d\n    style C fill:#f59e0b,stroke:#78350f\n    style D fill:#3b82f6,stroke:#0f172a\n    style E fill:#22c55e,stroke:#14532d\n```\n\nMoonshot optimized K3 for long‑horizon work, not quick chat. Internal results include: [1]\n\n- GPU training kernel latency cut from 283.6 ms to 114.4 ms over 24 hours  \n- MiniTriton compiler built from scratch to near‑Triton performance  \n- A working chip design for its nano model in 48 hours  \n- Reproduction of astrophysics I‑Love‑Q relations in ~2 hours (normally 1–2 weeks).\n\n⚡ **In the wild:** A fintech engineer let K3 run overnight on their entire trading‑risk codebase; it traced a subtle precision bug across dozens of files and produced a patch with regression tests—something their previous LLM stack could not reliably do without heavy orchestration. [2][9]\n\nKey technical features:  \n\n- **1‑million‑token context** and automatic **context caching** for whole monorepos, multi‑day logs, and large legal corpora in one session. [2][3]  \n- Cached segments can drop input cost from about $3.00 to $0.30 per million tokens. [2][4]  \n- Developer‑friendly API with:  \n  - Strict JSON output via `strict: true` in `json_schema`  \n  - Dynamic tool loading via `system` messages  \n  - Partial‑mode prefix control for constrained continuations  \n  - Native vision and video via base64 media. [2][4]\n\nOn external benchmarks, K3 ranks at or near the top on GPU kernel optimization, SWE Marathon, BrowseComp, and Automation Bench, often matching or exceeding Anthropic Fable 5 and GPT‑5.6 Sol in coding and agent tasks. [1][3][5]  \n\nPricing: around $3\u002FMTok input and $15\u002FMTok output, with cheaper cached tokens—aggressive versus US frontier models but costlier than DeepSeek V4 Pro or MiniMax M3. [4][8]\n\n📊 **Data point:** [Artificial Analysis](\u002Fentities\u002F69e69c736db79d4361e1e58a-artificial-analysis) estimates: [8]\n\n- K3: ~$0.94 per composite task  \n- GPT‑5.6 Sol: ~$1.04  \n- Fable 5: ~$2.75  \n- DeepSeek V4 Pro: ~$0.04  \n- MiniMax M3: ~$0.12.\n\n## Strategic Impact: What Kimi K3 Means for Open‑Weight AI\n\nK3’s near‑frontier, open‑weight status lands amid a heated policy debate on model openness. [3][9] US CEOs warn that powerful open‑weight models raise security risks, yet many of the strongest releases now come from Chinese labs those CEOs view as core competitors. [3][9]\n\nEconomically, K3 shows:\n\n- Open‑weight ≠ “cheapest”; it is China’s most capable and among its most expensive models  \n- A shift toward open systems competing on capability, control, and ROI, not just price. [8][10]\n\nHigh‑value use cases include: [1][2]\n\n- Enterprise knowledge management over millions of tokens  \n- Complex codebase refactoring and observability‑driven debugging  \n- Long‑running research and optimization agents (chip design, compilers, scientific workflows)  \n- Multimodal analytics across documents, diagrams, UIs, and video.\n\n⚠️ **Key point:** Risks remain. Access depends on Chinese infrastructure and regulation, and future export or safety controls—after the Fable\u002FMythos precedent—could rapidly alter availability. [3][5] Once weights are out, they can be repurposed for automated vulnerability discovery or influence operations, which concerns security analysts. [7][9]\n\n## Conclusion: A New Line in the Sand for Open‑Weight Frontier Models\n\nKimi K3 is a step change: a 2.8‑trillion‑parameter open‑weight model with 1‑million‑token context, multimodal reasoning, and benchmarks close to leading closed systems from OpenAI and Anthropic. [1][3][6] It narrows the US–China capability gap while keeping full weights—at least in principle—available to researchers, startups, and sovereign clouds outside major US labs. [6][8]\n\nFor technical leaders, the next step is empirical: benchmark K3 against your stack on long‑horizon workflows, strict JSON tooling, overnight agents, and costs with caching. [2][4] Over the next 12–18 months, the July 27 weight release and successors will show how open‑weight frontier systems ultimately compare with closed APIs on capability, control, and price. [2][6][9]","\u003Ch2>What Is \u003Ca href=\"\u002Fentities\u002F6a3e7cc4c460e8b42cde2c30-moonshot\">Moonshot\u003C\u002Fa> \u003Ca href=\"\u002Fentities\u002F6a5af275b336bdca17d22a28-kimi-k3\">Kimi K3\u003C\u002Fa> and Why It Matters Now\u003C\u002Fh2>\n\u003Cp>Moonshot’s Kimi K3 is a 2.8‑trillion‑parameter mixture‑of‑experts (MoE) large language model, currently the largest open‑weight AI system publicly announced. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003Cbr>\nIt is:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>~75% larger than \u003Ca href=\"\u002Fentities\u002F6a06ff641f0b27c1f4254b62-deepseek-v4-pro\">DeepSeek V4 Pro\u003C\u002Fa> (~1.6T)\u003C\u002Fli>\n\u003Cli>Far bigger than Zhipu’s 744B GLM‑5 series\u003C\u002Fli>\n\u003Cli>At the top of the open ecosystem by scale. \u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cstrong>Open‑weight\u003C\u002Fstrong> means developers can download, self‑host, and fine‑tune the model weights instead of using only a hosted API. \u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa> This differs from closed systems like \u003Ca href=\"\u002Fentities\u002F6939892d312dc892c4c1841a-openai\">OpenAI\u003C\u002Fa>’s GPT‑5.6 or \u003Ca href=\"\u002Fentities\u002F6939b254312dc892c4c1857e-anthropic\">Anthropic\u003C\u002Fa>’s Fable, where weights stay private. \u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa> For enterprises, this enables:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>On‑prem or sovereign‑cloud deployment\u003C\u002Fli>\n\u003Cli>Custom safety, governance, and logging\u003C\u002Fli>\n\u003Cli>Deep security and compliance review of the model itself.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Headline specs: a 2.8T sparse MoE, 1‑million‑token context window, and native multimodal support for text, images, and video in one system. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa> That supports:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Full‑repo and multi‑service code analysis\u003C\u002Fli>\n\u003Cli>Cross‑document legal or research review\u003C\u002Fli>\n\u003Cli>Mixed text–image–video investigations in a single session.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Geopolitically, K3 arrives just after Anthropic’s Fable and \u003Ca href=\"\u002Fentities\u002F69dc1d31dc9b12943743b5f6-mythos\">Mythos\u003C\u002Fa> were pulled back under US pressure, highlighting how quickly Chinese labs like Moonshot, \u003Ca href=\"\u002Fentities\u002F6a3ae887add847c9a8512ab7-zai\">Z.ai\u003C\u002Fa>, and \u003Ca href=\"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMinimax\" class=\"wiki-link\" target=\"_blank\" rel=\"noopener\">MiniMax\u003C\u002Fa> are closing on US frontier models from OpenAI and Anthropic. \u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-5\" class=\"citation-link\" title=\"View source [5]\">[5]\u003C\u002Fa>\u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa> Benchmarks show K3:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Competitive with Fable 5 on several tasks\u003C\u002Fli>\n\u003Cli>Ahead of some GPT‑5.x variants on GPU kernels and complex workflows. \u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-5\" class=\"citation-link\" title=\"View source [5]\">[5]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>💡 \u003Cstrong>Key takeaway:\u003C\u002Fstrong> Kimi K3 is a near‑frontier, open‑weight system that raises expectations for openness and performance at trillion‑parameter scale. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa>\u003C\u002Fp>\n\u003Ch2>Under the Hood: Architecture, Capabilities, and Benchmarks\u003C\u002Fh2>\n\u003Cp>K3’s 2.8T parameters use a Stable Latent MoE, activating only 16 of 896 experts per token for efficient inference. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa> Kimi Delta Attention (KDA) and Attention Residuals improve long‑range information flow, yielding ~2.5× better scaling efficiency than earlier Kimi models. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>The workflow below summarizes how K3 processes large multimodal contexts and turns them into long‑horizon outputs.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-mermaid\">flowchart TB\n    title Kimi K3 Architecture and Long-Context Workflow\n    A[Multimodal input] --&gt; B[Stable Latent MoE]\n    B --&gt; C[Delta Attention]\n    C --&gt; D[Context caching]\n    D --&gt; E[Structured outputs]\n    style A fill:#3b82f6,stroke:#0f172a\n    style B fill:#22c55e,stroke:#14532d\n    style C fill:#f59e0b,stroke:#78350f\n    style D fill:#3b82f6,stroke:#0f172a\n    style E fill:#22c55e,stroke:#14532d\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Moonshot optimized K3 for long‑horizon work, not quick chat. Internal results include: \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>GPU training kernel latency cut from 283.6 ms to 114.4 ms over 24 hours\u003C\u002Fli>\n\u003Cli>MiniTriton compiler built from scratch to near‑Triton performance\u003C\u002Fli>\n\u003Cli>A working chip design for its nano model in 48 hours\u003C\u002Fli>\n\u003Cli>Reproduction of astrophysics I‑Love‑Q relations in ~2 hours (normally 1–2 weeks).\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>⚡ \u003Cstrong>In the wild:\u003C\u002Fstrong> A fintech engineer let K3 run overnight on their entire trading‑risk codebase; it traced a subtle precision bug across dozens of files and produced a patch with regression tests—something their previous LLM stack could not reliably do without heavy orchestration. \u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Key technical features:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>1‑million‑token context\u003C\u002Fstrong> and automatic \u003Cstrong>context caching\u003C\u002Fstrong> for whole monorepos, multi‑day logs, and large legal corpora in one session. \u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>Cached segments can drop input cost from about $3.00 to $0.30 per million tokens. \u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>Developer‑friendly API with:\n\u003Cul>\n\u003Cli>Strict JSON output via \u003Ccode>strict: true\u003C\u002Fcode> in \u003Ccode>json_schema\u003C\u002Fcode>\u003C\u002Fli>\n\u003Cli>Dynamic tool loading via \u003Ccode>system\u003C\u002Fcode> messages\u003C\u002Fli>\n\u003Cli>Partial‑mode prefix control for constrained continuations\u003C\u002Fli>\n\u003Cli>Native vision and video via base64 media. \u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>On external benchmarks, K3 ranks at or near the top on GPU kernel optimization, SWE Marathon, BrowseComp, and Automation Bench, often matching or exceeding Anthropic Fable 5 and GPT‑5.6 Sol in coding and agent tasks. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-5\" class=\"citation-link\" title=\"View source [5]\">[5]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Pricing: around $3\u002FMTok input and $15\u002FMTok output, with cheaper cached tokens—aggressive versus US frontier models but costlier than DeepSeek V4 Pro or MiniMax M3. \u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>📊 \u003Cstrong>Data point:\u003C\u002Fstrong> \u003Ca href=\"\u002Fentities\u002F69e69c736db79d4361e1e58a-artificial-analysis\">Artificial Analysis\u003C\u002Fa> estimates: \u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>K3: ~$0.94 per composite task\u003C\u002Fli>\n\u003Cli>GPT‑5.6 Sol: ~$1.04\u003C\u002Fli>\n\u003Cli>Fable 5: ~$2.75\u003C\u002Fli>\n\u003Cli>DeepSeek V4 Pro: ~$0.04\u003C\u002Fli>\n\u003Cli>MiniMax M3: ~$0.12.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Strategic Impact: What Kimi K3 Means for Open‑Weight AI\u003C\u002Fh2>\n\u003Cp>K3’s near‑frontier, open‑weight status lands amid a heated policy debate on model openness. \u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa> US CEOs warn that powerful open‑weight models raise security risks, yet many of the strongest releases now come from Chinese labs those CEOs view as core competitors. \u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Economically, K3 shows:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Open‑weight ≠ “cheapest”; it is China’s most capable and among its most expensive models\u003C\u002Fli>\n\u003Cli>A shift toward open systems competing on capability, control, and ROI, not just price. \u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003Ca href=\"#source-10\" class=\"citation-link\" title=\"View source [10]\">[10]\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>High‑value use cases include: \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Enterprise knowledge management over millions of tokens\u003C\u002Fli>\n\u003Cli>Complex codebase refactoring and observability‑driven debugging\u003C\u002Fli>\n\u003Cli>Long‑running research and optimization agents (chip design, compilers, scientific workflows)\u003C\u002Fli>\n\u003Cli>Multimodal analytics across documents, diagrams, UIs, and video.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>⚠️ \u003Cstrong>Key point:\u003C\u002Fstrong> Risks remain. Access depends on Chinese infrastructure and regulation, and future export or safety controls—after the Fable\u002FMythos precedent—could rapidly alter availability. \u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-5\" class=\"citation-link\" title=\"View source [5]\">[5]\u003C\u002Fa> Once weights are out, they can be repurposed for automated vulnerability discovery or influence operations, which concerns security analysts. \u003Ca href=\"#source-7\" class=\"citation-link\" title=\"View source [7]\">[7]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa>\u003C\u002Fp>\n\u003Ch2>Conclusion: A New Line in the Sand for Open‑Weight Frontier Models\u003C\u002Fh2>\n\u003Cp>Kimi K3 is a step change: a 2.8‑trillion‑parameter open‑weight model with 1‑million‑token context, multimodal reasoning, and benchmarks close to leading closed systems from OpenAI and Anthropic. \u003Ca href=\"#source-1\" class=\"citation-link\" title=\"View source [1]\">[1]\u003C\u002Fa>\u003Ca href=\"#source-3\" class=\"citation-link\" title=\"View source [3]\">[3]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa> It narrows the US–China capability gap while keeping full weights—at least in principle—available to researchers, startups, and sovereign clouds outside major US labs. \u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-8\" class=\"citation-link\" title=\"View source [8]\">[8]\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>For technical leaders, the next step is empirical: benchmark K3 against your stack on long‑horizon workflows, strict JSON tooling, overnight agents, and costs with caching. \u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-4\" class=\"citation-link\" title=\"View source [4]\">[4]\u003C\u002Fa> Over the next 12–18 months, the July 27 weight release and successors will show how open‑weight frontier systems ultimately compare with closed APIs on capability, control, and price. \u003Ca href=\"#source-2\" class=\"citation-link\" title=\"View source [2]\">[2]\u003C\u002Fa>\u003Ca href=\"#source-6\" class=\"citation-link\" title=\"View source [6]\">[6]\u003C\u002Fa>\u003Ca href=\"#source-9\" class=\"citation-link\" title=\"View source [9]\">[9]\u003C\u002Fa>\u003C\u002Fp>\n","What Is Moonshot Kimi K3 and Why It Matters Now\n\nMoonshot’s Kimi K3 is a 2.8‑trillion‑parameter mixture‑of‑experts (MoE) large language model, currently the largest open‑weight AI system publicly anno...","trend-radar",[],915,5,"2026-07-18T03:29:24.109Z",[17,22,26,30,34,38,42,46,50,54],{"title":18,"url":19,"summary":20,"type":21},"Moonshot AI: Introducing Kimi K3 the best and largest OpenWeight model","https:\u002F\u002Fwww.linkedin.com\u002Fposts\u002Fraphaelmansuy_moonshot-ai-introducing-kimi-k3-the-best-activity-7483716125235564544-PiLh","Moonshot AI: Introducing Kimi K3 the best and largest OpenWeight model .... What does it take to convert raw compute into usable intelligence at trillion-parameter scale? Moonshot AI just answered tha...","kb",{"title":23,"url":24,"summary":25,"type":21},"Moonshot AI just dropped Kimi K3: An open weight frontier multimodal AI model (2.8T params) MoE with a 1M context window","https:\u002F\u002Fwww.reddit.com\u002Fr\u002FAIDeveloperNews\u002Fcomments\u002F1uyhbpm\u002Fmoonshot_ai_just_dropped_kimi_k3_an_open_weight\u002F","Moonshot AI has just launched Kimi K3. It is a 2.8 trillion-parameter Mixture of Experts (MoE) model built on a new Kimi Delta Attention (KDA) architecture. The Kimi API Platform is live right now, an...",{"title":27,"url":28,"summary":29,"type":21},"China's Moonshot unveils world's 'largest' open AI model, Kimi K3, closing in on US rivals","https:\u002F\u002Fwww.dawn.com\u002Fnews\u002F2016173","Chinese AI startup Moonshot on Friday unveiled Kimi K3, a 2.8 trillion-parameter model that it said is the world’s largest open-weight AI system and delivers performance approaching US giant Anthropic...",{"title":31,"url":32,"summary":33,"type":21},"Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3","https:\u002F\u002Fventurebeat.com\u002Ftechnology\u002Fchinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems","Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 — a 2.8-trillion-parameter model that the company says is now the largest open-source AI ...",{"title":35,"url":36,"summary":37,"type":21},"China’s Startup Moonshot Unveils World’s Largest Open AI Model","https:\u002F\u002Fstratnewsglobal.com\u002Ftrade-tech\u002Fchinas-startup-moonshot-unveils-worlds-largest-open-ai-model\u002F","China’s AI startup Moonshot on Friday unveiled Kimi K3, a 2.8 trillion-parameter model that it said is the world’s largest open-weight AI system and delivers performance approaching U.S. giant Anthrop...",{"title":39,"url":40,"summary":41,"type":21},"Moonshot AI unveils world’s largest open-source AI model as China narrows gap with US rivals","https:\u002F\u002Fwww.scmp.com\u002Ftech\u002Ftech-trends\u002Farticle\u002F3360885\u002Fmoonshot-ai-unveils-worlds-largest-open-source-ai-model-china-narrows-gap-us-rivals","Ben Jiang in Beijing and Minxiao Chang in Shenzhen\nPublished: 12:31pm, 17 Jul 2026 Updated: 2:26pm, 17 Jul 2026\n\nChinese start-up Moonshot AI has launched the world’s largest open-source artificial in...",{"title":43,"url":44,"summary":45,"type":21},"China's Moonshot unveils world's largest open-weight AI model, closing gap with US rivals","https:\u002F\u002Fwww.reuters.com\u002Fworld\u002Fchina\u002Fchinas-moonshot-unveils-worlds-largest-open-ai-model-closing-us-rivals-2026-07-17\u002F","BEIJING, July 17 (Reuters) - Chinese AI startup Moonshot on Friday unveiled Kimi K3, a 2.8 trillion-parameter model that it said is the world's largest open-weight AI system and delivers performance a...",{"title":47,"url":48,"summary":49,"type":21},"Moonshot unveils Kimi K3, largest open-weight AI model yet","https:\u002F\u002Fwww.siliconrepublic.com\u002Fmachines\u002Fmoonshot-unveils-kimi-k3-largest-open-weight-ai-model-yet","Moonshot announced a $2bn raise at a $20bn valuation in May.\n\nChina is showcasing its AI prowess despite restrictive policy efforts from Washington, with Moonshot AI’s new Kimi K3 boasting a performan...",{"title":51,"url":52,"summary":53,"type":21},"China's open-weight Kimi model stuns AI world with frontier-level results","https:\u002F\u002Fwww.axios.com\u002F2026\u002F07\u002F16\u002Fmoonshot-kimi-ai-china-model-openai-anthropic","Chinese AI startup Moonshot AI stunned developers on Thursday with a massive new model that may rival the best American systems at a fraction of the cost.\n\n Why it matters: Kimi K3's early performance...",{"title":55,"url":56,"summary":57,"type":21},"Kimi K3: China’s Most Capable and Most Expensive AI Model Yet","https:\u002F\u002Fwww.youtube.com\u002Fwatch?v=K2lcv0W-To8","Kimi K3 has just released, Moonshot's latest AI model. This is the first time that a Chinese labs AI model is extremely close to the benchmarks of the US frontier AI models like OpenAI's GPT 5.5 or An...",{"totalSources":59},10,{"generationDuration":61,"kbQueriesCount":59,"confidenceScore":62,"sourcesCount":59},249653,100,{"metaTitle":64,"metaDescription":65},"Moonshot Kimi K3 Open-Weight AI: 2.8T MoE Explained","Discover Moonshot Kimi K3, a 2.8T open-weight MoE with 1M token context and multimodal power. See why self-hosting, security and geopolitics will shift.","en","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1694432739753-ea35e4c26054?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxtb29uc2hvdCUyMGtpbWklMjB3b3JsZCUyMGxhcmdlc3R8ZW58MXwwfHx8MTc4NDM0NDg3NXww&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60",{"photographerName":69,"photographerUrl":70,"unsplashUrl":71},"Paul Harris","https:\u002F\u002Funsplash.com\u002F@paul_harris_coaching?utm_source=coreprose&utm_medium=referral","https:\u002F\u002Funsplash.com\u002Fphotos\u002Fa-close-up-of-a-helmet-on-a-motorcycle-IBIcqD_xZXQ?utm_source=coreprose&utm_medium=referral",true,"moonshot-kimi-k3-world-s-largest-open-weight-ai-model",{"score":62,"type":75,"sourceCount":76,"topSourceDomains":77,"detectedAt":81,"mentionsLast7Days":82},"spiking",67,[78,79,80],"unknown","fortune.com","forbes.com","2026-07-18T00:19:19.743Z",12,{"key":84,"name":85,"nameEn":86},"ia","Intelligence Artificielle","Artificial Intelligence",[88,90,92,94],{"text":89},"Moonshot Kimi K3 is a 2.8‑trillion‑parameter sparse mixture‑of‑experts (MoE) model with a 1‑million‑token context window and native multimodal (text, image, video) support.",{"text":91},"K3 activates 16 of 896 experts per token, delivers ~2.5× better scaling efficiency than prior Kimi models, and cuts GPU kernel latency from 283.6 ms to 114.4 ms in internal runs.",{"text":93},"K3 is open‑weight: developers can download, self‑host, and fine‑tune the full weights, enabling on‑prem\u002Fsovereign‑cloud deployment, custom safety and governance, and deep compliance reviews.",{"text":95},"Pricing targets roughly $3 per million input tokens and $15 per million output tokens (with cached segments reducing input cost to ~$0.30\u002FMTok), and estimated composite task cost of ~$0.94 versus $1.04 for GPT‑5.6 Sol.",[97,100,103],{"question":98,"answer":99},"What technical features make Kimi K3 stand out?","Kimi K3 is a 2.8T sparse MoE with a 1‑million‑token context and native multimodal reasoning. It uses a Stable Latent MoE that activates 16 of 896 experts per token, Kimi Delta Attention and Attention Residuals to improve long‑range flow, and a custom MiniTriton compiler that reduced kernel latency from 283.6 ms to 114.4 ms; these combine to prioritize long‑horizon workflows (whole monorepos, multi‑day logs, mixed text–image–video sessions). K3 also supports context caching to dramatically lower input costs (from about $3.00 to $0.30 per million tokens for cached segments) and provides developer controls like strict JSON output, dynamic tool loading, partial‑mode prefix control, and base64 media handling for vision and video.",{"question":101,"answer":102},"How can enterprises deploy and use K3 while managing security and compliance?","Enterprises can download and self‑host K3 weights on‑premises or in sovereign clouds to retain full control over data residency, logging, and governance; this open‑weight access allows organizations to run deep security audits, integrate corporate toolchains, and implement bespoke safety layers. Typical high‑value uses are long‑horizon codebase analysis and automated patch generation, multimodal legal and research review across millions of tokens, and running overnight optimization agents for chip design or scientific workflows; however, deployment requires substantial infrastructure, engineering to manage sparse MoE execution, and processes to mitigate misuse and vulnerability discovery.",{"question":104,"answer":105},"What are the main risks and geopolitical implications of K3’s release?","K3’s open‑weight availability significantly increases access to near‑frontier capabilities outside closed US labs, narrowing the capability gap and raising concerns about proliferation and dual‑use applications. Risks include potential repurposing for automated vulnerability discovery or influence operations, dependency on Chinese infrastructure and regulatory regimes, and the possibility of rapid export or safety controls (as seen with Fable\u002FMythos) that could restrict distribution; organizations must therefore balance capability advantages against operational, legal, and geopolitical risk when adopting an open‑weight frontier model.",[107,115,121,126,130,136,141,149,155,162,169,174,180,187],{"id":108,"name":109,"type":110,"confidence":111,"wikipediaUrl":112,"slug":113,"mentionCount":114},"6a5af341b336bdca17d22a89","Kimi Delta Attention (KDA)","concept",0.95,null,"6a5af341b336bdca17d22a89-kimi-delta-attention-kda",3,{"id":116,"name":117,"type":110,"confidence":118,"wikipediaUrl":112,"slug":119,"mentionCount":120},"6a5af341b336bdca17d22a88","Stable Latent MoE",0.93,"6a5af341b336bdca17d22a88-stable-latent-moe",1,{"id":122,"name":123,"type":110,"confidence":124,"wikipediaUrl":112,"slug":125,"mentionCount":120},"6a5af341b336bdca17d22a8b","SWE Marathon",0.78,"6a5af341b336bdca17d22a8b-swe-marathon",{"id":127,"name":128,"type":110,"confidence":124,"wikipediaUrl":112,"slug":129,"mentionCount":120},"6a5af341b336bdca17d22a8c","Automation Bench","6a5af341b336bdca17d22a8c-automation-bench",{"id":131,"name":132,"type":133,"confidence":111,"wikipediaUrl":112,"slug":134,"mentionCount":135},"69967e6c9aa9beba177c4cc4","BrowseComp","event","69967e6c9aa9beba177c4cc4-browsecomp",7,{"id":137,"name":138,"type":133,"confidence":139,"wikipediaUrl":112,"slug":140,"mentionCount":120},"6a5af342b336bdca17d22a8d","July 27 weight release",0.72,"6a5af342b336bdca17d22a8d-july-27-weight-release",{"id":142,"name":143,"type":144,"confidence":145,"wikipediaUrl":146,"slug":147,"mentionCount":148},"6939892d312dc892c4c1841a","OpenAI","organization",0.99,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenAI","6939892d312dc892c4c1841a-openai",899,{"id":150,"name":151,"type":144,"confidence":145,"wikipediaUrl":152,"slug":153,"mentionCount":154},"6939b254312dc892c4c1857e","Anthropic","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAnthropic","6939b254312dc892c4c1857e-anthropic",586,{"id":156,"name":157,"type":144,"confidence":158,"wikipediaUrl":159,"slug":160,"mentionCount":161},"69e69c736db79d4361e1e58a","Artificial Analysis",0.94,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FArtificial_intelligence","69e69c736db79d4361e1e58a-artificial-analysis",14,{"id":163,"name":164,"type":144,"confidence":165,"wikipediaUrl":166,"slug":167,"mentionCount":168},"6a3e7cc4c460e8b42cde2c30","Moonshot",0.98,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMoonshot_AI","6a3e7cc4c460e8b42cde2c30-moonshot",6,{"id":170,"name":171,"type":144,"confidence":165,"wikipediaUrl":172,"slug":173,"mentionCount":14},"6a3ae887add847c9a8512ab7","Z.ai","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FZ.ai","6a3ae887add847c9a8512ab7-zai",{"id":175,"name":176,"type":144,"confidence":177,"wikipediaUrl":178,"slug":179,"mentionCount":14},"6a3e7cc4c460e8b42cde2c31","MiniMax",0.9,"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMinimax","6a3e7cc4c460e8b42cde2c31-minimax",{"id":181,"name":182,"type":183,"confidence":165,"wikipediaUrl":184,"slug":185,"mentionCount":186},"69dc1d31dc9b12943743b5f6","Mythos","product","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FCthulhu_Mythos","69dc1d31dc9b12943743b5f6-mythos",52,{"id":188,"name":189,"type":183,"confidence":145,"wikipediaUrl":112,"slug":190,"mentionCount":191},"6a2cec93add847c9a84ed716","Fable 5","6a2cec93add847c9a84ed716-fable-5",51,[193,200,207,214],{"id":194,"title":195,"slug":196,"excerpt":197,"category":11,"featuredImage":198,"publishedAt":199},"6a5fc2ac366a05b9f721dbc4","Hugging Face Breached by an Autonomous AI Agent: What Happened and How to Respond","hugging-face-breached-by-an-autonomous-ai-agent-what-happened-and-how-to-respond","Hugging Face is the de facto hub for open-source machine learning, hosting over 45,000 models used by more than 50,000 organizations worldwide. [4] A compromise there is not just another vendor incide...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1499568509606-4f9b771232ed?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxodWdnaW5nJTIwZmFjZSUyMGJyZWFjaGVkJTIwYXV0b25vbW91c3xlbnwxfDB8fHwxNzg0NjYwNjUyfDA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-21T19:14:39.313Z",{"id":201,"title":202,"slug":203,"excerpt":204,"category":11,"featuredImage":205,"publishedAt":206},"6a5f0668366a05b9f721d5ed","Moonshot’s 2.8 Trillion-Parameter Kimi K3 Redraws the Open-Weight Frontier","moonshot-s-2-8-trillion-parameter-kimi-k3-redraws-the-open-weight-frontier","Moonshot’s Kimi K3 brings “near‑frontier” performance into a space enterprises can inspect, customize, and self‑host instead of renting via opaque APIs.[1][3] For technical and business leaders, this...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1459909633680-206dc5c67abb?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxtb29uc2hvdCUyMHVudmVpbHMlMjB0cmlsbGlvbiUyMHBhcmFtZXRlcnxlbnwxfDB8fHwxNzg0NjEyNDU2fDA&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-21T05:49:02.822Z",{"id":208,"title":209,"slug":210,"excerpt":211,"category":11,"featuredImage":212,"publishedAt":213},"6a5ef1854ead64f9f4e786c4","Chinese AI Model Kimi K3 Is Closing the Gap With Claude and ChatGPT","chinese-ai-model-kimi-k3-is-closing-the-gap-with-claude-and-chatgpt","For developers, CTOs, and policy teams, Kimi K3 is one of the first Chinese open‑weight LLMs that can seriously compete with the strongest versions of Claude and ChatGPT on coding and reasoning — not...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1607348595533-2eb150a869e3?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxfHxjaGluZXNlJTIwbW9kZWx8ZW58MXwwfHx8MTc4NDYwNzEwOXww&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-21T04:19:22.631Z",{"id":215,"title":216,"slug":217,"excerpt":218,"category":11,"featuredImage":219,"publishedAt":220},"6a5ebfac4ead64f9f4e783c9","How Moonshot AI’s Kimi K3 Surpassed US Frontier Models on Key Benchmarks","how-moonshot-ai-s-kimi-k3-surpassed-us-frontier-models-on-key-benchmarks","Moonshot AI’s Kimi K3 has turned what was a one‑sided US narrative on frontier models into a real contest, especially in coding and GPU efficiency.[1][6] For technical leaders, it shows that Chinese o...","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1739036868260-c26b292cd85d?ixid=M3w4OTczNDl8MHwxfHNlYXJjaHwxNnx8YXJ0aWZpY2lhbCUyMGludGVsbGlnZW5jZSUyMHRlY2hub2xvZ3l8ZW58MXwwfHx8MTc4NDU5NDM0OHww&ixlib=rb-4.1.0&w=1200&h=630&fit=crop&crop=entropy&auto=format,compress&q=60","2026-07-21T00:47:35.487Z",["Island",222],{"key":223,"params":224,"result":226},"ArticleBody_Z5TNxlCPgKejpwtXRJlz7AlDC2OMvev99r8U3MiFg",{"props":225},"{\"articleId\":\"6a5af12b08bc9b1b28e40d5c\",\"linkColor\":\"red\"}",{"head":227},{}]