Ranked #6 Local / Private AI — Your Brain, Your Machine, Your Rules
Moonshot AI

Kimi K3

A 2.8-trillion-parameter open-weight frontier model that proves private AI can be enormous and excellent. It also proves that a download can be free while the machine needed to run it is not.

Updated August 29, 2026 Open WeightsModified MIT2.8T MoE
8.5out of 10
Official Website
Best for

A 2.8-trillion-parameter open-weight frontier model that proves private AI can be enormous and excellent. It also proves that a download can be free while the machine needed to run it is not.

Why It Wins

Artificial Analysis scores Kimi K3 around 60 on its current Intelligence Index, tied with flagship GLM-5.3 at the open-weight frontier. It adds native multimodality, a 1M context, and strong Moonshot coding results.

Watch out

The MXFP4 checkpoint is roughly 1.4 TB, production serving can require dozens of accelerators, and the modified MIT license is not identical to standard MIT. More practical 75–306 GB challengers now sit above it in this Local ranking.

01

What It Actually Is

Kimi K3 is what happens when the phrase “run it yourself” is stretched to datacenter scale. The weights are downloadable, the model can remain inside an organization’s boundary, and its intelligence belongs near the frontier. Yet the official compressed checkpoint is roughly 1.4 terabytes. Privacy is possible; simplicity is not.

Artificial Analysis now places Kimi K3 around 60 on its current Intelligence Index, tied with flagship GLM-5.3 in the leading open-weight cluster. Earlier site copy described a score near 80 and a fifth-place rank on a previous scale. That photograph has expired. The new score is not worse intelligence; the ruler changed.

Moonshot’s task-specific story remains impressive. The company reports 88.3 on Terminal-Bench 2.1, 67.5 on DeepSWE, and 81.2 on FrontierSWE. Native image and video understanding and a million-token context turn K3 into more than a text engine. A private system can inspect visual material, large repositories, and long histories without shipping them to a third-party model provider.

Then physics sends the invoice. A mixture-of-experts architecture activates roughly 104 billion parameters per token, but the system still stores and routes the entire 2.8-trillion-parameter pool. Experimental installations begin with many high-memory accelerators; serious production can require dozens. The fact that only some specialists work each case does not let the hospital demolish the other offices.

The competitive staircase has also changed. GLM-5.3-Flash offers a 306 GiB official FP8 checkpoint, MIT rights, and slightly lower independent capability. Qwen3.8-Flash-Next descends toward 75 GB with aggressive quantization and only six billion active parameters, though its license is narrower. Qwen3.8-27B fits around 18 GB at four bits. None matches every Kimi strength, but each is far easier to own.

Kimi K3 therefore remains a milestone without remaining a practical top-two recommendation. Choose it when your organization already operates the required infrastructure and its multimodal profile wins a matched evaluation. For everyone else, the open-weight revolution now offers smaller doors into the same building.

02

Strengths and honest limitations

Key Strengths

  • Raw open-weight capability is exceptional: Current Artificial Analysis testing places Kimi K3 around 60, in the strongest downloadable cluster and close to the best closed models.
  • Multimodality and context are not afterthoughts: Native image and video understanding plus a million-token context suit private document, repository, and long-agent workflows.
  • Moonshot reports serious coding strength: Terminal-Bench 2.1 at 88.3, DeepSWE at 67.5, and FrontierSWE at 81.2 put it in frontier territory under the vendor’s disclosed setups.
  • The checkpoint enables sovereign deployment: Organizations with the infrastructure can keep model operation and sensitive context inside their own boundary instead of relying on a changing hosted endpoint.

Honest Limitations

  • The footprint is measured in terabytes: Roughly 1.4 TB of MXFP4 weights and production recommendations involving large accelerator clusters exclude nearly every individual and small team.
  • Active parameters do not erase storage: About 104B parameters participate per token, but the full 2.8T expert pool must still be stored and routed.
  • Most detailed task scores are first-party: Artificial Analysis validates broad ability; Moonshot’s exact coding and kernel results still depend on its harness and need matched reproduction.
  • The license is modified MIT, not standard MIT: Commercial use is broadly possible, but operators should read the attribution and product-name conditions rather than relying on a shorthand label.
03

Benchmark Snapshot

Artificial Analysis Intelligence Index — about 60

Current independent composite evidence ties Kimi K3 with GLM-5.3 near the open-weight frontier. Earlier site copy using an 80-point scale is obsolete.

Terminal-Bench 2.1 — 88.3 (Moonshot run)

Strong terminal-agent performance, close to other frontier systems in the vendor comparison. Harness and version must be named.

DeepSWE — 67.5 (Moonshot run)

Impressive long-horizon coding evidence that remains setup-dependent and should not be treated as an independent local-quant result.

Checkpoint footprint — roughly 1.4 TB MXFP4

The practical number that determines this ranking: even compressed official weights require datacenter-class storage and accelerator memory.

04

The Verdict

Kimi K3 shifts to #6 in Local / Private AI with an 8.5. It remains a landmark and one of the smartest models an organization can possess, but the category now has more useful stairs: Qwen3.8-27B for one-machine ownership, GLM-5.3-Flash for an MIT multimodal cluster, Qwen3.8-Flash-Next for high-RAM efficiency, DeepSeek for text-and-code serving, and flagship GLM-5.3 for a smaller maximum-capability cluster. Choose Kimi when its particular multimodal and coding profile justifies a very large installation—not because an old article still calls it #2.

05

Frequently Asked Questions