Ranked #5 Everyday Ecosystem — The Leading AI Assistants
Moonshot AI

Kimi K3

Kimi K3 is the first Kimi release that deserves a place beside the everyday AI platforms, not only on a model chart. It combines near-frontier reasoning with web and mobile chat, autonomous Agent workflows, Docs, Sheets, Slides, Deep Research, Kimi Work, and Kimi Code. The model already does impressive work, although some Agent workflows still come with practical limits.

Updated July 28, 2026 AgenticKnowledge WorkDocs & Sheets
9.6out of 10
Official Website
Best for

Kimi K3 is the first Kimi release that deserves a place beside the everyday AI platforms, not only on a model chart. It combines near-frontier reasoning with web and mobile chat, autonomous Agent workflows, Docs, Sheets, Slides, Deep Research, Kimi Work, and Kimi Code. The model already does impressive work, although some Agent workflows still come with practical limits.

Why It Wins

Artificial Analysis scores K3 at 57 overall, 1668 Elo on GDPval-AA v2, 1547 Elo on AA-Briefcase, and 53% on AutomationBench-AA. The product accepts text and images, offers a 1M model context window, and can produce websites, research reports, Word documents, spreadsheets, presentations, and code. Plans begin with a free tier, followed by paid tiers from $19 per month.

Watch out

Kimi's AI brain is already impressively capable. The catch is in the product details: standard Agent can forget early instructions during a long conversation, usually outputs only one file per task, and uses a smaller practical context than the model's 1M-token maximum. Artificial Analysis also measured a 51% hallucination rate, high token use, and slower-than-average generation. Kimi lacks the close ties to Workspace, Android, Chrome, and Search that keep Gemini above it for everyday use.

01

What It Actually Is

Most AI assistants are built around a simple chat window: you ask a question, and you get an answer. Kimi K3 goes further by combining chat with specialized tools for research, documents, spreadsheets, presentations, websites, coding, and automation. Moonshot did not merely train a larger model; it built a set of products that can turn the model’s reasoning into finished work.

The independent results justify taking those products seriously. Artificial Analysis gives K3 a score of 57 on its Intelligence Index, 1668 Elo on GDPval-AA v2, and 1547 Elo on AA-Briefcase. The last test asks a model to handle long, multi-step professional work, and K3 sits behind only Claude Fable 5. It also reaches 53% and first place on AutomationBench-AA, which tests everyday software tasks. In plain English: K3 is good not only at answering questions, but also at completing several connected steps to produce something useful.

Kimi’s product range is unusually broad. The general Agent can research, create websites, produce documents, analyze spreadsheets, and build presentations. Kimi Work brings those tasks to a desktop app. Kimi Code serves developers, while Agent Swarm divides large research and batch jobs among several agents. Kimi Claw can keep automated tasks running over time. The product names may sound dramatic, but they address a practical need: people often want the completed spreadsheet, presentation, or report—not instructions for making it themselves.

The one-million-token context window and native image input let the model consider a very large collection of source material at once. It can connect information across reports, screenshots, charts, and other documents. A large context window is not perfect memory, and adding more material can still introduce confusion. Even so, the extra capacity is genuinely useful for legal comparisons, research collections, financial analysis, and other projects with many documents.

Kimi is easy to try. It runs on the web and mobile apps, and the free tier includes a small number of Agent tasks. Paid membership begins at $19 per month, with higher tiers adding more Agent tasks, simultaneous jobs, Swarm usage, Kimi Code time, and professional-database calls. The API costs $3 per million input tokens and $15 per million output tokens, with cached input at $0.30. K3 is not a bargain model; it is priced for professional work.

Why, then, only #4? Because this category measures the whole ecosystem, not only the model. Gemini may be weaker on some independent tests, but Google already connects it to email, documents, calendars, browsers, phones, search, and workplace permissions. Kimi offers many tools, yet it is not as closely connected to the services that people already use every day. For someone choosing one assistant, that convenience genuinely matters.

Reliability is the other reason for caution. Artificial Analysis found that K3 became more accurate than K2.6 but also hallucinated more often, with a measured rate of 51%. In other words, the model may answer more questions correctly while still presenting some wrong answers with confidence. It can be extremely productive, but users should check citations, formulas, and important business claims.

But the million-token claim comes with a catch. The K3 model supports that context size, while Kimi’s documentation lists smaller practical limits for some Agent workflows. What the model can hold and what a particular app lets you use are not always the same. Check the limit in the exact tool before planning an enormous document job around it.

Finally, K3 is new. It launched at maximum reasoning effort, which can produce more careful answers but also makes them longer and slower. The full open weights arrived on July 27 under a Modified MIT license — but at 2.8 trillion parameters and roughly 1.4 terabytes in MXFP4, they require datacenter-grade hardware to run. For most users the hosted API and product suite remain the practical way to use K3.

For now, Kimi K3 deserves #4. Choose it when the assignment is larger than a chat: deep research, many documents, an editable report, a spreadsheet, a presentation, a website, or a long-running agent task. Gemini is still easier inside a workflow already built around Google. Kimi is the more adventurous choice: the model is ready for serious work, but you should check the limits of the exact Agent tool before planning a large job around it.

02

Strengths and honest limitations

Key Strengths

  • Independent knowledge-work performance is already near the top: K3 reaches 1668 Elo on GDPval-AA v2 and 1547 Elo on AA-Briefcase, where Artificial Analysis places it behind only Fable 5. This is evidence for reports, analysis, and multi-step professional work—not just academic question answering.
  • Kimi makes the file instead of merely explaining it: Kimi Agent can create and edit Word files, analyze and generate Excel workbooks, build presentations, produce PDFs, deploy websites, and deliver deep-research reports. For a non-technical user, a downloadable file is often more valuable than another paragraph explaining how to make one.
  • Its agents can do real work, not just chat: Artificial Analysis reports 53% and first place on AutomationBench-AA. Kimi’s Agent, Agent Swarm, Work, Code, and Claw tools carry that strength into multi-step web, desktop, coding, and scheduled tasks.
  • Large context and image input help with document-heavy work: The K3 model accepts text and images with a 1M-token context window. It can work across contracts, reports, screenshots, charts, source material, and large document collections without repeatedly losing important context.
  • Free and paid plans cover several levels of use: Kimi works on the web and on iOS, Android, and HarmonyOS. A free plan provides a small number of Agent tasks, while paid membership starts at $19 per month and increases Agent, Swarm, Code, and professional-database allowances.
  • The API price is reasonable for serious work: At $3/$15 per million input/output tokens and $0.30 for cached input, K3 is not the cheapest model. Artificial Analysis estimates about $0.94 per Intelligence Index task—close to GPT-5.6 Sol and far below Fable 5 in that evaluation.

Honest Limitations

  • Gemini still fits more naturally into everyday tools: Gemini is built into Workspace, Android, Chrome, and Search. Kimi has an impressive collection of its own tools, but far fewer people already use them as part of their daily work, so K3 stays below Gemini at #4.
  • Factual reliability remains a concern: Artificial Analysis measured K3’s hallucination rate at 51%, up from 39% for K2.6 even as accuracy improved. Sources, formulas, contract clauses, and business claims require verification.
  • The model limit and the app limit are not always the same: K3 supports a 1M-token context, while Kimi’s Agent documentation still describes smaller practical limits in some workflows. Check the exact Kimi tool before assuming that every million-token document job will fit.
  • Max reasoning is not instant conversation: K3 launched with maximum thinking effort as the default, and Artificial Analysis found it verbose and slower than average. That is appropriate for a report or difficult plan, but excessive for asking when the meeting starts.
  • The open weights are released — but running them is another matter: Moonshot published the full 2.8-trillion-parameter weights on July 27 under a Modified MIT license, confirming that K3 is genuinely self-hostable. But at roughly 1.4 terabytes in MXFP4, production serving needs dozens of H100-class GPUs across multiple nodes. For most users the hosted API remains the practical path; the weights matter most for enterprises and research labs that already operate GPU clusters.
03

Benchmark Snapshot

Artificial Analysis Intelligence Index — 57

Independent near-frontier composite result behind Fable 5 and GPT-5.6 Sol, and ahead of Grok 4.5. Multiple effort settings can change the displayed numerical rank.

GDPval-AA v2 — 1668 Elo

Professional agentic work covering realistic deliverables and decisions. K3 trails Fable 5 and GPT-5.6 Sol but leads Opus 4.8 and Grok 4.5 in the cited comparison.

AA-Briefcase — 1547 Elo (#2)

Private test of long, multi-step knowledge work. K3 is behind only Fable 5 and scores especially well on analytical quality.

AutomationBench-AA — 53% (#1)

Independent implementation of SaaS workflow automation tasks, supporting Kimi's claim that its Agent abilities extend beyond chat.

Intelligence Index cost — about $0.94 per task

Artificial Analysis estimate using measured token consumption and first-party API rates. Real cost depends heavily on reasoning length and cache reuse.

04

The Verdict

Kimi K3 enters the Everyday Ecosystem ranking at #4. The model is easily strong enough: independent tests place it near the frontier, and Kimi has built one of the broadest sets of agent tools outside the three dominant platforms. It remains below Gemini because an everyday ecosystem also depends on where your mail, documents, browser, phone, and colleagues already are. Choose Kimi when you want deep research, large document collections, autonomous work, or finished Office-style files in one service. Check important facts, and note that while the open weights are now downloadable, running them yourself requires datacenter-grade hardware.

05

Frequently Asked Questions