Ranked #5 Coding — AI That Writes Production Code
xAI

Grok 4.5

Grok 4.5 takes #4 for coding because it makes frontier-class agent loops economically normal. Kimi K3 now moves above it on raw independent intelligence and frontend preference, but Grok Build still ranks third on Artificial Analysis's Coding Agent Index, matches GPT-5.5's Codex result there, and works at a fraction of the per-task cost.

Updated July 10, 2026 Agentic CodingCursorGrok Build
9.7out of 10
Official Website
Best for

Grok 4.5 takes #4 for coding because it makes frontier-class agent loops economically normal. Kimi K3 now moves above it on raw independent intelligence and frontend preference, but Grok Build still ranks third on Artificial Analysis's Coding Agent Index, matches GPT-5.5's Codex result there, and works at a fraction of the per-task cost.

Why It Wins

Artificial Analysis scores Grok Build at 76 on its Coding Agent Index—third, on par with GPT-5.5 in Codex—and reports $2.49 per task versus $5.07 for GPT-5.5 and $11.80 for Fable 5. xAI reports Terminal-Bench 2.1 at 83.3%, SWE Marathon at 29.0%, 80 TPS, $2/$6 pricing, and 4.2× fewer output tokens per SWE-Bench Pro task than Opus 4.8 max. Grok 4.5 is available in Cursor on all plans, Grok Build, and the API.

Watch out

The #4 position rewards value and efficiency, not a claim that Grok beats every rival. Its published SWE-Bench Pro 64.7% trails Fable 5 and Opus 4.8; DeepSWE also trails the leaders. Kimi K3 now has the stronger raw-capability case, while cheap code still deserves tests, review, and security checks.

01

What It Actually Is

There are two ways to buy a coding agent. You can rent the smartest contractor in the city for every ticket, or you can hire a very good engineer who is fast enough and cheap enough to take another pass when the first one misses. Grok 4.5 is the second option, and that is not faint praise.

Artificial Analysis puts Grok Build third on its Coding Agent Index, at 76: level with GPT-5.5 in Codex, below Fable 5 in Claude Code. Then comes the number that changes the buying decision. The same analysis estimates $2.49 per agent task for Grok Build, versus $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code. Grok used far fewer tokens getting there.

That is why Grok 4.5 remains a strong #4 after Kimi K3 arrives. xAI’s own chart has it almost tied with the leaders on Terminal-Bench 2.1 at 83.3%, and leading the cited figures on SWE Marathon at 29.0% pass@1. It is also available in Cursor on every plan and is the default in Grok Build. Long agent loops, terminal work, multi-step debugging, repository exploration—these are suddenly cheap enough to repeat.

What not to claim

Grok 4.5 is not the new raw-score monarch. Its published 64.7% on SWE-Bench Pro falls behind Fable 5’s 80.4% and Opus 4.8’s 69.2%. Its DeepSWE results also trail Fable and GPT-5.5 in xAI’s comparison. If your workflow is a pure benchmark for resolving repository issues, the models above it still have the sharper resume.

What to do instead

Start Grok 4.5 on the work where a capable agent gets better by taking another loop: investigation, terminal tasks, refactors, tests, and app-building iterations. Spend the savings on verification. Escalate to GPT-5.6 or Fable 5 when the task has demonstrated that it needs their ceiling.

That is the real coding story: Grok 4.5 is not the most expensive hammer. It is the very good power tool you can afford to keep running.

02

Strengths and honest limitations

Key Strengths

  • Independent third place, not vendor-only applause: Artificial Analysis ranks Grok Build third on its Coding Agent Index at 76, on par with GPT-5.5 in Codex and below Fable 5 in Claude Code. That is the right shape of claim: a top-tier practical agent, not an invented sweep.
  • The cost-per-agent-task gap is enormous: Artificial Analysis reports $2.49 per Coding Agent Index task for Grok Build, versus $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code. It used 1.9M tokens per task versus 6.2M and 7.2M respectively.
  • Terminal work is genuinely competitive: xAI reports 83.3% on Terminal-Bench 2.1—just behind Fable 5’s 84.3% and GPT-5.5’s 83.4% in its published comparison. For tool use, shell work, and retry loops, that is a very usable neighborhood.
  • Fast and cheap encourages better engineering habits: xAI serves Grok 4.5 at 80 TPS and prices it at $2/$6 per 1M input/output tokens. That makes it easier to run another test, ask for a second implementation, or let an agent investigate before a human takes over.
  • Cursor and Grok Build make the model easy to deploy: Grok 4.5 is the default in Grok Build and is available in Cursor on all plans, alongside the API. It has a real developer home instead of requiring a bespoke harness before it can be useful.

Honest Limitations

  • Raw benchmark leadership belongs elsewhere: xAI reports 64.7% on SWE-Bench Pro, behind Fable 5’s 80.4% and Opus 4.8’s 69.2% in the same table. If repository-issue resolution is your only sport, do not discard those leaders.
  • DeepSWE is competitive, not dominant: On xAI’s cited comparison, Grok 4.5 scores 62.0% on DeepSWE 1.0 and 53% on DeepSWE 1.1, behind Fable 5 and GPT-5.5 on both. The value case does not erase the raw-score gap.
  • EU teams cannot use the launch today: xAI says Grok 4.5 is not yet in the EU through its products or API console. Expected mid-July is a plan, not availability.
  • Token efficiency does not guarantee correct code: Fewer steps and lower cost are advantages only after the change passes tests, review, security checks, and the messy context your benchmark did not include.
  • It is a fresh release: Use private holdout tasks before standardizing on it. The model is new, harness behavior matters, and a cost curve is not a substitute for your CI.
03

Benchmark Snapshot

Artificial Analysis Coding Agent Index — 76 (#3)

Independent result for Grok Build: on par with GPT-5.5 in Codex and below Fable 5 in Claude Code.

Coding Agent Index cost — $2.49 per task

Artificial Analysis reports Grok Build at $2.49 per task versus $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code.

Terminal-Bench 2.1 — 83.3%

xAI's published comparison puts Grok close to Fable 5 at 84.3% and GPT-5.5 at 83.4% on terminal-agent work.

SWE Marathon — 29.0% pass@1

xAI reports Grok ahead of the cited Opus 4.8 and Fable 5 figures on this long-running engineering evaluation.

SWE-Bench Pro — 64.7%

The important counterweight: xAI's chart places Grok below Fable 5 and Opus 4.8 on repository-issue resolution.

04

The Verdict

Grok 4.5 is #4 because the coding leaderboard should reward both capability and an agent teams can afford to run repeatedly. Kimi K3 now has the stronger raw and frontend evidence, while Grok remains the value-and-speed alternative: strong terminal behavior, third-place independent agent results, Cursor availability, and a startlingly low cost per task. Use K3, Fable 5, or GPT-5.6 when the problem demands a higher ceiling; use Grok 4.5 when you want a capable agent to keep trying without setting fire to the budget.

05

Frequently Asked Questions