There are two ways to buy a coding agent. You can rent the smartest contractor in the city for every ticket, or you can hire a very good engineer who is fast enough and cheap enough to take another pass when the first one misses. Grok 4.5 is the second option, and that is not faint praise.
Artificial Analysis puts Grok Build third on its Coding Agent Index, at 76: level with GPT-5.5 in Codex, below Fable 5 in Claude Code. Then comes the number that changes the buying decision. The same analysis estimates $2.49 per agent task for Grok Build, versus $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code. Grok used far fewer tokens getting there.
That is why Grok 4.5 remains a strong #4 after Kimi K3 arrives. xAI’s own chart has it almost tied with the leaders on Terminal-Bench 2.1 at 83.3%, and leading the cited figures on SWE Marathon at 29.0% pass@1. It is also available in Cursor on every plan and is the default in Grok Build. Long agent loops, terminal work, multi-step debugging, repository exploration—these are suddenly cheap enough to repeat.
What not to claim
Grok 4.5 is not the new raw-score monarch. Its published 64.7% on SWE-Bench Pro falls behind Fable 5’s 80.4% and Opus 4.8’s 69.2%. Its DeepSWE results also trail Fable and GPT-5.5 in xAI’s comparison. If your workflow is a pure benchmark for resolving repository issues, the models above it still have the sharper resume.
What to do instead
Start Grok 4.5 on the work where a capable agent gets better by taking another loop: investigation, terminal tasks, refactors, tests, and app-building iterations. Spend the savings on verification. Escalate to GPT-5.6 or Fable 5 when the task has demonstrated that it needs their ceiling.
That is the real coding story: Grok 4.5 is not the most expensive hammer. It is the very good power tool you can afford to keep running.