Ranked guide

Coding — AI That Writes Production Code

These are coding agents, not autocomplete toys. GPT-6 Astra takes #1 backed by Arena.ai's Code Arena (1,797 pts in WebDev), terminal and computer-use efficiency, and low cost per finished task; Opus 5 is the evidence-leading verifier at #2. GLM-5.3 and the cheaper multimodal GLM-5.3-Flash expand the open-weight end of the shortlist.

Decision first

Our ranking

Start with the winner, then compare the trade-offs that might change the answer for you.

#1 Coding

GPT-6 Astra

OpenAI

GPT-6 Astra takes the coding crown the same honest way GPT-5.6 held it: by finishing agent-shaped work — terminal sessions, failing pipelines, incident forensics, whole-desktop setups — now backed by Arena.ai's new #1 ranking on Code Arena (WebDev 1,797 pts) while burning roughly a third of Sol's tokens to do it.

Why It Wins

New #1 on Arena.ai's Code Arena: WebDev with 1,797 pts (+35 lead over Claude Fable 5.1 Max at 1,762, +109 over Opus 5 Max at 1,688, and +180 over Sol at #13), leading Data & Analytics, Consumer Product, and Content Creation Tools. Combined with Terminal-Bench 4.0 at 57.7–57.9% (Fable 5.1: 55.8%), table-best DeepSWE v1.1 at 74.1%, Terminal-Bench Science at 64.6%, SRE-Bench at 88% pass@1 / 99.2% pass@4, and 100% on ExploitBench with previously unknown 0-days found during evaluation. While Artificial Analysis' own agent harness still narrowly scores Fable 5.1 ahead (70 vs 67), Astra reshapes the Pareto frontier at $40/Mtoken blended and burns roughly a third of Sol's tokens.

The Catch

Even with Code Arena's #1 crown in WebDev and consumer tooling, Fable 5(.1) still holds the SWE-Bench Pro repo-issue record and leads on Artificial Analysis' own-harness agent index (70 vs 67). The API sticker is 2.5x Sol at $10/$50, cache reads cost four times Fable 5.1's $0.25, terse answers can skip report-polish steps, and stronger cyber safeguards add friction to exploit-adjacent work. Launch tables moved after publish — treat any single number as ±1–2 points.

9.9 Editorial score
Read review
Best for

GPT-6 Astra takes the coding crown the same honest way GPT-5.6 held it: by finishing agent-shaped work — terminal sessions, failing pipelines, incident forensics, whole-desktop setups — now backed by Arena.ai's new #1 ranking on Code Arena (WebDev 1,797 pts) while burning roughly a third of Sol's tokens to do it.

Why It Wins

New #1 on Arena.ai's Code Arena: WebDev with 1,797 pts (+35 lead over Claude Fable 5.1 Max at 1,762, +109 over Opus 5 Max at 1,688, and +180 over Sol at #13), leading Data & Analytics, Consumer Product, and Content Creation Tools. Combined with Terminal-Bench 4.0 at 57.7–57.9% (Fable 5.1: 55.8%), table-best DeepSWE v1.1 at 74.1%, Terminal-Bench Science at 64.6%, SRE-Bench at 88% pass@1 / 99.2% pass@4, and 100% on ExploitBench with previously unknown 0-days found during evaluation. While Artificial Analysis' own agent harness still narrowly scores Fable 5.1 ahead (70 vs 67), Astra reshapes the Pareto frontier at $40/Mtoken blended and burns roughly a third of Sol's tokens.

Watch out

Even with Code Arena's #1 crown in WebDev and consumer tooling, Fable 5(.1) still holds the SWE-Bench Pro repo-issue record and leads on Artificial Analysis' own-harness agent index (70 vs 67). The API sticker is 2.5x Sol at $10/$50, cache reads cost four times Fable 5.1's $0.25, terse answers can skip report-polish steps, and stronger cyber safeguards add friction to exploit-adjacent work. Launch tables moved after publish — treat any single number as ±1–2 points.

#2

Claude Opus 5

Anthropic

The evidence-leading frontier coder with unusually patient verification. Opus 5 remains our #2 only because this guide gives GPT-6 Astra an explicit tiebreak for Codex integration, terminal and computer-use efficiency, and cost per finished task.

9.9 Editorial score
Read review
#3

Claude Fable 5.1

Anthropic

The Mythos-class reasoning engine refined for agentic coding. Same weights as the restricted Mythos 5.1, but optimized with a 75% cache discount that transforms long-horizon engineering economics. Best deployed in Claude Code, where its AA Coding Agent score of 70 leads the field.

9.8 Editorial score
Read review
#4

Grok 4.6

xAI

Grok 4.6 is a serious coding-agent upgrade: better at repository discovery, long-running implementation, and visual first passes, while retaining $2/$6 pricing and immediate access in Cursor and Grok Build. It belongs in the frontier conversation, but it does not win every software-engineering test.

9.7 Editorial score
Read review
#5

GLM-5.3

Z.ai (Zhipu AI)

A heavyweight open-weight coding agent trained for the part that comes after the first answer: planning, editing, testing, recovering, and staying oriented through long repository jobs.

9.2 Editorial score
Read review
#6

Qwen3.8-Max

Alibaba / Qwen Team

A high-value visual and long-context coding challenger, now ranked #6 after independent testing placed flagship GLM-5.3 ahead. Its downloadable Max-tier checkpoint remains remarkable, but 2.4T parameters make self-hosting a datacenter project.

9.2 Editorial score
Read review
#7

GLM-5.3-Flash

Z.ai (Zhipu AI)

Frontier-adjacent coding and agent work at unusually low prices, with native vision and a 1M context. “Flash” means efficient and cheap here—not small, and not especially fast.

9.2 Editorial score
Read review
Questions, answered

Frequently Asked Questions