GPT-6 Astra takes the coding crown the same honest way GPT-5.6 held it: by finishing agent-shaped work — terminal sessions, failing pipelines, incident forensics, whole-desktop setups — now backed by Arena.ai's new #1 ranking on Code Arena (WebDev 1,797 pts) while burning roughly a third of Sol's tokens to do it.
New #1 on Arena.ai's Code Arena: WebDev with 1,797 pts (+35 lead over Claude Fable 5.1 Max at 1,762, +109 over Opus 5 Max at 1,688, and +180 over Sol at #13), leading Data & Analytics, Consumer Product, and Content Creation Tools. Combined with Terminal-Bench 4.0 at 57.7–57.9% (Fable 5.1: 55.8%), table-best DeepSWE v1.1 at 74.1%, Terminal-Bench Science at 64.6%, SRE-Bench at 88% pass@1 / 99.2% pass@4, and 100% on ExploitBench with previously unknown 0-days found during evaluation. While Artificial Analysis' own agent harness still narrowly scores Fable 5.1 ahead (70 vs 67), Astra reshapes the Pareto frontier at $40/Mtoken blended and burns roughly a third of Sol's tokens.
Even with Code Arena's #1 crown in WebDev and consumer tooling, Fable 5(.1) still holds the SWE-Bench Pro repo-issue record and leads on Artificial Analysis' own-harness agent index (70 vs 67). The API sticker is 2.5x Sol at $10/$50, cache reads cost four times Fable 5.1's $0.25, terse answers can skip report-polish steps, and stronger cyber safeguards add friction to exploit-adjacent work. Launch tables moved after publish — treat any single number as ±1–2 points.