Every developer has worked next to two kinds of colleague.
One is brilliant and thorough, and gives you a fifteen-minute lecture every time you ask a question. The other just leans over, types forty lines, runs the tests, and says “try that.”
In OpenAI and Anthropic’s line-ups, the careful, expensive colleague has the most titles. Grok 4.7 is the quick one who is cheap enough to ask a hundred times a day.
Released on September 21, 2026, Grok 4.7 sits at #5 on our Coding leaderboard. Its role is to be the best-value pair programmer in the top tier.
The improvements that matter to coders
Grok 4.6 was an interesting challenger. Grok 4.7 is a clear step up on every coding test xAI published:
- DeepSWE v1.1, complex engineering tasks in real codebases: from 65.2% to 71.0% (at high effort). In xAI’s table that’s slightly ahead of Claude Fable 5.1’s 70.0%.
- CursorBench 4.0, longer coding tasks inside the Cursor editor: from 40.4% to 46.3%.
- Terminal-Bench 4.0, multi-hour work in a command line: from 20.3% to 37.6%, nearly double.
xAI says the gains come from a larger base model, much longer training on tasks that take hours to finish, and specific work on checking its own answers and handling long context. In practice, that means fewer half-finished changes and fewer confident mistakes.
The economics of coding in the editor
When you code in Cursor or a similar editor, the AI isn’t called once a day. It’s called constantly. It reads your files, suggests edits, explains an error, rewrites a function, and runs again.
At $2 per million input tokens and $6 per million output, Grok 4.7 has the lowest output price among top-tier coding models. For comparison: GPT-6 Sol is $10, Claude Opus 5.5 is $20, and GPT-6 Astra is $50. If you work in the editor all day, that difference is real money.
There’s also a fast variant that produces output twice as fast at twice the price. It’s useful when you’re waiting on every answer.
Where it still falls short
Long unsupervised runs. Terminal-Bench 4.0 is the test that best predicts whether you can leave an agent alone in a terminal for hours. Grok 4.7’s 37.6% is a big improvement, but GPT-6 Astra and Claude Opus 5.5 both score 59.6% in independent testing, and Fable 5.1 scores 57.9% in xAI’s own table. For overnight autonomous work, those models are much less likely to get lost.
The best editor results. On CursorBench 4.0, Grok trails Claude Fable 5.1 (51.8%) and Anthropic’s reported 57.8% for Claude Opus 5.5.
Newest rivals. xAI compares Grok 4.7 with GPT-5.6 Sol and Fable 5.1. GPT-6 Sol launched the next day at a similar price, and we haven’t seen an independent head-to-head yet. All Grok numbers here are xAI’s own.
The verdict
For long, unsupervised agent runs, choose GPT-6 Astra. For deep architectural work across a large codebase, choose Claude Opus 5.5.
But for the hours you spend each day in your editor, fixing bugs, writing tests, and refactoring components, Grok 4.7 is a strong, fast, and very cheap partner. Test it against GPT-6 Sol on your own code; one of the two will likely become your daily driver.