Ranked #4 Coding — AI That Writes Production Code
OpenAI

GPT-6.1 Sol

GPT-6.1 Sol is the coding agent you can leave running all day without watching the bill. On independent tests it lands within a point or two of GPT-6 Astra, OpenAI's flagship, yet a typical task costs less than a quarter as much. The token price hasn't changed from GPT-6 Sol ($2 in, $10 out per million), but re-reading cached code now costs just $0.10 per million tokens, and big repositories are exactly where that saving shows.

Updated September 30, 2026 Value Coding AgentDeepSWE 75.2%$0.10 Cached Input
9.8out of 10
Official Website
Best for

GPT-6.1 Sol is the coding agent you can leave running all day without watching the bill. On independent tests it lands within a point or two of GPT-6 Astra, OpenAI's flagship, yet a typical task costs less than a quarter as much. The token price hasn't changed from GPT-6 Sol ($2 in, $10 out per million), but re-reading cached code now costs just $0.10 per million tokens, and big repositories are exactly where that saving shows.

Why It Wins

Artificial Analysis (September 29, 2026) measured it one point below GPT-6 Astra on its Intelligence Index (52 vs 53) at $0.72 per task versus $3.26. On the Coding Agent Index its xhigh setting beat Astra by a point for under 15% of the cost. OpenAI reports 75.2% on DeepSWE v1.1 at high effort, and on Cognition's FrontierCode 1.1 it scores 58.1% at low effort for $0.21 per task, the best result under $0.30 on that board.

Watch out

Its merge-quality score didn't improve: on FrontierCode 1.1 it scores 60.4%, the same as GPT-6 Sol. The gain is cost, not a higher ceiling. OpenAI's own tables show GPT-6 Astra still clearly ahead on scientific terminal work. Claude Sonnet 5.5 scores higher on Artificial Analysis's broad index, though it costs far more per task. In ChatGPT it lives only in Work and Codex on paid plans.

01

What It Actually Is

Picture a building site with two crane operators.

One is the veteran brought in for the lift nobody else will touch: the steel beam that has to thread between two towers in a crosswind. The other operator is nearly as good and charges a fraction of the day rate. For every ordinary lift, pallets of bricks, roof trusses, the day-in, day-out work, you’d be foolish to book the veteran.

GPT-6.1 Sol is the second operator, and this month the gap between the two got very small.

A release that came out of a cancellation

OpenAI launched GPT-6.1 Sol at its DevDay event on September 29, 2026, only a week after GPT-6 Sol. The timing had a backstory. The day before, OpenAI had scrapped the planned launch of GPT-6.1 Astra after internal tests found that model didn’t reliably stay within the scope and permissions of its tasks. Sol 6.1 is a separate, smaller model, and its own safety report looks cleaner: in OpenAI’s test of whether an agent admits that its search tool is broken instead of guessing, GPT-6.1 Sol failed to say so in 2.1% of cases, compared with 4.9% for GPT-6 Sol and 1.5% for GPT-6 Astra.

OpenAI’s pitch is short: near-Astra intelligence for a fifth of the price.

Is it really near Astra?

This is the part that matters, and for once we don’t have to take the vendor’s word for it.

Artificial Analysis, an independent testing firm, ran its benchmarks on launch day. On its Intelligence Index, a combination of ten hard evaluations, GPT-6.1 Sol at max effort scored 52, one point below GPT-6 Astra’s 53 and four above GPT-6 Sol. On its Coding Agent Index, the result is even closer: at max effort Sol sits 2 points behind Astra, and at the second-highest setting, xhigh, it beat Astra by a point for less than 15% of the cost per task.

Cognition, the company behind the Devin coding agent, tested it on FrontierCode 1.1, a benchmark that asks a strict question: would a senior engineer actually merge this change? GPT-6.1 Sol’s best score is 60.4%, essentially the same as GPT-6 Sol (60.7%). The difference is cost. At medium effort it spends $0.31 per task, 81% less than GPT-6 Sol spent at max to reach its best. At low effort it scores 58.1% for just $0.21, up from 50.5% for GPT-6 Sol at the same setting.

OpenAI’s own figures point the same way. On DeepSWE v1.1, complex engineering tasks in real codebases, its chart shows 75.2% at high effort for $0.65 per task. That’s 6.4 points above GPT-6 Sol’s best and a hair above GPT-6 Astra’s best (74.1%), which costs $4.43 per task. One oddity is worth knowing: at xhigh and max, GPT-6.1 Sol actually scores lower, 71.9%. More thinking isn’t always better thinking.

So the honest summary is: roughly Astra-class coding, and the same merge quality as last week’s Sol, at a much lower price per finished job.

Why the cache price is the real headline

A coding agent never reads your project just once. Every time it runs a test, reads the error, and tries again, it re-reads the same files, instructions, and history. Across a long session that can mean the same few hundred thousand tokens going back into the model dozens of times.

That’s what cached input covers: text the model has seen recently in this session. GPT-6.1 Sol charges $0.10 per million cached tokens, 95% off the normal $2 and half of GPT-6 Sol’s cache price. Normal input and output prices stay at $2 and $10 per million, a fifth of GPT-6 Astra’s $10/$50.

Add it up and you get the number that matters most: Artificial Analysis measured about $0.72 per task to run its index with GPT-6.1 Sol, against $3.26 for GPT-6 Astra, $5.98 for Claude Opus 5.5, and $7.60 for Claude Sonnet 5.5. One caution: prompts longer than 272,000 tokens cost double for input and 1.5 times for output, so truly enormous single requests get pricier.

Don’t turn the dial all the way up

GPT-6.1 Sol has five effort settings: low, medium (the default), high, xhigh, and max. It’s tempting to assume max is always best. On Artificial Analysis’s Coding Agent Index, it isn’t: xhigh scored higher than max, and it costs less. For routine fixes, low or medium is often enough, as Cognition’s 58.1%-for-$0.21 result shows. Save xhigh for the problems that really need the extra thinking.

The honest limits

  • Better value, not better code. On FrontierCode 1.1, merge quality is flat against GPT-6 Sol and GPT-5.6 Sol. If your problem was that Sol’s patches weren’t good enough, 6.1 doesn’t fix that. It makes the same patches much cheaper.
  • Science is still Astra’s job. On OpenAI’s Terminal-Bench Science 0.1, data analysis, simulations, and theorem proving in a terminal, GPT-6.1 Sol scores 57.0% at max effort. That’s more than double GPT-6 Sol’s 27.6%, but well behind GPT-6 Astra (68.1%) and Claude Opus 5.5 (63.3%). It does cost $5.47 a task, against about $23 for either rival.
  • Claude scores higher on the broad index. Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56) remain ahead of GPT-6.1 Sol (52) on Artificial Analysis’s Intelligence Index, and although Artificial Analysis measured a 12-point Terminal-Bench 4.0 gain for GPT-6.1 Sol over GPT-6 Sol, Sonnet’s 64% on that test is still higher. Sonnet gets there by burning far more tokens, so Sol wins on cost per task.
  • Not in regular Chat. In ChatGPT, GPT-6.1 Sol is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Developers use it in the API as gpt-6.1-sol. OpenAI says a faster “Ultrafast” mode for Codex is coming, but it isn’t live yet.

Who should use it

If your work looks like… Reach for Why
Bugs, tests, refactors, agents running in CI all day GPT-6.1 Sol Near-Astra coding at a fraction of the cost per task
Scientific computing, the hardest computer-use jobs GPT-6 Astra Still clearly ahead on science terminals
Open-ended architecture and judgment calls Claude Opus 5.5 Highest on the broad independent index
Long terminal sessions where tokens are no object Claude Sonnet 5.5 Stronger independent terminal score, higher token use

The coding verdict

GPT-6.1 Sol doesn’t raise the ceiling. What it does is bring the ceiling within reach of an everyday budget. Independent tests put it a point or two from OpenAI’s flagship at under a quarter of the cost per task, and its cache pricing is built for the way coding agents actually work. Hand it the backlog, set it to xhigh for the hard tickets, and call in Astra or Opus only when it gets stuck.

02

Strengths and honest limitations

Key Strengths

  • Near-flagship coding, independently checked: Artificial Analysis measured the Coding Agent Index on September 29, 2026. GPT-6.1 Sol at max effort sits 2 points below GPT-6 Astra, and its xhigh setting actually beats Astra by 1 point for less than 15% of the cost per task.
  • Cheap per finished task, not just per token: Running Artificial Analysis’s full index costs about $0.72 per task with GPT-6.1 Sol, against $3.26 for GPT-6 Astra, $5.98 for Claude Opus 5.5, and $7.60 for Claude Sonnet 5.5, all at max effort.
  • Cache price built for big codebases: Re-reading text the model has already seen, such as your repository, costs $0.10 per million tokens: 95% off the normal input price and half of GPT-6 Sol’s cache price. Coding agents re-read the same files constantly, so this is where most of the saving lands.
  • Strong even at low effort: On Cognition’s FrontierCode 1.1, which grades whether code changes are good enough to merge, GPT-6.1 Sol scores 58.1% at low effort for $0.21 per task, up from 50.5% for GPT-6 Sol at the same setting. Across every effort level it runs 44–57% cheaper per task than GPT-6 Sol.
  • A bigger jump in real engineering tasks: On DeepSWE v1.1, OpenAI’s chart shows 75.2% at high effort for $0.65 per task, 6.4 points above GPT-6 Sol’s best (68.8% at max, $2.74) and slightly above GPT-6 Astra’s best (74.1% at xhigh, $4.43).

Honest Limitations

  • Merge quality hasn’t moved: On FrontierCode 1.1, its best score of 60.4% matches GPT-6 Sol (60.7%) and GPT-5.6 Sol (60.6%). You get the same quality of mergeable code for much less money, not better code.
  • Not the model for science terminals: On OpenAI’s Terminal-Bench Science 0.1, which covers data analysis, simulations, and theorem proving, GPT-6.1 Sol scores 57.0% at max effort, more than double GPT-6 Sol (27.6%) but clearly behind GPT-6 Astra (68.1%) and Claude Opus 5.5 (63.3%).
  • Claude still leads the broad index: On Artificial Analysis’s Intelligence Index, Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56) score higher than GPT-6.1 Sol (52). Artificial Analysis measured a 12-point Terminal-Bench 4.0 gain over GPT-6 Sol, but Sonnet 5.5’s launch score of 64% is still higher, though Sonnet spends far more tokens to get there.
  • Max effort isn’t the best setting: On Artificial Analysis’s Coding Agent Index, xhigh scored higher than max. OpenAI’s own DeepSWE chart shows the same pattern: 75.2% at high effort, but 71.9% at xhigh and max, which cost more.
  • Many headline numbers are OpenAI’s: DeepSWE, OSWorld, GDP.pdf, and AutomationBench come from OpenAI’s launch tables. The Artificial Analysis and Cognition results are independent, and they broadly agree.
03

Benchmark Snapshot

Artificial Analysis Coding Agent Index — above GPT-6 Astra at xhigh

Independent, September 29, 2026. At xhigh effort GPT-6.1 Sol beats GPT-6 Astra by 1 point for under 15% of Astra's cost per task. At max it sits 2 points below Astra and 3 above GPT-6 Sol.

Artificial Analysis Intelligence Index — 52 (max effort)

Independent composite of 10 evaluations, including Terminal-Bench 4.0 and SciCode. One point below GPT-6 Astra (53), up from 48 for GPT-6 Sol, at $0.72 per task versus $3.26 for Astra.

FrontierCode 1.1 — 60.4% best; 58.1% at $0.21 (low effort)

Cognition's benchmark for merge-ready code, published September 29, 2026. Best score level with GPT-6 Sol. The low-effort result is the highest of any model under $0.30 per task on that board.

DeepSWE v1.1 — 75.2% (high effort, $0.65 per task)

OpenAI's chart for complex engineering tasks in real codebases. GPT-6 Astra's best is 74.1% at xhigh for $4.43; GPT-6 Sol's best is 68.8% at max for $2.74. GPT-6.1 Sol drops to 71.9% at xhigh and max.

Terminal-Bench Science 0.1 — 57.0% (max effort, $5.47 per task)

Scientific workflows in a terminal, per OpenAI's chart. GPT-6 Astra leads at 68.1% ($23.80 per task), Claude Opus 5.5 scores 63.3% ($23.21), and GPT-6 Sol 27.6% ($12.18).

Price — $2 / $10 per 1M tokens, $0.10 cached

Same token price as GPT-6 Sol and one-fifth of GPT-6 Astra's $10/$50. Prompts over 272K input tokens cost 2× input and 1.5× output; Batch and Flex are 50% off.

04

The Verdict

GPT-6.1 Sol is for teams that want near-flagship coding agents running all day at mid-tier prices. Independent tests put it within a point or two of GPT-6 Astra for under a quarter of the cost per task. It ranks just below Claude Sonnet 5.5 because Sonnet scores higher on the broad independent index and on terminal work, and because Sol’s merge-quality score didn’t move from GPT-6 Sol. It ranks above Claude Fable 5.1 for everyday coding on cost per finished task. Keep GPT-6 Astra for science terminals and the hardest computer-use jobs. Give GPT-6.1 Sol the backlog, and try xhigh before max.

05

Frequently Asked Questions