Ranked #5 Everyday Ecosystem — The Leading AI Assistants
OpenAI

GPT-6.1 Sol

GPT-6.1 Sol is OpenAI's everyday work model: the one you hand the reports, the PDFs, and the multi-step chores in ChatGPT Work. Independent tests put it one point below GPT-6 Astra, OpenAI's flagship, for less than a quarter of the cost per task. It reads long documents well and makes fewer factual mistakes than the model it replaces. It can also operate desktop apps nearly as well as Astra, according to OpenAI.

Updated September 30, 2026 Near-Astra ValueChatGPT WorkDocument Work
9.6out of 10
Official Website
Best for

GPT-6.1 Sol is OpenAI's everyday work model: the one you hand the reports, the PDFs, and the multi-step chores in ChatGPT Work. Independent tests put it one point below GPT-6 Astra, OpenAI's flagship, for less than a quarter of the cost per task. It reads long documents well and makes fewer factual mistakes than the model it replaces. It can also operate desktop apps nearly as well as Astra, according to OpenAI.

Why It Wins

Artificial Analysis (September 29, 2026) scored it 52 on its Intelligence Index, one point below GPT-6 Astra (53), at $0.72 per task versus $3.26. OpenAI reports it beats Claude Opus 5.5 on GDP.pdf, questions about complex professional PDFs, at under half the cost, and comes within 2.1 points of Astra on OSWorld 2.0, a test of operating a desktop, at about a seventh of the cost.

Watch out

It isn't in regular ChatGPT Chat yet. You need a paid plan and ChatGPT Work or Codex. On Artificial Analysis's knowledge test it still makes up an answer 54% of the time when it doesn't know. Claude Opus 5.5 and Sonnet 5.5 score higher on the same independent index, and the model reads images but doesn't create them itself.

01

What It Actually Is

Think of a good law firm. The senior partner is brilliant and bills accordingly. You’d never ask them to go through two hundred pages of supplier contracts looking for the cancellation clause. You give that to a trusted senior associate, someone almost as sharp and a lot cheaper.

GPT-6.1 Sol is that senior associate. It isn’t OpenAI’s most powerful model, but it is close, and it’s the one you can afford to use for most of the week’s work.

What it is, and where you’ll find it

OpenAI released GPT-6.1 Sol on September 29, 2026, at its DevDay event, one week after GPT-6 Sol. OpenAI describes it as near-Astra intelligence for coding, computer use, and professional work, at a fifth of the price of its flagship GPT-6 Astra.

Where you can use it matters for most readers. In ChatGPT, GPT-6.1 Sol lives in ChatGPT Work and Codex, the parts of the app built for longer tasks, files, and agents. It’s available on the Plus, Pro, Business, Enterprise, and Edu plans. It is not yet in regular Chat, and it isn’t on the free plan. Developers can use it through the API as gpt-6.1-sol, and coding tools such as Cognition’s Devin added it on launch day.

How close to the flagship?

Artificial Analysis, an independent testing firm, ran its benchmarks on launch day. Its Intelligence Index combines ten hard evaluations, from office work and business automation to scientific coding and factual knowledge. GPT-6.1 Sol scored 52. GPT-6 Astra scored 53, and last week’s GPT-6 Sol scored 48. That puts GPT-6.1 Sol 10th of the 222 models on the board.

Now look at the price. Artificial Analysis measured about $0.72 to run one task of its index with GPT-6.1 Sol, against $3.26 for Astra. That’s one point less capability for less than a quarter of the cost.

For fairness, Anthropic still leads this index. Claude Opus 5.5 scored 58 and Claude Sonnet 5.5 scored 56. But those scores cost much more to reach: $5.98 and $7.60 per task.

What it’s good at in a working day

OpenAI’s launch tests focus on the kind of work that fills an office week. Keep in mind that these are OpenAI’s own numbers.

  • Reading difficult PDFs. On GDP.pdf, the model answers questions about real documents from finance, law, healthcare, and seven other fields, with dense tables, charts, and fine print. OpenAI’s chart shows 32.0% at high effort for $0.35 per task. Claude Opus 5.5’s best is 28.8% at $0.83, and GPT-6 Astra’s best is 32.2% at $1.91, so Sol nearly matches the flagship for a fifth of the price.
  • Multi-step chores. On AutomationBench, an agent completes end-to-end workflows across 47 business tools, such as updating a CRM, filing a ticket, and sending a follow-up. At medium effort, GPT-6.1 Sol scores 31.7% for $0.19 per task, 2.2 points above Claude Opus 5.5 at the same setting. Read the whole chart, though: at max effort Opus 5.5 reaches 42.5% and Astra 41.4%, while Sol tops out at 36.1%. Sol wins on routine chores per dollar; the pricier models still finish more of the hardest ones.
  • Using a computer. On OSWorld 2.0, the model clicks and types its way through real desktop apps. OpenAI reports 71.4% at $1.27 per task, against 73.5% for Astra at $9.44.

Fewer made-up facts, but not none

OpenAI tested factual accuracy on deliberately hard questions taken from real conversations where users had flagged an earlier model’s mistake. At low effort, the share of answers containing a factual error fell from 11.4% to 7.7% compared with GPT-6 Sol, roughly a third fewer.

Artificial Analysis’s independent knowledge test gives a more sober picture. When GPT-6.1 Sol doesn’t know an answer, it still makes one up 54% of the time. That’s better than GPT-6 Sol’s 60%, but worse than Claude Sonnet 5.5’s 47%. The practical advice doesn’t change: for an obscure name, date, or figure, ask it to search the web or cite a source.

A word on the Astra that didn’t ship

GPT-6.1 Sol was not the only model OpenAI had planned for DevDay. The day before, the company cancelled GPT-6.1 Astra after internal tests found it didn’t reliably stay within the scope and permissions of its tasks. GPT-6.1 Sol is a separate model with its own safety report. In one of OpenAI’s tests, an agent’s search tool is quietly broken, and the model should say so rather than guess. GPT-6.1 Sol failed to disclose the problem in 2.1% of cases, against 4.9% for GPT-6 Sol and 1.5% for GPT-6 Astra. That’s a real improvement, though not a guarantee for every situation.

The honest catch

It’s behind a paywall and a tab. For a model aimed at everyday work, the biggest limitation is simply access: no free plan and no regular Chat, only ChatGPT Work and Codex on paid plans.

It reads images but doesn’t draw. GPT-6.1 Sol takes text and images as input and writes text back. It doesn’t generate images, audio, or video itself.

Claude is still ahead on the broad index. If you want the highest independent scores regardless of cost, Claude Opus 5.5 and Sonnet 5.5 are ahead. GPT-6.1 Sol’s advantage is doing nearly as well for much less.

Who should use it

If your day looks like… Reach for Why
Contracts, reports, and PDFs in ChatGPT Work GPT-6.1 Sol Near-Astra document work at a fraction of the cost
The hardest research and computer-use jobs GPT-6 Astra OpenAI’s strongest model, a point or two ahead
High-stakes judgment, obscure facts Claude Opus 5.5 Highest on the broad independent index
Everyday work on a free plan Claude Sonnet 5.5 or Gemini Strong models without a subscription

The everyday verdict

GPT-6.1 Sol won’t impress you by being the smartest model on the board. It impresses you by being almost as smart as OpenAI’s flagship at a fraction of the cost. If you already pay for ChatGPT and spend your day in Work, making it your default is the easy choice. Check the obscure facts, and switch to Astra only when a task truly needs the last couple of points.

02

Strengths and honest limitations

Key Strengths

  • Near-flagship brains at a mid-tier price: On Artificial Analysis’s Intelligence Index, measured September 29, 2026, GPT-6.1 Sol scores 52 against 53 for GPT-6 Astra. Running a task costs about $0.72 versus $3.26 for Astra, and it’s 10th of the 222 models tested.
  • Reads difficult documents well: On GDP.pdf, which asks questions about real finance, legal, and healthcare PDFs full of tables, charts, and fine print, OpenAI’s chart shows 32.0% at high effort for $0.35 per task, against Claude Opus 5.5’s best of 28.8% ($0.83) and GPT-6 Astra’s best of 32.2% ($1.91).
  • Good at routine office chores for little money: On AutomationBench, business workflows across 47 tools in sales, finance, support, and HR, OpenAI’s chart shows 31.7% at medium effort for $0.19 per task, 2.2 points above Claude Opus 5.5 at the same setting ($0.65).
  • Operates a computer nearly as well as Astra: On OSWorld 2.0, where an agent clicks and types through real desktop apps, OpenAI reports 71.4% at max effort for $1.27 per task, versus 73.5% for GPT-6 Astra at $9.44.
  • Fewer confident mistakes: On OpenAI’s toughest factuality test, the share of answers with a factual error fell from 11.4% to 7.7% at low effort compared with GPT-6 Sol. Artificial Analysis independently measured its hallucination rate falling from 60% to 54%.

Honest Limitations

  • Not in the main chat window: In ChatGPT, GPT-6.1 Sol runs only in ChatGPT Work and Codex, on Plus, Pro, Business, Enterprise, and Edu plans. Free users can’t pick it, and it isn’t in regular Chat yet.
  • Still guesses when unsure: On Artificial Analysis’s AA-Omniscience test, it made up an answer 54% of the time when it didn’t know, better than GPT-6 Sol’s 60% but worse than Claude Sonnet 5.5’s 47%. Check obscure facts.
  • Claude still scores higher on the broad index: Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56) both beat GPT-6.1 Sol (52) on Artificial Analysis’s Intelligence Index. Sol is cheaper per task, but not smarter.
  • Beats Opus at medium, not at max: OpenAI’s AutomationBench win over Claude Opus 5.5 is at medium effort. At max effort, its own chart shows Opus 5.5 at 42.5% and GPT-6 Astra at 41.4%, well above GPT-6.1 Sol’s 36.1%. For the hardest multi-step workflows, the pricier models still finish more of the job.
  • Reads pictures, doesn’t make them: The model accepts text and images and produces text. Image generation in ChatGPT comes from a separate tool, and the model doesn’t handle audio or video.
  • Many comparisons are OpenAI’s: GDP.pdf, AutomationBench, OSWorld, and the factuality test come from OpenAI’s launch post. The Artificial Analysis index results are independent and point the same way.
03

Benchmark Snapshot

Artificial Analysis Intelligence Index — 52 (max effort)

Independent composite of 10 evaluations, measured September 29, 2026. One point below GPT-6 Astra (53) and up from 48 for GPT-6 Sol. Claude Opus 5.5 scores 58 and Claude Sonnet 5.5 56.

Cost per index task — $0.72

Artificial Analysis's cost to run one task of its index: GPT-6 Astra $3.26, Claude Opus 5.5 $5.98, Claude Sonnet 5.5 $7.60, all at max effort.

GDP.pdf — 32.0% (high effort, $0.35 per task)

Questions about complex professional PDFs in finance, legal, healthcare, and seven other fields, per OpenAI's chart. Claude Opus 5.5's best is 28.8% ($0.83 per task) and GPT-6 Astra's best 32.2% ($1.91).

OSWorld 2.0 offline — 71.4% (max effort)

Operating real desktop apps, partial-credit scoring, per OpenAI. GPT-6 Astra scores 73.5% at $9.44 per task and GPT-6 Sol 64.4% at $3.37; GPT-6.1 Sol costs $1.27.

AA-Omniscience hallucination rate — 54%

Independent. How often the model makes up an answer on questions it gets wrong; lower is better. GPT-6 Sol: 60%. Claude Sonnet 5.5: 47%.

Price — $2 / $10 per 1M tokens, $0.10 cached

One-fifth of GPT-6 Astra's API price. Cached input is 95% off, and Batch and Flex processing are 50% off.

04

The Verdict

GPT-6.1 Sol is for people who already pay for ChatGPT and want its Work agent to handle documents, spreadsheets, and multi-step office jobs without the flagship’s cost. Independent tests put it one point below GPT-6 Astra. It ranks below GPT-6 Astra and the Claude models because Opus 5.5 and Sonnet 5.5 score higher on the same independent index and Sonnet is on Claude’s free plan, while Sol isn’t in regular Chat or on any free plan. It ranks above Gemini 3.1 Pro and Grok 4.7 on current independent results, although Gemini remains free to use. If you use ChatGPT Work every day, it’s the sensible default model.

05

Frequently Asked Questions