Ranked #2 Everyday Ecosystem — The Leading AI Assistants
OpenAI

GPT-6 Astra

GPT-6 Astra is OpenAI's new flagship for the part of work that happens outside the chat box. It drives the browser, clicks through desktop apps, runs multi-step professional workflows to the end, and checks its own facts roughly twice as well as GPT-5.6 Sol — inside the same ChatGPT, Codex, and API ecosystem you already know.

Updated September 5, 2026 AgenticChatGPT WorkComputer Use
9.8out of 10
Official Website
Best for

GPT-6 Astra is OpenAI's new flagship for the part of work that happens outside the chat box. It drives the browser, clicks through desktop apps, runs multi-step professional workflows to the end, and checks its own facts roughly twice as well as GPT-5.6 Sol — inside the same ChatGPT, Codex, and API ecosystem you already know.

Why It Wins

Independent Artificial Analysis testing scores Astra 61.2 on its Intelligence Index — statistically tied with Sol — while the launch tables show 72.6% on OSWorld 2.0 computer use in about half of Sol's time, 59.3% on Agents' Last Exam, 62.7% on the standard ARC-AGI-3 harness (99.9% only with OpenAI's own adapter), 64.6% on Terminal-Bench Science, and hallucinations roughly halved at higher accuracy. Available across ChatGPT, Codex, the API, Azure, and Bedrock from day one.

Watch out

This is a specialist upgrade, not a new IQ record. Independent knowledge-work scores (GDPval-AA) dipped below Sol, Claude Fable 5.1 still leads broad intelligence composites, and the API sticker is 2.5x Sol at $10/$50 per million tokens with a $1 cache read. Astra's terse answers lost ground on legal, finance, and presentation-quality checks, and the day-one rollout left some paying users locked out.

01

What It Actually Is

Every flagship launch arrives with the same poster: the most intelligent model ever built. GPT-6 Astra’s poster is different, and more useful. It says: this one takes the keyboard. Astra is OpenAI’s new flagship for work that happens outside the chat box — operating a browser, clicking through desktop apps, carrying a multi-step professional workflow to the end — and it checks its own facts far better while doing so.

The measured story matches the pitch. On OSWorld 2.0, the desktop-operations benchmark, Astra scores 72.6% and finishes in roughly 40 minutes where GPT-5.6 Sol needed about 75; Opus 5 sits at 70.2%. On Agents’ Last Exam, a long-horizon professional-workflow evaluation across 55 fields, it reaches 59.3%. And Artificial Analysis’ factuality tracking shows the hallucination rate roughly halved versus Sol at higher accuracy. Fewer confident mistakes, faster finished work: that is the upgrade you can feel in the first week.

Here is the honest asterisk. On Artificial Analysis’ Intelligence Index, Astra scores 61.2 — a statistical tie with Sol’s 60.9 — while Claude Fable 5.1 leads at 65.7. On GDPval-AA v2 knowledge work, Astra’s roughly 1629 trails Opus 5 (1824) and Fable 5.1 (1853) — and Sol. If your day is diligence, memos, and judgment, the new flagship is not automatically the better tool; it is a different tool. And the API sticker — $10/$50 per million tokens, 2.5x Sol, with $1 cache reads — prices that difference plainly.

Route the work

Job Route it to Why
“Do it on my machine”: browser, files, desktop apps, finished end to end Astra Best measured computer use, about twice as fast per task
Judgment-heavy knowledge work, diligence, long documents Opus 5 Higher GDPval at a quarter of the API price
The hardest pure reasoning and messy long context Fable 5 / Fable 5.1 Leads independent intelligence composites
Volume, routine drafts, the daily conveyor belt Sol / Terra / Luna Still available, still the cheapest seats in the house

The honest catch

Astra is terse by temperament. That speed-first style wins coding evaluations and loses legal, finance, and tax rubrics that reward listing every step; its presentation polish dropped versus Sol even as its analysis quality jumped. The day-one rollout was messy — paying users hit an access lottery, and Enterprise accounts shipped with Astra off by default. There is no native image generation and no audio or video output. And its much stronger cyber capabilities (100% on ExploitBench, with previously unknown 0-days surfaced during evaluation) come with stricter safeguards that can slow legitimate security-adjacent work.

So the winning habit: Astra for the driver’s seat, Opus for judgment, Fable for the hard thinking, Sol for the conveyor belt. OpenAI did not build a bigger brain; it built better hands. For a lot of real work, that is the upgrade that matters.

02

Strengths and honest limitations

Key Strengths

  • Computer use that finishes: Astra posts 72.6% on OSWorld 2.0 against Sol’s 65.7% and Opus 5’s 70.2%, and clears the task set in roughly 40 minutes where Sol needed about 75. In ChatGPT Work that becomes an agent that browses, clicks, types, and moves files until the job is done — with you as the approver.
  • Roughly half the hallucinations: Artificial Analysis’ factuality tracking shows Astra’s hallucination rate cut to about half of Sol’s while accuracy rose by roughly four points. For professional work, fewer confident mistakes is worth more than another exam point.
  • A serious science colleague: Astra scores 64.6% on Terminal-Bench Science — against Fable 5.1’s 52.6% and Sol’s 22.4% — and about 97.6% on FrontierMath Tier 4. It even produced proofs for two of 68 open Erdős problems during evaluation, which is research color, not a product promise.
  • Novel-problem muscles: On ARC-AGI-3’s standard harness Astra reaches 62.7%, roughly double Opus 5’s 30.2%. The 99.9% figure you may have seen requires OpenAI’s provider adapter and compaction; treat 62.7% as the honest, comparable number.
  • Ecosystem continuity: Astra drops into ChatGPT Work, Codex, the desktop app, and Sites on day one, with the API model ID gpt-6-astra, Azure and Bedrock listings, Zero Data Retention for eligible API work, a roughly 1.05M-token context, and day-one support in tools such as Devin.

Honest Limitations

  • Not a general-intelligence leap: Artificial Analysis puts Astra at 61.2 versus Sol’s 60.9 — a tie — while Claude Fable 5.1 leads at 65.7. On GDPval-AA v2 knowledge work, Astra’s roughly 1629 trails Opus 5 (1824) and Fable 5.1 (1853) — and Sol. If your day is memos, diligence, and judgment, Astra is not automatically the better buy.
  • Premium economics: The API lists at $10/$50 per million input/output tokens — 2.5x Sol — with a $1 cache read (Fable 5.1 charges $0.25), a $12.50 cache write, doubled input pricing for prompts over 272k tokens, and 2x pricing in Fast mode. Artificial Analysis computes Astra about 75% more expensive per intelligence-index task than Sol despite using about 10% fewer tokens.
  • Style and access friction: Astra is terse, so it misses rubric items on legal, finance, and tax evaluations, and its presentation polish dropped versus Sol even as analysis quality jumped. The day-one rollout was messy — Enterprise access shipped off by default — and there is no native image generation or audio/video output. Stronger cyber safeguards slow some legitimate security-adjacent work, and OpenAI’s own tables moved slightly after publish, so treat any single cell as ±1–2 points.
03

Benchmark Snapshot

OSWorld 2.0 — Astra 72.6% in ~40 min

OpenAI's computer-use table, against Sol at 65.7% (~75 min) and Opus 5 at 70.2%. Measured on a partial, offline task set — a strong directional signal, not a guarantee for your specific apps.

Artificial Analysis Intelligence Index — 61.2

Independent composite: tied with GPT-5.6 Sol at 60.9 and behind Claude Fable 5.1 at 65.7. The honest headline is 'more agentic,' not 'smarter.'

AA factuality tracking — hallucinations roughly halved

Independent measurement showing about half of Sol's hallucination rate at roughly four points higher accuracy — the reliability change professionals actually feel.

ARC-AGI-3 — 62.7% standard / 99.9% provider adapter

Novel interactive problem solving. 62.7% is the comparable standard-harness number (Opus 5: 30.2%); 99.9% includes OpenAI's Responses harness and compaction and should always carry that footnote.

API pricing — $10/$50 · cache read $1 · write $12.50

Per 1M input/output tokens, plus 2x input / 1.5x output for prompts over 272k tokens and 2x Fast mode. Compare finished-task cost and cache behavior, not sticker prices.

04

The Verdict

GPT-6 Astra takes the #2 everyday seat from GPT-5.6 because it is the strongest computer-use and science agent inside the broadest consumer platform — ChatGPT Work, Codex, desktop, Sites, and the API. Opus 5 keeps #1 on judgment per dollar for pure knowledge work, and Fable 5 remains the intelligence specialist. Choose Astra when the job is ‘do it on my machine and finish it’; choose Opus 5 when the deliverable is judgment; keep Sol, Terra, and Luna for volume routing at the old prices.

05

Frequently Asked Questions