Every flagship launch arrives with the same poster: the most intelligent model ever built. GPT-6 Astra’s poster is different, and more useful. It says: this one takes the keyboard. Astra is OpenAI’s new flagship for work that happens outside the chat box — operating a browser, clicking through desktop apps, carrying a multi-step professional workflow to the end — and it checks its own facts far better while doing so.
The measured story matches the pitch. On OSWorld 2.0, the desktop-operations benchmark, Astra scores 72.6% and finishes in roughly 40 minutes where GPT-5.6 Sol needed about 75; Opus 5 sits at 70.2%. On Agents’ Last Exam, a long-horizon professional-workflow evaluation across 55 fields, it reaches 59.3%. And Artificial Analysis’ factuality tracking shows the hallucination rate roughly halved versus Sol at higher accuracy. Fewer confident mistakes, faster finished work: that is the upgrade you can feel in the first week.
Here is the honest asterisk. On Artificial Analysis’ Intelligence Index, Astra scores 61.2 — a statistical tie with Sol’s 60.9 — while Claude Fable 5.1 leads at 65.7. On GDPval-AA v2 knowledge work, Astra’s roughly 1629 trails Opus 5 (1824) and Fable 5.1 (1853) — and Sol. If your day is diligence, memos, and judgment, the new flagship is not automatically the better tool; it is a different tool. And the API sticker — $10/$50 per million tokens, 2.5x Sol, with $1 cache reads — prices that difference plainly.
Route the work
| Job | Route it to | Why |
|---|---|---|
| “Do it on my machine”: browser, files, desktop apps, finished end to end | Astra | Best measured computer use, about twice as fast per task |
| Judgment-heavy knowledge work, diligence, long documents | Opus 5 | Higher GDPval at a quarter of the API price |
| The hardest pure reasoning and messy long context | Fable 5 / Fable 5.1 | Leads independent intelligence composites |
| Volume, routine drafts, the daily conveyor belt | Sol / Terra / Luna | Still available, still the cheapest seats in the house |
The honest catch
Astra is terse by temperament. That speed-first style wins coding evaluations and loses legal, finance, and tax rubrics that reward listing every step; its presentation polish dropped versus Sol even as its analysis quality jumped. The day-one rollout was messy — paying users hit an access lottery, and Enterprise accounts shipped with Astra off by default. There is no native image generation and no audio or video output. And its much stronger cyber capabilities (100% on ExploitBench, with previously unknown 0-days surfaced during evaluation) come with stricter safeguards that can slow legitimate security-adjacent work.
So the winning habit: Astra for the driver’s seat, Opus for judgment, Fable for the hard thinking, Sol for the conveyor belt. OpenAI did not build a bigger brain; it built better hands. For a lot of real work, that is the upgrade that matters.