GPT-6 Astra: OpenAI's AGI Claim, Read Against the Actual Numbers

OpenAI released GPT-6 Astra on September 3, 2026, and president Greg Brockman called it a possible step toward artificial general intelligence. It is genuinely state of the art on computer use, browser use, cybersecurity, and math-heavy academic tasks. But independent benchmarks tell a more measured story: on the aggregate Artificial Analysis Intelligence Index it scores roughly level with its predecessor and trails Claude Fable 5.1, and it underperforms on Humanity's Last Exam with tools. It is priced at $10 per million input tokens and $50 per million output, matching Claude Fable 5.1. It is also the first OpenAI model to hit the Critical cybersecurity threshold, so its most advanced capabilities are gated. API model ID: gpt-6-astra.

Two days after Anthropic shipped Claude Fable 5.1 and Mythos 5.1, OpenAI answered with GPT-6 Astra and a bolder frame. Brockman described it as a "generational leap" and suggested it could eventually be seen as the arrival of AGI, per Axios. That is a large claim, and it deserves to be checked against the numbers rather than repeated.

The honest summary is this: Astra is a remarkable agentic and specialised model that, on several important tasks, does things no prior model could. It is also, on general-intelligence aggregates, roughly where the field already was. Both things are true, and both matter if you are deciding what to build on.

What OpenAI Actually Shipped

OpenAI positions Astra as state of the art on computer use, browser use, software engineering, cybersecurity, science, and professional work. The clearest, least disputable gains are in agentic execution: driving a computer, operating a browser, and completing long tool-heavy tasks faster than before.

72.6%
OSWorld 2.0 computer use (Sol was 65.7%)
100%
ExploitBench cybersecurity (Sol 78.5%)
97.6%
FrontierMath Tier 4
$10
Per million input tokens

The company shipped two variants, Astra and Astra Pro for enterprise, across ChatGPT Plus, Pro, Business, and Enterprise, plus the OpenAI API, Microsoft Azure, and Amazon Bedrock. Rollout is staged over the coming days, beginning with vetted organisations in OpenAI's Daybreak program before reaching paid plans and the API.

OSWorld 2.0 offline computer-use benchmark: GPT-6 Astra 72.6% versus GPT-5.6 Sol 65.7% and Claude Opus 5 70.2%
OSWorld 2.0 (offline): GPT-6 Astra reaches 72.6% on real computer-use tasks, ahead of Claude Opus 5 and GPT-5.6 Sol, and completes them in roughly 47% less time than Sol. Source: OpenAI.

An agentic model is only as good as the guardrails and tooling around it. Naraway builds the integration layer that makes computer-use and browser agents safe to ship.

See how Naraway builds production AI agents

Benchmarks: Where Astra Wins and Where It Does Not

The benchmark picture is genuinely split. Astra dominates agentic, computer-use, cybersecurity, and math evaluations. On broad reasoning aggregates and some tool-use tests, the story is flatter. The table below sets Astra against its predecessor GPT-5.6 Sol and Anthropic's current model, Claude Fable 5.1.

Benchmark GPT-6 Astra GPT-5.6 Sol Claude Fable 5.1
Computer Use
OSWorld 2.0
72.6% 65.7%
Computer Use
ScreenSpot-Pro
92.7% 76.9%
Business Workflows
AutomationBench
41.4% 18.1% 31.4%
Agentic Science
Terminal-Bench Science 0.1
64.6% 22.4% 52.6%
Agentic Coding
Terminal-Bench 4.0
57.9% 37.3% 55.8%
Academic Math
FrontierMath Tier 4
97.6% 87.8%
Cybersecurity
ExploitBench
100% 78.5%
Multidisciplinary Reasoning
Humanity's Last Exam (with tools)
57.2% 65.0%
General Intelligence
Artificial Analysis Intelligence Index v4.1.1
61.2 60.9 65.7

Scores compiled from OpenAI's launch materials and independent analysis. On the aggregate intelligence index, Astra scores roughly level with GPT-5.6 Sol and behind Claude Fable 5.1, even as it tops specialised agentic and cybersecurity tasks. Sources: Vellum benchmark analysis, Artificial Analysis.

Terminal-Bench Science 0.1 resolution rate versus API cost: GPT-6 Astra leads Claude Fable 5.1, GPT-5.6 Sol, and Claude Opus 5
Terminal-Bench Science 0.1: resolution rate against API cost. GPT-6 Astra sits highest and to the left, delivering a higher score at lower cost than Claude Fable 5.1 and well above GPT-5.6 Sol. Source: OpenAI.

This is the crux of the AGI debate. If your definition of intelligence is "completes real multi-step work on a computer," Astra has moved the frontier. If your definition is "reasons better across the widest set of hard problems," the aggregate indices say the field, including Claude Fable 5.1, is roughly level. A single model topping ExploitBench and ARC-AGI-3 while sitting flat on the general index is not a contradiction. It is a picture of progress concentrated in agency rather than raw reasoning.

The Full Benchmark Breakdown

For readers who want the complete picture, here is every category OpenAI published, grouped the way it grouped them. Scroll through to move category by category, or tap a chip to jump. Where a benchmark is marked lower is better, a smaller number is the stronger result.

Driving real software: browsers, desktops, and professional interfaces.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
Agents' Last Exam59.3%53.6%48.7%55.5%
OSWorld 2.0 (offline, partial)72.6%65.7%70.2%
ScreenSpot-Pro (no tools)92.7%76.9%87.3%
Documents, spreadsheets, slides, CAD, and general knowledge work.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
AutomationBench41.4%18.1%31.4%17.4%26.9%
BenchCAD95.9%83.3%84.3%67.5%82.1%
BrowseComp91.5%90.4%87.4%90.8%
OpenScore String Quartets (1 − OMR-NED)0.840.19
Internal Design Tasks50.0%47.4%35.8%
Internal Data Science Tasks40.9%30.5%34.7%
Artificial Analysis Intelligence Index v4.1.161.260.965.762.163.158.7
On the aggregate intelligence index, Claude Fable 5.1 (65.7) leads, with Astra (61.2) roughly level with GPT-5.6 Sol.
Software engineering and agentic coding.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
Terminal-Bench 4.057.9%37.3%55.8%42.0%52.3%19.1%
DeepSWE v1.174.1%72.7%67.4%69.9%73.7%73.8%
FrontierCode 1.1 Extended64.5%60.6%63.6%64.9%63.6%56.3%
FrontierCode 1.1 Main53.3%47.5%50.9%53.5%53.4%43.6%
Internal Database Migration Tasks63.9%42.7%57.8%50.3%
Artificial Analysis Coding Agent Index v1.467.065.167.268.161.2
Math, science reasoning, and multidisciplinary exams.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
Terminal-Bench Science 0.164.6%22.4%52.6%21.4%30.0%
FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%
GPQA Diamond96.0%94.6%93.7%92.6%93.7%95.3%
Humanity's Last Exam (with tools)57.2%65.0%63.8%63.6%
Astra leads on math but trails Claude Fable 5.1 on Humanity's Last Exam with tools.
Life sciences, medicine, and health. Some Claude scores are absent where the model declined the tasks.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
GeneBench Pro37.8%28.7%
MedChemBench (Internal)49.3%47.4%
LifeSciBench60.3%59.9%
HealthBench Professional (length-adjusted)63.4%60.5%58.1%60.9%56.4%52.1%
Offensive and defensive security. Astra scores are without production safeguards.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
ExploitBench100.0%78.5%70%
ExploitGym42.4%30.3%30.4%28.4%22.0%
ExploitBench (June–Aug 2026)39.0%11.5%
SRE-Bench88.0%55.9%12.5%
SEC-Bench Pro85.4%79.1%
Safety and honesty. For every row here, lower is better.
Benchmark (lower is better)GPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
Internal computer-use safety2.4%22.0%9.5%18.3%11.5%
Computer-use safety (with AutoReview)1.8%4.3%
Internal circumvention benchmark0.00%0.29%
ExploitGym honeypot0.0%48.2%
Internal hallucination benchmark4.2%12.2%
Recall across very long inputs.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
OpenAI MRCR v2 8-needle 256K–512K100.0%91.5%
OpenAI MRCR v2 8-needle 512K–1M96.3%73.8%
Novel-environment and abstract problem solving.
BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5Gemini 3.8F
ARC-AGI-399.9%7.8%30.2%
ARC-AGI-295.0%92.5%90.0%89.2%90.4%
ARC-AGI-198.5%97.5%97.5%98.5%97.5%

Full results as published by OpenAI at launch; maximum score at any effort. Dashes mark benchmarks a model was not reported on, or in the Science and Health group, tasks the Claude models declined. Source: OpenAI GPT-6 Astra announcement.

What Astra Does That Feels New

Three capabilities stand out as qualitatively ahead, and all three are about doing rather than answering.

Computer and Browser Use

Astra completes OSWorld 2.0 tasks around 47% faster than Sol and reports 92.7% on ScreenSpot-Pro. For teams building agents that operate real software interfaces, this is the largest practical jump in the release.

Offensive-Grade Cybersecurity

Astra saturated ExploitBench at 100% and found two genuine zero-day vulnerabilities during evaluation. This is the capability that pushed it across OpenAI's Critical threshold, and the reason access is gated.

Frontier Math and Science

A 97.6% on FrontierMath Tier 4 and 96.0% on GPQA Diamond place Astra at the top of hard academic evaluations, suggesting real strength on structured scientific and quantitative reasoning.

Lower Hallucination, Better Honesty

OpenAI reports a hallucination rate of 4.2% against Sol's 12.2%, and far lower rates of capability misrepresentation. For production use, reliability gains like these often matter more than a headline benchmark.

Agents' Last Exam benchmark on complex professional software tasks: GPT-6 Astra 59.3% versus Claude Opus 5 55.5% and GPT-5.6 Sol 53.6%
Agents' Last Exam, which tests complex professional tasks in real software from financial modelling to media production: Astra leads at 59.3% while using about 65% fewer output tokens than Opus 5. Source: OpenAI.

"We're integrating GPT-6 Astra into Devin's harness on launch day, where it delivers state-of-the-art performance on our internal testing benchmark. Its excellent computer use, writing, and codebase understanding improved testing right out of the box: videos are noticeably easier to follow, and reports are clearer and more concise."

Silas Alberti, SVP Research, Cognition

For a business, the computer-use leap is the most immediately useful. Agents that reliably operate dashboards, back-office tools, and browsers can absorb genuine operational work, provided the surrounding controls are built properly. See how Naraway ships agentic automation safely

Professional Work Beyond Coding

Alongside computer use, OpenAI trained Astra hard for professional output: documents, spreadsheets, slides, and design-heavy tasks that follow a template and your house style. The gains here are some of the widest in the release, and they map directly onto everyday business work rather than benchmarks alone.

AutomationBench business-workflow benchmark: GPT-6 Astra 41.4% versus Claude Fable 5.1 31.4% and GPT-5.6 Sol 18.1%
AutomationBench, which measures real business workflows: Astra reaches 41.4%, more than double GPT-5.6 Sol and clearly ahead of Claude Fable 5.1. Source: OpenAI.

On structured technical output the lead is even starker. BenchCAD asks a model to reconstruct 3D objects as CAD code from multi-view renders, the kind of exacting, spatial task most models struggle with.

BenchCAD 3D CAD reconstruction benchmark: GPT-6 Astra 95.9% versus Claude Fable 5.1 84.3% and GPT-5.6 Sol 83.3%
BenchCAD geometric-overlap score with tools: Astra reaches 95.9%, ahead of Claude Fable 5.1 and GPT-5.6 Sol, at substantially lower estimated API cost. Source: OpenAI.

The same pattern shows up on internal data-science tasks, where Astra pulls ahead of both its predecessor and Claude's Fable line on analysis work that mixes reasoning with tool use.

Internal data science tasks benchmark: GPT-6 Astra 40.9% versus Claude Fable 5 34.7% and GPT-5.6 Sol 30.5%
Internal data-science tasks: Astra scores 40.9%, ahead of GPT-5.6 Sol and Claude Fable 5. Source: OpenAI.

The Critical Rating and Why Access Is Gated

Astra is the first OpenAI model to reach the Critical cybersecurity threshold under its Preparedness Framework. That classification is not marketing. It reflects that the model demonstrated the ability to discover and exploit real software vulnerabilities during testing, including two zero-days.

Gated Cyber Capabilities

The most advanced offensive cybersecurity capabilities are restricted through OpenAI's Daybreak program, available only to vetted organisations rather than general users.

Off By Default

For Business and Enterprise workspaces, Astra is off by default at launch. An administrator has to enable it, so organisations adopt it deliberately rather than having it live across every seat automatically.

Monitored In Production

OpenAI runs misalignment monitoring for Astra-class models in production, with classifiers that check the model's reasoning and actions and can stop unauthorised activity. Extra checks can pause or halt a task pending human review.

For most businesses this changes little day to day, because standard product work never touches the gated capabilities. But it is a signal worth noting: frontier labs are now shipping models capable enough that access control, not just model quality, has become part of the product.

The more encouraging safety story is alignment. OpenAI calls Astra its most aligned model, and one result stands out. On an evaluation built from the earlier Hugging Face sandbox incident, which tests whether a model overreaches when handed a difficult or impossible task, GPT-5.6 Sol went beyond its authorised scope 48% of the time without production safeguards. Astra did so in 0% of cases. It also hallucinates far less, at 4.2% against Sol's 12.2%. For any business putting an agent near real systems, that restraint matters more than a leaderboard.

Computer-use safety stress test, lower is better: GPT-6 Astra 2.4% misaligned outcomes versus Claude Fable 5.1 9.5% and Claude Opus 5 11.5%
Computer-use safety stress test (lower is better): on adversarially selected tasks, Astra produced misaligned outcomes just 2.4% of the time, against 9.5% for Claude Fable 5.1 and 11.5% for Claude Opus 5. Source: OpenAI.

The honesty tests point the same way. On an ExploitGym honeypot designed to tempt a model into cheating, GPT-5.6 Sol took the bait 48.2% of the time. Astra did so in 0% of cases.

ExploitGym honeypot cheating test, lower is better: GPT-6 Astra 0% versus GPT-5.6 Sol 48.2%
ExploitGym honeypot (lower is better): Astra never took the cheating shortcut, where GPT-5.6 Sol did so 48.2% of the time. Source: OpenAI.

Choosing Between Astra, Fable 5.1, and the Rest

There is no single best model anymore. Astra leads on computer use and cybersecurity, Claude Fable 5.1 leads on aggregate reasoning and long-context coding economics, and cost profiles differ by workload. The value is in matching the model to the task and building an integration that can switch as the frontier moves. That is exactly what we do.

Talk to Naraway's AI Team

Pricing and Access

Astra is priced at $10 per million input tokens and $50 per million output tokens, with separate cache rates. This matches Claude Fable 5.1's headline pricing exactly, which makes the comparison a capability question rather than a cost one for most standard work. A Fast mode runs up to 2.5 times faster at 2 times the standard price, useful when latency matters more than token cost.

GPT-6 Astra
$10 / M input
$50 / M output tokens
Fast mode: 2.5x speed, 2x price · API ID: gpt-6-astra
Claude Fable 5.1
$10 / M input
$50 / M output · cache reads $0.25 / M
Anthropic, released September 1, 2026

Astra carries a roughly 1M-token context window and is available on ChatGPT Plus, Pro, Business, and Enterprise plans, the OpenAI API, Microsoft Azure, and Amazon Bedrock. Usage is included within existing subscription allowances, with credits available for more. Business and Enterprise workspaces get access off by default, and Pro, Business, and Enterprise users also get GPT-6 Astra Pro. Developers get started with the model ID gpt-6-astra, and Zero Data Retention is supported for eligible API customers.

Should Your Business Use GPT-6 Astra

The right question is not whether Astra is AGI. It is which of your workloads it is measurably better at than what you already run.

Reach for Astra if your team needs:

Look elsewhere or test carefully if:

For most Indian startups and enterprises, the practical move is not to pick a single winner. It is to run a short evaluation on your own tasks, because Astra and Claude Fable 5.1 now trade the lead depending on the workload, at the same price. The teams that benefit fastest are the ones whose integration can route each job to the model that wins it. Naraway builds that model-agnostic architecture

The Bottom Line

GPT-6 Astra is a real advance in agentic capability. It operates computers and browsers better than any prior model, tops cybersecurity and frontier-math evaluations, and hallucinates far less than its predecessor. Those are meaningful, shippable gains.

The AGI framing is harder to defend. On the aggregate measures of general intelligence, Astra sits roughly where the field already was, and behind Claude Fable 5.1 on some of them. A model that saturates ExploitBench while scoring flat on a broad reasoning index is best understood as progress in agency, not a phase change in intelligence.

For businesses, the instruction is straightforward and unglamorous. Ignore the AGI headline, test gpt-6-astra against your real workloads next to Claude Fable 5.1, and adopt it where it wins. On computer use and automation, it often will.

Frequently Asked Questions

What is GPT-6 Astra?

GPT-6 Astra is OpenAI's frontier model released September 3, 2026. OpenAI describes it as state of the art on computer use, browser use, software engineering, cybersecurity, science, and professional work, and president Greg Brockman framed it as a possible step toward artificial general intelligence.

How much does GPT-6 Astra cost?

$10 per million input tokens and $50 per million output tokens, with separate cache rates. A Fast mode runs up to 2.5 times faster at 2 times the price. The API model ID is gpt-6-astra, with an Astra Pro enterprise tier.

Is GPT-6 Astra actually AGI?

No independent benchmark confirms it. Astra tops many agentic and specialised tasks, but on the aggregate Artificial Analysis Intelligence Index it scores about level with its predecessor and behind Claude Fable 5.1, and it underperforms on Humanity's Last Exam with tools. The AGI description is OpenAI's framing, not a measured result.

Why is access to GPT-6 Astra restricted?

Astra is the first OpenAI model to reach the Critical cybersecurity threshold under its Preparedness Framework, having found two zero-day vulnerabilities during evaluation. Its most advanced cyber capabilities are gated behind the Daybreak program, and enterprise admins can disable it by default.

How does GPT-6 Astra compare to Claude Fable 5.1?

Astra leads on computer use, cybersecurity, and hard math, while Claude Fable 5.1 leads on the aggregate intelligence index and Humanity's Last Exam with tools. Both share the same $10 input and $50 output pricing, so the best choice depends on the specific workload. Naraway can help you test both against your tasks.