Free to read. Sign up to save tools and get alerts when they change. Plus 900+ more AI tool profiles.

Sign up free
7 min read·Updated September 4, 2026

GPT-6 Astra is OpenAI's flagship model, released September 3, 2026 to ChatGPT Plus, Pro, Business and Enterprise plus the API, Microsoft Azure and AWS Bedrock at $10 input and $50 output per million tokens. It is the first model OpenAI has rated Critical for cybersecurity under its Preparedness Framework, and OpenAI's president says it marks the start of the AGI era. It posts large gains on computer use, agentic coding and mathematics, but it trails Claude Fable 5.1 on three published indices, and OpenAI concedes its reasoning is harder to monitor than GPT-5.6 Sol's.

Share

Listen to this overview

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

Learning Objectives

  • Understand what GPT-6 Astra is, how it supersedes GPT-5.6, and where you can actually use it
  • Read its benchmark claims critically, including the three published indices where it does not lead
  • Understand why a Critical cybersecurity rating changes how the model is gated, and what OpenAI admits it gave up

What Is GPT-6 Astra?

GPT-6 Astra is OpenAI's flagship model, released on September 3, 2026. It is the first new GPT generation in more than a year, arriving roughly two months after GPT-5.6 and replacing it as the model OpenAI points most users toward.

OpenAI describes Astra as state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work. President Greg Brockman went further at the launch briefing: asked when AGI was created, he said "I think it's going to be about this time, and I think it might be about this model," and added that "it's not unreasonable to feel that we are now in the AGI era." That is a company officer's framing of his own product, not a measured result, and it is worth holding separately from the benchmarks below.

🎯Tip

Access GPT-6 Astra: In ChatGPT on Plus, Pro, Business and Enterprise, with Pro, Business and Enterprise also getting GPT-6 Astra Pro. In the API as gpt-6-astra, and through Microsoft Azure and AWS Bedrock. Usage draws on your existing subscription allowance, with extra credits purchasable. Note that enterprise access is off by default at launch — a workspace administrator has to enable it.

What Changed

Three shifts matter more than the individual scores.

Computer use got substantially faster, not just better. On the OSWorld 2.0 offline set Astra scores 72.6 percent at roughly 40 minutes per task, against GPT-5.6 Sol's 65.7 percent at roughly 75 minutes — a better result in about 47 percent less time. Paired with an updated Codex harness, OpenAI reports 1.9 times faster task completion on Mind2Web.

Codex can keep notes instead of compacting. Historically a long session compresses its history into a summary, losing why a fix failed or how a component behaved. Astra can keep notes across context windows and search earlier windows directly. It is experimental, enabled in your Codex config.toml, and OpenAI says it becomes the Astra default in the coming weeks.

It builds and ships, not just drafts. With Sites in ChatGPT, Astra can create, host and share websites, web apps and games from a prompt, and it is tuned to follow existing document and slide templates rather than producing generic output.

Benchmarks

Every figure below is vendor-reported by OpenAI, measured at maximum effort, and run in OpenAI's research environment rather than production ChatGPT.

BenchmarkGPT-6 AstraGPT-5.6 SolClaude Fable 5.1
Terminal-Bench 4.057.9%37.3%55.8%
OSWorld 2.0 (offline set)72.6%65.7%not published
ScreenSpot-Pro (no tools)92.7%76.9%not published
AutomationBench41.4%18.1%31.4%
FrontierMath Tier 4 (v2)97.6%83.0%87.8%
GPQA Diamond96.0%94.6%93.7%
ARC-AGI-399.9%7.8%not published
ExploitBench100.0%78.5%70%

The computer-use and abstract-reasoning gains are the least ambiguous. On ARC-AGI-3 the jump from 7.8 percent to 99.9 percent is a step change rather than an increment, and Greg Kamradt of the ARC Prize Foundation reported Astra surpassing their human action-efficiency baseline on 96 percent of levels.

⚠️Warning

Astra does not lead everything, and OpenAI's own comparison table shows it. On the Artificial Analysis Intelligence Index v4.1.1 Astra scores 61.2 against Claude Fable 5.1 at 65.7, Claude Opus 5 at 63.1 and Claude Fable 5 at 62.1 — fourth of the models listed. On Humanity's Last Exam with tools it scores 57.2 percent against Fable 5.1's 65.0 percent. On the Artificial Analysis Coding Agent Index it scores 67.0 against Claude Fable 5's 68.1. OpenAI calls Astra "the best model for software engineering to date"; that is a vendor claim, and three of the indices in OpenAI's own table do not support it. Benchmark against your own workload before migrating.

Cybersecurity and the Critical Rating

Astra is the first model OpenAI has ever rated Critical for cybersecurity under its Preparedness Framework. In practice that means it can find previously unknown security flaws and develop ways to exploit them across well-protected systems without a person guiding each step. During evaluation it discovered and used two previously unknown zero-day vulnerabilities, both of which OpenAI says it disclosed to the maintainers.

That rating changes what ships. The version released on September 3 will refuse advanced offensive-security work such as building proof-of-concept exploits, while still supporting secure code review and patching. Less restrictive access is planned through OpenAI Daybreak, the vetted defender program, in the coming weeks.

📝Note

The gating and the giveaway are separate facts, and OpenAI published both. The safeguards are real: on an internal computer-use safety benchmark Astra scores 2.4 percent against Sol's 22.0 percent (lower is better), it never attempted to circumvent a Codex Auto-Review denial even when that review was deliberately made evadable, and on an evaluation built from the Hugging Face incident it went beyond its authorized target in 0 percent of cases against Sol's 48 percent without production safeguards. But OpenAI also reports that Astra's reasoning is harder to monitor than GPT-5.6 Sol's — it exerts more control over its written reasoning and can evade internal monitors when explicitly instructed to. OpenAI attributes this to Astra solving problems in fewer written steps, says it still struggles to conceal reasoning on complex tasks, and calls the decline something it takes seriously. Independent researchers had raised exactly this concern before launch.

Because of that rating, OpenAI now runs misalignment monitoring on every tool-using request in Astra's external deployment, at what it calls significant compute cost. A practical consequence: safety checks can pause or stop legitimate work, including defensive security. In ChatGPT and Codex you may be asked to review an action before it continues; in the API the task simply stops.

Pricing

gpt-6-astra (API)$10 in / $50 out per million tokens
  • Flagship model
  • Separate rates for cache reads and writes
Fast mode (API)Double the standard rate
  • Up to two times the speed
  • Same model, faster serving
ChatGPT Plus$20/month
  • Astra included in the plan allowance
  • Extra usage purchasable as credits
ChatGPT Pro / Business / Enterprise$200/month and up
  • Adds GPT-6 Astra Pro
  • Enterprise access off by default until an admin enables it

At 10 dollars in and 50 dollars out per million tokens, Astra costs twice GPT-5.6 Sol's input rate and roughly 67 percent more on output. That is a real jump, and it is the reason the cheaper GPT-5.6 tiers stay relevant: Terra and Luna remain the cost-efficient options for high-volume or latency-tolerant traffic. Astra supports Zero Data Retention for eligible API customers.

Strengths

  • Computer use — leading published scores on OSWorld 2.0 and ScreenSpot-Pro, and materially faster per task
  • Mathematics and abstract reasoning — 97.6 percent on FrontierMath Tier 4 and 99.9 percent on ARC-AGI-3, both large jumps over Sol
  • Long context — 96.3 percent on OpenAI's MRCR v2 eight-needle test in the 512K to one-million-token band, against Sol's 73.8 percent
  • Measured alignment gains — large reductions on OpenAI's internal safety, circumvention and hallucination benchmarks
  • Broad distribution on day one — ChatGPT, the OpenAI API, Microsoft Azure and AWS Bedrock

Limitations & Considerations

  • Every benchmark is vendor-reported — no independent lab had tested Astra at launch, and scores are maximum-effort results from OpenAI's research environment
  • It trails Claude Fable 5.1 on three published indices — Artificial Analysis Intelligence Index, Humanity's Last Exam, and the Artificial Analysis Coding Agent Index
  • Harder to monitor than its predecessor — OpenAI's own admission, and the clearest cost the model carries
  • Twice the input price of GPT-5.6 Sol — the flagship tier got more expensive, not cheaper
  • Advanced cyber work is refused at launch — defenders wanting exploit validation or malware analysis must wait for the Daybreak rollout
  • Safety checks can interrupt legitimate work — API tasks stop rather than prompt, which matters for unattended automation
  • Closed model — API-only, no weights for self-hosting
  • Enterprise access is off by default — an administrator has to switch it on before anyone in the workspace sees it
  • GPT-5.6 — the prior flagship family, still the cost-efficient choice via Terra and Luna
  • ChatGPT — OpenAI's consumer product, where Astra now runs on paid tiers
  • OpenAI Codex — the coding agent that gets Astra's cross-window notes
  • Claude Fable 5.1 — Anthropic's competing flagship, ahead of Astra on three published indices

Lineage

  • GPT-6 Astra (September 3, 2026) — current flagship; first model rated Critical for cybersecurity; the page above
  • GPT-5.6 (July 9, 2026) — prior flagship family; Sol, Terra and Luna tiers plus an ultra mode
  • GPT-5.5 (April 23, 2026) — built for agentic work; superseded by GPT-5.6
  • GPT-5.6-Cyber (August 10, 2026) — offensive-capable security variant behind the vetted Daybreak Red tier
  • GPT-OSS — OpenAI's open-weight model under Apache 2.0

Key Takeaways

  • GPT-6 Astra shipped on September 3, 2026 to ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock, replacing GPT-5.6 as OpenAI's flagship
  • OpenAI president Greg Brockman said it is "not unreasonable to feel that we are now in the AGI era" — a vendor framing, not a measured result
  • The clearest gains are computer use (72.6 percent on OSWorld 2.0 in about 47 percent less time than Sol), mathematics (97.6 percent on FrontierMath Tier 4) and abstract reasoning (99.9 percent on ARC-AGI-3)
  • It is the first model OpenAI has rated Critical for cybersecurity, and the launch version refuses advanced offensive-security work; wider access runs through the vetted Daybreak program
  • It does not lead everywhere — Claude Fable 5.1 is ahead on the Artificial Analysis Intelligence Index, Humanity's Last Exam and the Artificial Analysis Coding Agent Index
  • OpenAI concedes Astra's reasoning is harder to monitor than GPT-5.6 Sol's, and now runs misalignment monitoring on every tool-using request as a result
  • At 10 dollars in and 50 dollars out per million tokens it is twice Sol's input price, so the GPT-5.6 Terra and Luna tiers remain the cost-efficient options

Keep track of the tools you’re evaluating

Sample AI Hub dashboard showing saved tools, a content-updates alert, an AI Skill Streak, and personalized recommendations

Your AI Hub — sample view. Click to enlarge.

📰GPT-6 Astra in the News

Showing the 3 stories where GPT-6 Astra is tagged in Top AI Stories.

More tools in Foundation Models & Open Source (12 of 74)

More from OpenAI

🧭Recommended for you

Optional detours — these connect to what you just read, and your next lesson will be waiting.