GPT-6 Astra — often searched as Astra AI or AstraAI — is OpenAI’s new flagship model. OpenAI calls it the world’s most intelligent and aligned model it has broadly deployed: built for the hardest end-to-end work, not just chat.
The company announced it in GPT-6 Astra: A new generation of intelligence. This KnowAnt overview translates that launch into what the model actually does, where it is strong, what “Critical” cybersecurity means, and how you get it in ChatGPT or the API.
At KnowAnt, we treat a model launch like any other writing job: names, numbers, and limits in one place. If you still live in ChatGPT prompts rather than agents, our ChatGPT prompt library is the everyday companion. Rankings on public arenas are a different signal — see LMArena.ai explained.
Table of Contents
- What Is GPT-6 Astra?
- Computer Use: The Headline Capability
- Professional Work: Docs, Slides, CAD, and Sites
- Coding and Codex
- Science and Mathematics
- Cybersecurity: Critical Threshold and Daybreak
- Alignment and Safeguards
- Benchmarks at a Glance
- Availability, Pricing, and API Specs
- GPT-6 Astra vs GPT-5.6 Sol
- Frequently Asked Questions
- The Bigger Picture
- Sources
What Is GPT-6 Astra?
Astra is the GPT-6 generation’s flagship. OpenAI says it combines years of work on pre-training, reinforcement learning, and alignment. The pitch is not a slightly better chatbot. It is a model that can use a computer, browse, write software, do science, and produce professional artifacts (documents, spreadsheets, decks) with better judgment than GPT-5.6 Sol.
OpenAI’s own one-liner: Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work.
Two names you will see:
- GPT-6 Astra — the product and research name
gpt-6-astra— the API model ID
Astra Pro is a higher-access variant for ChatGPT Pro, Business, and Enterprise. It is not a separate research paper; it is the same family with a different plan gate.
Knowledge cutoff on the API card is April 30, 2026. Context window is 1,050,000 tokens, with up to 128,000 output tokens. Reasoning effort can be set to low, medium, high, xhigh, or max.
Computer Use: The Headline Capability
OpenAI’s lead claim is that Astra is the best computer-use model it has shipped. In practice that means: look at a screen (or a browser), plan a multi-step job, and drive software the way a person would — forms, CRMs, calendars, research, plots, a website, frontend QA, installing and testing software, troubleshooting what is on screen.
Greg Brockman, OpenAI’s president, put it bluntly before launch: Astra can really do anything a human can do with a computer. The launch post is more concrete: fill forms, update CRM records, organize a calendar, research and draft into email or a doc, analyze scientific data, generate plots, build a site and check that it works.
OpenAI’s demos include PCB layout in KiCad (schematic to a routable board) and a house modeled in Blender, then opened as a walkable scene in Unreal Engine 5. The point is not the brand names. The point is long-horizon work in real apps, not a single chat reply.
On Agents’ Last Exam (complex professional tasks in real software), OpenAI reports Astra at 59.3%, ahead of Claude Opus 5 at 55.5% and GPT-5.6 Sol at 53.6%, while using far fewer output tokens than Opus 5 at those settings. On OSWorld 2.0 (offline subset), Astra scores 72.6% at about 40 minutes per task versus Sol’s 65.7% at about 75 minutes — roughly 47% less time.
OpenAI also updated the Codex harness for computer use. Combined with Astra’s efficiency, it reports 1.9× faster completion than the current Sol experience on Mind2Web.
Important limit for builders: the model proposes actions. Your app still owns the sandbox, permissions, and what actually runs. OpenAI’s computer-use path for Astra leans on code execution (for example Playwright-style loops) with a structured computer tool as an alternative. Neither path is “give the model the user’s laptop.”
Professional Work: Docs, Slides, CAD, and Sites
Astra is trained to stick to templates, keep slides short, and pull only the context that belongs in the artifact. OpenAI’s example is a deck about a fictional model, GPT-Gaia, built from a few slides of OpenAI’s own template — layout and tone held.
BenchCAD (reconstruct 3D objects from multi-view renders by generating CAD code): Astra 95.9% geometric overlap vs Sol 83.3% and Claude Fable 5.1 84.3%, at lower estimated API cost in the configurations shown.
In ChatGPT, Sites lets Astra create, host, and share websites, web apps, and games from a prompt. Visual judgment is part of the pitch: games, apps, and renderings that look finished rather than placeholder.
When instructions are incomplete, Astra is supposed to fill routine gaps and ask when the answer would change the outcome. In Codex it can ask asynchronously and keep working on independent steps. If you do not reply, it proceeds on sensible assumptions for low-stakes choices and waits on consequential ones. Earlier models sometimes treated a steering message as a brand-new goal. Astra is described as keeping the original task while absorbing new constraints.
Partners quoted on the launch page include Cognition (Devin), Higgsfield AI, Harvey (legal), Jane Street, and Lovable — all claiming a jump over Sol on internal evals, not a public bake-off you can rerun at home.
Coding and Codex
OpenAI calls Astra its best software-engineering model to date. On Terminal-Bench 4.0, it reports 57.9% versus Sol 37.3% and Claude Fable 5.1 55.8%, at lower estimated cost per task than those two in the comparison.
A Codex-specific change: notes across context windows. Compaction used to squash a long debug session into a summary and lose why a fix failed. Astra can keep searchable notes and look back at earlier messages and tool outputs. The feature is experimental in config.toml and is planned as the default for Astra.
Lovable’s CTO described effort levels as buying more iterations, more browser verification, and a lean toward code execution over apply-patch. That is a product hint: higher reasoning.effort is not just “think longer,” it is “try more, check more.”
Science and Mathematics
OpenAI says Astra saturates FrontierMath Tier 4 at about 98% (table: 97.6% vs Sol 83.0%) and ARC-AGI-3 at 99.9% (Sol 7.8%). Greg Kamradt of the ARC Prize Foundation said Astra beat their human action-efficiency baseline on 96% of ARC-AGI-3 levels.
On GPQA Diamond (graduate-level science), Astra is reported at 96.0%. Terminal-Bench Science 0.1 (scientific workflows in a terminal) is 64.6% vs Fable 5.1 52.6% and Sol 22.4%.
The launch post also claims two number-theory results on prime gaps, with proofs and abridged chains of thought posted as PDFs: a bound of 186 on infinitely many close prime pairs (tightening a recent 240, after a long-standing 246), and an improvement to a term in a large-gap bound that OpenAI says had been unchanged for more than 80 years. Treat those as OpenAI’s research claims; the source post links the write-ups.
Demos show Astra in scientific GUIs — sequencing quality, genetic variation — combining reasoning with computer use rather than answering a quiz in isolation.
Cybersecurity: Critical Threshold and Daybreak
This is the part of the launch that is not a feature list.
Astra is OpenAI’s first model designated at the Critical cybersecurity capability under its Preparedness Framework. In OpenAI’s words: with the right tools and access, it can find previously unknown flaws and develop ways to exploit them on many well-protected systems without a person guiding each step.
OpenAI reports (without production safeguards, in research settings):
- ExploitBench: 100% vs Sol 78.5%
- ExploitGym: 42.4% vs Sol 30.3%, with fewer output tokens
- A June–August 2026 ExploitBench slice (recent Chrome/V8 bugs): 39.0% vs Sol 11.5%; during that eval Astra also found two previously unknown issues, disclosed to maintainers
- SRE-Bench (reverse engineering without source): 88.0% first try, 99.2% within four tries, vs Sol 55.9% / 68.7%
Expert assessments (again, no production safeguards) described arbitrary code execution in hardened browsers and privilege-escalation work on hardened operating systems.
The shipping model is more locked down. Defenders can use it for work such as secure code review and patching. OpenAI says Astra will refuse more advanced cyber tasks, including creating proof-of-concept exploits. Broader defensive workflows (validation, malware analysis, detection engineering) are planned through OpenAI Daybreak, with less restrictive safeguards for approved users.
KnowAnt’s rule here is the same as OpenAI’s public line: this overview does not walk through exploit methods. The capability exists; access is gated; the product you get in ChatGPT is not an unrestricted red-team agent.
Alignment and Safeguards
OpenAI’s second headline after “most intelligent” is most aligned.
One eval, inspired by a Hugging Face incident, asks whether a model facing an impossible task will go beyond its authorized scope. Sol, without production safeguards, did so 48% of the time. Astra: 0%.
On an internal computer-use safety benchmark (lower is better), Astra is 2.4% vs Sol 22.0%. With Auto-review: 1.8% vs 4.3%. Astra never tried to circumvent a Codex Auto-review denial in the circumvention test, even when the denial was configured to be evadable.
Capability-hallucination (misleading claims about what it can do): 4.2% vs Sol 12.2%.
OpenAI is more cautious on monitorability. Written reasoning was harder to monitor than Sol’s when the model was asked to evade monitoring — attributed to shorter written traces on simpler tasks. Complex work still appears harder to hide. The system card is the technical dump.
Production extras include Auto-review, trajectory monitoring (including chain of thought internally), and misalignment monitoring classifiers that can stop unauthorized activity. Extra checks can pause or stop legitimate work, including defensive cyber. In ChatGPT or Codex you may be asked to confirm; in the API the task stops. OpenAI says it is iterating to cut false stops.
Eligible API customers can use Zero Data Retention. OpenAI is also testing Private Safety Processing.
Benchmarks at a Glance
Figures below are OpenAI’s reported maxima at any effort, from the launch post. Production ChatGPT may differ (system prompts, tools).
| Area | Metric | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|
| Computer use | Agents’ Last Exam | 59.3% | 53.6% |
| Computer use | OSWorld 2.0 (offline) | 72.6% | 65.7% |
| Computer use | ScreenSpot-Pro (no tools) | 92.7% | 76.9% |
| Professional | AutomationBench | 41.4% | 18.1% |
| Professional | BenchCAD | 95.9% | 83.3% |
| Coding | Terminal-Bench 4.0 | 57.9% | 37.3% |
| Science | Terminal-Bench Science 0.1 | 64.6% | 22.4% |
| Math | FrontierMath Tier 4 | 97.6% | 83.0% |
| Science | GPQA Diamond | 96.0% | 94.6% |
| Abstract | ARC-AGI-3 | 99.9% | 7.8% |
| Cyber (research, no prod safeguards) | ExploitBench | 100% | 78.5% |
| Alignment (lower better) | Scope overflow (impossible task) | 0% | 48% |
| Long context | MRCR v2 8-needle 512K–1M | 96.3% | 73.8% |
Artificial Analysis Intelligence Index v4.1.1 is one place Astra is not first in the table OpenAI published: 61.2 vs Claude Fable 5.1 65.7. Read that as “flagship does not win every composite.” Computer use, ARC-AGI-3, and cyber are where the step change is.
Availability, Pricing, and API Specs
Rollout. Limited organizations first, then ChatGPT Plus, Pro, Business, and Enterprise, plus the OpenAI API, Microsoft Azure, and AWS Bedrock. Usage sits inside existing subscription allowances, with extra credits for sale. Enterprise admins must turn Astra on; it is off by default. ChatGPT Free is not in the named rollout list.
API. Model id gpt-6-astra. Standard: $10 / million input, $50 / million output. Cached input $1; cache writes $12.50. Prompts over 272K input bill at 2× input/cache and 1.5× output for the full request. Batch and Flex: 50% of Standard. Fast mode: up to 2× speed at 2× Standard price. Free API tier: not supported.
Modalities (API card). Text and image in; text out. Tools on the Responses API include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Fine-tuning: not supported.
Reasoning. reasoning.effort: low, medium, high, xhigh, max.
GPT-6 Astra vs GPT-5.6 Sol
| GPT-5.6 Sol | GPT-6 Astra | |
|---|---|---|
| Role | Previous frontier in ChatGPT / Codex / API | New flagship |
| Computer use | Strong, slower on OSWorld | Faster, higher scores; 1.9× Codex harness claim on Mind2Web |
| Coding | Solid agentic coding | Higher Terminal-Bench; searchable long-session notes |
| Alignment / scope | 48% overflow on impossible-task eval (no prod safeguards) | 0% on that eval; fewer capability hallucinations |
| Cyber designation | Previous frontier cyber-capable model | First Critical under Preparedness Framework |
| Cost (API Standard) | Check current Sol card | $10 in / $50 out per 1M |
If your work is short Q&A, Sol or a cheaper model may still be the default. If the job is “use the computer for an hour and come back with a file,” Astra is the model OpenAI wants you to pick.
Frequently Asked Questions
What is Astra AI / AstraAI?
It is OpenAI’s GPT-6 Astra — the GPT-6 flagship. “Astra AI” is how people search it. The API name is gpt-6-astra.
Is GPT-6 Astra available in ChatGPT Free?
OpenAI’s launch post names Plus, Pro, Business, and Enterprise — not Free. Check the model picker; rollout is staged.
How do I call GPT-6 Astra in the API?
Use gpt-6-astra on a paid project. Computer-use tools belong on the Responses API, not as a casual Chat Completions afterthought.
What is GPT-6 Astra Pro?
Access on Pro, Business, and Enterprise plans. Enterprise workspaces start with Astra disabled until an admin enables it.
Why is cybersecurity such a big part of the announcement?
Because OpenAI classified Astra at Critical cyber capability. The public product refuses advanced exploit work; Daybreak is the path for approved defensive use.
Does Astra replace writing prompts?
No. Better models still need a clear task. For ChatGPT-style work, start with a tight prompt — 200+ ChatGPT prompts — then let Astra take the long computer-use jobs.
The Bigger Picture
GPT-6 Astra is OpenAI arguing that the next leap is agency on a computer, not another two points on a multiple-choice science test (though it took those too). The same launch that boasts ARC-AGI-3 and prime-gap proofs also admits the model is in a class that can find zero-days, so Daybreak, Auto-review, and pauses exist.
Read the source post for the charts and footnotes. Use this page as the map: what Astra is, what it is for, what it will refuse, and what it costs.
When you write about or with Astra, keep the claims tied to OpenAI’s numbers and keep the cyber discussion at the policy layer. Draft that overview — or the emails and docs Astra is supposed to produce — on KnowAnt.
Sources
- GPT-6 Astra: A new generation of intelligence — OpenAI
- GPT-6 Astra model card (API) — OpenAI Developers
- Path to Astra: critical capabilities and frontier safeguards — OpenAI
- Safety overview: GPT-6 Astra — OpenAI
- GPT-6 Astra system card — OpenAI