GPT-6 Astra Capabilities and Benchmarks :
GPT-6 Astra is OpenAI’s flagship model, released in early September 2026. It is the engine powering always-on agents like Dots and sits at the top of the company’s current lineup for complex, multi-step work.
Here’s what the model actually delivers on paper and in practice.
Core Capabilities
Astra was built for long-horizon tasks that previous models struggled to finish cleanly. OpenAI positions it as state-of-the-art in six areas:
- Computer use and browser navigation
- Software engineering and agentic coding
- Cybersecurity (first model to hit “Critical” under OpenAI’s Preparedness Framework)
- Advanced mathematics and scientific reasoning
- Professional knowledge work
- Abstract reasoning in novel environments
It supports a 1.05-million-token context window, image and file inputs, tool use (including computer-use tools), and multiple reasoning-effort settings (low through max). Fast mode is available at roughly 2× the standard price for higher throughput.
In real workflows the model can drive a desktop environment, write and debug code across large repositories, run scientific analysis pipelines, and handle multi-step professional tasks with fewer interruptions than earlier GPT-5.x and GPT-6 Sol releases. That same capability set is what makes it suitable as the backbone for persistent agents—exactly the reason it underpins how does OpenAI dots agent work after scrapping GPT-6.1 Astra over safety concerns.
Key Benchmark Results
OpenAI published extensive numbers at launch. Independent evaluators and secondary sources have largely confirmed the direction of the gains, with the usual caveats about harness differences on abstract-reasoning tests.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Notes / Context |
|---|---|---|---|
| OSWorld 2.0 (offline, partial) | 72.6% | 65.7% | ~40 min/task vs ~75 min (47% faster) |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | UI element targeting |
| Agents’ Last Exam | 59.3% | 53.6% | Professional software tasks |
| FrontierMath Tier 4 (v2) | 97.6–98% | 83.0% | Research-grade math; near saturation |
| GPQA Diamond | 96.0% | 94.6% | Graduate-level science |
| Terminal-Bench 4.0 | 57.7–57.9% | 37.3% | Shell / CLI agentic coding |
| DeepSWE v1.1 | 74.1% | ~72.7% | Complex software engineering |
| ExploitBench | 100% | 78.5% | Known vulnerability → working exploit |
| ARC-AGI-3 (provider adapter) | 99.9% | 7.8% | Stateful harness; standard harness lower |
| ARC-AGI-2 | 95.0% | 92.5% | — |
| BrowseComp | 91.5% | 90.4% | Web browsing competence |
| AutomationBench | 41.4% | 18.1% | Office / workflow automation |
Astra also recorded strong results on Terminal-Bench Science (64.6%), SRE-Bench, and several internal design and data-science task suites. On long-context retrieval (MRCR v2 at 512K–1M tokens) it scored 96.3%.
A few caveats matter for honest evaluation. The 99.9% ARC-AGI-3 figure uses OpenAI’s provider-adapter harness that preserves reasoning state across steps. Independent runs under the standard ARC Prize harness land substantially lower (around the low-to-mid 60s depending on reasoning effort). Cyber scores such as the perfect ExploitBench result were measured without full production safeguards in some evaluations; the shipped model includes additional refusal and monitoring layers.

Pricing and Access
API pricing sits at $10 per million input tokens and $50 per million output tokens (cached input is lower). Fast mode doubles the price for higher speed. The model is available in ChatGPT (Plus and above), the OpenAI API, Microsoft Azure, and AWS Bedrock. Enterprise access for the highest cyber capabilities remains restricted by default.
Alignment and Safety Improvements
OpenAI reports that Astra is its most aligned model to date. Internal evaluations showed fewer unintended outcomes on computer-use safety benchmarks and better adherence to scope boundaries than GPT-5.6 Sol. Hallucination rates on certain factuality tests dropped noticeably. These alignment gains are one reason the company felt comfortable shipping Dots on Astra even after canceling the more aggressive GPT-6.1 Astra update.
Practical Takeaways
GPT-6 Astra is strongest where previous models were weakest: sustained computer use, research-grade math, and multi-step professional work that requires tool use and judgment. It is not a universal winner on every independent intelligence index (some Anthropic models still edge it on certain composite scores), but the combination of computer-use speed, coding reliability, and alignment improvements makes it the current default for agentic systems.
If you are evaluating it for production agents or heavy research workflows, start with the OSWorld and Terminal-Bench numbers—they map most directly to real desktop and coding tasks. Pair those scores with the safety controls that ship around Dots, and you get a clearer picture of both capability and operational risk.
For the full picture of how this model is being used in always-on agents after the GPT-6.1 Astra cancellation, see the linked analysis of Dots.
3 FAQs
What are the top GPT-6 Astra capabilities and benchmarks?
GPT-6 Astra leads in computer use (72.6% on OSWorld 2.0), research math (97.6–98% on FrontierMath Tier 4), abstract reasoning (99.9% on ARC-AGI-3 with provider harness), and cybersecurity (100% on ExploitBench). It also scores 96% on GPQA Diamond and 74.1% on DeepSWE v1.1.
How does GPT-6 Astra compare to GPT-5.6 Sol on key benchmarks?
Astra improves significantly: OSWorld 2.0 jumps from 65.7% to 72.6% (and finishes tasks ~47% faster), FrontierMath Tier 4 rises from 83% to 97.6%, and ExploitBench goes from 78.5% to 100%. Alignment metrics also improved, with fewer unintended outcomes on safety tests.
Is GPT-6 Astra available for API use and what does it cost?
Yes. API pricing is $10 per million input tokens and $50 per million output tokens. It is available through the OpenAI API, Azure, AWS Bedrock, and ChatGPT (Plus and higher plans). Fast mode is offered at roughly double the standard rate.