GPT-6 Astra is about to be released to the public—find out what’s new with this powerful AI model.
GPT-6 Astra is rolling out today to a limited number of organizations and, over the coming days, will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock.

A new era begins
OpenAI is launching its most powerful AI model today, GPT-6 Astra, and its performance is as impressive as it is daunting. Based on the tests shown, there is a significant leap in performance and efficiency, meaning a more affordable price for a product of surprising quality.
“On ARC-AGI-3, Astra surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark. Not only is this the best model we’ve ever tested, but it also represents a meaningful step change in frontier-model performance—not only in its ability to navigate and solve novel environments, but also in how efficiently it learns to do so.”
Greg Kamradt, ARC Prize Foundation
A More Reliable Platform
In tests, GPT-6 also demonstrates greater confidence in performing tasks within the predetermined scope compared to GPT-5.6 SOL, the new model showed that it performs tasks while respecting the authorized limits in 100% of cases—a major leap compared to its predecessors, which exceeded the established limits in 48% of the tested cases.

A Faster and More Accurate Model
GPT-6 Astra can perform tasks more quickly and accurately—tasks such as filling out forms, organizing calendars, analyzing scientific documents, creating websites, conducting quality tests, running and testing programs, and much more—all at a much lower cost to the consumer. In mind2web’s tests, Astra showed a 1.9x increase in speed when performing tasks compared to the previous version, GPT-5.6 Sol.

Agents’ Last Exam tests agents on complex professional tasks in real software, from financial modeling to engineering and media production. GPT‑6 Astra reaches a new high in the comparison shown, scoring 59.3%, compared with 55.5% for Claude Opus 5 and 53.6% for GPT‑5.6 Sol. At these highest-scoring settings, Astra also uses approximately 65% fewer output tokens than Opus 5.Image Credit: OpenAI
Major Improvements for the Professional Environment
Astra also provides better visual judgment for building websites, games, applications, and rendered designs, resulting in outputs that more closely match what was specified in the prompt. With its advanced multitasking capabilities and ability to solve complex problems, it significantly enhances professional workflows. By making better use of the codex, Astra also offers improved “memory,” capable of keeping track of issues and messages from a previous context, thereby assisting with the resolution of long and complex problems.

Terminal-Bench 4.0 tests agents on complex terminal-based tasks, including software engineering, system configuration, and data analysis. GPT‑6 Astra reaches a new high at 57.9%, compared with 37.3% for GPT‑5.6 Sol2 and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task, respectively.
Image Credit: OpenAI
Greater security, but also greater vulnerability
With its improved ability to detect vulnerabilities in software and websites, GPT-6 presents a critical threshold in cybersecurity. While developers can use this capability to identify and resolve these vulnerabilities, it also creates a need for more robust and rapid security measures, since this capability can be used to exploit these weaknesses. OpenAI has been working to create safeguards against misuse of the model, but nothing is guaranteed yet, which is a major concern for companies and developers.

tested the model without production safeguards on ExploitBench, which evaluate whether models can turn known software vulnerabilities into working exploits. On ExploitBench, Astra achieved a perfect score of 100%, compared with 78.5% for GPT‑5.6 Sol.
Image Credit: OpenAI
“The story is: end of one era, start of another.”
Greg Burnham, EpochAI
