OpenAI Launches GPT-6 Astra

OpenAI Launches GPT-6 Astra: Price, ChatGPT Rollout and 99.9% ARC-AGI-3 Score

SAN FRANCISCO — OpenAI has launched GPT-6 Astra, a new frontier AI model built to do more than answer prompts. Astra is designed to operate computers, browse the web, write and test code, work across professional software and complete longer tasks with less supervision — while posting a headline 99.9% result on one ARC-AGI-3 evaluation.

The September 3 launch matters to ChatGPT users, developers and businesses because OpenAI is pushing its flagship technology deeper into agent-style work. Astra is rolling out in stages across eligible ChatGPT products, the OpenAI API and enterprise deployments, meaning availability may not appear for everyone at the same time.

GPT-6 Astra price and ChatGPT rollout

OpenAI says Astra usage is included within applicable subscription allowances, while additional usage can be purchased. Eligible Pro, Business and Enterprise customers also get access to a higher-capability GPT-6 Astra Pro offering, and Enterprise administrators can control availability for their workspaces.

For developers, the API model is gpt-6-astra. Standard pricing is $10 per million input tokens and $50 per million output tokens, with separate cache rates. Fast processing is available at up to twice the Standard speed for twice the Standard price.

According to OpenAI’s GPT-6 Astra documentation, the model has a roughly 1.05-million-token context window and supports up to 128,000 output tokens, making it suitable for large codebases, documents and long-running workflows.

What can GPT-6 Astra do?

Computer use is one of Astra’s biggest upgrades. OpenAI says the model can conduct online research, fill out forms, organize calendars, work with spreadsheets and documents, analyze scientific data, generate plots, build websites and test software interfaces.

OpenAI has demonstrated Astra using specialized applications including engineering and design software. In Codex, the model can also preserve useful notes across context windows and retrieve earlier requirements or test results during lengthy development projects.

The shift builds on the growing use of AI for complete workflows rather than isolated answers. Practical ChatGPT workflows for real-world tasks have increasingly involved research, planning and multi-step assignments where maintaining context is essential.

What does the 99.9% ARC-AGI-3 score mean?

Astra’s most eye-catching result is 99.9% on ARC-AGI-3, a benchmark that asks AI agents to explore unfamiliar environments, discover rules, identify goals and execute solutions.

There is an important qualification. The 99.9% result was achieved at high reasoning effort using ARC Prize’s Provider Adapter harness, which can take advantage of provider-specific context management. Astra’s best result using ARC Prize’s provider-neutral Standard harness was 62.7%.

In Provider Adapter testing, ARC Prize also found that Astra used fewer actions than its median human baseline on 96% of levels it completed and averaged 51.7% fewer actions per level.

That does not mean GPT-6 Astra has proved it is AGI. ARC Prize explicitly cautions that ARC-AGI-3 uses bounded, deterministic environments and that saturating the benchmark is not proof of artificial general intelligence.

GPT-6 Astra vs GPT-5.6 Sol

OpenAI reports substantial improvements in several practical evaluations. Astra scored 57.9% on Terminal-Bench 4.0 compared with 37.3% for GPT-5.6 Sol. On OSWorld 2.0, Astra reached 72.6% against Sol’s 65.7%.

Science results include 97.6% on FrontierMath Tier 4 and 96% on GPQA Diamond. Astra also reached 59.3% on Agents’ Last Exam, an evaluation covering complex professional tasks performed inside real software.

Those gains arrive as professionals increasingly depend on competing AI tools for coding and everyday work. Recent disruption involving ChatGPT, Claude, Grok and Cursor services illustrated how closely AI platforms are becoming tied to developer and workplace workflows.

Why Astra’s cybersecurity power matters

Cybersecurity is the most sensitive part of the release. OpenAI says Astra reaches the Critical cybersecurity capability threshold under its Preparedness Framework.

Without normal production safeguards, Astra scored 100% on ExploitBench, compared with 78.5% for GPT-5.6 Sol. OpenAI also says Astra discovered and used two previously unknown zero-day vulnerabilities during controlled evaluation and is disclosing them to their maintainers.

OpenAI is therefore restricting advanced offensive uses. Astra will refuse certain requests such as developing proof-of-concept exploits while still supporting defensive tasks including secure code review and patching.

More capable — but not without new risks

OpenAI says Astra is better at respecting authorization boundaries and avoiding unintended actions. However, the company also disclosed that Astra’s written reasoning proved harder to monitor than GPT-5.6 Sol’s in tests specifically designed to examine monitoring evasion.

That combination explains why this launch is more consequential than another benchmark race. Astra pairs stronger reasoning with the ability to act inside real software. The question now is whether those improvements remain dependable when millions of users begin delegating longer, more consequential tasks to the model.

Add Swikblog as a preferred source on Google

Make Swikblog your go-to source on Google for reliable updates, smart insights, and daily trends.

Get the Swikblog App
Stay updated with breaking news, trending stories and the latest updates—all in one place.
Get the Swikblog app on Google Play
Free download for Android devices

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *