Features, Benchmarks, Pricing, and What’s New

0
1
Features, Benchmarks, Pricing, and What’s New


OpenAI has released GPT-6 Astra, its newest frontier model, less than a week after Anthropic’s Claude Fable 5.1. OpenAI calls Astra the world’s most intelligent and aligned model yet.

What actually makes Astra different?

The simplest way to put it is this: Astra is built to do more, not just answer more. It can use a computer to complete tasks instead of telling you how to do them. It can create finished documents instead of giving you a rough first draft. During long coding sessions, Astra can also remember what happened earlier instead of starting from scratch. And it’s better at knowing when to take action and when to stop and ask.

The bigger shift isn’t just smarter answers. AI is getting closer to actually getting the work done.

Some important caveats remain, though. Astra doesn’t beat every competing model, and some of its biggest benchmark numbers come with asterisks. Let’s look at what’s new, what Astra does well, and where the caveats are.

What Actually Changed in GPT-6 Astra

Drive a Computer End-to-end

GPT-6 Astra can drive a computer end-to-end

Astra can fill out forms, update CRM records, run frontend QA checks on a website, and troubleshoot software by watching what happens on screen, without step-by-step hand-holding. OpenAI reports a score of 72.6% on OSWorld 2.0, a benchmark for real desktop computer use, narrowly ahead of Claude Opus 5’s 70.2% and well ahead of GPT-5.6 Sol’s 65.7%.

The more interesting number sits next to the accuracy score: OpenAI reports Astra completes these tasks in about 40 minutes on average, versus roughly 75 minutes for Sol. A 2-point accuracy gain is nice to have. Cutting the time to finish a task nearly in half is the more commercially meaningful claim, because in agentic work, time is cost.

Knows When to Ask and When to Guess

GPT Astra can drive a computer end-to-end

Earlier models tended to either guess wrong on ambiguous instructions or interrupt with unnecessary questions. Astra is trained to fill in routine gaps on its own and pause only when the answer would meaningfully change the outcome.

OpenAI’s own side-by-side demo shows the difference: GPT-5.6 Sol built a personal career website on its own in about 13 minutes. Astra paused after 20 seconds to ask what career the user was actually moving into. That’s a small moment, but it’s a good illustration of the difference between a model that acts and a model that uses judgment about when to act.

Produces Finished Documents

GPT 6 produces finished, on-brand documents, not drafts

There’s a meaningful difference between an AI that generates content and one that completes a deliverable. A generic model hands you text for a presentation. OpenAI says Astra is trained to match your existing templates, tone, and structure, and to pull in only the context that’s relevant rather than padding the output with everything it knows.

In one OpenAI demo, Astra builds a slide deck from just a handful of template slides while keeping the tone and layout consistent throughout. For teams that currently spend time reformatting AI output into a house style, that’s the part worth testing first.

Remembers Across Long Coding Sessions

In Codex, Astra can now keep searchable notes across context windows instead of repeatedly compressing long debugging sessions into a single summary, a change OpenAI says preserves details that compaction tends to lose, like why an earlier fix failed. It’s opt-in for now through Codex’s config file, and OpenAI says it will become the default in the coming weeks.

Cybersecurity

Astra in  Cybersecurity

Astra’s biggest jump isn’t a productivity feature at all. OpenAI reports Astra reaches the “Critical” threshold for cybersecurity under its own Preparedness Framework, its highest risk tier, meaning the model can independently identify and develop working exploits for previously unknown vulnerabilities. OpenAI reports a 100% score on ExploitBench and says Astra solved 88% of SRE-Bench reverse-engineering tasks on the first attempt.

Because of that risk tier, OpenAI is gating the more dangerous parts of this capability at launch. Astra will help with defensive work like secure code review and patch validation, but it refuses to create proof-of-concept exploits until OpenAI expands access through its Daybreak program.

GPT-6 Astra Benchmarks

The table below pulls together the headline comparisons OpenAI published against GPT-5.6 Sol, Claude, and Gemini. All figures are self-reported by OpenAI in its launch materials.

Benchmark GPT-6 Astra GPT-5.6 Sol Claude Fable 5.1 Claude Opus 5 Gemini 3.8 Flash
OSWorld 2.0 (computer use) 72.6% 65.7% 70.2%
FrontierMath Tier 4 97.6% 83.0% 87.8% 73.2%
GPQA Diamond 96.0% 94.6% 93.7% 93.7% 95.3%
Terminal-Bench 4.0 (coding) 57.7% 37.3% 55.8% 52.3% 19.1%
ExploitBench 100.0% 78.5% 70.0%
Humanity’s Last Exam (w/ tools) 57.2% 65.0% 63.6%

What this actually tells us:

  • Astra’s clearest strength is computer use. It leads Claude Opus 5 on OSWorld 2.0 and does it in meaningfully less time per task.
  • Its coding scores have improved substantially over its own predecessor, though the lead over Claude Fable 5.1 on Terminal-Bench 4.0 is narrow rather than decisive.
  • It is not universally ahead of Claude. On Humanity’s Last Exam with tools, Astra scores 57.2%, behind both Claude Fable 5.1 (65.0%) and Claude Opus 5 (63.6%).
  • Some headline numbers rely on evaluation setups that don’t reflect normal usage. Astra’s marketed 99.9% on ARC-AGI-3 depends on an expensive, stateful evaluation harness. Independent testing by the ARC Prize Foundation found that a standard, stateless API call scores far lower, somewhere between 17% and 63% depending on the reasoning tier used. Anyone calling the model through a normal API integration should expect the lower end, not the headline figure.

Also Read: GPT-5.6 Sol vs Claude Fable 5: Benchmarks, Pricing & Hands-On

GPT-6 Astra vs Claude Fable 5.1 Comparision

As neither model is publicly available yet, I turned to X to see what people with early access are sharing about their experiences. In the following section, I’ll highlight some of the most interesting examples and head-to-head videos I came across.

Building a 3D Villa from a Single Prompt

Developer Karan Kendre posted a head-to-head Blender comparison, asking Claude Fable 5.1 and GPT-6 Astra to build the same villa scene. Astra comes out ahead in the visual comparison, with a more polished and realistic-looking result, particularly in the interior details and overall presentation.

Source: KaranKendre on X

Designing a Travel App from the Same Prompt

In this task GPT-6 Astra and Claude Fable 5.1 were prompted to design an app based on the same prompt and goal. Fable went with a sky-themed design, while Astra took an astro-inspired approach. Astra comes out ahead here, with a more distinctive visual identity and stronger thematic consistency across the design, while Fable’s version feels cleaner but more conventional. It’s a good example of how two models can interpret the same creative brief in very different ways.

Source: jaimin on X

Video Prompt Generation

GPT-6 Astra and Claude Fable 5.1 were each used to generate a prompt for the same video concept, which was then created using Higgsfield AI. Astra comes out ahead here, with a more coherent sequence and stronger visual storytelling. Its version maintains the monk, giant koi, and movement more consistently across shots, while Fable’s version feels more dramatic but has some less consistent elements toward the end. It’s a good example of how the quality of the prompt can influence the final AI-generated video.

Source: HiggsField AI on X

Cost of GPT Astra

GPT-6 Astra is rolling out in stages: starting with a limited set of organizations, then expanding to all ChatGPT Plus, Pro, Business, and Enterprise users over the following days. Enterprise admins need to manually enable it for their workspace, since it’s off by default at launch. Pro, Business, and Enterprise users also get access to a separate GPT-6 Astra Pro tier.

For developers, Astra is available as gpt-6-astra through the OpenAI API, Microsoft Azure, and Amazon Bedrock.

API pricing:

  • Input: $10 per million tokens
  • Output: $50 per million tokens
  • Fast mode: roughly 2.5x the speed of standard processing, at 2x the price

That’s well above GPT-5.6 Terra’s $2/$12 rates and Claude Opus 5’s $5/$25 rates. The pricing itself is a signal: this is built for high-value autonomous work, not routine bulk-text generation. Teams paying $10/$50 per million tokens are likely doing so because a task gets finished reliably with minimal supervision, not because they need more words generated per dollar.

The model also supports zero data retention for eligible API customers, and OpenAI says it’s testing private safety processing to strengthen monitoring w hile preserving customer privacy.

Learn more about GPT-6 Astra here.

Conclusion

The real test for Astra won’t be whether it can top another benchmark. It will be whether companies can hand it a messy, multi-step task and trust it to get the job done with minimal supervision.

On paper, Astra looks like a significant step forward, particularly in computer use, coding, and agentic workflows. But the benchmarks also show that it doesn’t lead everywhere. Claude’s models still have an edge on broader reasoning benchmarks, which is why the right model ultimately depends on the work you need it to do.

The bigger question is what happens outside the leaderboard. Once Astra is widely available, real-world testing will show whether its strengths translate into reliable, everyday workflows. That’s the test that matters most, and one only hands-on use can answer.

Frequently Asked Questions

Q1. What is GPT-6 Astra?

A. GPT-6 Astra is OpenAI’s newest frontier model, succeeding GPT-5.6 Sol. OpenAI positions it around three strengths: computer use, producing finished professional work, and a major jump in cybersecurity capability.

Q2. How is GPT-6 Astra different from GPT-5.6 Sol?

A. OpenAI reports Astra outperforms Sol on most measured benchmarks, including a substantial cut in time per task on computer use, a large jump on math benchmarks, and a perfect self-reported score on ExploitBench. It’s also OpenAI’s first model to cross the Critical cybersecurity threshold under its Preparedness Framework.

Q3. How does GPT-6 Astra compare to Claude Opus 5 and Sonnet 5?

A. Astra reports a lead over Claude Opus 5 on computer use and a narrow one on coding, but trails Claude Fable 5.1 and Opus 5 on Humanity’s Last Exam with tools. It’s a strong model with clear specializations rather than one that wins across every metric.

Q4. Where can I access GPT-6 Astra?

A. It’s rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, along with the OpenAI API, Microsoft Azure, and Amazon Bedrock. Enterprise workspaces need to manually enable it, since it’s off by default at launch.

Q5. How much does GPT-6 Astra cost?

A. Standard API pricing is $10 per million input tokens and $50 per million output tokens, with fast mode available at roughly 2.5x speed for 2x the price. That’s notably more expensive than Claude Opus 5’s $5/$25 rates.

Q6. Is GPT-6 Astra safe to use?

A. OpenAI classifies Astra as reaching the Critical threshold for cybersecurity risk, so its most advanced exploit-creation capabilities are gated behind the company’s Daybreak program at launch. It ships with additional alignment training and safeguards, though OpenAI has flagged a regression in how easily its reasoning can be monitored, and lists that as an ongoing research priority.

Hello, I am Nitika, a tech-savvy Content Creator and Marketer. Creativity and learning new things come naturally to me. I have expertise in creating result-driven content strategies. I am well versed in SEO Management, Keyword Operations, Web Content Writing, Communication, Content Strategy, Editing, and Writing.

Login to continue reading and enjoy expert-curated content.