GPT-6 Astra Review: Is It Worth the Hype?
ChatGPT's GPT-6 Astra is strong at using apps and hard maths, but contested on coding and hungry on usage. What the launch claims and the testers found.

OpenAI launched GPT-6 Astra on 3 September 2026 and called it the start of the “AGI era”. A month on, the launch claims have been checked against what independent testers actually found, and the picture is more interesting than either the hype or the backlash. This review covers what ChatGPT’s new model is, what it is good at, where it falls short, what it costs, and whether it is worth switching on.
The short version: Astra is a real step forward at using software and at hard maths, and it is not the clean sweep the headline numbers suggest.
Key takeaways
- Astra’s real strength is operating apps. OpenAI reports 72.6% on OSWorld 2.0 in about 40 minutes per task, against 65.7% in about 75 minutes for the previous model, GPT-5.6 Sol.
- It is far ahead on hard maths: 97.6% on FrontierMath Tier 4 against 87.8% for Claude Fable 5.1.
- Coding is contested. Gemini 4 Argon scored 77.9% on DeepSWE against Astra’s 74.1%, and Claude Fable 5.1 is reported to lead coding by a modest margin.
- The famous ARC-AGI-3 score of about 99% depends on a special harness. Plain API calls reportedly score between 17% and 63%.
- It uses your allowance faster than the last model, and Plus users get a small number of messages per window.
- Creative writing is a weak spot, with one reviewer calling it “boring and visibly machine-generated”.
What GPT-6 Astra is
Astra is the successor to GPT-5.6 Sol and the new flagship behind ChatGPT. It is a large language model with a one-million-token context window and up to 128,000 tokens of output. What sets it apart is the job description. Earlier models were mostly judged on how well they answered. Astra is built to do things: fill in online forms, update records, organise calendars, work in spreadsheets, build and test websites, and produce documents and presentations. OpenAI stresses that it can work through applications that have no API, by reading the screen and clicking like a person would.
In practice that means a loop like this one, repeated until the job is done:
Look
Reads what is on the screen
Decide
Works out the next step
Act
Clicks, types or opens an app
Check
Confirms it worked, then repeats
It also arrived with an unusual safety footnote. Reporting around the launch said OpenAI had warned about its advanced cyber capabilities and gated the most sensitive uses behind an access programme called Daybreak. DataCamp notes that secure code review and patching are available, while exploit creation stays restricted.
What it does well
Using a computer. This is the headline. On OSWorld 2.0, a test of completing tasks in a real desktop environment, Astra beat its predecessor on both accuracy and speed. MindStudio’s analysis calls this the most practical gain, since a model that finishes in half the time is cheaper to run as an agent.
Hard maths and science. Astra scores 97.6% on FrontierMath Tier 4, which DataCamp describes as close to saturation, and 96% on GPQA Diamond. Epoch AI’s independent maths runs are cited as putting it first.
Staying on task. OpenAI says it is significantly better than Sol at keeping hold of context during long work, and that it asks clarifying questions only when the answer would materially change the result. In OpenAI’s own internal tests without safeguards, Sol went beyond its authorised scope 48.2% of the time and Astra 0%. That is OpenAI’s measurement of its own model, so treat it as a claim rather than a finding.
Where it falls short
The benchmark story is messier than the launch. VentureBeat noted that OpenAI did not publish results on GDPval, its own benchmark for real-world work, which is a strange omission for a model launched on the strength of professional work. Artificial Analysis, an independent aggregator, ranked Astra only around fifth overall at launch, even though testers said it felt like a much bigger leap in practice. The two sources that quote its exact scores disagree with each other, so the ranking is more reliable than the numbers.
The ARC-AGI-3 caveat. The roughly 99% score was achieved using OpenAI’s own stateful harness, which can give the model more help than a bare call. DataCamp warns that people using the API should not expect that figure out of the box. VentureBeat adds that NVIDIA reached 100% on the same test with a Claude model inside its own agent system, which suggests the score measures the whole setup as much as the model.
Humanity’s Last Exam. Astra scores 57.2% against 65.0% for Claude Fable 5.1, so it does not lead everywhere on reasoning.
Writing. One reviewer reported that its creative writing was a step backwards. Anyone who wants a model mainly for prose should test it before trusting the launch.
Transparency and certification. Layer3 Labs points out that Astra’s reasoning is less interpretable by design, and that no HIPAA or SOC 2 certifications were published at launch.
Hallucinations remain possible. Nothing about the launch removes the need to check output. A confident wrong answer, what people call a hallucination, is still a risk, and agent-style work raises the stakes because the model now acts instead of only suggesting.
Astra against its predecessor
The fairest comparison is with the model it replaces. GPT-5.6 Sol was OpenAI’s flagship until September, and most of Astra’s claimed gains are measured against it, which is also why they are the easiest to trust. A new model beating its own predecessor on the maker’s chosen tests is the least surprising result in the field.
The computer-use numbers are the clearest improvement. Astra finished OSWorld 2.0 tasks with 72.6% accuracy in about 40 minutes, where Sol managed 65.7% in about 75 minutes. That is roughly 47% less time per task and a gain in accuracy at the same time, which is unusual, because speed and accuracy normally pull against each other. MindStudio also reports a jump on the ARC-AGI-3 reasoning test, from 7.8% for GPT-5.6 to roughly 99% for Astra, but that is the figure with the harness caveat above, so it says more about the setup than about the model on its own.
Staying on track over long jobs is the other change OpenAI stresses. It says Astra is significantly better than Sol at maintaining context during complex work. DataCamp adds a detail about how: in Codex, OpenAI’s coding tool, Astra keeps notes in searchable records that carry across context windows, instead of squeezing the earlier work into a summary. That matters for any task long enough to overflow even a million-token window, because a summary loses detail and a note you can search does not.
The downsides arrive with the upgrades. Astra uses more of your allowance, costs more per task in Layer3 Labs’ week-one testing, and gives up some readability in its reasoning. OpenAI describes the architecture as using “recurrent depth”, which lets it think harder on difficult steps, but it also makes the reasoning harder for outsiders to follow. Whether that trade is worth it depends on whether you need the extra capability or the extra transparency.
The cyber question
Astra came with a warning attached. Around the launch, CNBC reported that the company flagged its advanced cyber capabilities and began rolling Astra out to a limited set of organisations before opening it to everyone. The most sensitive uses sit behind an access programme called Daybreak.
The numbers explain the caution. VentureBeat lists Astra at 100% on ExploitBench, a test of whether a model can work out how to exploit software weaknesses. DataCamp reports that the public version can do secure code review and patching, but that exploit creation stays gated. For an ordinary user this changes little day to day. It is still worth knowing about, because it is a signal that the company itself thinks the model has crossed a line that earlier ones had not, and it explains why access was staged instead of switched on for everyone at once.
It also sits oddly with the other safety claim. OpenAI says that in its own tests without safeguards Astra stayed within its authorised scope every time, where Sol went outside it nearly half the time. Both can be true, since a model can be very capable and well behaved, but both are OpenAI grading its own homework, and independent testing will say more over the coming months.
Price and access
Through the API, Astra is billed per million tokens, and output costs five times as much as input. Cached input is a tenth of the standard input rate. A fast mode runs about 2.5 times quicker for roughly double the rate. Layer3 Labs found its cost per task in the first week was about 2.5 times that of the prior version, so a single hard task can cost noticeably more. OpenAI publishes the current figures on its own pricing page. MindStudio points the other way: one tester said Claude Fable 5.1 consumed 40% of a weekly allowance on a single task, so the cheapest model on paper is not always the cheapest in use.
In ChatGPT there is no separate plan. Astra is included in existing limits, and OpenAI warns that it can use your allowance faster than GPT-5.6 Sol. Estimated Plus allowances are 5 to 45 messages per five-hour window, rising to 25 to 225 on the lower Pro plan and 100 to 900 on the higher one. Plus users reach it through Work and Codex first, with regular Chat arriving gradually, and free users are excluded. OpenAI also notes that even when Work runs locally, task context may still be stored in the cloud.
Microsoft has made Astra generally available in Foundry, with its standard
identity, encryption and access controls, and states that prompts and outputs
are not used to train the models.
Replit’s CTO described it as moving beyond code generation to active software creation, which is the kind of endorsement a launch partner is expected to give.
How it compares
Each chart below sets Astra against the rival that was reported on that test. The figures come from different sources and testing setups, so read them as a rough picture, not a league table.
Hard maths (FrontierMath Tier 4). Astra’s biggest lead over Claude.
Computer use (OSWorld 2.0). Astra against the model it replaces, and nearly twice as fast.
Coding (DeepSWE v1.1). Here Astra is beaten.
Humanity’s Last Exam. Claude comes out ahead on broad reasoning.
Who it suits
Astra suits people who give an AI long, fiddly jobs and want it to get on with them. Think of the sort of task that would otherwise mean an afternoon of clicking between a spreadsheet, a form and a calendar. This is agentic AI at its most literal: an AI agent that does the work in the same apps you use, not a chatbot that tells you how. It also suits anyone doing serious maths or science, where its lead is the widest and the independent testing agrees.
It suits people less if the main use is writing, because the early verdict on its prose is poor, or if the main use is coding and you already have a model you trust, because the coding results are a draw at best. It also suits casual users less than the launch suggests. A Plus subscriber gets a small number of messages per window, and a model that burns through them quickly is frustrating if you only want a quick answer.
Whatever you use it for, keep a person in the loop. An agent that clicks through real apps can make a real mistake, such as sending the wrong email or changing the wrong record, and the speed that makes Astra attractive is also what makes a mistake travel. A human in the loop who reads the result before it is final turns a risky tool into a useful one.
Verdict
Astra deserves some of the hype and not all of it. If you want a model that operates software, handles long multi-step jobs and does serious maths, it is the strongest option reported so far. If you mostly want writing, or you want the best coder, the evidence does not back it, and the independent overall ranking puts it behind Claude. The sensible way to judge it is to take one task you actually do, run it on Astra and on its rivals, and compare the results yourself. Treat any single benchmark, including the ones in this review, as a hint and not a verdict.
Further reading
- Claude vs ChatGPT: Which fits your business best?: how the two assistants differ on data, admin and access.
- Gemini vs ChatGPT for business: the other main rival, compared.
- How to check AI output before it goes to a customer: the review step no model removes.
Sources
- GPT-6 Astra: the next generation in intelligence for work (OpenAI): the launch announcement, capability claims and rollout.
- OpenAI begins rolling out Astra after warning of its advanced cyber capabilities (CNBC): the cyber warning and gated access.
- “Welcome to the AGI era”: OpenAI launches GPT-6 Astra (VentureBeat): benchmark scores, the missing GDPval result and the ARC-AGI-3 caveat.
- ChatGPT Astra is now rolling out to Plus subscribers (BleepingComputer): plan availability and limits.
- How to use GPT-6 Astra when it rolls out to you (Engadget): rollout by plan and message allowances.
- GPT-6 Astra: frontier intelligence for work, now generally available (Microsoft Azure blog): Foundry availability, governance features and customer quotes.
- GPT-6 Astra: features, benchmarks and pricing (DataCamp): benchmark detail, fast-mode pricing and caveats.
- GPT-6 Astra benchmarks: how it really compares (MindStudio): comparison against Claude and Gemini.
- GPT-6 Astra review (Layer3 Labs): week-one test results, cost per task and compliance gaps.
FAQ
Questions we get asked
What is GPT-6 Astra?
GPT-6 Astra is OpenAI's flagship model, launched on 3 September 2026 and available in ChatGPT, Codex and the API. OpenAI describes it as its most intelligent and aligned model, built for computer use, browsing, coding, science and longer pieces of work.
Is ChatGPT Astra Free?
No. Free and Go users are excluded. Plus subscribers can use it inside ChatGPT Work and Codex, with regular Chat rolling out gradually. Pro, Business and Enterprise users get it too, subject to their workspace settings.
How much does GPT-6 Astra cost?
In ChatGPT it is included in your existing plan limits, though OpenAI warns it can use your allowance faster than GPT-5.6 Sol. Through the API it is billed per million tokens, with output costing five times as much as input, and there is a fast mode that is about 2.5 times quicker at roughly twice the rate. Current figures are on OpenAI's own pricing page.
Is GPT-6 Astra better than Claude?
It depends on the task. Astra leads on hard maths and on computer use. On coding the results are mixed, with Claude Fable 5.1 and Gemini 4 Argon each ahead on some benchmarks, and on Artificial Analysis's overall ranking Astra sits behind Claude.
How many messages do I get with GPT-6 Astra?
OpenAI's estimate for Plus is 5 to 45 messages per five-hour window, rising to 25 to 225 on the lower Pro plan and 100 to 900 on the higher one. Astra can use your allowance faster than GPT-5.6 Sol, depending on how complex the task is, how much context it carries and how much reasoning it needs, so heavy jobs use up the window sooner.
Ready to put this to work?
Tell us where your team is with AI and we will tell you honestly what would make the biggest difference.

