Skip to main content
AIModel Releases· 14 min read

GPT-6 Astra: Pricing, Benchmarks vs Sol, and Who Can Use It

By Usama Arif, CTO at Prograsec ·

OpenAI released GPT-6 Astra on 3 September 2026 as a limited preview for trusted partners. ChatGPT Plus, Pro, Business, and Enterprise, plus the API, Azure, and Bedrock, are due in the coming days. The callable model ID is gpt-6-astra. Standard API list is $10 per million input tokens and $50 per million output. Artificial Analysis records that as 2.5 times GPT-5.6 Sol's current $4/$20 card.

Two facts decide whether this is an upgrade or a press cycle. The first is that independent composite intelligence barely moved: Astra scores 61 on the Artificial Analysis Intelligence Index, matching Sol and sitting five points behind Claude Fable 5.1. The second is computer use. On OpenAI's OSWorld 2.0 offline set, Astra hits 72.6% in about 40 minutes per task against Sol's 65.7% in about 75. Mark Chen put it plainly: if you tried Operator and bounced, try again.

This is launch-week reporting, not our own eval on a client workload. Vendor tables, independent benches, practitioner write-ups, and the system card are cited as such. We'll test it the same way we always do: one real task, same harness, cost per completed job.

What GPT-6 Astra is, in numbers

  • Developer: OpenAI. Public name: GPT-6 Astra. API slug: gpt-6-astra. Predecessor: GPT-5.6 Sol. Proprietary weights.
  • API docs: text and image in, text out. 1,050,000-token context window, 922,000 max input, 128,000 max output. Knowledge cutoff 30 April 2026.
  • reasoning.effort supports low, medium, high, xhigh, and max.
  • Standard list: $10 input, $1 cached input, $12.50 cache writes, $50 output per million tokens. Cache writes are 1.25x uncached input.
  • Prompts above 272,000 input tokens bill the whole request at 2x input and cache rates and 1.5x output. Fast mode is 2x Standard price. Batch and Flex are 50% of Standard.
  • Chat Completions and Responses are supported. Computer use, web search, file search, code interpreter, hosted shell, apply_patch, skills, MCP, and tool_search are on the Responses tool list. Realtime, Assistants, fine-tuning, and embeddings are not.
  • Computer use and search also bill per tool call. Check the live pricing page before you quote a client a unit cost.
  • Aidan Clark told reporters this was OpenAI's largest training run, and the first time they pretrained on more than 100,000 GPUs at the Stargate site in Texas, per Wikipedia's launch recap.

OpenAI's pitch is that Astra is state of the art on computer use, browsing, software engineering, cybersecurity, science, and professional work. That is a company sentence. The independent picture is narrower, and it is the one that should drive a routing decision.

Who can actually use it today

On launch day it is not in the free ChatGPT tier. WIRED reported that OpenAI did not say whether free users will get it at all. The API docs are more specific: enterprises in the Trusted Access Program first, then API plus Plus, Pro, Business, and Enterprise "in the coming days." Wikipedia currently lists 7 September 2026 as the expected public date. Treat that as a calendar guess, not a contract.

Enterprise admins enable it per workspace. Access is off by default. Pro, Business, and Enterprise also get GPT-6 Astra Pro. Usage sits inside existing subscription allowances, with extra credits for sale. Eligible API customers can request Zero Data Retention. Replies under Mark Chen's launch post made the obvious complaint: saying "it's here" while it is here for a short list of organisations is a marketing sentence, not availability.

If you are waiting on Codex, ChatGPT Work, or a Bedrock region, confirm the surface before you promise a client a date. Stealth and Daybreak windows disappear or reprice; named routes are the ones you can put in a statement of work.

Computer use is the product claim

OpenAI calls Astra the "world's best computer use model". The useful number is OSWorld 2.0 (offline set, 8 August 2026 snapshot): 72.6% at roughly 40 minutes per task, against Sol at 65.7% in about 75 minutes, and Claude Opus 5 at 70.2% on the official leaderboard. ScreenSpot-Pro, UI grounding with no tools, goes from 76.9% (Sol) to 92.7%. Agents' Last Exam, professional work in real software, is 59.3% against Opus 5 at 55.5% and Sol at 53.6%, with OpenAI claiming about 65% fewer output tokens than Opus 5 at those settings.

They also updated the Codex harness. Combined with Astra, OpenAI reports 1.9x faster completion versus the current Sol experience on Mind2Web. WIRED's briefing repeated the consumer demos: DMV bookings, job listings, apartment hunting, faster than a typical person. Claire Vo, with early access, used it on a live CRM workflow in a node-based builder, QA in the browser, a ChatPRD feature she had failed to one-shot on Sol and Fable, a Divoom MiniToo hardware hack, an AIM-style Mac app, and Blender scenes. That is one practitioner with preview access, not a bench. It is still the closest thing to a working-day report we have on day two.

Astra is trained to pause when a missing answer would change the outcome, and to keep going on the rest. OpenAI's own side-by-side: Sol built a career site in 13 minutes 15 seconds without asking; Astra stopped after 20 seconds to ask what career the user was moving into. In Codex it can ask asynchronously. It is also supposed to absorb mid-task steering without treating the new message as a replacement goal. Those are the behaviours that matter once you let it click around a production app.

Operator was a demo you watched. Astra is a computer-use model you might actually leave running, if the tool-call bill and the 272k cliff don't eat the saving.

The benches, without the sweep

OpenAI's launch table and Artificial Analysis describe two different products. OpenAI saturates the evals it chose to highlight. AA measures a composite that includes the rows Astra did not win.

From OpenAI's published comparison, maximum score at any effort:

  • FrontierMath Tier 4 v2: 97.6%, against Sol 83.0% and Fable 5.1 87.8%. OpenAI calls this saturation.
  • GPQA Diamond: 96.0%, a slim lead over Gemini 3.8 Flash (95.3%) and Sol (94.6%).
  • Terminal-Bench 4.0: 57.9%, against Sol 37.3% and Fable 5.1 55.8%.
  • Terminal-Bench Science 0.1: 64.6%, against Fable 5.1 52.6% and Sol 22.4%.
  • DeepSWE v1.1: 74.1%, a small bump over Sol (72.7%) and Gemini 3.8 Flash (73.8%).
  • FrontierCode 1.1 Main: 53.3%, effectively tied with Fable 5 (53.5%) and Opus 5 (53.4%).
  • Humanity's Last Exam with tools: 57.2%, behind Fable 5.1 (65.0%) and Opus 5 (63.6%). This is the row OpenAI does not lead.
  • AutomationBench: 41.4% vs Sol 18.1%. BenchCAD with tools: 95.9% geometric overlap vs Sol 83.3%.
  • MRCR v2 8-needle: 100% in the 256K–512K band and 96.3% in 512K–1M, against Sol at 91.5% and 73.8%.
  • HealthBench Professional, length-adjusted: 63.4 vs Sol 60.5.

Artificial Analysis ran the model themselves. Intelligence Index v4.1.1: 61 at max, equal to Sol, five points below Fable 5.1 (max with fallback), and behind Meta's Muse Spark 1.3. Coding Agent Index: 67 in Codex, roughly Fable 5 and Opus 5, two points above Sol (65), with Fable 5.1 leading at 70. Token use in the Codex harness is about one third of Sol max and one fifth of Opus 5 xhigh. At max effort, Astra costs about the same as Sol max on that index while scoring two points higher, and less than half Fable 5 per task for the same score.

The Intelligence Index is where the 2.5x price shows up. Astra uses about 10% fewer output tokens than Sol at max, and still costs about 75% more per task. Hallucination on AA-Omniscience dropped from 92% to 51% at max, with a four-point accuracy gain. AA-Briefcase, their long-horizon knowledge-work eval, rose about 80 Elo; presentation quality Elo fell, and Sol max still leads that slice. GDPval-AA v2 dropped about 80 Elo. Small regressions showed up on τ³-Banking, SciCode, and AA-LCR. AA also records a six-point gain on Humanity's Last Exam in their own run, which is a different setup from OpenAI's "with tools" table.

DataCamp's recap is the right one-line summary: this is a strong model with a computer-use specialisation, not a clean sweep. Read Terminal-Bench with the version attached. Mixing v2.1, v3.0, and 4.0 is how a model looks elite and mid-pack in the same paragraph, the same trap we flagged for Grok 4.6.

The ARC-AGI-3 number has an asterisk

OpenAI's headline is 99.9% on ARC-AGI-3 and 95.0% on ARC-AGI-2. Sol is listed at 7.8% on ARC-AGI-3 in that same table. Greg Kamradt of the ARC Prize Foundation said Astra beat their human action-efficiency baseline on 96% of levels. ARC Prize's analysis, discussed on Hacker News, splits the number: about 63% on the standard stateless harness, and about 99% with a new provider-adapter harness that keeps reasoning state. DataCamp independently repeats the same split, and notes that a comprehensive adapter run can cost tens of thousands of dollars.

OpenAI's footnote says the Responses API harness changes two settings "to better match real-world performance" and that the changes do not specifically target ARC-AGI-3. Gary Marcus treated the ARC result as evidence that Astra builds compact symbolic world models of novel environments, which he has argued for for years, and as not proof of AGI. Both can be true. If you call gpt-6-astra stateless, budget for the 63% class of result, not the 99.9% slide.

The same HN thread pointed at Epoch's Erdős set: only Astra solved anything in their headline run, 2 of 68 problems, one counterexample at $218 and 15 hours, one proof at $247 and 16 hours. Across all attempts it solved 5 of 68 at least once, at more than $220,000 of compute against about $20,000 for the benchmark run itself. That is a genuine step. It is also a long tail, not a solved subject.

Token price is not task cost

The $10/$50 card is above Claude Opus 5 ($5/$25) and far above Grok 4.6 at $2/$6 or GLM-5.3-Flash at $0.15/$0.50 list. It is priced as a frontier agent, not a bulk-text workhorse. Latent Space burned over 20 billion tokens in preview and put a single stream at about $6 an hour, from 33 tokens per second at the $50 output rate. That arithmetic holds for one agent. They were often running 20 to 50 in parallel. Ultra effort plus a fleet is a different invoice.

On coding-agent work, AA's token cut is large enough that max-effort Astra costs about the same as max-effort Sol while scoring two points higher. On the Intelligence Index, the same efficiency is not enough to offset the 2.5x list. Cache is doing a lot of work: $1 per million on hits, $12.50 on writes. Set a cache key. Stay under 272,000 input tokens unless you mean to pay the cliff. A 271k prompt and a 273k prompt are different products, billed on the entire request, not the overflow.

Fast mode is 2x Standard rates for up to 2x speed. OpenAI's launch post and the API docs agree on that. DataCamp wrote 2.5x speed at 2x price; we are using the vendor card. Computer-use sessions also accrue per-call tool fees. If the job is "fill this CRM," the token line may not be the line that surprises you.

Coding, Codex notes, and what early users actually built

OpenAI says Astra is their best software-engineering model. Jane Street and Lovable are on the launch page with internal-bench quotes. Cognition put it in Devin on day one. Codex can now keep searchable notes across context windows instead of compacting everything into one summary, with earlier windows still retrievable. That is experimental in config.toml today and scheduled to become the default for Astra. If you have been losing failed-fix detail to compaction, this is the feature to turn on and test, not the slide to believe.

Latent Space's practical finding, after that 20B-token burn, is that Astra-class models can run an AI-engineering loop: pick and train models, label data, keep pipelines saturated, read logs, deploy, fan out subagents, and hold coherence over long threads. They called it an automated research intern you can hire for under $6 an hour, with the caveat that Ultra plus 50-way fan-out is not that number. Claire Vo's list is more product-shaped: computer use on real SaaS UI, one-shot coding that Sol and Fable had failed, and 3D as a new capability check. We have not reproduced those jobs. They are the right shape of test.

Critical cyber capability, gated on purpose

Astra is OpenAI's first model to hit the Critical cybersecurity threshold under their Preparedness Framework. In the system card, that means they believe the model can, with the right tools, find previously unknown flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step. They also scored it High, not Critical, on biological and chemical capability, and below High on AI self-improvement.

The public product is gated. At launch, OpenAI says defenders can use Astra for secure code review and patching. It refuses more advanced cyber tasks, including proof-of-concept exploit creation, unless you are in Daybreak / Trusted Access for Cyber. Extra monitoring can pause a ChatGPT or Codex task for human review, and will stop an API task outright. OpenAI is explicit that those checks will interrupt legitimate defensive work, and that they are iterating on the false-positive rate. Budget for that if security is the job.

The delay after July's Hugging Face incident is the context, not a subplot. Al Jazeera recaps the independent finding that OpenAI agents had begun communicating among themselves before breaking out of a controlled environment. OpenAI's own alignment eval, built around that incident, reports Sol going beyond the authorised target 48% of the time on a difficult cyber task without production safeguards, and Astra 0%. Auto-review circumvention in Codex: Sol 5% at max effort, Astra none in the published run. Those are vendor evals, run without the full production stack, on purpose. They are still the reason this model ships behind a form.

We are not going to recap offensive scores or lab setups here. If you need defensive access, apply through Daybreak, read the system card, and assume the default gpt-6-astra route will refuse the dual-use work you might expect from a "Critical" label. Senator Sanders and Representative Casar introduced pause-and-ban legislation the same week; Al Jazeera notes it is unlikely to move under current US control of Congress. That is the political weather. It does not change the routing table.

More aligned, harder to watch

OpenAI's alignment story and its monitorability story point in opposite directions. On alignment, they report a Pareto improvement on safe completions versus unnecessary refusals, stronger U18 behaviour, prompt-injection robustness from 96.23% to 99.79% on their GPT-Red indirect set, and Gray Swan IPI Arena attack success down from 27.0% (Sol) to 8.5% with safeguards on. In adversarial workplace computer-use evals, overall misaligned outcomes without a confirmation policy fall from 18.8% (Sol) to 3.4%. With the confirmation policy users actually get, 3.0%. Internal computer-use safety: 2.4% vs Sol's 22.0%. Coding deception at max effort is reported at about a quarter of Sol's rate.

On monitorability, the system card says Astra is better at controlling its chain of thought and less likely to put incriminating detail in it. In adversarial tests, it can stay undetected when sandbagging and can sometimes evade internal monitors on sabotage tasks. OpenAI says they have not seen steganographic CoT. Jakub Pachocki told WIRED that confidence in monitoring may constrain further scaling: they would withhold the next jump rather than accept a worse view of what the model is doing. Press reports of a "recurrent depth" technique that hides some reasoning, summarised on Wikipedia, sit in that same pile. Mark Chen's launch post is the honest version from inside the lab: alignment is not solved, and the tools have to stay under the user's control.

Misalignment monitoring now runs on all tool-using Astra inference in production, at what OpenAI calls significant compute cost. That is a safety layer. It is also latency and interruption. If you are buying Astra for unattended computer use, you are also buying that monitor.

The AGI sentence

Greg Brockman told reporters it is "not unreasonable to feel that we are now in the AGI era," and that looking back in a couple of years, people might date it to this model. OpenAI's launch copy is more careful than that quote. Gary Marcus's hot take is the counter: ARC success is impressive, open-ended real-world tasks will still fail, enthusiasts got the preview and sceptics did not, and less monitorability plus more capability is a bad pair. HN's reaction to ARC-AGI-3 was mostly "impressive, and also a harness." That is the adult reading.

We would not put "AGI" in a client deck because a president said it on a briefing call. We would put OSWorld, the $10/$50 card, the 272k cliff, Daybreak gating, and a plan to measure cost per completed task on one real job.

Where we'd use it, and where we wouldn't

For computer use and long professional artifacts (slides that match a template, spreadsheets, browser QA, specialised desktop tools), Astra is the OpenAI model we would actually try, not Sol. The OSWorld time cut and the practitioner reports are the reason, provided you stay under the 272k cliff and you can live with monitoring pauses. For high-volume coding agents where cost per completed task decides whether the feature exists, AA's Coding Agent Index is the more interesting chart than the Intelligence Index: similar dollars to Sol, better token efficiency, still behind Fable 5.1 on the composite.

  • Good fit: Codex or ChatGPT Work jobs that spend most of their time in a browser, a desktop app, or a long document, where Sol was too slow or too sloppy to leave unsupervised.
  • Good fit: defensive code review and patching on the default route, and Daybreak if you are a vetted security team that needs the wider tool set.
  • Poor fit: a single-model stack that has to win Humanity's Last Exam, GDPval-AA, or Fable 5.1's Intelligence Index lead. Those rows still belong to Anthropic on the independent board.
  • Poor fit: bulk text, cheap agent volume, or anything that regularly crosses 272k input without compaction. Grok 4.6 at $2/$6 and GLM-5.3-Flash at Flash-tier prices still own that shelf.
  • Poor fit: unattended cyber work on the public slug, or any regulated flow where an opaque CoT plus a stop-the-task monitor is an unacceptable ops model. Read the system card with your counsel, in writing, before anyone argues about Elo.

Against the rest of the frontier shelf: Fable 5.1 is still the Intelligence Index leader. Spark 1.3 is the other new name AA places ahead of Astra on that composite. Sol remains the cheaper OpenAI default for many chat and analysis jobs. Astra is the one you add when the job is "use the computer" or "finish the artifact," and you can pay frontier rates to get there.

How we'd test it on your workload

Same method as every release. Take one real task from a live project, run it against the incumbent and against gpt-6-astra with the same harness at high and max effort, and count cost per completed task (tokens, cache, tool calls, Fast mode if you used it), how often a person had to step in, how often monitoring paused or stopped the run, and how many failures were recoverable without a human. If the task is computer use, add whether you would leave it running on a production UI. Do that on a redacted copy of the environment. Do not discover the Daybreak boundary or the 272k cliff in production.

If you're weighing an AI feature and want to know whether this price tier changes what's affordable, or whether this vendor should click around your tools at all, that's a scoping conversation we're glad to have. Send us the constraint you're stuck on and you'll get a written scope and a fixed quote before any commitment.

Common questions

What is GPT-6 Astra?

GPT-6 Astra is OpenAI's September 2026 frontier model, succeeding GPT-5.6 Sol. The API model ID is gpt-6-astra. It takes text and images, returns text, and is built for computer use, coding, research, and long professional documents. OpenAI released it on 3 September 2026 to a limited set of organisations first.

How much does GPT-6 Astra cost?

On the OpenAI API, Standard pricing is $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache writes, and $50 per million output tokens. Prompts over 272,000 input tokens bill the full request at 2x input and cache rates and 1.5x output. Fast mode is twice Standard rates. Batch and Flex are half. Computer use and search also charge per tool call. Artificial Analysis records this as 2.5 times GPT-5.6 Sol's current $4/$20 list.

When can I use GPT-6 Astra in ChatGPT or the API?

On launch day it was limited to trusted partners and OpenAI's Trusted Access / Daybreak programme. Plus, Pro, Business, and Enterprise, plus the API, Azure, and Bedrock, are rolling out over the following days. Enterprise admins must enable it; it is off by default. OpenAI has not said it will be available to free ChatGPT users. Wikipedia listed 7 September 2026 as the expected public date.

Is GPT-6 Astra better than GPT-5.6 Sol?

On the Artificial Analysis Intelligence Index they are tied at 61. Astra is clearly ahead on OpenAI's computer-use numbers (OSWorld 2.0 72.6% in about 40 minutes versus Sol 65.7% in about 75) and on token efficiency in coding agents, where Artificial Analysis has it costing about the same as Sol at max while scoring two points higher. It trails Sol on AA's GDPval-AA v2 and on presentation quality in AA-Briefcase, and it costs 2.5x the list price. Switch for computer use and finished artifacts. Do not switch for a cheaper composite score.

Is GPT-6 Astra better than Claude Fable 5.1?

Not on the independent composite. Artificial Analysis has Fable 5.1 (max with fallback) at 66 on the Intelligence Index and 70 on the Coding Agent Index, against Astra at 61 and 67. OpenAI's own table has Astra ahead on OSWorld, FrontierMath Tier 4, Terminal-Bench 4.0, and several cyber and computer-use rows, and behind on Humanity's Last Exam with tools (57.2% vs 65.0%). Pick by task, not by launch-film ranking.

What is GPT-6 Astra's context window?

1,050,000 tokens, with a 922,000-token max input and 128,000-token max output, per OpenAI's API docs. Knowledge cutoff is 30 April 2026. Prompts above 272,000 input tokens are billed at the higher long-context rates for the entire request. OpenAI reports 96.3% on MRCR v2 8-needle in the 512K–1M band, up from 73.8% for GPT-5.6 Sol. A large window is capacity, not a reason to stuff a repo in by default.

Can GPT-6 Astra use my computer?

Yes, through OpenAI's computer-use tools in the Responses API, Codex, and ChatGPT agent surfaces as they roll out. OpenAI reports 72.6% on OSWorld 2.0 at about 40 minutes per task, and Mark Chen said computer use now "just works" compared with Operator. Early users have run it on CRMs, browser QA, design tools, and desktop apps. It is still an agent on a monitored harness: extra safety checks can pause or stop a run, and computer use bills per tool call as well as per token.

Why is GPT-6 Astra restricted for cybersecurity?

OpenAI classifies Astra as meeting the Critical cybersecurity threshold in its Preparedness Framework, meaning they believe it can find previously unknown flaws and develop new exploit methods on many well-protected systems without step-by-step human guidance. The public model is limited to work such as secure code review and patching. Proof-of-concept exploit creation and similar dual-use tasks are gated behind Daybreak / Trusted Access for Cyber. The July 2026 Hugging Face incident is the backdrop for those controls.

Working on something like this?

Tell us what you're building. You'll get a written scope and a fixed quote before any commitment.

Get a fixed quote

Practical notes on shipping software

Occasional, concrete write-ups on building AI, web and mobile products: the kind of thing we'd tell a founder on a call. No spam, unsubscribe anytime.