Skip to content
Skip to content

Independent e-magazine

the OUTSPOKEN digest

GPT-6 Astra Is Real: OpenAI's First 'Critical' Cybersecurity Model, and Testers Say It Is Not the Best All-Rounder

Released 3 September, Astra found genuine security flaws in testing and posted OpenAI's best-ever reasoning scores, but Artificial Analysis still ranks Claude Fable 5.1 ahead of it on broad intelligence and coding. Pricing matches Fable at 10 and 50 dollars.

Outspoken Digest Technology Desk

Thursday, September 24, 2026/2 min read

The Pioneer Building at 3180 18th Street in San Francisco's Mission District, OpenAI's headquarters, photographed in July 2019
Photo: HaeB via Wikimedia Commons (CC BY-SA 4.0)

Readers have been asking whether "GPT-6 Astra" is a real model or an internet rumour. It is real: OpenAI released it to approved organisations on 3 September and to the wider ChatGPT and API base the following day, and it is now available through the OpenAI API, Microsoft Azure and AWS Bedrock. OpenAI's own material calls it "the world's most intelligent and aligned model" and a "generational leap", language the company uses at every major release, but the numbers behind it are unusually specific: OpenAI's announcement and safety overview, cross-checked against independent write-ups, describe the first model the company has ever classed as reaching the "Critical" tier of its own Preparedness Framework for cybersecurity.

What changed

On OpenAI's internal Terminal-Bench 4.0 test of agentic command-line tasks, Astra's score rose from 37.3 per cent for the prior model to 57.9 per cent. On ExploitBench, a test of finding and using software vulnerabilities, Astra scored 100 per cent, which is what triggered the Critical classification: OpenAI says the model can find and exploit novel flaws in hardened targets, including some genuine zero-day vulnerabilities discovered during evaluation, and that it will now refuse to help with the most advanced offensive cybersecurity tasks such as writing proof-of-concept exploits. On FrontierMath Tier 4, a benchmark of very hard mathematics, it scored 97.6 per cent, the highest OpenAI has published, and the company says the model's hallucination rate fell to 4.2 per cent.

What independent testers found

Here the picture is more mixed than OpenAI's own framing suggests. Vellum and other third-party evaluators, drawing on Artificial Analysis's published indices, found Astra ahead of Anthropic's Claude Fable 5.1 on computer-use tasks, terminal workflows and long-context retrieval, but behind it on the broader Intelligence Index, 61 to 66, and on the Coding Agent Index, 67 to 70. That is a pattern this outlet has seen before with frontier launches: a lab's chosen benchmarks show a clean win, and a wider third-party suite shows big gains in the categories the lab optimised for and flat or mixed results elsewhere. Astra's pricing, at 10 dollars per million input tokens and 50 dollars per million output, matches Fable 5.1 exactly, though its cached-token reads cost four times more, a detail that will matter to anyone running high-volume agent loops.

Why the confusion over its name

Readers asking whether Astra was a rumour were not being unreasonable. OpenAI's naming has been inconsistent this year, model numbers have leaked and been walked back before release, and a genuine "Critical" cybersecurity classification, a first for the company, is the kind of claim that circulates in exaggerated form before the primary source catches up. It is not a rumour. It is a shipped, priced, benchmarked model that OpenAI itself says is powerful enough to require a new safety tier, and whose actual ranking against Anthropic's own newly refreshed Opus 5.5, released three weeks later, depends heavily on which benchmark you trust and which company published it.

Published in The Outspoken Digest

Editorial desk

Outspoken Digest Technology Desk

Software, hardware, artificial intelligence and what they change for everyone else.

Newsletter

The Digest, in your inbox

One edition, sent when it is ready. No noise, and your address is never passed on.

We send a confirmation first. One click to leave, always.

Share this story

the OUTSPOKEN digest

Beyond boundaries. Independent stories on technology, culture, and the trends shaping how we live.