Reading note. This article draws on the public announcement of GPT-6 Astra on September 3, 2026, on ARC Prize's publication of the ARC-AGI-3 results, and on the reporting available at the time of writing. Benchmark and pricing figures come from OpenAI or from identified third parties; we cite them as given and flag what still needs independent verification.
In one sentence
On September 3, 2026, OpenAI unveiled GPT-6 Astra, successor to GPT-5.6 Sol, and its president Greg Brockman closed the press briefing with "Welcome to the AGI era" — AGI standing for artificial general intelligence. We believe the real news is not that sentence, which no institution can validate, but a soberer shift: Astra no longer merely produces text, it drives a computer. And the launch's most-quoted figure — 99.9% on ARC-AGI-3 — collapses to 62.7% as soon as you change the scaffolding around the model. That gap, more than the word AGI, is what we find instructive.
1. What was announced, without the varnish
Let us restate the facts before discussing what they mean.
The model. Astra succeeds GPT-5.6 Sol, whose tiering we analyzed in July. OpenAI describes it as its most capable model to date in computer use, software engineering and scientific work. It comes, the company says, from its largest training run ever — on the order of 100,000 graphics processors at its Stargate site in Texas — and it is the first of its models whose training was reportedly supervised, in part, by other models.
The capabilities on show. Filling in online forms, updating a customer record, organizing a calendar, running research, producing documents, spreadsheets and presentations. Demonstrations reported by the press range from tax preparation to architectural rendering, job hunting and the formatting of legal documents.
The figures, as published:
- OSWorld 2.0 (operating-system control): 72.6%, at roughly 47% less time per task than GPT-5.6 Sol.
- ExploitBench (offensive cybersecurity): 100%, with Astra discovering two zero-day vulnerabilities along the way — flaws unknown to the vendors, and therefore unpatched.
- ARC-AGI-3: 62.7% with ARC Prize's standard harness, 99.9% with an OpenAI-specific adapter. We return to this at length in section 3.
Distribution. First a restricted set of organizations through the Daybreak Access program, then ChatGPT Plus, Pro, Business and Enterprise, the API — the interface through which a third-party program calls the model remotely — under the name gpt-6-astra, and Amazon Bedrock. Announced API pricing is $10 per million input tokens and $50 per million output tokens, with a fast mode at twice the price.
That last point deserves immediate attention, because it breaks with the past twelve months. GPT-5.6 Sol was $5 / $30. Astra doubles the input and raises the output by two thirds. After a year of steady declines in the price of intelligence, the top tier is going back up. That is not a contradiction: it is a sign that OpenAI is no longer selling an answer, but working time.
2. The real shift: from AI that answers to AI that executes
Since ChatGPT, we have used artificial intelligence as a conversational interface. It writes, translates, summarizes, codes, analyzes. The user asks: "How do I do this?"
Astra pushes another mode of use, the one known as agentic AI: the system receives a goal, breaks the problem down, uses several tools, moves between applications, checks its own work and continues across many steps. The question becomes: "Do it."
That is a change in kind, not in degree. A model that answers produces text someone else must then act on. A model that acts produces effects — a form submitted, a record modified, a file written, a message sent. An error is no longer fixed by re-reading: it is fixed by repairing.
For a company, this potentially turns the model into a general-purpose digital colleague. Desk research, file preparation, financial analysis, development, report layout: all delegable. Economically, this shift strikes us as more important than the argument over the word AGI — and considerably easier to verify.
3. The 99.9% on ARC-AGI, and why it should be read twice
This is the figure that traveled fastest, and the one that most deserves attention.
ARC-AGI is a family of tests designed to measure something other than memorization: a system's ability to understand and solve environments it has never seen. Its third version, ARC-AGI-3, drops the model into unknown interactive worlds whose rules it must infer by playing.
The results published by ARC Prize are as follows:
- With the standard harness — the minimal interface, identical for every model — Astra scores 62.7%.
- With a provider adapter, an infrastructure tailored to OpenAI's own mechanisms, which preserves the model's reasoning state between requests and compacts long exchanges, it reaches 99.9%.
For comparison, GPT-5.6 Sol scored 7.8% and Claude Opus 5 scored 30.2%. Even at 62.7%, the jump is substantial. And ARC Prize notes that Astra used fewer actions than the median tested human on 96% of levels.
But the gap between the two numbers is not a methodological footnote. It says something important:
"The intelligence of a modern system no longer lives in the model alone. It lives in what surrounds it — memory, context, tools, persistence."
The same model, given better working memory and better context handling, goes from "very good" to "saturated." What is being measured, then, is the capability of the model plus scaffolding pair, not of the model alone. ARC Prize, for its part, claims no AGI.
We draw a working hypothesis from this, one that matters for small outfits too: if a large share of performance comes from the scaffolding, then the gap between a frontier model and a well-equipped modest one is narrower than it looks. That is good news for anyone building under constraints.
4. The "critical" cybersecurity threshold
Astra crosses another threshold, a less cheerful one. It is the first model OpenAI rates at the highest level of its own risk framework for autonomous cyber capability. The 100% on ExploitBench, obtained by discovering two previously unknown flaws, is the illustration.
The difficulty is structural, not moral: the skill that finds a vulnerability in order to patch it is exactly the skill that exploits it. There is no defense-only version of that capability.
OpenAI's answer is restriction: advanced cyber uses are reserved for users and professionals judged trustworthy. That is a reasonable answer, and it is also an admission — that control now runs through access management rather than through properties of the model. Jakub Pachocki, OpenAI's chief scientist, acknowledges that pinning down precisely what these systems can do gets harder as they improve.
We note the consistency with what we wrote about Fable 5: the more sensitive a capability is judged to be, the more conditional its access becomes. The tap exists, and someone has their hand on it.
5. AGI: a word nobody can certify
The problem with the announcement is not that it is overblown. It is that it is unverifiable.
There is no consensus definition of artificial general intelligence. In its charter, OpenAI has historically described it as highly autonomous systems that outperform humans at most economically valuable work — a strictly economic definition, requiring neither consciousness, nor emotion, nor thinking "the human way." By that criterion, the line does become hard to draw.
But no authority issues an AGI certificate. No single benchmark settles it — and the previous section shows that the most symbolic benchmark gives two opposite answers depending on the tooling used.
The irony is that the best-phrased caution comes from OpenAI itself: Sam Altman has been repeating for a while that AGI is "not a super useful term," precisely because everyone uses it with their own definition. So a company can find itself producing a genuinely consequential technology while being unable to state scientifically: there, at this instant, AGI exists.
Our reading: "AGI era" belongs here to the register of conviction and communication, not measurement. That does not disqualify Astra. It simply moves the burden of proof toward what can be measured — work actually completed.
6. The question that can be measured: how much work?
So let us replace "is Astra an AGI?" with an auditable question:
What share of human work can a system like Astra complete end to end, without rework?
If an agent takes a brief in the morning, uses several pieces of software, runs research, produces documents and returns usable work a few hours later, the economic consequences need no AGI label to be considerable. They would progressively touch development, administration, consulting, finance, communications, legal, research and engineering.
Three caveats, however, before concluding anything:
- The rework rate is unknown. An agent that succeeds on 72.6% of desktop tasks fails on more than one in four. In production, a failed task is not free: it must be detected, understood and redone.
- The price went up. At $10 / $50 per million tokens, a long agentic task — one that consumes context and produces a lot — is not trivial. Profitability will hinge on how much rework it generates.
- The demonstrations are in-house. They are built to succeed. General availability, in the coming weeks, will show what survives outside the lab.
7. Signals to watch
- Independent measurement on OSWorld and ExploitBench. Do the launch figures hold up against third-party evaluation, on tasks nobody anticipated?
- The gap between the two ARC-AGI harnesses. If the community converges on comparable open scaffolding, the 62.7 / 99.9 gap should shrink. If it persists, we will have to accept that frontier performance is partly proprietary through infrastructure, not only through weights.
- The price trajectory. Astra raises the top-tier rate. The question is whether a cheaper middle tier follows, as Terra did for Sol. That tier, not the flagship, decides what is within reach of a small outfit.
- The access regime for cyber capability. Temporary restriction or durable norm? The answer will shape the governance of the models that follow.
- The first real agentic deployments. Not demonstrations: accounts from companies that handed Astra an actual assignment, and the rework rate they observed.
8. A situated word
We write from Réunion Island, 9,000 km from Silicon Valley. From here, the announcement of an "AGI era" reads not as a historic event but as a change in operating conditions.
What concerns us directly comes down to three points.
First, the scaffolding matters as much as the model. That is the most actionable lesson of the launch. We will never have 100,000 graphics processors; we can, however, take care of the memory, the context, the tooling and the verification loop around a more modest model. Part of the road from 62.7% to 99.9% runs through there, and that part is within our reach.
Second, the top-tier price is rising again. The steady fall in the cost of a token does not hold at the frontier. For an island player, the sensible strategy is not to chase Astra, but to know exactly which tasks justify a model at $50 per million output tokens — and to do everything else elsewhere.
Third, a rented capability remains revocable. The more powerful and sensitive a tool becomes, the more conditional its access. We still believe one should consume the abundance without becoming its prisoner: use the best tools, and keep alongside them a modest capability nobody can unplug.
We long imagined the arrival of artificial general intelligence as a scene: a lab switches on a machine, and the world understands. Reality looks less cinematic. Models write better, then reason better, then use tools, then a computer, then work for hours. When we finally ask when AGI arrived, it may already be impossible to locate the line.
It is too early to claim OpenAI has built an artificial general intelligence. It is becoming hard to treat these systems as mere chatbots. Between the two, one thing remains to be done, and it is not spectacular: measure the work actually completed, and build accordingly. 汎
Sources and further reading
- OpenAI — "GPT-6 Astra: A new generation of intelligence" — Official announcement, capabilities, benchmarks, availability and pricing.
- ARC Prize — "OpenAI's GPT-6 Astra on ARC-AGI-3" and the results page — Detail on the 62.7% (standard harness) and 99.9% (provider adapter) scores, and methodology.
- Axios — "'Welcome to the AGI era,' OpenAI says as GPT-6 Astra debuts" — Greg Brockman's remarks and the press-briefing context.
- Fortune — "OpenAI debuts GPT-6 Astra and touts its ability to use your computer" — Computer-use capability, the critical cyber threshold, training scale.
- CNBC — "Sam Altman now says AGI is 'not a super useful term'" — Altman's caution about the definition itself.
- Ryuzaki Labs — "GPT-5.6, in Three Moons", "The Collapse of Token Cost" and "Fable 5, Unplugged in 72 Hours" — Our earlier analyses of OpenAI's tiers, the price of intelligence, and the revocability of rented capability.
This document is updated if new elements emerge. Last revision: 文 September 4, 2026.