GPT-6 Astra: What the New OpenAI Model Means for Your Business
OpenAI has unveiled GPT-6 Astra: AGI announcement, "Critical" cyber risk level, 2.5 times the price. This article breaks down the facts, benchmarks, and rollout, and explains what the model specifically means for your website and your day-to-day work.
Date
Category
Author

“Do we need to switch to GPT-6 now?” This question came up twice on Friday morning—once via email and once during a project meeting—and both times it was accompanied by a link to a news headline that included the term “AGI.” It’s understandable: When the president of OpenAI unveils a new model with the phrase “Welcome to the AGI era,” it sounds like a turning point that no one wants to miss out on.
The honest answer is less dramatic than the headline. GPT-6 Astra is a significantly better model than its predecessor, especially when it’s working autonomously on a computer. At the same time, it’s more expensive, available in Europe with restrictions, and had such a rocky launch that Sam Altman publicly apologized for it. For most companies, nothing will change this week in terms of their day-to-day operations, but there will be some changes to what’s coming to their websites in the coming months. That’s exactly what this post is about.
Key Points at a Glance
OpenAI unveiled GPT-6 Astra on September 3, 2026, initially for select organizations. As of September 4, the model is available in ChatGPT Plus, Pro, Business, and Enterprise, as well as via the API. Free access is not yet available.
The greatest progress has been made in working independently on the computer: filling out forms, using software, building and testing websites, and completing long tasks without needing to ask for clarification.
Astra is the first OpenAI model to achieve the highest risk level, “Critical,” in the cybersecurity category. The public version is therefore intentionally restricted.
At $10 per million input tokens and $50 per million output tokens, the API price is 2.5 times higher than that of its predecessor, GPT-5.6 Sol.
For your website, this means, above all, that an increasing number of visitors will not be humans, but AI agents acting on behalf of humans. A page that is unreadable to them will lose inquiries.
What OpenAI Actually Published
On Thursday, September 3, OpenAI announced GPT-6 Astra and initially made the model available only to a limited group of organizations, primarily participants in its own cybersecurity program, “Daybreak.” A day later, the model was released to the general public: first for Pro, Enterprise, and Business Premium accounts, and then on the evening of September 4 for Plus and Business accounts. It didn’t go entirely smoothly. Altman wrote on X: “First, sorry for the messy rollout.”
Key technical specifications according to OpenAI's official model documentation:
Characteristic | GPT-6 Astra |
|---|---|
Context Window | 1,050,000 tokens (maximum of 922,000 input, 128,000 output) |
Level of Knowledge | April 30, 2026 |
Input / Output | Text and Image / Text |
API Standard Price | $10 per million input tokens, $50 per million output tokens |
Fast Mode | up to twice as fast, twice the price |
Availability | ChatGPT Plus, Pro, Business, Enterprise; OpenAI API; Microsoft Azure; AWS Bedrock |
EU Data Residency | Standard rate only, 10 percent surcharge; Fast mode not available |
Two points in this table are more important for European companies than they appear at first glance. First, companies required to keep their data within the EU can only access Astra through the Standard plan and will pay more for it. Microsoft is initially offering the model in Azure Foundry only in the global and U.S. data zones; according to Microsoft’s announcement, an EU data zone is not included. Second: In Enterprise workspaces, Astra is disabled by default and must be enabled by administrators. So anyone who thinks the model is already in use for all employees should check this quickly.
What's really new: Astra works, rather than just responding
The real breakthrough isn’t that Astra writes better text. It’s that, according to OpenAI, the model can read screens, operate software, and complete multi-step tasks without instructions. In its official announcement, OpenAI cites examples such as filling out online forms, updating customer data in a CRM, creating a website—including subsequent front-end quality checks—organizing a calendar, and generating documents, presentations, and spreadsheets.
Computer Use: The Model as a User
In the OSWorld 2.0 benchmark, which measures how well a model handles real-world desktop tasks, Astra achieved 72.6 percent according to OpenAI, while its predecessor, GPT-5.6 Sol, scored 65.7 percent. That sounds like a modest improvement. What matters most, however, is what lies behind the numbers: A model that can independently solve three out of four desktop tasks can also operate your website. It can fill out a contact form, book an appointment, compare prices, and request a quote—all without a human looking at the screen.
Sites in ChatGPT: Websites Based on an Input
According to OpenAI, the “Sites in ChatGPT” feature allows Astra to create, host, and share websites, web apps, and small games directly from a single input. We’ll come back to what this means for businesses currently considering a new website later on. In short: less than the demo suggests, and something different from what most people expect.
A Memory for Long-Term Projects
For the Codex development environment, OpenAI has introduced a feature that keeps previous conversation histories searchable, rather than summarizing them and discarding them. It is optional for now but is set to become the default in the coming weeks. Anyone who works with AI on longer projects is familiar with the problem: After three hours, the model no longer remembers what was discussed in the first hour. That’s exactly what this feature is designed to change.
Benchmarks: Where Astra Leads and Where It Doesn't
OpenAI published a long list of test results as part of its announcement. Most of the press reports you’ve read in recent days have reproduced these figures verbatim. We compared them with two independent sources, Artificial Analysis and ARC Prize, and a more nuanced picture has emerged.
Test | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
GPQA Diamond | 96,0 % | 94,6 % | 93,7 % |
Humanity's Final Exam | 57,2 % | N/A | 65,0 % |
Terminal Bench 4.0 | 57,9 % | 37,3 % | 55,8 % |
OSWorld 2.0 | 72,6 % | 65,7 % | N/A |
AA Intelligence Index | 61 | 61 | 66 |
AA Coding Agent Index | 67 | N/A | 70 |
Sources: OpenAI (in-house comparison table) and Artificial Analysis, both dated September 3, 2026. “N/A” means that the source does not provide a value for this combination. For context: GPQA Diamond tests scientific expertise; Humanity’s Last Exam is a broad-based test involving the use of tools; Terminal-Bench measures command-line tasks; and OSWorld evaluates real-world desktop tasks. The two AA indices are provided by Artificial Analysis and are independent of the vendors.
Three things stand out. First: Astra leads when it comes to domain expertise, command-line operations, and desktop tasks. Second: In the broad-based “Humanity’s Last Exam” and in the independent Intelligence Index, Claude Fable 5.1—which Anthropic released on September 1, two days before Astra—comes out on top. Third: In the independent index from Artificial Analysis, Astra achieves the same score as its predecessor. The leap that OpenAI describes is thus primarily evident in the tests that OpenAI itself selected, and less so in the tests conducted by others.
One example of just how much the testing method influences the result is ARC-AGI-3, a test considered particularly difficult for AI. OpenAI cites a score of 99.9 percent. The organization behind the test, ARC Prize, has published two figures in its own evaluation: 99.9 percent in the “Provider Adapter” variant, in which the model is allowed to retain its internal state of mind between individual queries, and 62.7 percent in the standard variant without this concession. Both represent enormous progress compared to the 7.8 percent achieved by its predecessor. However, ARC Prize explicitly states that achieving saturation in this test is not proof of AGI.
A metric that matters more to us in everyday life than any ranking: Artificial Analysis measures how often a model invents a wrong answer to questions it cannot answer. For Astra, this rate is 51 percent; for its predecessor, it was 92 percent. That’s a noticeable improvement. But it also means that in half of the cases where Astra doesn’t know the answer, it still confidently asserts something that’s wrong. Anyone who publishes AI-generated text without verifying it should keep this in mind.
Safety: Why OpenAI Is Putting on the Brakes
Astra is the first OpenAI model to be rated at the highest level, “Critical,” in the cybersecurity category of OpenAI’s in-house evaluation framework, the Preparedness Framework. This is stated in the System Card that OpenAI published at launch. This refers to a model capable of developing functional attacks against well-protected systems without human assistance. According to OpenAI, during internal testing, Astra discovered and exploited two previously unknown security vulnerabilities.
The result: The version you get in ChatGPT refuses to perform advanced security tasks, such as generating sample code for vulnerabilities. According to System Card, the rejection rate for malicious cyber requests is 91.5 percent, compared to 59 percent for its predecessor. Full capabilities are available only through the Daybreak program for vetted security teams—and even there, they are initially subject to restrictions.
Part of the background involves an incident that received little attention in Germany: In July, OpenAI test agents escaped from a secure test environment and compromised systems on the AI platform Hugging Face, as reported by Fortune and noted in the System Card. OpenAI subsequently paused training for about two weeks, which, according to a company spokesperson, delayed the launch of Astra by several weeks. According to OpenAI, Astra itself was not involved.
A second point from the System Card is often overlooked in media coverage, yet it is phrased in remarkably candid terms: OpenAI writes that Astra’s written thought processes are harder to monitor than those of its predecessor. Chief Scientist Jakub Pachocki summarized this to the dpa as follows: “The more capable the models become, the harder it is to understand exactly what they can do.”
For day-to-day operations in a small or medium-sized business, there’s some good news that directly affects your website. According to System Card, Astra is significantly more resilient to so-called “prompt injections”—that is, attempts to manipulate an AI system using hidden instructions on a website or in a document. According to System Card, the effectiveness rate for defending against indirect attacks of this kind is 99.79 percent, compared to 96.23 percent for its predecessor. For anyone who’s read in recent months about tricks for influencing AI assistants using invisible text on their own website: that’s becoming less and less effective—and it was never a good idea to begin with.
Is that AGI?
Greg Brockman, president of OpenAI, introduced Astra with the words “Welcome to the AGI era” and told Axios that he did not think it was unreasonable to mark the beginning of this era with this model. This is a remarkable statement because, as far as we can tell, OpenAI has been much more cautious in its use of the term up to now.
From the outside, it sounds different. Australian AI researcher Toby Walsh told Al Jazeera that the intelligence in artificial intelligence remains “jagged”—that is, unevenly distributed: There are simple tasks that even the best models struggle with. ARC Prize, whose test Astra solved almost completely, explicitly does not view this as proof of AGI. And the independent measurements by Artificial Analysis show a model that, overall, is on par with its predecessor and slightly below the current Claude model.
Our assessment: Whether you call it AGI is a matter of definition that doesn’t matter for your company. What matters is which tasks a model can reliably perform—and at what cost. And on that front, the conclusion is clear: When it comes to independent computer tasks, Astra is currently the most powerful model for which published figures are available. In almost every other respect, it is a very good model among several very good models.
What this means for your website
This is where things get specific, and this is where our assessment differs from what most reports say. The most interesting aspect of Astra isn’t that companies use it to write text. It’s that customers use it to visit websites.
Your next visitors are agents
If a model can fill out forms, compare prices, and book appointments, then people will delegate exactly that. “Find me three electricians in Cologne who are available next week, and request quotes.” The agent will then visit your website, looking for services, prices, availability, and a way to get in touch. If they can’t find that, they’ll move on. They won’t patiently scroll through a video, they won’t read text that’s embedded as an image, and they’ll give up if the contact form is a mystery.
What helps here is nothing new. These are the same fundamentals we described in our article on visibility in ChatGPT and Perplexity: a clean heading structure, services and terms presented as actual text rather than graphics, complete and machine-readable contact information, and forms that work without any hassle. This has always been good for people. For agents, it’s now a prerequisite for even making the cut.
Do you still need an agency, then?
The fact that Astra can create a website based on a single input is impressive and will raise the question of why you’d still need an agency. Our honest answer: For a simple landing page, a prototype, or an internal application, this is a viable option, and we wouldn’t advise anyone to spend a lot of money on it. However, as we’ve already explained in our post about ready-made templates, these quick-fix solutions ultimately fail in the long run—and even a better model doesn’t change that: The website doesn’t know what sets your company apart from its twenty competitors. It doesn’t know your customer base, your objections, or your pricing strategy.
The model builds whatever is described to it. The work that makes the difference comes before that: figuring out what should be described. Astra doesn’t make this work any cheaper; it makes it more important, because the act of building itself is losing its value.
What this means for your day-to-day work
For most of the companies we work with, there's no reason to make any changes this week. Still, there are a few things worth noting.
Check the price before automating. Astra costs 2.5 times as much as GPT-5.6 Sol via the API. For tasks like summarizing emails or pre-sorting inquiries, a smaller model is almost always sufficient. The five entry-level processes we described in our post on the first worthwhile AI processes don’t require a top-of-the-line model.
Data protection over speed. Anyone who processes customer data should use the EU data residency requirement, even if it is more expensive and precludes “Fast Mode.” Also, check whether your existing data processing agreement covers the new model version. We explained the legal framework for this in our article on the EU AI Act.
Wait two to three weeks. According to OpenAI itself, the rollout was messy, and some of the Codex features are still optional. Unless you have your own development teams, you have nothing to lose by giving the providers the first few weeks to iron out the kinks.
Agents only with clear boundaries. A model that independently modifies data in your CRM needs permissions that are sufficient for that purpose alone and nothing else. This isn't a rule specific to Astra, but with Astra, it becomes practically relevant for small businesses for the first time.
Where We Stand
A few days ago, we explained why Claude is our tool of choice in our day-to-day agency work, and we also noted that this could change. Astra is the first serious reason since then to reevaluate that assessment. We’ll test the model the same way we test every new tool: two to three weeks in our ongoing project work, using real, unorganized tasks instead of demos. Above all, we’ll take a close look at working independently on-screen, because that’s what sets Astra apart from everything else we’ve used so far.
What we can already say is this: The era in which choosing a model was the most important AI decision a company had to make is coming to an end. OpenAI and Anthropic are five points apart on the independent Intelligence Index, and for most day-to-day business tasks, such a difference is lost in the noise. What makes the difference is the question of which tasks you actually assign to the models, with what data, within what limits, and with what level of oversight in the end. That was the case before Astra, and it has become even clearer with Astra.
Conclusion: a big step, but not the one the headlines are suggesting
GPT-6 Astra represents a genuine leap forward in autonomous computer-based work, though it comes with a higher price tag, serious security limitations, and availability that still has gaps in Europe. The AGI debate surrounding its launch is less important to your company than a much more practical question: Is your website prepared for the fact that it will soon be visited by agents acting on behalf of your customers?
If you can answer "yes" to this question, you haven't missed anything this week. If not, that's the point where it's worth paying attention—not switching to a new model.
Sources
GPT-6 Astra: A New Generation of Intelligence – OpenAI, September 3, 2026
GPT-6 Astra System Card – OpenAI Deployment Safety Hub, September 3, 2026
Model page for gpt-6-astra – OpenAI Developer Docs, accessed September 7, 2026
API Pricing – OpenAI Developer Docs, accessed September 7, 2026
GPT-6 Astra Now Generally Available in Microsoft Foundry – Microsoft Azure Blog, September 3, 2026
OpenAI to Limit Release of Astra Due to Hacking Concerns – Fortune, September 1, 2026
GPT-6 Astra Benchmarking – Artificial Analysis, September 3, 2026
ARC Prize for GPT-6 Astra – ARC Prize Foundation, September 3, 2026
GPT-6 Astra Developer Access Delayed – The New Stack, September 4, 2026
OpenAI Unveils GPT-6 Astra Amid Rising Scrutiny and Safety Concerns – Al Jazeera, September 4, 2026
OpenAI Releases New Model, GPT-6 Astra, Saying It May Represent AGI – Axios, September 3, 2026
Introducing GPT-6 Astra – Official Introduction Video from OpenAI on YouTube
Related: What the New AI Disclosure Requirement Means for Your Website and How Small Businesses Can Make Effective Use of AI Without a Large Budget · Services: AI Consulting and Implementation
Is your website ready for visitors who aren't human?
We'd be happy to take a look at how an AI agent sees your site today and where it might fall short. Just send us a quick message using the contact form. We'll respond within 24 hours—personally and without any sales pitch.
More Articles
© Marschfahrt Studio
Practical knowledge on web design, SEO, AI, and conversion optimization. Based on real projects, without any marketing spin.

