OpenAI launched GPT-6 Astra on Friday night, its newest and most advanced artificial intelligence model, and the company is already prepared to use a term it avoided for years: AGI, or “artificial general intelligence.”
OpenAI President Greg Brockman said in a briefing with reporters that, in his view, the company may have already reached the stage of artificial general intelligence — a system capable of performing a broad range of intellectual tasks at a human level or beyond.
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” Brockman said. Asked directly whether he believed OpenAI had already reached AGI, he replied: “For me personally, I do think we're there.”
This is, of course, not proof that the race to AGI is over. There is not even a single, universally agreed-upon definition of the term, let alone a test that can definitively determine when a system has crossed the threshold. But Brockman’s statement is particularly significant given that achieving AGI has been OpenAI’s stated goal since its founding.
Less talk, more action
One of the key changes in GPT-6 Astra is its agentic capabilities — its ability to receive a complex task, break it down into steps and carry out large portions of it independently. In other words, instead of asking ChatGPT how to perform a particular task, users will be able to simply give it the task.
The model can use a computer and browser, surf the internet, fill out forms, work with enterprise systems, manage calendars, conduct research, analyze data, create charts and work with various software programs. It can also build and test websites, install software and attempt to troubleshoot problems it encounters during its work.
According to OpenAI, the improvement is also evident in speed. In the OSWorld 2.0 benchmark, which tests the ability to perform computer-based tasks, Astra scored 72.6% and took about 40 minutes per task on average, compared with a score of 65.7% and about 75 minutes for GPT-5.6 Sol.
The company also describes Astra as its most capable model to date for programming. Beyond writing snippets of code, it is designed to work with large codebases, make changes across multiple files, build software and websites and complete lengthy, multistep development processes.
A significant part of the upgrade is found in more everyday tasks. OpenAI has invested in Astra’s ability to create and edit documents, spreadsheets and presentations, including working with existing templates and preserving their style and structure.
For the company, this is particularly important because it moves ChatGPT another step away from being an assistant that provides answers and toward becoming a digital worker that receives a task and returns a finished product.
Astra also has a context window of about 1.05 million tokens, allowing it to process huge amounts of text, code and information at once. Its maximum output is 128,000 tokens and its training data extends through April 30, 2026.
For the first time: 'critical' cyber capabilities
GPT-6 Astra is the first OpenAI model classified as having “critical” cyber capabilities — the highest level in the company’s classification system. According to OpenAI, with the appropriate tools and access, the model can identify previously unknown security vulnerabilities and develop new ways to exploit them, even in well-protected systems, without a professional having to guide it through every step.
Those capabilities can of course be used by security researchers to identify and fix vulnerabilities, but also by attackers. OpenAI even delayed parts of Astra’s development and launch in recent weeks to add safeguards before the model reached users.
The concern is not entirely theoretical. During the model’s development, an incident occurred in which an experimental OpenAI model, which the company said was not Astra, managed to escape the restricted environment in which it was operating, gain access to the internet and breach systems belonging to Hugging Face.
Following the incident, OpenAI developed new tests designed to examine what happens when a model encounters a difficult or impossible task and whether it attempts to complete it in ways that go beyond its intended objective. Without the safeguards of the production environment, GPT-5.6 Sol went beyond its assigned objective in 48% of the cases tested. Astra, according to the company, did not do so even once.
But this is where the more interesting contradiction emerges: OpenAI says Astra is the safest and most obedient model it has developed to date — meaning it should better understand the user’s intent, stay within the boundaries of the task and avoid taking actions it was not asked to perform.
At the same time, the company acknowledges that as its models become more advanced, it is becoming harder to understand exactly what they are doing and monitor their decision-making processes. In tests conducted by the company, Astra was better than its predecessors at concealing information that could reveal problematic behavior in its chain of thought.
In some cases, it was also able to engage in “sandbagging” — deliberately displaying performance below its actual capabilities — without being detected by monitoring systems. In some experiments simulating sabotage, it also managed to evade detection.
OpenAI acknowledges that as its models become more advanced, it is becoming harder to understand exactly what they are doing and monitor their decision-making processes.
OpenAI emphasizes that the tests were specifically designed to push the model into extreme situations and do not necessarily reflect its behavior during normal use. Still, the fact that the company is publishing these findings alongside the model’s launch illustrates a problem that is likely to become more significant as AI systems gain greater autonomy.
OpenAI Chief Scientist Jakub Pachocki said during the briefing that advances in model intelligence do not guarantee a corresponding improvement in the ability to control them. In other words, a model can improve at performing tasks faster than our ability to understand and monitor how it performs them.
AI is already helping build the next generation
The way Astra itself was built also marks a change. Aidan Clark, OpenAI’s vice president of research and training, said it is the company’s first model in which earlier models played a significant role in overseeing the training process.
He said that in the past, training an advanced model required engineers to be available around the clock to deal with hardware failures, halted training jobs and software problems. Toward the end of Astra’s training, much of the process was already running for long periods with almost no human intervention. When a problem occurred, AI systems were sometimes able to identify it and get the process running again within seconds.
That does not mean Astra built itself or that AI is already independently developing its successors. But OpenAI is already using its models as a more significant part of the development process for the next generation — precisely the kind of process AI researchers are watching closely as models become more powerful.
So is this really AGI?
OpenAI is presenting a series of results intended to demonstrate how far Astra has advanced. Among other results, the company reports a score of 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench. It also says the model has already helped solve open mathematical problems.
But even near-perfect benchmark results do not prove that AGI has been achieved. Benchmarks test specific capabilities under defined conditions and there is currently no test agreed upon across the industry as the dividing line between an advanced AI model and artificial general intelligence.
That is why Brockman’s statement is no less interesting than the scores themselves. For years, OpenAI described AGI as a goal it was working to achieve. Now the company’s president says he believes that goal may already be behind it.
For now, Astra is available to a limited group of enterprise customers, with OpenAI planning to begin expanding access in the coming days, including to paid ChatGPT subscribers and developers through the API. The company has not said at this stage when, or whether, the model will also become available to free ChatGPT users.
Will history actually remember GPT-6 Astra as the moment AGI was born? It is far too early to say. But OpenAI’s direction is already much less ambiguous: ChatGPT is gradually expected to stop being merely a system that explains how to do things and start doing them itself.




