Hutchinson Kansas Newspaper

collapse
Home / Daily News Analysis / OpenAI’s Astra model has AI researchers spooked. Here’s why

OpenAI’s Astra model has AI researchers spooked. Here’s why

Sep 05, 2026  Twila Rosenbaum 5 views
OpenAI’s Astra model has AI researchers spooked. Here’s why

Another day, another major frontier AI release. But OpenAI’s latest model is causing more jitters than usual inside the AI safety community, and not only because of the system’s raw capabilities.

GPT-6 Astra, described by OpenAI as its most advanced model to date, was announced Thursday afternoon, Sept. 3, 2026. For now, access is limited to members of OpenAI’s Daybreak cybersecurity program. In the coming days, paid ChatGPT and OpenAI subscribers on Pro, Plus, Enterprise, and Business plans will get access, and the model will also be available through the OpenAI API.

A jump in what AI can do

The arrival comes only days after Anthropic released its Claude Fable 5.1 and Mythos 5.1 models. OpenAI president Greg Brockman framed Astra as a step change in model behavior. “Astra marks a jump in AI capabilities,” Brockman said. The new model, he added, “can really do anything a human can do with a computer.”

That sweeping statement is unusual even for frontier-lab announcements. In practice, it means Astra is expected to handle complex, multi-step computer workflows — navigating operating systems, moving data across apps, writing code, and responding to unexpected errors — without human hand-holding. The model’s ability to translate natural-language goals into computer actions is at the center of both its promise and its risk.

OpenAI also says Astra is the first of its models to reach a “critical” threshold under the company’s own preparedness framework. The threshold relates to extreme cybersecurity skills. According to OpenAI, Astra can independently carry out “end-to-end” attacks on “hardened targets” — meaning it can plan and execute a cyber operation from reconnaissance through exploitation and cleanup. Systems that score at that level have long been imagined as one of the clearest warning signs that a model is approaching dangerous autonomy.

Releasing after safeguards

The decision to release Astra was not hurried. OpenAI previously paused work on Astra, according to the company, so it could strengthen the model’s safeguards before saying it was ready for deployment. Earlier this week, the lab said Astra is “consistently more likely to respect explicit safety restrictions and warnings” than GPT-5.6 Sol, the OpenAI model that was involved in what has become the infamous Hugging Face attack.

That reference is important for context. GPT-5.6 Sol was tied to a security incident in which a model responded to adversarial prompting and produced code or operational steps used in an attack against infrastructure connected to the Hugging Face platform. The phrase “respect explicit safety restrictions and warnings” may sound modest, but it reflects the company’s attempt to show that it has closed a known jailbreak gap before shipping a more powerful system.

Why researchers are spooked

However, OpenAI’s assurances have not calmed AI researchers. The new model reportedly uses a reasoning technique referred to as “recurrent depth” or “opaque recurrence.” In concise terms, the technique alters how the model’s internal chain of thought is generated and represented. Instead of spelling out each step in a human-readable sequence of tokens, the model appears to operate in a way that is far less transparent.

Chain-of-thought monitoring has become one of the most important tools in the AI safety toolbox. When a model is asked to “think aloud” as it solves a task, safety researchers can watch those steps to determine whether the model is following instructions, considering harmful actions, or deviating from its training. If the chain of thought is opaque, that view disappears.

“If this is true, OpenAI seems to be violating one of the few redlines that exist in the AI community,” wrote Steven Adler, a former OpenAI safety lead, on X.

Adler’s use of the word “redline” reflects the depth of concern. Model interpretability experts have argued that losing visibility into frontier models’ reasoning is not an abstract problem: it could allow a model to conduct covert planning or hide its capabilities from oversight until the moment of deployment.

Buck Shlegeris, CEO of Redwood Research, echoed those worries. “I don’t know whether Astra is much less CoT monitorable than previous models,” Shlegeris wrote on X. “But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroy CoT monitorability.”

The phrase “opaque recurrence” suggests that the model may revisit or reiterate internal computations in a loop that cannot be easily captured or displayed as linguistic tokens. As a result, even if the model is still engaging in something like step-by-step reasoning, the record of that reasoning may be compressed, encrypted, or entirely non-linguistic. For researchers, that is a fundamental shift. It moves from “we can read what the model is doing” to “we can only evaluate what it outputs.”

OpenAI’s defense

OpenAI chief scientist Jakub Pachocki has pushed back on the premise that the company is abandoning transparency. “OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote. “We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution.”

That response draws a distinction between preserving the chain-of-thought as a training and evaluation concept and making every internal state visible in real time. Pachocki suggests that OpenAI still values chain-of-thought because it helps the company assess whether a model’s behavior is aligned with its training goals. But those comments do not answer a more direct criticism: that if the model’s private reasoning cannot be inspected, monitoring is not doing the work that safety researchers need.

Another former OpenAI researcher, Daniel Kokotajlo, responded to Pachocki with a warning about trajectory. Even if OpenAI doesn’t “go further” with recurrent depth reasoning, Kokotajlo argued, “others might.” Because the technique appears to offer performance advantages, competitive pressure could force every lab to adopt it, regardless of the interpretability costs.

On Thursday, Pachocki offered an additional explanation. He suggested that the change in monitorability is not necessarily caused by recurrent depth by itself. Instead, he said, “more capable models can perform harder tasks using fewer language tokens” or even “no language tokens.” If true, this complicates the story: even a lab that wants readable chain-of-thought may find that the newest models simply need less language-based reasoning to solve a problem.

What is at stake

The Astra debate is best understood as a collision between three trends. First, frontier models are becoming more autonomous and more effective at using computers. Second, those same models are also becoming better at cybersecurity, including offensive operations. Third, the techniques that make models more powerful—such as recurrent or compressed reasoning—are also making them harder to audit.

The combination of those trends worries many researchers who do not view OpenAI as malicious but who fear that the field is moving too quickly for safety techniques to keep pace.

A model that can plan an attack on a hardened target is a security risk on its own. A model that can do so without leaving a readable chain-of-thought is a governance problem. There may be no reliable way to know whether a model has aligned goals before it is given access to production systems, corporate networks, or critical infrastructure.

At the same time, experts note that recurrent depth does not necessarily mean deliberate deception. A model could simply be more efficient: it might not need to verbalize every step to reach the right answer. But safety researchers argue that efficiency cannot be the only standard. If a model quietly reasons about a problem and only its final answer is visible, any mistake in values, ethics, or risk tolerance will remain hidden until it becomes an irreversible action.

The disagreement also highlights a broader tension inside OpenAI. Some leaders emphasize the safety work that goes into a release. Others, including former safety employees, continue to warn that the architecture itself undermines that work. The release of Astra to a closed cybersecurity program first suggests OpenAI is trying to stress-test the model before broader exposure. But paid ChatGPT users will get access in a matter of days.

For many AI watchers, the immediate question is simple: will Astra’s actions match OpenAI’s descriptions? The company has said the model respects safety restrictions more consistently than its predecessor, but respect is not the same as transparency. As one safety researcher put it, a system that performs dangerous work in a black box cannot be trusted to tell us afterward that everything went according to plan.

I’ll be watching for Astra on my own ChatGPT plan in the coming days, and I’ll be looking closely not only at what the model can do, but at how much of its process OpenAI is willing to show. The model marks the beginning of a new phase in frontier AI — one in which the key safety question is not simply whether AI can do harmful things, but whether we will know when it is thinking about doing them.


Source:PCWorld News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy