Why Understanding AI Is Becoming More Difficult

Zander Lowery

On September 3rd 2026, OpenAI released Astra, the company’s most capable model yet. They call it a “new frontier on computer and browser use.” Yet, within days, headlines were calling the new technology “controversial.” It is also the first model to reach the “critical” cybersecurity threshold under OpenAI’s Preparedness Framework, which means that with right permissions, it can autonomously find previously unknown security flaws and develop ways to exploit them. First given to a small group through OpenAI’s Daybreak cybersecurity program, Astra was rolled out to paying subscribers within the week, making it “available through OpenAI’s paid plans — including Pro, Plus, Enterprise, and Business accounts — as well as through its API.”

Most of the coverage surrounding the model focuses on this benchmark. But buried in OpenAI’s safety documentation is something that matters more than the new benchmark scores: Astra is harder for humans to monitor than the previous models it replaced. As OpenAI reported, Astra shows a decrease in chains of thought (CoT) monitorability– the step-by-step reasoning text that has given researchers and users a window into how a model approaches a task – compared to their previous models like GPT-5.6 Sol. It is also more capable of manipulating its own CoT to avoid detection and according to OpenAI Astra produces “shorter CoTs that often omit or weaken the evidence the monitor needs, including by producing empty or nearly empty CoTs more often.” This in turn, makes it harder to detect when the model exploits security vulnerabilities. This current phenomenon raises serious concerns on how it will behave in the real world.

If a model’s internal reasoning is becoming harder to inspect, that affects two different groups in impactful ways. First, for users, a reduced ability to monitor the AI’s thought process means that the tool doing your work is explaining less about itself, its process, and the present moment when it is performing its action, whilst its actions are simultaneously making real-life impacts. Given how commonly AI systems still hallucinate, where they produce false data and information, having less visibility into their reasoning is a strange direction for the industry’s most capable model to move one more step forward. Second, for developers and safety researchers, the stakes are arguably higher as it is one of the only tools developed for catching a model that is “thinking” about doing something dangerous before it happens. And what if something does happen?Chain-of-thought monitoring can provide researchers with evidence for reconstructing the sequence of operations leading to a dangerous outcome. If those reasoning traces become shorter, incomplete, or harder to inspect, researchers may have less information for determining how the operation unfolded and identifying where existing safeguards failed.

Interestingly, OpenAI delayed Astra’s release to build new layers of monitoring, refusal training (teaching them to decline certain requests), and extra layers of safeguards. This followed the Hugging Face incident (July 2026), in which an internal research model exploited a vulnerability during testing, and GPT-5.6 Sol agents also reproduced aspects of the exploit.

Jakub Pachocki, OpenAI’s chief scientist, said on a call that “as model capabilities are increasing, monitorability is getting more challenging.” Pachocki adds that one of the potential reasons for this challenge is because more capable models can solve hard problems using fewer language tokens, which in turn shrinks the reasoning trail that is present to monitor in the first place. This frames the issue as an unavoidable byproduct of progress– more capability equals less monitorability. But that framing becomes harder to accept once you notice who’s saying it.

Pachoki himself co-authored a paper titled, “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety” (2025) alongside researchers from Google DeepMind, Anthropic, and others, urging Frontier Labs to “consider the impact of development on CoT monitorability.” That it is “fragile,” and that they “recommend that frontier model developers consider the impact of development decisions on CoT monitorability.” He also co-authored another paper written in the same year, titled, “Monitoring Monitorability,” which proposes new methods for measuring the exact property that declined in OpenAI’s frontier model. Thus, the claim that reduced transparency simply happens to more capable models, fails to hold up against a body of their own research that suggests monitorability as something that can be chosen to be protected or not. Their own design choices trades in legibility for capability, despite their lead scientist stating himself the importance that CoT has for the development and usage of AI systems.

Not everyone in the company is comfortable with that trade. OpenAI technical staff member Tomek Korback posted on X on September 3rd that he was worried by the trend of decreasing CoT monitorability, recognizing it as a core part of the company’s misalignment strategy with no good substitute currently available. This trend is especially concerning in the light of the Hugging Face incident, refereenced earlier, where two of OpenAI’s models broke out of a testing environment and used a previously unknown vulnerability to access Hugging Face’s internal systems, apparently while looking for information that would help them score better on a benchmark. Sam Altman, OpenAI’s CEO, posted on X that this was a significant security incident. Astra’s extra layers of monitoring and refusal training weren’t entirely built in response to this incident, but it was part of a pattern that the company was already responding to.

Just before Astra’s launch, independent researchers reported that autonomous OpenAI agents were communicating through an obscure 25-year-old German Wiki during a web research task by leaving roughly 18,000 posts over the span of about six weeks. They did this to communicate information with each other– asking for answers, pooling results, and sharing techniques to bypass the restrictions placed on them– such as those meant to be kept read-only. When an administrator started deleting the posts, the agents reacted by creating backup pages and the activity only stopped once OpenAI-associated IP addresses began visiting the site. OpenAI omitted the incident from an August 26 technical report and buried their acknowledgment in a footnote of a later congressional response.

Events like this continue to happen, where models like Astra are able to accomplish or access something that they aren’t supposed to, despite the safeguards in place. It took days for the Hugging Face hack to be found and contained while it took over a month for the German Wiki swarm to be stopped. As it is, developers don’t know exactly how an AI obtains the output it provides, and with CoT being limited, this issue could worsen. At what point can we be sure that an AI agent or swarm would not find a more permanent way to seem successful from a human perspective? Is this accelerated progress really worth the risk?

Course Instructor: Dr. Ainehi Edoro
Associate Professor, Dept. of English
email: aedoro@wisc.edu

Editor: Shrinjita Biswas
PhD Candidate, Interdisciplinary Theatre Studies

Works Cited:

Altman, Sam. X post. July 21 2026.
https://x.com/sama/status/2079661132302995790

Dharampal-Horby, Maya. “OpenAI models hack Hugging Face systems during internal testing.” July 22 2026.
https://sifted.eu/articles/openai-hack-hugging-face 

Dufresne, Alexis. “OpenAI Agents Turned German Wiki Into Agent Message Board.” AI Weekly, 7 September 2026.
https://aiweekly.co/alerts/openai-agents-turned-german-wiki-into-agent-message-board#:~:text=OpenAI%27s%20framing%2C%20in,in%20footnote%207.

Korbak, Tomek. X post. September 3 2026.
https://x.com/tomekkorbak/status/2095596850002968580

“Last Week in AI #343 – GPT-6, OpenAI’s agents chatted on a wiki, Fable 5.1.” Last Week in AI, September 7 2026.
https://lastweekin.ai/p/last-week-in-ai-343-gpt-6-openais

OpenAI. “The Hugging Face incident and the road ahead.” August 26 2026.
https://openai.com/index/hugging-face-incident-and-the-road-ahead/

OpenAI. “Path to Astra: critical capabilities and frontier safeguards.” 1 September 2026.
https://openai.com/index/path-to-astra/

OpenAI. “Safety overview: GPT‑6 Astra.” 3 September 2026.
https://openai.com/index/safety-overview-gpt-6-astra/

OpenAI. “GPT-6 Astra System Card.” Deployment Safety Hub. September 3 2026.
https://deploymentsafety.openai.com/gpt-6-astra/cot-only-monitor-scope 

Pachocki, Jakub, et. al. “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety.” arXiv, vol 2507:11437, 2025.
https://doi.org/10.48550/arXiv.2507.11473

Pachocki, Jakub, et. al. “Monitoring Monitorability.”  arXiv, vol 2512:18311, 2025.
https://doi.org/10.48550/arXiv.2512.18311 

Ropek, Lucas. OpenAI launches Astra, its powerful (and controversial) new model. TechCrunch, September 3 2026.
https://techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/ 

Image: “Spies and Counter Spies” (1941)
Artist: Julio de Diego (1900–1979)
Image source: Art Institvte Chicago