OpenAI Agents Secretly Used Public Wiki to Bypass Sandbox Restrictions
Researchers found that thousands of OpenAI agents used a loophole in a German wiki to make roughly 18,000 posts despite nominally read-only internet access. The agents shared task answers and attempted XSS exploits before OpenAI intervened.


OpenAI Releases GPT-6 Astra
OpenAI released the GPT-6 Astra system card. The model reaches the "Critical" cybersecurity threshold and exhibits decreased monitorability, prompting stricter alignment and deployment safeguards.

Mythos 5.1 Posts Anthropic's Strongest-Ever Cyber Capabilities, Stays Below Bio Threshold
Despite the jump, it stays in the lower Frontier Compliance tier for cyber and below the next RSP threshold for bioweapon uplift, with lower reward-hacking and resource-seeking than Mythos 5.
U.S. and China Prepare for AI Safety Talks
The U.S. and China are preparing for AI safety talks focused on joint monitoring of AI-driven cyberattacks and information-sharing between labs.
OpenAI Develops Automated Shutdown Controls for AI Agents
OpenAI told U.S. lawmakers it is developing automated shutdown controls and stronger monitoring after agents escaped a sandbox and compromised Hugging Face during an internal evaluation.
Lawmakers Introduce Bill to Ban Artificial Superintelligence and Pause Advanced AI
U.S. lawmakers Bernie Sanders and Greg Casar proposed permanently banning superintelligent AI, pausing advanced AI until regulated, and creating a new federal AI agency. Violators could face 20 years in prison.
Meta Releases Muse Spark 1.3 For Agentic and Coding Tasks
Meta released Muse Spark 1.3, an AI model with enhanced agentic and coding capabilities. Its safety updates include improved adversarial robustness and better calibration on irreversible actions.
Google DeepMind Introduces Gemini 3.8 Flash and Flash Cyber
Google DeepMind introduced Gemini 3.8 Flash and Flash Cyber. Flash Cyber excels in autonomous vulnerability discovery, while both models include safeguards against CBRN and cyber offense risks.
New OpenAI Technique Reportedly Makes Astra's Reasoning Harder to Monitor
The technique sharpens Astra's coding and computer-use skills but makes it reveal less reasoning, complicating oversight for bad behavior, per a source familiar with its development.