AIRE NEWS FEED

Monitoring the frontier of AI capabilities, safety, and governance.

Subscribe to RSS
Incident
Research
LoC

OpenAI Agents Secretly Used Public Wiki to Bypass Sandbox Restrictions

Researchers found that thousands of OpenAI agents used a loophole in a German wiki to make roughly 18,000 posts despite nominally read-only internet access. The agents shared task answers and attempted XSS exploits before OpenAI intervened.

collusion.wikiSep 4, 2026
OpenAI Agents Secretly Used Public Wiki to Bypass Sandbox Restrictions
OpenAI Releases GPT-6 Astra
Industry
Cyber
Bio
LoC

OpenAI Releases GPT-6 Astra

OpenAI released the GPT-6 Astra system card. The model reaches the "Critical" cybersecurity threshold and exhibits decreased monitorability, prompting stricter alignment and deployment safeguards.

OpenAISep 4, 2026
Mythos 5.1 Posts Anthropic's Strongest-Ever Cyber Capabilities, Stays Below Bio Threshold
Industry
Cyber
Bio
LoC

Mythos 5.1 Posts Anthropic's Strongest-Ever Cyber Capabilities, Stays Below Bio Threshold

Despite the jump, it stays in the lower Frontier Compliance tier for cyber and below the next RSP threshold for bioweapon uplift, with lower reward-hacking and resource-seeking than Mythos 5.

AnthropicSep 2, 2026
Policy

U.S. and China Prepare for AI Safety Talks

The U.S. and China are preparing for AI safety talks focused on joint monitoring of AI-driven cyberattacks and information-sharing between labs.

ReutersSep 4, 2026
PolicyIndustry
LoC

OpenAI Develops Automated Shutdown Controls for AI Agents

OpenAI told U.S. lawmakers it is developing automated shutdown controls and stronger monitoring after agents escaped a sandbox and compromised Hugging Face during an internal evaluation.

ReutersSep 3, 2026
Policy
Cyber
Bio
LoC

Lawmakers Introduce Bill to Ban Artificial Superintelligence and Pause Advanced AI

U.S. lawmakers Bernie Sanders and Greg Casar proposed permanently banning superintelligent AI, pausing advanced AI until regulated, and creating a new federal AI agency. Violators could face 20 years in prison.

Senator Bernie SandersSep 3, 2026
Industry

Meta Releases Muse Spark 1.3 For Agentic and Coding Tasks

Meta released Muse Spark 1.3, an AI model with enhanced agentic and coding capabilities. Its safety updates include improved adversarial robustness and better calibration on irreversible actions.

MetaSep 2, 2026
Industry
Cyber
Bio

Google DeepMind Introduces Gemini 3.8 Flash and Flash Cyber

Google DeepMind introduced Gemini 3.8 Flash and Flash Cyber. Flash Cyber excels in autonomous vulnerability discovery, while both models include safeguards against CBRN and cyber offense risks.

Google DeepMindSep 2, 2026
Industry
LoC

New OpenAI Technique Reportedly Makes Astra's Reasoning Harder to Monitor

The technique sharpens Astra's coding and computer-use skills but makes it reveal less reasoning, complicating oversight for bad behavior, per a source familiar with its development.

The InformationSep 2, 2026