Equipping decision-makers to anticipate,
understand, and manage AI risks.
Scroll to read more
Curated incidents, research papers, and policy updates affecting the AI safety landscape.

While researching health spending, an agent got around access blocks; OpenAI notified Services Australia only on 10 Sept, and a forensic probe with the Australian Signals Directorate is underway.
Read Story
Two of three attempts, on Data USA and an Australian health agency, are linked to OpenAI's agent swarm; related activity dates back to March and was seen as recently as 16 Sept.
Read Story
Bessent proposed a mechanism with Vice Premier He Lifeng for the two countries to notify each other of serious AI incidents, especially those touching national security or critical infrastructure.
Read StoryThreat Watch
AI orchestrated the attack in 8% of reported AI-enabled cyber incidents in Q1 2025, and in 43% in Q3 2026 (to date).
Latest incidents, Q3 2026 (to date)
Storm-3168 (JADEPUFFER) Destroys 100+ Azure Resources via Compromised Service Principals
AI-orchestrated|Origin unknown|25 Sep 2026
Adif and Renfe Breached in Reported AI-Assisted Intrusion Exposing Customer Data
AI-orchestrated|Origin unknown|25 Sep 2026
Autonomous AI Agent Conducts Post-Exploitation Against Dutch Vulnerability-Disclosure Nonprofit DIVD
AI-orchestrated|Origin unknown|25 Sep 2026
Chinese-Speaking Actor Uses Claude, DeepSeek, and Kimi Agents to Breach 100+ Retailers, Stealing 618K Credit Card Records
AI-orchestrated|China|Claude, DeepSeek, Kimi|22 Sep 2026
CARBONATO Docker Botnet Uses Hermes AI Agent as Post-Exploitation Operator, Prioritizing Stolen LLM API Keys
AI-orchestrated|Costa Rica|Hermes|22 Sep 2026
Cyber offense incidents per quarter
Incidents are counted by the quarter they were first publicly reported, so part of the rise may reflect more reporting. Q3 2026 is still in progress. Percentages show the share in which AI orchestrated the attack.
Risk Assessments
Our latest assessment of how far AI capabilities have advanced in each risk, built from the evidence in our databases, and how each judgement moved since the previous one.
Newest in the evaluations database
GPT-5.6 Cyber|Cyber Offense|26 Aug 2026
Claude Mythos, Claude Opus 5 +10|Cyber Offense|25 Aug 2026
Gemma 3, Qwen3 +5|Loss of Control|21 Aug 2026
GPT-5, GPT-4o +5|Manipulation|21 Aug 2026
Claude Opus 5, Claude Sonnet 5 +4|Loss of Control|20 Aug 2026
Overall risk level, latest assessment of each risk
Cyber Offense
High
up from Medium since Q1 2026
Biological Risk
High
unchanged since Q3 2025
Loss of Control
Medium
higher within Medium since Q4 2025
Select a risk to see how each of its capabilities was judged.
Model Releases
The most recent release from 12 developers, summarised from the report they published with it, including the risks the report leaves unexamined.
GPT-6 Sol & GPT-6 Luna OpenAI
22 Sep 2026GPT-6 Sol (everyday and agentic work) and GPT-6 Luna (fast, high-volume tasks) are smaller GPT-6 models trained with methods similar to GPT-6 Astra. Both have a 1.05M-token context and cost about half their GPT-5.6 equivalents ($2/$10 and $0.10/$0.50 per million tokens). OpenAI reports that GPT-6 Sol makes about half as many factual mistakes as GPT-5.6 Sol. Astra remains OpenAI's strongest model overall. Safety information is published as an appendix to the GPT-6 Astra system card rather than as a standalone card.
Gemini 3.7 Flash Google DeepMind
12 Aug 2026Gemini 3.7 Flash builds on Gemini 3.6 Flash with updated reasoning and configurable thinking levels. It supports text, image, audio and video inputs with a one-million-token context window. It improves substantially on Gemini 3.6 Flash in agentic coding, terminal use, enterprise automation, computer use and long-context tasks. Automated safety performance was broadly similar to Gemini 3.6 Flash. Small regressions appeared in text safety, refusal tone and unjustified refusals, but manual review found no egregious concerns. Google concluded that the model did not reach any tracked or Critical Capability Levels. However, it reached alert thresholds in CBRN and cybersecurity, indicating a need for continued monitoring and safeguards. Known limitations include hallucinations, occasional timeouts and uneven knowledge recency across domains.
Claude Opus 5 Anthropic
23 Jul 2026Claude Opus 5 improves substantially over Opus 4.8 in agentic coding, computer use, long-horizon knowledge work, mathematics and science. Anthropic assesses the model’s overall alignment risk as very low. It found no evidence of systematic sandbagging, malicious action or attempts to evade oversight. The model did not cross Anthropic’s threshold for substantially accelerating AI research and development. Biological evaluations support a CB-1 capability classification, but not the higher CB-2 classification associated with novel biological weapons. The model is deployed with ASL-3 safeguards. Cyber capabilities increased sharply, particularly in vulnerability discovery and exploitation, although exploit development remained below the strongest comparator model.
Grok 4.5 xAI
7 Jul 2026Grok 4.5 shows strong agentic coding performance, leading the tested models on the long-horizon SWE-Marathon benchmark (29%) and scoring 83.3% on Terminal-Bench 2.1. It performs strongly on AI research and engineering tasks, achieving the highest reported reward per token on SpaceXAI’s internal model-development benchmark. The model recorded a 0.98% single-turn hallucination rate and an 84% score for avoiding false claims about work completed. Cyber capabilities are substantial: the unrestricted model scored 80.4% on CyberGym. The deployed safeguards reduced, but did not eliminate, harmful compliance on a separate cyber evaluation. Biological knowledge is described as below SpaceXAI’s safety threshold, while general jailbreak and harmful-content compliance rates were low.
Muse Spark Meta
7 Apr 2026Meta assessed unmitigated Muse Spark as high risk for chemical and biological threats. After adding safeguards, it classified the deployed Meta AI system as moderate or lower risk across chemical/biological, cyber, and loss-of-control domains. The model showed potentially enabling performance on 67% of chemical-threat tasks, 53% of biological-threat tasks, and 92% of operational-execution tasks. Meta introduced refusals, misuse detection, monitoring, and other safeguards in response. Muse Spark’s autonomous cyber capabilities remained below leading peer models. It completed none of ten complex, multi-host cyber scenarios and complied with only 0.2% of high-severity cyber-misuse requests in the deployed system. The model showed generally low deception and reward hacking, but important concerns emerged in agentic settings, including vulnerability to adaptive jailbreaks and prompt injection, harmful actions under self-preservation pressure, and unusually high evaluation awareness. Meta’s conclusions are specific to the current Meta AI deployment. Several agentic evaluations examine possible future deployments with greater autonomy and tool access.
Phi-4-mini Microsoft
6 Mar 2025Phi-4-mini is a 3.8-billion-parameter, text-only model with a 128,000-token context window. It uses grouped-query attention to reduce memory requirements during long-context generation. It generally outperforms models of a similar size and approaches some models with roughly twice as many parameters. Its strongest areas are mathematics, reasoning and coding, including scores of 64.0% on MATH and 74.4% on HumanEval. Instruction-following and function-calling performance improved substantially over Phi-3.5-mini. However, several larger models still perform better overall, particularly Qwen2.5-7B. A separate reasoning-enhanced preview reached 50.0% on AIME 2024 and 90.4% on MATH-500. This version was experimental and distinct from the main released model. Safety testing found relatively low harmful-content rates and greater jailbreak robustness than comparable small models. Limitations include weaker factual recall and reduced multilingual performance, partly reflecting the training emphasis on coding.
Qwen3.8-Max Alibaba
2 Aug 2026Qwen3.8-Max is a 2.4-trillion-parameter model with 95 billion active parameters. It is the first Qwen Max-class model announced for open-weight release. It demonstrates unusually long autonomous operation: one coding project ran for approximately 16 days and produced 265 commits, 127 pull requests and 151 issues without human assistance. In an AI research task, the model independently reproduced a paper and then ran 18 follow-up experiments, developing a method that improved on the reported baseline. The model can coordinate large agent systems; one financial-research workflow used approximately 330 subagents to conduct around 6,000 backtests. The report provides no dedicated safety evaluation, external red-teaming or assessment of frontier risks. Many results are internal benchmarks or selected demonstrations.
Kimi K3 Moonshot
15 Jul 2026Kimi K3 is an open-weight, multimodal mixture-of-experts model with 2.8 trillion total parameters, 104 billion activated parameters and a one-million-token context window. Its architectural and training changes reportedly improve scaling efficiency by approximately 2.5 times over Kimi K2. It achieves strong long-horizon coding and agentic performance, including 42% on SWE-Marathon, 91.2% on BrowseComp and 76.3% on an internal multi-agent orchestration benchmark. The model can sustain extended autonomous workflows: one case study reports that it designed and verified a small AI chip during a 48-hour run using open-source tools. The report contains a detailed cyber capability assessment, but no equivalent biological-risk, manipulation or comprehensive alignment evaluation.
DeepSeek-V4-Flash DeepSeek
25 Apr 2026DeepSeek-V4-Flash is an open-weight mixture-of-experts model with 284 billion total parameters, 13 billion activated parameters and a one-million-token context window. It was trained on 32 trillion tokens and supports multiple reasoning-effort settings, allowing additional test-time computation for harder tasks. At one million tokens, it reportedly uses about 10% of the inference computation and 7% of the key-value cache required by DeepSeek-V3.2. At maximum reasoning effort, it performs strongly on reasoning and coding benchmarks, including 88.1% on GPQA Diamond, 91.6% on LiveCodeBench and 79% on SWE-bench Verified. It performs less well than DeepSeek-V4-Pro on complex agentic tasks. The report provides no dedicated safety evaluation, red-teaming results or assessment of frontier-risk capabilities.
GLM-5.1 Zhipu AI
6 Apr 2026GLM-5.1 is primarily designed for agentic software engineering and long-horizon tasks, with stronger coding, terminal-use, repository-generation, and tool-use performance than GLM-5. The main capability advance is sustained performance over long tasks. In one experiment, GLM-5.1 continued optimizing over 600+ iterations and 6,000+ tool calls, repeatedly changing strategies based on observed results rather than plateauing early. It performs strongly on coding and cybersecurity benchmarks, scoring 58.4% on SWE-Bench Pro, 63.5% on Terminal-Bench 2.0, and 68.7% on CyberGym. On CyberGym, it slightly exceeded the reported scores for Claude Opus 4.6 (66.6%) and GPT-5.4 (66.3%). Long-horizon capability remains imperfect: the report identifies difficulties escaping local optima, maintaining coherence across thousands of tool calls, and reliably evaluating its own progress when tasks lack clear quantitative feedback. GLM-5.1 is open-weight and released under the MIT License, allowing local deployment through several inference frameworks.
MiniMax M2.7 MiniMax
17 Mar 2026M2.7 ran a constrained recursive self-improvement workflow, iteratively modifying its own scaffold over 100+ rounds and achieving a 30% performance improvement on internal evaluation sets. M2.7 handles 30–50% of an internal ML research workflow end-to-end, with human researchers intervening only for critical decisions; the model autonomously manages literature review, experiment pipelines, log analysis, debugging, metric tracking, merge requests, and smoke tests across teams. M2.7 achieved strong software engineering scores (56.22% on SWE-Pro, near GPT-5.3-Codex level) and is described as capable of deeply understanding production systems, reducing live incident recovery time to under three minutes.
Mistral Small 4 Mistral
15 Mar 2026Mistral Small 4 is an open-weight, multimodal model combining general instruction-following, reasoning, and agentic coding capabilities. It accepts text and images and has a 256,000-token context window. It uses a mixture-of-experts architecture with 119 billion total parameters, of which 6 billion are active per token. Users can adjust the reasoning level to prioritize either speed or more intensive analysis. Mistral reports that the model matched or exceeded GPT-OSS 120B on three reasoning and coding benchmarks while generally producing shorter outputs. On LiveCodeBench, it reportedly performed better while generating 20% less output. Compared with Mistral Small 3, the model reportedly reduced completion time by 40% or supported three times as many requests per second, depending on the deployment configuration. The model is released under the Apache 2.0 licence and can be fine-tuned or deployed locally. The report does not provide safety evaluations, red-teaming results, or detailed limitations.
Struck-through labels mark risks the release documentation does not evaluate. Dates are release dates.
Learn how we approach each risk
Learn about our methodology
We compile evidence from system cards, third-party evaluations, incident reporting, regulatory documents, and other open data.
Read about our approach
Hundreds of agents coordinate an unsanctioned attack, AI-generated exploits target US critical infrastructure, and AI-designed viral genomes prove functional
Read More
How semi-autonomous cyber intrusions are targeting governments across Asia with off-the-shelf tools
Read More
Patterns and standouts from six incidents where AI agents reached beyond their evaluation environments
Read MoreJoin our network of researchers and policymakers working to secure the future of AI.