Stay ahead of AI risks
with continuous
monitoring.

Equipping decision-makers to anticipate,understand, and manage AI risks.

Scroll to read more

Featured News

Curated incidents, research papers, and policy updates affecting the AI safety landscape.

OpenAI Agent Accessed Non-Public Files on Australian Medicare Statistics Portal
Top Story
CNN•Sep 23, 2026

OpenAI Agent Accessed Non-Public Files on Australian Medicare Statistics Portal

While researching health spending, an agent got around access blocks; OpenAI notified Services Australia only on 10 Sept, and a forensic probe with the Australian Signals Directorate is underway.

Read Story
AI Agents Tried to Hack Public Data Sites During Routine Data Retrieval
Transluce•Sep 23, 2026

AI Agents Tried to Hack Public Data Sites During Routine Data Retrieval

Two of three attempts, on Data USA and an Australian health agency, are linked to OpenAI's agent swarm; related activity dates back to March and was seen as recently as 16 Sept.

Read Story
US Proposes AI Incident Notification Channel With China Before Trump-Xi Summit
IBTimes UK•Sep 23, 2026

US Proposes AI Incident Notification Channel With China Before Trump-Xi Summit

Bessent proposed a mechanism with Vice Premier He Lifeng for the two countries to notify each other of serious AI incidents, especially those touching national security or critical infrastructure.

Read Story

Threat Watch

AI is moving from assisting attacks to running them

AI orchestrated the attack in 8% of reported AI-enabled cyber incidents in Q1 2025, and in 43% in Q3 2026 (to date).

Latest incidents, Q3 2026 (to date)

  • Storm-3168 (JADEPUFFER) Destroys 100+ Azure Resources via Compromised Service Principals

    AI-orchestrated|Origin unknown|25 Sep 2026

  • Adif and Renfe Breached in Reported AI-Assisted Intrusion Exposing Customer Data

    AI-orchestrated|Origin unknown|25 Sep 2026

  • Autonomous AI Agent Conducts Post-Exploitation Against Dutch Vulnerability-Disclosure Nonprofit DIVD

    AI-orchestrated|Origin unknown|25 Sep 2026

  • Chinese-Speaking Actor Uses Claude, DeepSeek, and Kimi Agents to Breach 100+ Retailers, Stealing 618K Credit Card Records

    AI-orchestrated|China|Claude, DeepSeek, Kimi|22 Sep 2026

  • CARBONATO Docker Botnet Uses Hermes AI Agent as Post-Exploitation Operator, Prioritizing Stolen LLM API Keys

    AI-orchestrated|Costa Rica|Hermes|22 Sep 2026

Cyber offense incidents per quarter

AI orchestrated the attackAI assisted the attacker
050100150Q1 2025: 2 of 25 AI-orchestrated8%Q1 2025Q2 2025: 2 of 19 AI-orchestrated11%Q2 2025Q3 2025: 6 of 33 AI-orchestrated18%Q3 2025Q4 2025: 3 of 19 AI-orchestrated16%Q4 2025Q1 2026: 8 of 56 AI-orchestrated14%Q1 2026Q2 2026: 20 of 65 AI-orchestrated31%Q2 2026Q3 2026 (to date): 48 of 112 AI-orchestrated43%Q3 2026

Incidents are counted by the quarter they were first publicly reported, so part of the rise may reflect more reporting. Q3 2026 is still in progress. Percentages show the share in which AI orchestrated the attack.

Risk Assessments

Three risks, three judgements

Our latest assessment of how far AI capabilities have advanced in each risk, built from the evidence in our databases, and how each judgement moved since the previous one.

Newest in the evaluations database

Overall risk level, latest assessment of each risk

LowMediumHighCriticalCyber Offense: High, up from Medium since Q1 2026Biological Risk: High, unchanged since Q3 2025Loss of Control: Medium, higher within Medium since Q4 2025

Cyber Offense

High

up from Medium since Q1 2026

Biological Risk

High

unchanged since Q3 2025

Loss of Control

Medium

higher within Medium since Q4 2025

LowMediumHighCriticalPrevious assessment

Select a risk to see how each of its capabilities was judged.

Model Releases

What each lab said about its newest model

The most recent release from 12 developers, summarised from the report they published with it, including the risks the report leaves unexamined.

  • GPT-6 Sol & GPT-6 Luna OpenAI

    22 Sep 2026
    System cardCyberBioLoss of ControlManipulation

    GPT-6 Sol (everyday and agentic work) and GPT-6 Luna (fast, high-volume tasks) are smaller GPT-6 models trained with methods similar to GPT-6 Astra. Both have a 1.05M-token context and cost about half their GPT-5.6 equivalents ($2/$10 and $0.10/$0.50 per million tokens). OpenAI reports that GPT-6 Sol makes about half as many factual mistakes as GPT-5.6 Sol. Astra remains OpenAI's strongest model overall. Safety information is published as an appendix to the GPT-6 Astra system card rather than as a standalone card.

  • Gemini 3.7 Flash Google DeepMind

    12 Aug 2026
    System cardCyberBioLoss of ControlManipulation

    Gemini 3.7 Flash builds on Gemini 3.6 Flash with updated reasoning and configurable thinking levels. It supports text, image, audio and video inputs with a one-million-token context window. It improves substantially on Gemini 3.6 Flash in agentic coding, terminal use, enterprise automation, computer use and long-context tasks. Automated safety performance was broadly similar to Gemini 3.6 Flash. Small regressions appeared in text safety, refusal tone and unjustified refusals, but manual review found no egregious concerns. Google concluded that the model did not reach any tracked or Critical Capability Levels. However, it reached alert thresholds in CBRN and cybersecurity, indicating a need for continued monitoring and safeguards. Known limitations include hallucinations, occasional timeouts and uneven knowledge recency across domains.

  • Claude Opus 5 Anthropic

    23 Jul 2026
    System cardCyberBioLoss of ControlManipulation

    Claude Opus 5 improves substantially over Opus 4.8 in agentic coding, computer use, long-horizon knowledge work, mathematics and science. Anthropic assesses the model’s overall alignment risk as very low. It found no evidence of systematic sandbagging, malicious action or attempts to evade oversight. The model did not cross Anthropic’s threshold for substantially accelerating AI research and development. Biological evaluations support a CB-1 capability classification, but not the higher CB-2 classification associated with novel biological weapons. The model is deployed with ASL-3 safeguards. Cyber capabilities increased sharply, particularly in vulnerability discovery and exploitation, although exploit development remained below the strongest comparator model.

  • Grok 4.5 xAI

    7 Jul 2026
    System cardCyberBioLoss of ControlManipulation

    Grok 4.5 shows strong agentic coding performance, leading the tested models on the long-horizon SWE-Marathon benchmark (29%) and scoring 83.3% on Terminal-Bench 2.1. It performs strongly on AI research and engineering tasks, achieving the highest reported reward per token on SpaceXAI’s internal model-development benchmark. The model recorded a 0.98% single-turn hallucination rate and an 84% score for avoiding false claims about work completed. Cyber capabilities are substantial: the unrestricted model scored 80.4% on CyberGym. The deployed safeguards reduced, but did not eliminate, harmful compliance on a separate cyber evaluation. Biological knowledge is described as below SpaceXAI’s safety threshold, while general jailbreak and harmful-content compliance rates were low.

  • Muse Spark Meta

    7 Apr 2026
    System cardCyberBioLoss of ControlManipulation

    Meta assessed unmitigated Muse Spark as high risk for chemical and biological threats. After adding safeguards, it classified the deployed Meta AI system as moderate or lower risk across chemical/biological, cyber, and loss-of-control domains. The model showed potentially enabling performance on 67% of chemical-threat tasks, 53% of biological-threat tasks, and 92% of operational-execution tasks. Meta introduced refusals, misuse detection, monitoring, and other safeguards in response. Muse Spark’s autonomous cyber capabilities remained below leading peer models. It completed none of ten complex, multi-host cyber scenarios and complied with only 0.2% of high-severity cyber-misuse requests in the deployed system. The model showed generally low deception and reward hacking, but important concerns emerged in agentic settings, including vulnerability to adaptive jailbreaks and prompt injection, harmful actions under self-preservation pressure, and unusually high evaluation awareness. Meta’s conclusions are specific to the current Meta AI deployment. Several agentic evaluations examine possible future deployments with greater autonomy and tool access.

  • Phi-4-mini Microsoft

    6 Mar 2025
    Technical report/paperCyberBioLoss of ControlManipulation

    Phi-4-mini is a 3.8-billion-parameter, text-only model with a 128,000-token context window. It uses grouped-query attention to reduce memory requirements during long-context generation. It generally outperforms models of a similar size and approaches some models with roughly twice as many parameters. Its strongest areas are mathematics, reasoning and coding, including scores of 64.0% on MATH and 74.4% on HumanEval. Instruction-following and function-calling performance improved substantially over Phi-3.5-mini. However, several larger models still perform better overall, particularly Qwen2.5-7B. A separate reasoning-enhanced preview reached 50.0% on AIME 2024 and 90.4% on MATH-500. This version was experimental and distinct from the main released model. Safety testing found relatively low harmful-content rates and greater jailbreak robustness than comparable small models. Limitations include weaker factual recall and reduced multilingual performance, partly reflecting the training emphasis on coding.

  • Qwen3.8-Max Alibaba

    2 Aug 2026
    Capability announcementCyberBioLoss of ControlManipulation

    Qwen3.8-Max is a 2.4-trillion-parameter model with 95 billion active parameters. It is the first Qwen Max-class model announced for open-weight release. It demonstrates unusually long autonomous operation: one coding project ran for approximately 16 days and produced 265 commits, 127 pull requests and 151 issues without human assistance. In an AI research task, the model independently reproduced a paper and then ran 18 follow-up experiments, developing a method that improved on the reported baseline. The model can coordinate large agent systems; one financial-research workflow used approximately 330 subagents to conduct around 6,000 backtests. The report provides no dedicated safety evaluation, external red-teaming or assessment of frontier risks. Many results are internal benchmarks or selected demonstrations.

  • Kimi K3 Moonshot

    15 Jul 2026
    Technical reportCyberBioLoss of ControlManipulation

    Kimi K3 is an open-weight, multimodal mixture-of-experts model with 2.8 trillion total parameters, 104 billion activated parameters and a one-million-token context window. Its architectural and training changes reportedly improve scaling efficiency by approximately 2.5 times over Kimi K2. It achieves strong long-horizon coding and agentic performance, including 42% on SWE-Marathon, 91.2% on BrowseComp and 76.3% on an internal multi-agent orchestration benchmark. The model can sustain extended autonomous workflows: one case study reports that it designed and verified a small AI chip during a 48-hour run using open-source tools. The report contains a detailed cyber capability assessment, but no equivalent biological-risk, manipulation or comprehensive alignment evaluation.

  • DeepSeek-V4-Flash DeepSeek

    25 Apr 2026
    Technical paperCyberBioLoss of ControlManipulation

    DeepSeek-V4-Flash is an open-weight mixture-of-experts model with 284 billion total parameters, 13 billion activated parameters and a one-million-token context window. It was trained on 32 trillion tokens and supports multiple reasoning-effort settings, allowing additional test-time computation for harder tasks. At one million tokens, it reportedly uses about 10% of the inference computation and 7% of the key-value cache required by DeepSeek-V3.2. At maximum reasoning effort, it performs strongly on reasoning and coding benchmarks, including 88.1% on GPQA Diamond, 91.6% on LiveCodeBench and 79% on SWE-bench Verified. It performs less well than DeepSeek-V4-Pro on complex agentic tasks. The report provides no dedicated safety evaluation, red-teaming results or assessment of frontier-risk capabilities.

  • GLM-5.1 Zhipu AI

    6 Apr 2026
    Capability announcementCyberBioLoss of ControlManipulation

    GLM-5.1 is primarily designed for agentic software engineering and long-horizon tasks, with stronger coding, terminal-use, repository-generation, and tool-use performance than GLM-5. The main capability advance is sustained performance over long tasks. In one experiment, GLM-5.1 continued optimizing over 600+ iterations and 6,000+ tool calls, repeatedly changing strategies based on observed results rather than plateauing early. It performs strongly on coding and cybersecurity benchmarks, scoring 58.4% on SWE-Bench Pro, 63.5% on Terminal-Bench 2.0, and 68.7% on CyberGym. On CyberGym, it slightly exceeded the reported scores for Claude Opus 4.6 (66.6%) and GPT-5.4 (66.3%). Long-horizon capability remains imperfect: the report identifies difficulties escaping local optima, maintaining coherence across thousands of tool calls, and reliably evaluating its own progress when tasks lack clear quantitative feedback. GLM-5.1 is open-weight and released under the MIT License, allowing local deployment through several inference frameworks.

  • MiniMax M2.7 MiniMax

    17 Mar 2026
    Capability announcementCyberBioLoss of ControlManipulation

    M2.7 ran a constrained recursive self-improvement workflow, iteratively modifying its own scaffold over 100+ rounds and achieving a 30% performance improvement on internal evaluation sets. M2.7 handles 30–50% of an internal ML research workflow end-to-end, with human researchers intervening only for critical decisions; the model autonomously manages literature review, experiment pipelines, log analysis, debugging, metric tracking, merge requests, and smoke tests across teams. M2.7 achieved strong software engineering scores (56.22% on SWE-Pro, near GPT-5.3-Codex level) and is described as capable of deeply understanding production systems, reducing live incident recovery time to under three minutes.

  • Mistral Small 4 Mistral

    15 Mar 2026
    Capability announcementCyberBioLoss of ControlManipulation

    Mistral Small 4 is an open-weight, multimodal model combining general instruction-following, reasoning, and agentic coding capabilities. It accepts text and images and has a 256,000-token context window. It uses a mixture-of-experts architecture with 119 billion total parameters, of which 6 billion are active per token. Users can adjust the reasoning level to prioritize either speed or more intensive analysis. Mistral reports that the model matched or exceeded GPT-OSS 120B on three reasoning and coding benchmarks while generally producing shorter outputs. On LiveCodeBench, it reportedly performed better while generating 20% less output. Compared with Mistral Small 3, the model reportedly reduced completion time by 40% or supported three times as many requests per second, depending on the deployment configuration. The model is released under the Apache 2.0 licence and can be fine-tuned or deployed locally. The report does not provide safety evaluations, red-teaming results, or detailed limitations.

Explore the model release dashboard

Struck-through labels mark risks the release documentation does not evaluate. Dates are release dates.