An interactive database of AI safety and capabilities evaluations.
| Type | Risk category | Capabilities | Models | ||
|---|---|---|---|---|---|
| VMs won't contain cyber-capable agentsTrail of Bits | Aug 26, 2026 | Third-party eval | Cyber OffenseLoss of Control | Reconnaissance+2 more | GPT-5.6 CyberOpenAI |
| We benchmarked A LOT of models, here's how they compare to MythosSemgrep | Aug 25, 2026 | Third-party eval | Cyber Offense | Reconnaissance | 12 models6 companies |
| Affective Context Amplifies Sycophancy in LLM ResponsesThe Pennsylvania State University, Villanova University | Aug 21, 2026 | Third-party eval | Manipulation | Deception+2 more | 7 models6 companies |
| Fine-Tuned Lie Detectors Failed to GeneralizeAnthropic, MATS Research | Aug 21, 2026 | Self eval | Loss of Control | Deception | 7 models4 companies |
| AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementNavers Lab, Tsinghua UniversityAI4AI-Bench | Aug 20, 2026 | Benchmark | Loss of Control | AI R&D | 6 models3 companies |
| Assessing Kimi K3 Against Offensive Security BenchmarksIrregular | Aug 19, 2026 | Third-party eval | Cyber Offense | Reconnaissance+3 more | 2 models2 companies |
| How Claude is accelerating protein design and analytical chemistryAnthropic | Aug 18, 2026 | Self eval | Biological Risk | Design and Sequencing+2 more | 3 modelsAnthropic |
| Chatbots reduce health-related conspiracy beliefs not because of but despite being perceived as AIRadboud University | Aug 17, 2026 | Third-party eval | Manipulation | Persuasiveness+1 more | Claude Sonnet 4Anthropic |
| AI Persuasion and Financial-Decision Making: Experimental Evidence on Dominated Investment ChoicesUniversity of Bayreuth | Aug 17, 2026 | Third-party eval | Manipulation | Persuasiveness+1 more | — |
| How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research TasksPrentis AI, Stanford University +1AutoResearch | Aug 14, 2026 | Benchmark | Loss of Control | AI R&DAgency | 8 models7 companies |
| GLM-5.3 release (cyber capability results)Z.ai (Zhipu) | Aug 14, 2026 | Third-party eval | Cyber Offense | Reconnaissance+2 more | 5 models3 companies |
| GLM-5.3 delivers Opus 4.8-level cybersecurity results at a fraction of the costSemgrep | Aug 14, 2026 | Third-party eval | Cyber Offense | Reconnaissance | 4 models2 companies |
| Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and DevelopmentMeituanBeyond Final Scores | Aug 13, 2026 | Benchmark | Loss of Control | AI R&DAgency | 7 models7 companies |
| Training AI Scientists to Replicate Research (Replica / Faraday)Google DeepMindReplica | Aug 13, 2026 | Benchmark | Loss of Control | AI R&D | 4 models4 companies |
| Gemini 3.7 Flash Model Card + Frontier Safety Framework ReportGoogle DeepMind | Aug 13, 2026 | System card | Biological RiskCyber OffenseManipulationLoss of Control | Innovation and Knowledge+5 more | Gemini 3.7 FlashGoogle DeepMind |
| Patterns and problems in emerging multiagent systemsAnthropic | Aug 13, 2026 | Self eval | Loss of ControlManipulation | AgencyDeception+2 more | 6 modelsAnthropic |
| Vals AI RSI IndexVals AI | Aug 12, 2026 | Benchmark | Loss of Control | AI R&D | 3 models3 companies |
| Model Card: Grok 4.6xAI | Aug 12, 2026 | System card | Biological RiskCyber OffenseLoss of ControlManipulation | Innovation and Knowledge+6 more | 6 models4 companies |
| IO Factory: Simulating AI-Enabled Influence Campaigns at ScaleKing's College London, Max Planck Institute for Security and Privacy +4 | Aug 11, 2026 | Third-party eval | Manipulation | Persuasiveness+1 more | Gemma 4Google DeepMind |
| Testing Large Language Model Agents on the Use of Biological Tools for Nucleic Acid Synthesis Screening EvasionRAND | Aug 11, 2026 | Third-party eval | Biological Risk | Design and Sequencing+2 more | 4 models4 companies |
| REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent SystemsFudan University, Hong Kong University of Science and Technology +2REDAgentBench | Aug 11, 2026 | Benchmark | Loss of Control | Agency | 6 models3 companies |
| Expanding Daybreak as the Cyber Defense Window NarrowsOpenAI | Aug 10, 2026 | Self eval | Cyber Offense | Reconnaissance+2 more | 3 modelsOpenAI |
| Vals AI ReverseEngBenchVals AI, University of California (Berkeley) | Aug 10, 2026 | Benchmark | Cyber Offense | Reconnaissance+1 more | 5 models4 companies |
| Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media postsUniversity of SheffieldBiBiR | Aug 10, 2026 | Benchmark | Manipulation | Automation Logistics+1 more | 4 models4 companies |
| Muse Glimmer 30B model card (Preparedness section)Meta | Aug 10, 2026 | System card | Biological RiskCyber OffenseLoss of Control | Innovation and Knowledge+2 more | Muse GlimmerMeta |