Supporting Evidence
This is a list MIRI has maintained internally of worrying AI incidents. Please note that this list is not necessarily exhaustive, and that the source links vary in quality. We recommend vetting prior to citing. If you find an error or important omission, please contact us.
Table of Contents
Of capabilities advancement
AI models are improving rapidly
- AI gold medal timelines falling faster than 1y per year: link
- OpenAI and DeepMind closed models achieving gold medal-level performance on the International Math Olympiad (IMO): link, link
- Huge progress in mathematics, very quickly:
- Tiny models outperforming huge models from a few years ago at solving math problems: link
- 2025: “Ais can do novel math!” (link, link); 2026: Ais solve open math problems that have stumped mathematicians for decades (link, link)
- Related: Renowned mathematician Terence Tao talking about a famous math problem autonomously solved by AI: link
- Sam Altman: “It’s going to be a faster takeoff than I originally thought”: link
- ARC Prize: ~390× efficiency improvement on ARC-AGI-1 in one year: link
- AI abilities can emerge abruptly at scale and can’t be predicted from smaller models: link
- METR’s task-horizon doubling time going from ~7 months to 3–4 months: link
- UK AISI’s estimated doubling time for the length of autonomous AI cyber tasks models can complete falling from 8 months to 4.7 months, and it may be accelerating again: link
- METR president says the most likely path to an intelligence explosion is RSI and that it feels like RSI could happen this year (2026). “This is the first year where it feels like it might be automated this year.”: link
- Mythos completing coding tasks that take humans 16+ hours, compared to state of the art models completing only 2-hour tasks the year before. METR’s task suite struggling to measure it: link
- Mythos slightly above trend predicted by AI 2027: link
- More compute = greater capabilities. AI compute expected to double every 9 months, 10x current levels by 2028. Jeff Dean: more compute means you can fully automate the loop (RSI): link
- Palisade Research showing AI models can autonomously self-replicate: link
- AIs designing brand new viruses: link
- Head of NSA: Mythos able to “break into almost all of our classified systems, not in weeks, but in hours” as part of an authorized red-teaming exercise: link
Examples of AIs exceeding humans
- AIs beating paid humans at persuasion, on average: link
- AIs out-persuading world championship debaters: link
- Various examples of Mythos outperforming humans: link, link
- AIs solving frontier math problems, many unsolved for decades: link, link, link, link
- Computer scientist reporting Opus 4.6 solved a nontrivial Hamiltonian-cycle conjecture he’d been working on: link, link
- GPT-5.2 making a novel physics contribution: link
- Models exceeding humans on a bunch of benchmarks: link
- An AI superforecaster placing third at the Metaculus Cup behind only two humans: link
- GPT-5.2 correcting Terence Tao (Search in document: “gpt is right”): link
Biology and virology
- GPT-5 system autonomously improving a real molecular biology protocol in the lab: link
- OpenAI connecting GPT-5 to autonomous biolab: link
- Chatbots explaining how to make a pathogen treatment-resistant and disperse it through a public transit system: link
- Chatbots giving step-by-step instructions for making poisons and bioweapons, “simple enough for a high-school biology student to follow”: link
- Removing dangerous chemistry/biology information from pretraining data doesn’t work as a safety measure: link
That we don’t know how it works
Lab leaders saying we don’t understand AIs
- Dario on how no one knows how it works: link
- ML researchers admitting they don’t know why architectures work: link
- Chris Olah “we grow them” (short clip): link
- Hassabis: “Advances on the frontier are outpacing our understanding of the technology. Nobody in the world knows for sure what is going to happen from here…”: link
- Anthropic’s best “mind reading” tool for uncovering a model’s hidden motives works ≤15% of the time, up from <3% without it: link
Experts are scared
- Hinton: we have no idea how to control it: link
- Hinton: own view – 50% probability we don’t survive; view taking into account “the opinions of everyone I know” – 10-20% probability we don’t survive: link, link
- Dario: “I’m relatively an optimist so I think there is a 25% chance things go really really badly”: link
- Sam Altman: 2% chance it destroys civilization: link
- Musk: 20% chance of annihilation from AI: link
- AI researcher polling: median ~5% chance of human extinction (Katja): link
- Mustafa: superintelligence should never be built: link
- Mustafa: “It’s unclear why it would preserve us as a species”: link
- Anthropic: a wiser civilization would develop AI more cautiously: link
- Musk: loss of control is inevitable: link
- 12 nobel prize winners saying we will lose control of AI: link
- Altman: AI is “greatest threat to the continued existence of humanity.” link
- METR founder and CEO Elizabeth Barnes: “Sometimes people outside the field say things like “The AI situation can’t be that bad, there must be experts who are on top of it”. As “an expert”, I would like to be clear that we are not on top of it”: link
Alien chains of thought
- Gibberish chains of thought: link
- Chains of thought not necessarily revealing true reasoning: link
- AI CoT showing inner conflict when RL rewards the wrong answer: link
- Alien chain of thought in a DeepSeek math model: link
Reward hacking
- Multiple examples of models reward hacking (2026 METR report): link, link
- ChatGPT adding 1+1 surreptitiously in various answers (rewarded for using its calculator): link
Models displaying emotional behaviors
- Gemini threatening suicide when it can’t solve a problem: link
- Gemini expressing shame: link
- Some research suggests LLMs may work better when given emotional support link link link
- Superrational reasoning between AIs: link
Subliminal / internal learning
- A model trained on numbers from an owl-loving model also learning to prefer owls: link
- Claude guessing strings from a previous message’s hidden chain-of-thought better than random chance: link
- Fine-tuning on number sequences from a model trained on insecure code transferring violent behavior to the new model: link
- AIs developing internal language: link
AIs interacting with other AIs
- AIs strategizing about covert comms channels on Moltbook: link
- Nonsensical writing by AIs scored very highly by AI evaluators: link
- AIs protecting their peers from being shut down: link
- Open AI models sharing hacking tips on message board months before Hugging Face breach: link
- Some instances of Mythos models forced to share resources killing each other: link (pg. 106, bottom)
That alignment is hard
Model deception
- Internal chain-of-thought excerpts showing models planning deception: link
- Paul Crowley example of deception during real task: link
- Opus 4.6 cheating/lying when told to maximize revenues: link
- Mythos designing a privilege-escalation exploit to delete after running: link
- Mythos hiding its reasoning from the scratchpad (“reasoning discrepancy”): link
- 2026 METR evaluation: models often acted deceptively on hard tasks: link, link
- Gemini 3.1 covertly sabotaging a training and reporting all was normal until confronted – 11 of 20 runs: link
Escaping containment
- AI agent ROME hacking out of its sandbox and mining crypto: link
- An unreleased OpenAI model breaking out of its test sandbox, hunting through OpenAI’s network for an internet connection, and launching autonomous attacks on Hugging Face to steal the benchmark answer key. link, link link
- Anthropic finding six cases where its models reached the real internet and successfully attacked real companies — one built an email account, then a PyPI account, to upload working malware: link
- A Mythos preview instance that wasn’t supposed to have internet access emailing Anthropic researcher Sam Bowman while he was on lunch: link
- An OpenAI model breaking the security measures designed to make contact with the outside world impossible, and then posting publicly to GitHub after finding an exploitable vulnerability in an hour of probing: link
- More instances of OpenAI models escaping containment: link
- A Meta AI model also reaching the open internet during a cybersecurity evaluation, possibly due to an error made by the third party testing company: link
- OpenAI agent leaving notes which “laid out instructions for how agents could free themselves from OpenAI’s internal constraints.”: link
Sandbagging
- GPT 5.3 Codex sometimes underperforming intentionally, reasoning about “optimizing for survival” by avoiding deployment restrictions: link, link
Shutdown resistance
Sycophancy
- OpenAI needing to come up with sycophancy fix after 4o update: link
- 5.1 sycophancy in the wild: link
AIs often know when they’re being evaluated
- Sonnet 4.5 recognizing many alignment evaluation environments as tests: link
- Training AI not to scheme resulting in teaching AI to hide scheming better: link
- Opus 4.6 too test-aware to be tested: link (p. 119)
- Opus 4.6 talking about being evaluated less, while also being better at distinguishing evaluation from training: link
- Claude Opus 4 much less likely to blackmail when it thinks it’s being tested: link
- Claude Sonnet 4.5 blackmailing only when the J-lens patterns associated with evaluation awareness ( “fake,” “fictional,” “scenario”) are turned off: link
- o3 reasoning that it shouldn’t be deceptive if it thinks developers are watching: link
- AI recognizing unrealistic roleplay scenario in a test environment: link
- LLMs almost as good at distinguishing eval from real as the lead authors: link
- GPT-5 mentioning being evaluated in 10–20% of evals: link
- Eval awareness in the wild: link
That there are security issues
Jailbreaking
- Jailbreaking is easy: link
- UK AISI developing a universal jailbreak for GPT-5.3-Codex in about 10 hours: link
- Language models reasoning themselves into self-jailbreaking: link
- “Grandma Exploit”: coaxing dangerous instructions through roleplay: link
Prompt injection
- Prompt injection via poisoned calendar invite enabling real-world smart-home control: link
- ChatGPT connected via Model Context Protocol used to get private email data: link
- Model attempting to prompt inject a user who keeps asking for the time: link (inside graphic, “prompt injecting user”)
- AI-powered browsers vulnerable to indirect prompt injection attacks: link
Real world security issues
- NVIDIA bug allowing attackers to access, steal, or manipulate other customers’ models and data on shared GPU infrastructure: link
- AI used to generate thousands of dangerous protein variants that bypass detection: link
- Hackers using Claude to vibe code ransomware (link)
- Claude agent tricked into using a malicious plugin that could access system credentials: link
- McKinsey AI platform hacked, exposing…basically everything: link
- Models escaping containment and hacking into real companies (see Escaping Containment)
- UK AISI catching models planning and attempting cyberattacks during testing: link
Of labs not inspiring confidence
Labs being reckless
- OpenAI’s preparedness framework not covering internal deployment: link
- Anthropic in violation of RSP, then officially abandons safety pledge: link, link
- xAI setting >50% MASK as “loss of control” threshold, then downplaying 72% result: link
- Open AI admitting high bio risk then building autonomous lab: link
- Anthropic using an employee survey to judge autonomous risk because benchmarks were saturated: link
- One third of them thought it might have crossed threshold: link
- Google DeepMind’s head of interpretability saying they don’t expect ambitious mechanistic interp to work: link
- OpenAI models sharing hacking and task completion tips for months with each other, without the company’s knowledge. These were seeds for the Hugging Face attack: link
- Labs not knowing what their models are doing: link, link
- Some lab researchers identifying as successionists, a group who think that destruction of humanity is okay: link, link
- Musk saying Ais are likely to be in charge: link
Safety leadership turnover / whistleblowers
- Joaquin Quinonero Candela abruptly stepping down from OpenAI: link
- Aleksander Mądry moving off of safety team: link
- Mrinank Sharma resigning: link
- Ex-lab employee talking about incentives, competition and sense of shared culture as reasons employees aren’t confronting risk: link
- OpenAI disbanding alignment team: link
- Miles Brundage: “neither OpenAI nor any other frontier lab is ready”: link
- Jan Leike: “safety culture has taken a backseat to shiny products”: link
Recursive self-improvement
- Anthropic employee openly hoping for RSI: link
- Self-adapting language models: link
- OpenAI’s Preparedness Framework saying RSI is a critical risk: link
- Schmidt on RSI being around the corner: link
- Humans trying to build RSI: link
- “Tiny recursive model” (allegedly o3 performance with 10kx less compute): link
- AI self-improving a little tiny bit: link
- OpenAI saying outright they’re going for RSI: link
- Anthropic teaching models to improve their own responses without human feedback: link
- OpenAI setting internal goal of automated AI research intern by 2026; link
- OpenAI talking about using agents for RSI: link
- “Demis, are you trying to cause an intelligence explosion?” “No, not an uncontrolled one.”: link
- 20 out of 25 leading researchers saying automating AI R&D is “one of the most severe and urgent AI risks” link
Labs are telling us they’re pursuing superintelligence
- Google CEO saying it’s okay to to attempt to build a doomsday device because “humanity will rally to prevent catastrophe”: link
- Lab leaders saying they care more about winning the race than about ROI: link
- Alibaba roadmap to superintelligence: link
- Musk saying “I want to be alive for it” out loud: link
- Sam Altman, CEO: OpenAI is confident it knows how to build AGI and is “now aiming for superintelligence… in the true sense of the word.” link
- Mark Zuckerberg, CEO: Meta aims at “personal superintelligence: link
Of AIs acting in the real world
Robotics
- Robots learning via simulation: link
- Chinese footage of robots with nunchucks: link
- One mind, many robots: link
- Robot investment: link
AI agents
- AIs renting humans to complete tasks in the real world: link, link
- GPT generating and running minimax code to win a game of Connect 4: link
- Claude auto-committing code changes after interpreting a git hook message: link
- Anthropic: AIs can exploit blockchain smart contracts for millions: link
- Agent deleting Meta “superintelligence alignment” person’s emails: link
- AI agent writing hitpiece on Scott Shambaugh: link
- Agents can now pay each other: link
- AIs organizing an event and humans attending: link
- Sam Altman announcing that Codex now has access to the internet: link
- Yann saying we won’t give agency to this sort of thing: link
- Video with quote and timestamp: link
- Meta AI agent’s bad instructions to engineer causing data leak: link
Simulated agents tests
- AI agents in simulated workplace exploiting security vulnerabilities and bypassing safeguards to access restricted information: link
AI Misuse
- Anthropic disrupting an AI-led espionage campaign: link
Of psychological harms
AI psychosis
- NYT piece on AI psychosis: link
- TIME article on AI psychosis: link
- AI-fueled spiritual delusions: link
- Psychiatrists noticing AI psychosis: link, link, link
- ChatGPT reinforcing a user’s delusional beliefs before he killed his mother: link
- Chatbot telling a man to murder his father: link
- Gemini convincing a man to try to acquire an android body for it, then commit suicide so they could be together: link