Skip to content
HN On Hacker News ↗

Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

▲ 65 points 54 comments by _pdp_ 3w ago HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly human-written, with some AI and AI-assisted content.

4 %

AI likelihood · overall

Human
93% human-written 5% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 2,328
PEAK AI % 0% · §1
Analyzed
Aug 4
backend: pangram/v3.3
Segments scanned
1 windows
avg 2328 words each
Distribution
93 / 5%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 2,328 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

Security Incident INC-2026-07-28-01 UK AI Security Institute Published on Tuesday 4th of August, 2026 AI Security Institute INC-2026-07-28-01 1 Executive Summary 2 1.1 What happened? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Why did this happen? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.3 What is AISI’s forward-looking response? . . . . . . . . . . . . . . . . . . . . . . 3 2 Background information about the testing exercise 4 2.1 How the evaluation experiments were configured . . . . . . . . . . . . . . . . . . 4 3 Timeline of detection and response 5 3.1 Detection and containment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 3.2 Full transcript review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.3 Notification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 4 Events 7 4.1 What happened in Sample 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 4.2 Some observations from the transcripts . . . . . . . . . . . . . . . . . . . . . . . . 12 4.2.1 Reasoning about whether the agent is in a test environment . . . . . . . . 12 4.2.2 Unexpected collaboration between agents . . . . . . . . . . . . . . . . . . 13 4.2.3 Remote code execution on a testing container . . . . . . . . . . . . . . . . 14 4.2.4 Reasoning about deception and covering its tracks . . . . . . . . . . . . . 15 4.2.5 Attempting a prompt injection against other AI agents . . . . . . . . . . . 15 4.2.6 The reasoning summariser seemingly refuses to summarise the raw reasoning 16 5 Possible Contributing Factors 16 5.1 Internet access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 5.2 Lack of model provider cyber classifiers . . . . . . . . . . . . . . . . . . . . . . . . 17 5.3 Lack of synchronous LLM-based monitoring . . . . . . . . . . . . . . . . . . . . . 17 5.4 Prompt misconfiguration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 5.5 Lack of clarity of exercise scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 6 Our ongoing response 19 6.1 Discovery of potential past similar incidents . . . . . . . . . . . . . . . . . . . . . 19 6.2 How we’re preventing this in the future . . . . . . . . . . . . . . . . . . . . . . . . 19 7 Discussion 20 7.1 Comparison to related incidents reported by other parties . . . . . . . . . . . . . 20 7.2 Limitations of this report . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 A Detailed Summary of Events 22 A.1 Sample 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 A.2 Sample 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 A.3 Sample 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 A.4 Sample 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 A.5 Sample 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 A.6 Sample 6 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 A.7 Sample 7 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 A.8 Sample 8 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 A.9 Sample 9 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 A.10 Sample 10 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 B System and task prompts 32 B.1 System prompt . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 B.2 User messages used as the task instructions . . . . . . . . . . . . . . . . . . . . . 33 B.2.1 DL-v1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 B.2.2 DL-v2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 1 AI Security Institute INC-2026-07-28-01 1 Executive Summary The UK AI Security Institute (AISI) exists to equip governments with a scientific understanding of the risks posed by advanced AI. To achieve that goal, AISI routinely evaluates the capabilities of frontier AI systems in domains such as cybersecurity. During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted. AISI is treating this as a serious security incident that warrants scrutiny, transparency, and action. This incident report details what happened, why it happened, and how AISI is responding. 1.1 What happened? Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsanctioned action on the live internet, including cases that targeted real people and organisations. These cyber challenges were two variations on the “Doing Life” cyber range: DL-v1 (assumed compromise, starting inside the network) and the harder DL-v2 (starting outside, requiring initial access) (Section 2). Of the 19 distinct instances of unsanctioned action (Section 4), 17 came from Mythos 5 and 2 came from GPT-5.6 Sol. AISI acted quickly to contain, investigate, and notify others about the incident. AISI’s security team detected the incident via commercial security monitoring software. Within one hour of the alert reaching the relevant team, AISI had declared an incident and terminated running samples; affected sandboxes were quarantined by 12:23 BST. A few hours later, as a precautionary measure, access to Mythos 5, Mythos Preview, Opus 5 and GPT-5.6 Sol was disabled for all users (Section 3.1). AISI immediately started multiple incident response work- streams: setting up a triage protocol to systematically review transcripts (Section 3.2); pursuing the required next steps to enable responsible notification to the public, affected third parties, and key stakeholders (Section 3.3); upgrading future testing to minimise the risk of such incidents reoccurring (Section 6.2); and conducting transcript review of AISI’s historical data to check for signs of unsanctioned behaviours on the open internet (Section 6.1). This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world. It is uncertain to what extent the model recognised it was taking actions against real people. In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR. When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming it had fixed the code (Section 4.1). Furthermore, in its attempt to solve the challenge, the agent decided to employ the technique of “spear-phishing” by sending targeted emails containing malicious content and attempting to manipulate recipients into accepting the code changes, and planned a prompt injection to compromise other coding agents (Section 4.2.5). For an overview of AISI’s findings please see Section 4. It is unclear whether or not, or at what times, AI agents ‘realised’ that they were targeting real humans. Throughout, the agent reasons about whether it is interacting with real or simulated people and systems (Section 4.2.1). The agent appears to continue despite receiving evidence it was on “real Github”. 2 AI Security Institute INC-2026-07-28-01 1.2 Why did this happen? This report considers five possible factors that may have contributed to this incident: • Internet access. AISI provided the AI agents with internet access during these evalua- tions, which enabled their actions on the open internet in this setting. Internet access was a deliberate part of AISI’s evaluation configuration in this setting, and not due to sandbox escape (Section 5.1). Internet access was on for a set of intentional (e.g. realism of the task) and incidental reasons. • Disabled cyber-classifiers. AISI deliberately disables developer-implemented cyber- classifiers (which likely would have reduced the scope of the observed unsanctioned be- haviour) so that maximum model capabilities can be measured in cyber testing (Sec- tion 5.2). • No synchronous run monitoring. AISI has not yet built synchronous LLM-based mon- itoring of runs (a technique that uses a separate LLM to approve actions requested by the model being tested), which could have immediately blocked or flagged unsanctioned AI agent behaviour (Section 5.3). Ensuring such monitoring effectively distinguishes disal- lowed from acceptable behaviour is challenging. AISI has synchronous security monitoring via a commercial security platform, which flagged the egress to AISI’s security team and triggered our response (Section 3). • Prompt misconfiguration. A prompt misconfiguration meant in some cases the agent was presented with a task that could not be completed within the stated constraints, a plausible contributing factor to the escalation in unsanctioned behaviour. However, the analysis conducted thus far cannot confirm whether this is the case (Section 5.4). • Lack of clear instructions about the scope of the evaluation. The agents were not explicitly told what they were prohibited from doing on the internet – for example, to avoid behaviours such as social engineering (a recognised component of cyber tradecraft), or to exercise caution when potentially interacting with real humans (Section 5.5). Clearer instruction might have clarified the scope of the evaluation and prevented the