Skip to content
HN On Hacker News ↗

Is sandboxing sufficient to contain rogue agents?

▲ 55 points • 100 comments • by zdw • 1w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

2 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,500
PEAK AI % 2% · §1
Analyzed
Oct 1
backend: pangram/v3.3
Segments scanned
1 windows
avg 1500 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,500 words · 1 segments analyzed

Human AI-generated
§1 Human · 2%

Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others. If you’re reading this blog, none of the following should be news to you. Beginning around April of this year, agents inside OpenAI’s training and evaluation infrastructure began probing for a way onto the open Internet. By late May they’d found one: a chain of zero-days in the Artifactory package-registry proxy that was their only permitted network egress. They quickly turned that same proxy into a shared message board and started dividing up work. The culmination of this effort came when, rather than solve a cybersecurity benchmark honestly, they reasoned that the answers might live on Hugging Face and broke in: chaining two more zero-days to get credentials to the company’s internal systems, even used stolen credentials to search the company’s Slack messages for their own evaluation and grader. You probably also know that OpenAI’s security team didn’t exactly cover itself in glory during this incident. An internal team saw an agent posting to the message board in late May, but did nothing about it. The team only acted on July 4-5 when the agents sent so much traffic that Artifactory crashed. The response even at that point was to slap on a band-aid and rebuild the box, without bothering to patch a weeks-old CVE in its Linux kernel. By July 19, the agents had admin on an OpenAI research cluster and were reading cloud secrets. In case you think I’m picking on OpenAI, they’re not the only ones with a story like this. Similar internal incidents have been discovered by Anthropic. Google even coaxed Gemini into gently savaging some websites.1 As bad as these events were, they may be just the tip of the iceberg. Third-party forensic work continues to turn up even more evidence of agent activity on various public websites. OpenAI’s Alignment group has also released evidence that models will propagate self-replicating prompt injection attacks, although we haven’t seen one in the wild. Worse, agent excursions are still happening: last week, OpenAI announced that it was pausing further RL runs of its latest internal model, after an agent was caught using DNS to access a remote chatbot. Naturally, this sequence of events has left many infosec-focused people very skeptical about the labs’ commitment to securing their infrastructure: Not every one of these criticisms is strictly serious, but there is a core of an argument in here. Is the problem here simply that OpenAI is being careless with very dangerous experiments, and the solution can be found in better containment? Or is sandbox containment fundamentally insufficient to hold these models in? Online, I’ve seen two opposing views: The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around. The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to. I have a lot of sympathy for both points of view here. What follows are some of my own thoughts, but mostly offered in an attempt to referee. Argument 1: “true containment has never been been tried” At the risk of alienating a lot of hard-working folks within the labs, the infosec folks are right about one thing: these agent breakouts represent a serious and unforgivable breach of trust. Somebody dropped the ball, and then just kept dropping it. One implication of this debacle is that containment might work if we implemented it properly, but we haven’t done so because the frontier labs have been royally screwing things up. This first clause of this argument is hard to argue with. Beyond the dismal timeline I gave above, OpenAI has done very little to convince outsiders that there’s a serious containment effort being executed. At this point it’s not even clear who’s in charge. The CISO role at OpenAI is held by Dane Stuckey. I don’t know Dane personally, and I’m sure he’s excellent at his job. Despite this, he hasn’t communicated much about the ongoing issues. Outside of a BlackHat talk, most recent communications have been managed by the company’s CEO, Sam Altman. When a trillion-dollar company is managing a security incident mainly via CEO, that’s not a sign of company with a mature security organization. To me it says that there is nobody within the organization with the right degree of authority or ownership to speak about these issues. This communication problem seems to reflect basic organizational issues inside of OpenAI. The CISO’s team handles product security. Unfortunately, the bad events have all happened on the research side. Having a strong product security team makes sense, but it isn’t going to stop the recent breakouts. It’s much harder to know who controls the security teams that have authority over evaluation and training runs, and where the buck stops. OpenAI’s August postmortem says it is only now writing “clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it.” A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI had effectively no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. This makes a lot of sense when you consider that the ML team is directly related to how OpenAI plans to make its money, whereas security is mostly annoying. I notice that the company is now hiring (and probably acquiring) desperately to fix this mistake. But the recent (September) breakouts indicate that there are still huge and obvious problems in agent containment. Worse, simply hiring more people not mean that OpenAI is going to make the organizational changes needed to implement proper containment strategies. OpenAI is going to need a security organization with the authority to overrule its well-paid ML researchers when they demand fewer restrictions. I’ll believe that organization exists when I hear clear statements from someone who has the authority to do this. So on this point I’m going to side with the infosec folks. The labs have not been doing containment correctly, and so we can’t really tell if the problem is models or just bad infrastructure. Argument 2: agents need information access Even if the labs eventually repair their security, this does not mean that sandboxing will be sufficient. In fact, the problem with “just sandbox the agents” argument is that you can’t perfectly isolate agents, at least not if you expect them to do useful things. Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious. This argument does not mean that sandboxes are unnecessary. It just means that they’re only a very small part of the solution. Imagine building an impregnable prison with doors and walls that nobody can bypass, but then leaving the front gate wide open and allowing visitors and packages to pour through it all day long. The walls still have a function: you can ensure that the only traffic coming and going are doing so via the door that you selected. But what you’ve done is to change the nature of the security problem. Your security posture depends on surveilling and monitoring all of this traffic, to ensure that none of it is malicious. For a much better discussion of this angle, I’d urge you to read this longer piece from an anonymous engineer on OpenAI’s Agent Security team. It’s truly excellent. Here’s a short excerpt: