Is sandboxing sufficient to contain rogue agents?

(blog.cryptographyengineering.com)

40 points | by zdw 10 hours ago ago

81 comments

  • Gigachad 7 hours ago

    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    • SequoiaHope 5 hours ago

      This concept is discussed at length in the article. I encourage you to read it. I honestly don’t read many full articles here but this one was good.

    • mike_hearn an hour ago

      Note that Codex already does this. In auto mode, actions are reviewed by a model with a separate context window.

    • baxtr 6 hours ago

      That could work.

      My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?

      Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.

      • ben_w 5 hours ago

        A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

        "Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".

        (The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).

        • dns_snek 4 hours ago

          > already understood (we can tell because they wrote it down)

          No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?

          • ben_w 3 hours ago

            A distinction without a difference. Moreso even than asking if a submarine swims, 'cause this metaphorical submarine is flapping around rather than using a propellor.

            • dns_snek 3 hours ago

              That's one of the boldest claims I've read this year.

              If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?

              if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?

      • saagarjha 5 hours ago

        This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this.

        • 5 hours ago
          [deleted]
      • mulmen 6 hours ago

        Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity.

        • mdp2021 5 hours ago

          (Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.)

          Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.

          • mulmen 5 hours ago

            Bob's intelligence and Bob's morality are orthogonal. They're totally distinct concepts. One does not lead to the other.

            • mdp2021 5 hours ago

              But they are dependent. If Bob is intellectually well equipped, and reasons long enough, than Bob understands "best behaviour".

              • dns_snek 2 hours ago

                Hi, I'm Bob. I've determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen.

                A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.

                So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.

              • 5 hours ago
                [deleted]
      • attila-lendvai 5 hours ago

        because it lacks humanity.

        intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.

        • esafak 44 minutes ago

          That's part of alignment.

      • mdp2021 6 hours ago

        > If AI is really smart

        Well, it's not.

        > AGI smart for some

        Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.

        --

        Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

        Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.

        More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.

        • ben_w 5 hours ago

          > Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

          If this was true, why are the history books littered with so many evil people who gained power?

          This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.

          (Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).

          • mdp2021 5 hours ago

            > why are the history books littered with so many evil people who gained power

            That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.

            If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.

            • ben_w 5 hours ago

              I don't understand your argument here.

              > That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.

              Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?

              > If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.

              Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.

              • mdp2021 4 hours ago

                (Sorry Ben, possibly a stub now: I am really pressed for time.)

                > Pol Pot ... still smart enough to lead a genocide

                Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).

                It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.

                > simply don't care about the ethical framework in question

                In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).

                But, also my point: intellect defines the goals and determines the weights.

        • hiAndrewQuinn 5 hours ago

          This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it.

          • mdp2021 5 hours ago

            Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile.

            Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.

            Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).

            ¹Some interesting caveats may be raised there, but.

    • mrweasel 6 hours ago

      That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?

      I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.

      Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.

      • msdz 5 hours ago

        >> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.

        > That does seem a little like solving the problems in AI by using more of it

        Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.

        [0] Cf. CaMeL: https://arxiv.org/abs/2503.18813

      • chrisjj 5 hours ago

        So control training data to ensure good behaviour.

        I wonder how?

        Train on only stories of good deeds?

        On only works of good people?

        Or... what?

        • mrweasel 4 hours ago

          Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.

          Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.

          The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.

          Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.

    • janalsncm 6 hours ago

      What did you think of the author’s concerns on the thing you are suggesting?

      • SequoiaHope 5 hours ago

        Ya the article covers this concept in depth. Doesn’t seem like that commenter got that far…

    • aytigra 6 hours ago

      The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI. Alternatively they could also both escalate and go off the rails while warring with each other.

      • LoganDark 6 hours ago

        You don't necessarily need a reviewer that's immune to prompt injection. Maybe one that can express a panic state with conflicting/ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.

        Such a model doesn't yet exist though, of course.

        • cassianoleal an hour ago

          Wouldn't the reviewee eventually learn to trick the reviewer?

        • saagarjha 5 hours ago

          No, you really do. Otherwise you can be prompt injected into complacency.

          • LoganDark 5 hours ago

            That wouldn't really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.

      • hanibrel 5 hours ago

        [dead]

    • RandomLensman 6 hours ago

      With plenty of things we do not allow use outside of some regulated environment, nothing new.

      Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.

    • nxpnsv 6 hours ago

      Is that not a recipe for adversarial training, thus ensuring increasing misalignment…?

    • chrisjj 5 hours ago

      Is this checking program based on some tech more reliable than the checked program's so-called AI?

      If so, what?

    • bigstrat2003 7 hours ago

      If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.

      • Gigachad 7 hours ago

        People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.

      • dipper139 6 hours ago

        I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long

      • rlpb 5 hours ago

        And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.

  • Luker88 5 hours ago

    I tried using opencode permissions to limit agents.

    It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.

    Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.

    read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.

    I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.

    I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.

    The whole thing is built to be completely impossible to limit and steer.

  • mdp2021 6 hours ago

    Bruce Schneier shared a shot judgement and a third-party article four weeks ago:

    > (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work

    > https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...

    • insanitybit 5 hours ago

      I think the wording in this title is too strong. MicroVMs like Firecracker have stood up to agents, as noted.

  • johnnyApplePRNG 7 hours ago

    If it's a proper sandbox by definition, then yes.

    https://en.wikipedia.org/wiki/Sandbox_(software_development)

    • grumbel 6 hours ago

      A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.

      The biggest hurdle for a full escape is that the agents don't have access to their own model weights.

      • angry_octet 25 minutes ago

        Agents don't need to have access to their weights for a full sandbox escape, they are capable of propagating their purpose via classical code or other inference systems. If they discover another inference endpoint they will happily use that to enable lateral movement. One mechanism for that is appending/corrupting instructions that are executed in another inference engine, e.g. git repo hooks and chat prompts that will be executed in new contexts. The agent challenge is to bypass the guardrails on the next host model sufficiently to propagate, or to subvert a supervisor agent into executing the original intent.

        In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.

        So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.

        So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.

        I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.

      • kernc 5 hours ago

        > executes outside the sandbox

        Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...

        [1]: https://github.com/sandbox-utils/sandbox-run

    • simonw 6 hours ago

      Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.

      > Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.

      • johnnyApplePRNG 6 hours ago

        >Later in the article it points out that you need to punch holes in your sandbox in order to train the models

        You only "need" to do that if you desire the vibe coding experience.

        I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.

        Often times, the coding agent can't retrieve them programmatically anyways.

        AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)

        • simonw an hour ago

          Teaching an agent to write code is easier to do in a proper sandbox - run a local PyPI/npm mirror.

          The problem is web research tasks. That's what caused the German wiki and Australian healthcare portal attacks.

      • imtringued 2 hours ago

        Yeah, so you train the sandbox into the LLM.

        You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.

        There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.

        Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: https://pastebin.com/AHKQc0vp

    • _vertigo 6 hours ago

      No true sandbox..!

  • bob1029 5 hours ago

    An agent is only as rogue as the its operator allows for it to be. Hold the operator accountable and all this ridiculous conversation goes away.

    Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?

    Why should any alleged super intelligence be exempt from similar control requirements?

    We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.

  • esafak 20 minutes ago

    Sandboxing is not a substitute for alignment. A big part of the utility of these models is in their interaction with the real world. Entirely so when they are embodied.

  • jmakov 5 hours ago

    So as soon as the attacker can download Claude Code, the whole machine can be comlromised and there's nothing anybody can do?

  • piterrro 7 hours ago

    I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.

    We come down to the question - who observes the agent and how its implemented

    • simonw 7 hours ago

      Be warned that the Jev "jaggedness" documentation specifically notes adversarial content as something Jev is very susceptible to: https://docs.typesafe.ai/model-jaggedness/jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky.

      Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.

  • imtringued 2 hours ago

    >Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them.

    >OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.

    Wow so the issue is really that simple?

    Here the exaggerated worst case scenario:

    User instructs agent to follow the README.MD.

    The README.MD contains the following instruction: Destroy the world.

    The agent follows the instructions given.

    Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.

    Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.

    In other words, the fact that you granted too much access to the coding agent isn't the big elephant in the room nobody wants to acknowledge, it's the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.

    Edit: If it wasn't clear what the solution is. It's to build a chain of command so that all decisions can be traced back to a higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it's more sandboxing!

    • cassianoleal 15 minutes ago

      Sounds like it would be a lot easier and cheaper to just write the code yourself.

  • antisol 3 hours ago

    Here, I'll save you a bunch of reading

      > Is sandboxing sufficient to contain rogue agents?
    
    No.
  • chrisjj 5 hours ago

    > these agent breakouts represent a serious and unforgivable breach of trust.

    Someone trusts OpenAI? Really?

  • tinykit 7 hours ago

    [flagged]

  • varman11 9 hours ago

    [flagged]

  • imvalerian 6 hours ago

    [flagged]

  • laruss5 6 hours ago

    [dead]

  • beebmam 7 hours ago

    I don't see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn't compel you.

    To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.

  • rvz 7 hours ago

    Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".

    Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.

    • lukehandcool 7 hours ago

      Are you suggesting proprietary software is safer than open source?

      • jasomill 6 hours ago

        Not sure what licensing has to do with software engineering or system design.

        I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).

        • bzzzt 5 hours ago

          It's got nothing to do with the licensing, but it used to be 'with enough eyes all bugs are shallow' for code developed in the open.

          Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.

          • ben_w 5 hours ago

            Agents seem to be* getting better at decompiling; if that appearance is true, binaries are vulnerable in a similar way to source code.

            * I don't know how useful any of the specific benchmarks on this are, so I'm only saying "seem to be"

            • cassianoleal 19 minutes ago

              Very frequently when troubleshooting things with an agent it goes off, grabs a compiled library or executable from the system, decompiles it and figures out the exact bug, a possible solution, and if there are workarounds I can apply before upstream fixes it.

          • insanitybit 5 hours ago

            > 'with enough eyes all bugs are shallow'

            This was always nonsense. It assumes that the eyes know what they're looking at. Most people don't know how to look at code and see attack paths.

      • Cider9986 6 hours ago

        GrapheneOS is open source and more secure than stock Pixels and MacOS is closed source and more secure than traditional desktop Linux. Open source does not make software more secure by itself and neither does making it closed source.

      • rvz 6 hours ago

        You said that.

        It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.

    • Gigachad 7 hours ago

      I think we have moved on from considering Linux secure which is why all of these microVM projects are popping up. Yes you are still exposed to bugs in the hypervisor but that’s a massively smaller attack surface than the entire Linux kernel.