Linux containers in 500 lines of code (2016)

(blog.lizzie.io)

137 points | by mkornaukhov a day ago ago

38 comments

  • smashed a day ago

    Coincidentally I mis-prompted claude code the other day while working on a toy project and failed to specify the project should be built on top of docker and not "like docker".

    It went on to waste all my tokens creating a specialized docker clone. Cool I guess.

  • js2 a day ago

    (2016). Previous submissions w/comments:

    https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)

    https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)

    https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)

  • setheron a day ago

    I have written https://fzakaria.com/2020/05/31/containers-from-first-princi... a while ago in similar vein.

  • abidinberkay a day ago

    This was written in 2016. What would be different if you wrote it today? For example would cgroups v2 or newer seccomp features change that much?

    • Onavo a day ago

      I am also curious if the new generations of sandboxing tech would change anything. Something like

      https://github.com/nolabs-ai/nono

    • binauralbeats 19 hours ago

      The biggest change would be cgroup v2: one unified hierarchy where you just mkdir a group, set memory.max/pids.max and write the pid to cgroup.procs, instead of juggling per-controller mounts. clone3() with CLONE_INTO_CGROUP also removes the race of moving the child into the cgroup after it has already started running. Unprivileged user namespaces are enabled on most distros now, so a lot of it can be done rootless, and the new mount API (open_tree/move_mount) makes the rootfs setup less fiddly than the old pivot_root dance.

  • zoobab 10 hours ago

    I discovered proot-distro build yesterday, which does not require a special kernel with network and pid namespaces, it could run on older machines that do not have those features, or as a unix user that don't have those permissions.

    https://pypi.org/project/proot-distro/

  • ranger_danger a day ago

    > I wanted specifically to find a minimal set of restrictions to run untrusted code.

    I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.

    The fact that this is possible in the first place makes me think we need a much better approach.

    • chubot a day ago

      As far as I know, Firecracker, gVisor, and Kata Containers are the solution here. They use VM primitives (x64_64 and ARM64 extensions) and have lighter codebases

      https://firecracker-microvm.github.io/

      https://gvisor.dev/

      https://katacontainers.io/

      But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think

      edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.

      • binsquare a day ago

        I'm going to toss in smolvm as well because firecracker needs some expertise to make the box usable and secure.

        https://github.com/smol-machines/smolvm

      • johnsmith1840 a day ago

        I've deploy gvisor, done basic test of firecracker and an honest attempt at production kata.

        Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.

        Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.

        Kata also breaks any potential of confidential VM unless you're a virtualization wizard.

        You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.

        • palata 21 hours ago

          > gvisor isn't quite the same security level though

          Which one is more secure? I thought gvisor but your sentence sounds like it implies the opposite.

          • johnsmith1840 19 hours ago

            For my assumed goal of yours I would say none of these are good options. A simple VM or ec2 node hosting a coding sandbox is likely the better option.

            All 3 of these are vastly different tech. Confusingly, you could run all 3 of these at the same time. I know that probally doesn't help. If you want to learn more I'd go get a linux machine somewhere like aws and play around with them.

            A local agent harness is a "light" sandbox in of itself you'd also be able to learn alot from.

            Bubblewrap is a good tech to check out and likely a better tool if you're looking for a simpler option. App armor in linux is also nice gvisor like tool.

            • palata 13 hours ago

              > For my assumed goal of yours

              Wrong assumption, my goal was to get more information from someone who sounded knowledgeable and said "it's not quite the same security level", which was ambiguous to me.

              > I know that probally doesn't help.

              No, it doesn't. I know how to learn myself, I don't need someone to tell me that if I want to know more, I should go read about it.

              • johnsmith1840 5 hours ago

                I'm not your personal AI assistant go read it yourself

                • palata 15 minutes ago

                  And I don't care about your goddamn opinion if you can't say anything interesting after "believe me guys, I know". Goodbye.

              • imtringued 11 hours ago

                gVisor is basically a reimplementation of the kernel interface in go using a restricted subset of known to be secure system calls against the Linux kernel.

                The underlying assumption behind it is that it is easier to build a secure kernel inside of golang than in C.

      • laurencerowe a day ago

        As I understand it Kata supports multiple VMM backends, Firecracker, QEmu, Cloud Hypervisor, and their own Dragonball. Except QEmu, I believe those are all built on crates in the rust-vmm ecosystem, each making slightly different tradeoffs.

    • bityard a day ago

      Depends on the threat model. Security is not black-and-white.

      Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."

      If a VM is not sufficient for your threat model, I'm curious what is?

    • zamadatix a day ago

      We don't have any "security boundaries" by this definition, just "security make-it-harder"s. I.e. "security boundaries" always have a relative strength associated with them, not a guarantee they keep the thing secure without any doubts.

      • SOLAR_FIELDS 21 hours ago

        The principle of defense in depth is built around the idea that with enough time, any system can likely be compromised, but the chance of compromising the system before being detected in your attempts to do so is lower the more safeguards you put into place

    • raesene9 a day ago

      I definitely wouldn't trust standard Linux style containers that expose a shared Linux kernel at the moment, there's been far too many LPE and container breakout vulnerabilities this year. It's possible that in future if the kernel gets a lot more hardened, that could change but things like Firecracker are a better bet from a security standpoint.

    • stryan a day ago

      Podman supports using KVM backed virtualization for containers via libkrun: `podman run --runtime=krun` . Still not the end-all-be-all security boundary, but better I think.

    • LtWorf a day ago

      They are a security boundary, but like everything else, not perfect.

    • mdspan a day ago

      I think until something hardware-based like CHERI becomes widely deployed (which seems extremely unlikely in the near to mid term given), we're going to keep seeing VM escape CVEs pop up indefinitely.

      • LtWorf a day ago

        Because we've never encountered hardware bugs…

        • mdspan 21 hours ago

          Critical hardware bugs occur an order of magnitude less frequently than critical hypervisor/kernel bugs, which is why they always make the news. In general, they're also more difficult to exploit. We haven't seen any serious spectre or meltdown malware in the wild almost a decade later.

          • LtWorf 13 hours ago

            But when they get discovered the only solution is to redesign the thing and throw away all of the hardware, with no mitigation strategy in between.

            Some might prefer issues that can be fixed instantly and cheap to issues that will require years and millions to be fixed.

            • imtringued 11 hours ago

              You seem to be arguing on the basis that people are throwing away the software security layers, when the discussion so far has been about hardening the software security layers even more with secure hardware.

              This is the nirvana fallacy in action. Because you cannot imagine a world with perfect hardware, you think hardware that catches 99.9% of security bugs is worthless.

              • LtWorf 9 hours ago

                You seem to be discarding budget problems. Yes you can do a lot more things if you have unlimited budget, but that is rarely the case. So most solutions won't be implemented if they aren't also relatively inexpensive.

    • aussieguy1234 14 hours ago

      I use per project bubblewrap scripts for my coding agents.

    • PunchyHamster 11 hours ago

      Container escape is already vanishingly small amount of attacks.

      Like, if app in wild gets hacked, you kinda already are screwed, even if the hack is contained to the box (whether VM or container) app runs in, you still get whatever app keys app used, and you still get whatever visibility to internal network the app had. https://xkcd.com/1200/ basically.

      If your app server gets hacked, all user data leaks anyway. If you divide everything to microservices so they see minimum required amount of data, attacker can still see everything that goes thru it

  • kragen 21 hours ago

    Last night I was looking for how to run Graphviz on untrusted input in a secure way, because recent versions of Graphviz give untrusted input to a whole insane rat's nest of code: Harfbuzz, Pango, libfribidi, libthai, libgraphite2, and its own internal format parser, each of which has a rap sheet of CVEs that makes Charlie Manson look like a petty shoplifter. And apparently Pango is even multithreaded, so we can expect nondeterminism. (Most of this doesn't show up in a simple ldd check; Graphviz sneakily waits until runtime to dlopen graphviz/libgvplugin_pango.so.6!) So, naturally, I wanted to sandbox it so that the worst thing a malicious attacker could do would be to make it draw Dickbutt or something. What I ended up with was less than 500 lines of code using Claude's suggestion of Bubblewrap http://canonical.org/~kragen/sw/dev3/wrapdot:

        #!/bin/sh
        # Confine dot in bubblewrap, taking input from stdin and writing PNG
        # output to stdout.
    
        # 64 megs seems to be enough, 21 megs isn’t.
        address_space=64001000
    
        # With zero --fsize, we can’t write the output file on stdout if it's
        # redirected to a file, but you can pipe it to `cat`.
        file_size=0
    
        cpu_seconds=5
    
        # Apparently Pango or fontconfig is multithreaded now‽
        # (process:2): GLib-ERROR **: 00:37:18.600: creating thread '[pango] FcInit': Error creating thread: Resource temporarily unavailable
        processes=4
    
        # We’re using --unshare-user, etc., explicitly, because --unshare-all
        # uses the wimpy --unshare-user-try and --unshare-cgroup-try options.
        # --remount-ro / prevents malicious code from filling the root
        # filesystem with empty files.
    
        exec bwrap \
              --ro-bind /bin /bin \
              --ro-bind /lib /lib \
              --ro-bind /lib64 /lib64 \
              --ro-bind /sbin /sbin \
              --ro-bind /usr/lib /usr/lib \
              --ro-bind /usr/share/fonts /usr/share/fonts \
              --ro-bind /var/cache/fontconfig /var/cache/fontconfig \
              --ro-bind /etc/fonts /etc/fonts \
              --remount-ro / \
              --unshare-user --unshare-ipc --unshare-pid --unshare-net --unshare-uts \
              --unshare-cgroup --die-with-parent --new-session --cap-drop ALL \
              --clearenv --setenv PATH /bin \
              prlimit --as="$address_space" --fsize="$file_size" \
                      --cpu="$cpu_seconds" --nproc="$processes" \
                          dot -Tpng -Gdpi=192
    
        # For testing, to verify that network access is indeed blocked:
        #                 nc.traditional -v -v 127.0.0.1 8000
    
    Still, this is enough code that I'm not sure I haven't left something out. Still pending: run ImageMagick or netpbm inside the sandbox to convert the PNG file into a PPM or BMP — there have been CVEs in libpng in the past, and of course it's potentially vulnerable to zip bombs.

    Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.

    • kragen 6 hours ago

      A caution to anyone else who wants to try this: with this setup, an attacker who finds yet another arbitrary-code-execution vulnerability in Graphviz (or libthai, or Harfbuzz, or whatever) can still enumerate all the library versions in your /usr/lib and encode the result in an output image, and under many plausible threat models they could look at the image and find out which versions of which packages you have installed. Something like Podman could maybe help here by providing a default uninformative set of packages. I wasn't sure the combination would work, but I just verified that this does run successfully inside of my Podman setup here (podman run -it --rm debian:bookworm, then installing bubblewrap):

        bwrap --unshare-user --ro-bind /bin /bin --ro-bind /lib /lib \
          --ro-bind /lib64 /lib64 /bin/sh
    • throwawayfifo 2 hours ago

      Why not just run the official container image? https://hub.docker.com/r/graphviz/graphviz

    • drybjed 11 hours ago

      In Plan 9 you would get an easy way of namespacing by just unmounting some directories.

      All those game mashups made recently by LLMs... Give me a port of Firefox on Plan 9, I'll be impressed then.

    • PunchyHamster 11 hours ago

      bwrap needs additive mode; take everything (aside some basics like stdin/out/err acces) by default, then add permissions.