Solo – a .so loader for static Linux binaries

(github.com)

33 points | by zX41ZdbW 2 hours ago ago

43 comments

  • pg83 25 minutes ago

    How this differs (is better!) from prior art - https://github.com/pg83/solo#how-this-differs-from-prior-wor...

  • simonask 36 minutes ago

    It is a testament to the complete failure of the GNU/Linux userland that something like this seems at all attractive to spend time on (or, it seems, LLM tokens).

    Actually, scratch that, because Windows and macOS have historically struggled with ABI compatibility as well (macOS less so, due to not caring about backward compatibility in the first place).

    How did we get to the point where people feel they need to go to the length of embedding an ELF loader in their binary (!!) rather than just linking with glibc?

    • diabllicseagull 2 minutes ago

      I'm mostly taken aback all the solutions devised to go around the issue, especially the container-based ones. I really disliked it when I grabbed the flatpak version of Blender only to find out that it can't have HIP support. (they might have fixed it by now but you get the point)

    • pg83 34 minutes ago

      Glibc has a terrible history of binary incompatibility. If that's so hard to believe, try running binaries built on one distribution on other distributions. Linux has two stable ABIs: the kernel ABI for static programs, and, ironically, WINE.

      • vlovich123 17 minutes ago

        I haven’t heard of this and I don’t think you’re right. Glibc, for all its faults, as a general rule does backward compatibility well. The problem is if you compile against a newer glibc (common in CI by default) and try to run on a distro with an older (common in the wild). If your CI uses an older glibc you should be fine AFAIK.

        • pg83 10 minutes ago

          https://bugzilla.redhat.com/show_bug.cgi?id=638477 is the most "famous" example.

          There are also much less well-known "little things" that regularly pop up here and there.

          > If your CI uses an older glibc you should be fine AFAIK.

          In any case, my binaries work not only under glibc, but also under Alpine, and (work in progress) under android/bionic.

      • diabllicseagull 9 minutes ago

        according to appimage recommendations as long as you build against glibc with an earlier version than the system it's run on it should be fine.

        https://docs.appimage.org/reference/best-practices.html

        I hear you about WINE though.

  • nomel an hour ago

    I don't know much about musl.

    > GPU: Vulkan and OpenGL drivers are supplied by the host as shared objects, usually built against glibc, and a fully static musl binary cannot normally dlopen() them.

    Why? Have people managed to break the ancient concept of shared libraries, and this is a fix for that?

    • okanat an hour ago

      Because glibc and GNU set a terrible precedent. On GNU/Linux systems the shared binary interpreter / loader, GCC compiler, the C library and the system C/C++ ABI all depend into each other. You cannot change any of them independently. All shared libraries depend on the specific glibc version to load them into memory to be able to use that specific glibc version as their C library and make calls like dlopen.

      Shared libraries have always been broken in Linux. Unfortunately many things like GPU drivers, graphics libraries and NSS need shared libraries to dynamically load certain runtimes (because you don't want to load all possible GPU drivers in existence to your RAM). So an ecosystem has been developed on top of terrible ABI and architecture GNU/glibc provided.

      • asveikau 27 minutes ago

        > system C/C++ ABI

        C++ abi should not be included in this. It is independent from the other pieces and historically a source of incompatibility on its own.

        Saying "C/C++ abi" as if they are the same is looney tunes, the former is very simple and stable and the latter is very complex.

      • uecker 36 minutes ago

        In what sense do binary interpreter / loader, GCC compiler, C library and system C/C++ ABI dependent on each other? I have certainly mixed different versions of all these components without problems so far.

        • okanat 27 minutes ago

          When you compile GCC you need to provide a full glibc installation as your target. It is also a dependency of libstdc++.

          C++ global/static variable initialization depends on the specific version of glibc (they don't usually break compat, but they can and they did in the past) which also provides ld-linux.so that loads those global variable placeholders in the correct manner such that glibc and libstdc++ can initialize them correctly.

          This is just one example. Thread local variables and behavior of things like pthreads with signal, fork etc all depend on glibc.

          • uecker 16 minutes ago

            I can't comment on the C++, I can imagine there plenty of issues, but for C I don't see this. You need some libc if you compile with gcc, but this generally does not introduce a hard version dependency on the specific version (there may be a minimum requirement if you compile against a new version that a symbol with a different ABI).

        • mananaysiempre 16 minutes ago

          I don’t imagine that you’re unaware of any of this, but: ld.so and libc.so are heavily interdependent in deliberately undocumented ways with Glibc and outright the same file with shared Musl. And while you might usually get away with using any old GCC with the right architecture and ABI (especially for C; cf the musl-gcc hack), technically it needs to be built to target a specific libc version (particularly via symbol versioning; I’ve long wanted to gather a set of patches to build an old Glibc and subsequently a cross-compiler using a new GCC so I could avoid PyPA’s manylinux monster or its moral equivalent for compatible dynamic binaries in simple cases). The C compiler of course is tied to the C ABI, and this wouldn’t be really worth mentioning except for the time where the GCC devs accidentally the whole SysV i386 ABI and pretended that the stack was always 16-byte aligned, why do you ask, except on RHEL. The C++ parts I can’t really comment on.

          • uecker 8 minutes ago

            I am not really sure. For ld.so and libc.so I may believe this. The C ABI is very stable, and if you use a new symbol from a newer glibc, you certainly depend on it, but this can also be avoided. In any case, I do not see what is fundamentally misdesigned here. I can't quite image how it could work differently. If you upgrade something so that the e ABI changed you natually need to update other components. Static linking certainly seems a very poor replacement for this.

        • pg83 21 minutes ago

          For example, the itanium unwind ABI implementation lies between these three entities.

        • duped 21 minutes ago

          The interpreter/loader is glibc and a key part of bootstrapping an executable built against glibc is loading libc itself before continuing on to load the program. Versioning is a problem when distributing binaries linked against a newer glibc to distros that ship an older one. The C compiler doesn't really care as much.

          • okanat 3 minutes ago

            > The C compiler doesn't really care as much.

            Until you define a thread local variable (C11) or use atomics (also C11) or define a global with an initial value. Then it happily generates code that depends on "whatever my target glibc + ld-linux.so needs".

      • b5n 24 minutes ago

        While I don't disagree with some of the pain you describe, you conveniently gloss over the fact that gnu developed a system that worked, and then made it free to everyone to consult and use.

        • okanat 24 minutes ago

          BSD also did it. They did it better. Maybe more modern but AOSP also did it but at a different level of binary: instead of ELF, using compiled Java bytecode archives.

      • duped 24 minutes ago

        > all shared libraries depend on the specific glibc version to load them

        Not really, though. glibc uses symbol versions that are forward but not backward compatible. If you got an error that said "this program was built for a newer version of <distro>" would you say the same thing?

        Note this is the same (if not worse) on MacOS, and on windows you used to distribute the CRT with your application just to deal with the same problem.

        • okanat 14 minutes ago

          See https://news.ycombinator.com/item?id=49355262 .

          Yes glibc has some backwards compat but you cannot load a binary compiled with a newer version of glibc using an older ld-linux.so. That's because the interdependency. Nor you can load binaries that depend on different libc.so files with glibc systems

          I cannot comment on macOS, I have never used it. However this is not a problem with Windows. You can ship a newer CRT or you can install it as a system component using Microsoft's MSI. The dependency is one way on Windows. CRT purely depends on Win32. Moreover the loader is completely independent and DLLs are loaded into their own unique scoped namespace unlike Linux that loads them in global symbol namespace. That's why you can mix and match DLLs compiled for different CRT versions.

          • duped 9 minutes ago

            That's what I'm saying though, glibc-linked binaries are forwards (but not backwards) compatible.

    • akerl_ 36 minutes ago

      musl has no problem building and using shared libraries.

      What you can't do is build something statically with musl and then reliably dlopen shared libraries built with glibc.

      • pg83 33 minutes ago

        Well, now it's possible! Furthermore, SoLo binaries can run, without modification, on glibc-based distros, alpine, and soon on android/bionic (not committed yet).

    • ranger_danger an hour ago

      musl does not perfectly emulate all aspects of glibc, so trying to use libraries that assume glibc can sometimes lead to problems.

      • pg83 36 minutes ago

        On the one hand, this is technically true, but on the other, what serious issues do you know that will cause problems in practice? I run tests on 1,000 of the most popular Debian packages.

  • j16sdiz an hour ago

    > backed by its own ELF loader (x86-64 and aarch64) and a glibc ABI bridge

    Yacks

    • mananaysiempre 29 minutes ago

      If you want to load the OpenGL/Vulkan vendor driver then unfortunately you don’t have much of a choice: those are linked against glibc, and I believe generally also against libwayland so screw off if you want a different protocol library (I might be wrong about the latter part). If you instead want to load plugins or whatnot into your statically linked executable, then personally I’d argue that you shouldn’t be emulating Linux dynamic linking semantics at all, because the whole late-bound global namespace thing is silly and wrong. (Solaris, which is where Glibc took this model from, moved away from it[1] as much as compatibility allowed, and so did Darwin[2], whereas Windows never made the mistake to begin with, but Glibc persisted and Musl copied it.)

      [1] https://www.linker-aliens.org/blogs/rie/entry/direct_binding...

      [2] https://web.archive.org/web/20011004090044/http://developer....

    • fabiensanglard an hour ago

      Please elaborate and explain to people with less knowledge why this is bad.

      • arjvik 44 minutes ago

        Mapping parts of files into executable memory, and then executing them, had better be bulletproof! Exploiting this seems like a direct path to RCE, and it's likely that this sort of library is used by privileged code.

        Purely academically, this is a very cool piece of code! Just hoping that it gets a thorough vetting before used by privileged/security-critical software :)

    • lunixbochs an hour ago
      • pg83 38 minutes ago

        Oh, cool, another prior art I didn't know :)

  • catlifeonmars an hour ago

    So not completely static, since it must link against a libc :P

    • pg83 39 minutes ago

      The binary itself is completely static; the link even provides commands on how to check this!

  • jeffbee an hour ago

    How are we supposed to take this stuff seriously if the author (sic) isn't even willing to write the readme? Claude exists! If I want some slop I can push the button myself.

    • pg83 24 minutes ago

      Tell me, are there any substantive comments on the text, on what has been done, and on the technical implementation, and not on the form?

    • pg83 40 minutes ago
      • WD-42 37 minutes ago

        Parents point is the readme was written by Claude. Signals low effort.

        • pg83 29 minutes ago

          README.md was written by me, and I, of course, used claude/codex for it. In general, I do everything through claude/codex, the reasons are described in https://github.com/pg83/solo/blob/main/CONTRIBUTING.md . And no, it's not low effort, and no, I don't see the point in wasting time de-claude-ifying the text just to avoid it looking like I didn't spend enough time on it.

          • Barbing 11 minutes ago

            I believe usually when someone complains about text written by a language model they are hoping to read human-written text instead of human-laundered LLM output.

      • jlebar 36 minutes ago

        Implicit assumption of gp is that the README is llm authored. (Which, I agree is how it reads to me.)