17 comments

  • ziyadb 2 minutes ago

    I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collectively. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.

    From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of that knowledge.

    Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.

  • ZoomZoomZoom 3 minutes ago

    The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?

  • eulgro 10 minutes ago

    We've been seeing that headline for a few weeks now and I really don't understand the problem.

    Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.

    Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.

    So what's the problem here exactly?

    Also from the article:

    > It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.

    I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.

    • xandrius 5 minutes ago

      The problem is that you probably do little research or read very few old books.

      There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.

      You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.

      Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.

      • brainwad 2 minutes ago

        Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.

    • josem 7 minutes ago

      I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.

    • FartyMcFarter 7 minutes ago

      The problem is we don't know what we're losing, due to lack of transparency.

  • warkdarrior 18 minutes ago

    > Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.

    Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.

    • throwatdem12311 16 minutes ago

      At least the knowledge will be available to everyone instead of mashed together and regurgitated poorly through proprietary LLMs.

    • torh 16 minutes ago

      At least there will be a copy left for us. The AI companies won't share these books in their original form.

      • brainwad 13 minutes ago

        Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things.

        Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.

        • Paratoner 8 minutes ago

          The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place

        • embedding-shape 3 minutes ago

          > because copyright law forces them to do stupid things

          This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?

          Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?

          Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.

        • xandrius 11 minutes ago

          Google probably wanted to sell the whole of Google Books.

        • naasking 9 minutes ago

          Yes, government regulations are almost always behind commercial entities making seemingly irrational choices.

    • voidhorse 11 minutes ago

      Yeah, which is completely fine. There's a major difference between:

      A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:

      - This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.

      - The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.

      - The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?

      B: Company uses freely available scanned copy of the text:

      None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.

      I much prefer B.