Follow-Up Thoughts on Watermarking Schemes for AI-Generated Text

(daringfireball.net)

10 points | by tambourine_man 13 hours ago ago

2 comments

  • sushiburps 11 hours ago

    > Contra a bunch of idiots at Hacker News and elsewhere, I understand that popular LLMs do not just pick the “best” token (word) at each decision point.

    A tad defensive, considering he generally doesn't shy away from dressing people down when they get it wrong. If you included the "Temperature" section of this follow-up above the original post it'd be an awfully confusing read.

    I agree in parts with his more principled arguments against watermarking, and I don't think he needed to tread water in the technicalities, especially given how close he comes in this post to abandoning the quality degradation argument all together anyway.

  • no-name-here 9 hours ago

    > The only thing worth evaluating is what we human readers are naturally good at determining: whether it is good or bad. If it’s good, read it. If it’s not, don’t.

    1. To determine whether something is good or bad requires reading it, which is the thing he says you should not do if it’s not good? But a big difference with AI text is that now even a single reader reading it, to determine if it’s good or bad (and therefore should not have been read per OP) takes exponentially more time than it did for the “author” to create it. (Kind of like spam?)

    2. Also, Gruber claims that output quality will not be preserved, but apparently without basing that on data about actual watermarked examples. Meanwhile, years ago Google released a study on 20 million examples showing quality was preserved: https://www.nature.com/articles/s41586-024-08025-4

    > My advice is not to care whether anything was written by an AI or a human.… If you’ve got a job where you’re surrounded by colleagues filling your inbox with AI-generated messages that you can’t abide, get a new job or learn to live with it.

    3. That assumes those exposing you to AI text is a limited number of colleagues, as opposed to, for example, a virtually unlimited number of usernames who could use AI to post articles to somewhere like Hacker News or any other site you may use. Would the corollary then be “Well then stop using Hacker News or reading other such sites” (as OP similarly says he doesn’t read LinkedIn).

    > Anyone in a situation where “getting caught” would matter — students, say — is going to use non-watermarking LLMs

    4. Is his argument that people don’t neglect to take such steps for existing watermarks such as on photos?