> “I didn’t think anyone would care” prevented me from writing, though
A few years ago I started writing Twitter threads [0]. A few weeks ago I passed 200 total threads.
When I started writing them, my thought process was: "Is anyone going to be interested in my stories/ideas??"
Dear HN comment reader, I can 100% assure you of two things:
1. If you write things, at least one person will read them.
2. It is VERY hard to predict what people will find interesting
e.g. some of the threads I thought people would find the least interesting got the most traction and vice versa. The only way to find out is to write it down.
I would also add that just writing, a LOT, helps you become a better thinker and writer. Twitter threads in particular are great as they force you to distill a story down into bite sized chunks.
One additional benefit: you meet amazing people when you write about what you are interested in. Why? Because if someone likes your writing, they would probably like talking to you and you to them.
I recall a similar quote from Elton John that I'll paraphrase:
"I've had a lot of hits so you'd think I'd know in advance which ones will become hits. Songs that I was sure would become hits went nowhere and some songs that I didn't think anything of became my biggest hits"
> how do you prove a human wrote something? Forget about the why or the value in it, just: how?
> Semoi is a plugin (currently only available for Obsidian) which tracks the length of time it took for a document to be written up
Trying to mechanistically prove that a human created some content as opposed to ai, in the age of LLMs and style transfer when you can just ask for something to be written in the style of Mark Twain or drawn in a style of van Gogh and get a great output, is a fool's errand.
All solutions to this end are going to be some form of attestation.
Even the proposed approach of tracking keystrokes and timing as a form of mechanical attestation, is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication. May not even need an ai for this, a stochastic program could conceivably reproduce this behavior.
Ok so what if I just rephrase the essay that the LLM gave me, while keeping all the overall message and ideas and arguments, just slightly rewritten in my own words, but without having any idea about what it all really means and whether it's correct. Is that something worth reading now, but the LLM output is not worth reading?
No, still not worth reading. Do you have any ideas? Tell me about them. If you don't have any idea what it all really means and whether it's correct, then don't waste my time, whether or not an AI is involved.
Right. But you'll never have a checker for that. How do you check for the overall idea being AI-shaped? It's a much deeper check than style, even typing patterns or anything like that. If I consult the chatbot, then memorize the overall arguments and ideas, and then record myself typing out an essay regurgitating it in my own words, how can you distinguish that from writing down my own ideas?
Fundamentally, I don't actually care about whether AI wrote it or not. I care about whether the ideas are well-thought-out or half-baked. I care whether the writing is slop (whether AI or corporate PR or politician-speak or whatever), or whether the writing is actually something a human would want to read.
But as you say, we don't know how to write a checker for that...
> is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication
You can remotely attest the input devices. You can do it anonymously (long story, but doable) and without requiring some kind of pre-signed image for the while OS. (The OS passes through recent-input attestations.)
I have a slightly different approach on my blog https://ezeugo.dev where my entire thought process (edits, original ideas, rewrites are a part of the actual essay). By making the event stream part of the product, it shows the process and output as one.
I think we are reaching a point where some kind of solid attestation that things are not AI generated would be very valuable, for all kinds of different media (printed word, photo, video, etc).
I'm not sure how you would actually do any of that attestation, if it's even possible. Text seems especially difficult. Maybe photo/video could be achieved with specialized hardware and cryptographic signing though I don't know much about either so I'm not sure how it would work. Maybe all that attestation would also tie in to some kind of universal personal identifier online, so that bad actors can be tracked or excluded and can't repeatedly spin up new accounts.
It might mean a big reduction in privacy for certain online spaces that opt in to such a system... but the alternative of all trust being eroded and voices drowned out by a sea of bots or generated content seems potentially worse.
Just like we trained on human language to create LLMs, we can train on human keystrokes with a similar algorithm and spit out believable (at least statistically) "human" keystrokes generated by machine.
Yeah I think the "how can you prove text was handwritten" question is a subset of the larger "how can you prove that a computer is being driven by a human" problem that all of the work around captcha, attestations, biometrics, and government-id auth has been aimed at. The fundamental issue, it seems to me, is that any signal that a human can provide to a computer (keystroke, camera frame, mouse click, etc) is inherently only parsable by code because a sensor has translated the analog signal into a digital one. That same requirement also ensures that the input can be digitally spoofed or automated. There's a similar problem on the output side: how can an analog user trust a digital certificate? What's stopping me from copying the certificate HTML or taking a screenshot and using it to trick people into thinking my AI content is handwritten?
I don't have any suggestions. I worry that the only strong solutions require a lot of power to be given to a centralized authority.
> One could work around Semoi by, for example, typing out a bunch of gibberish, leaving their editor open, and then pasting in an LLM generated texting and minting the proof. To which I would respond: why? That’s really pathetic.
Even before LLMs, proving that a specific human wrote some piece of text was difficult or imprecise due to coauthoring, editing, plagiarism, etc. The important question then, as now, is rather to determine whether a specific human approved some piece of text for publication under their name.
This is what puzzles me about the whole topic as well. Why do we need to treat AI texts as anything special? As long as there is a name next to the text, we should care about the message instead of the way it was written. It reminds me of the situation with political ads, when after the main part of a commercial we hear: "I'm John Doe and I approve this message." It doesn't matter who wrote the script of the commercial - John Doe confirms that it is his position, and that is what matters most.
For example, English is not my native language. I can speak, read, write - I have no issue using it for work or everyday life. My own kid only speaks English. But when we talk about writing an article, I would want to polish it. I would want to put my thoughts into a better form, so people may enjoy reading a well-written text which may have some fragments written or edited by AI so it will be simply better. I would use it for additional fact-check. Communication is not a competition in language skills.
I once heard a story told by a journalist. He used to write articles for the NYT from time to time, and the process was like this: he knows English, but the NYT asked him to write in his native language, a very experienced translator produced the English text, and then they polished it together with the editors. Once the article was published, it mentioned only his name — no mentions of editors or translators. Why? Because creating a text is not just writing or typing. Often it is a more complicated process which may or may not include other people or systems. What matters most is whether the author puts their signature at the end or not.
This means that we must hold humans responsible for every stupid thing they say, whether or not AI told it to them. And in fact, that's the road we should have been on from day one. "Computers do not lie" is an idea that should have died in the 1970s.
As others have pointed out, it's relatively a lot of effort to create an artifact that realistically current systems can pretty well forge.
I don't know that there is a scalable and comfortable solution to this problem (or at least one that is scalable and comfortable proportional to the demand for it).
AI detection has gotten a lot better recently, and I've found that Pangram is "good enough" for a lot of use cases. A lot of people are uncomfortable with opaque detection methods, but waiting for a perfect solution (if such a solution even exists) will cause communication channels to be filled with bad actors who turn your community/platform/etc... into something resembling LinkedIn.
Detecting Ai use based on "how many keystrokes did you type and when" doesn't solve the problem, because someone can just write a program to mimic human typing in the words.
So the issue isn't "did a human write something", it's what the actual content is
I'll be honest, I think this is a social issue that you're trying to apply a process-solution to solve.
If someone writes "this was completely hand-written with no AI assistance", I'll just believe them. I'm already committed to letting your words fill my brain for a bit, so I don't know why I would NEED a cryptographic signature to PROVE you aren't lying about WHO wrote it.
Being called out as a lier will be a LOT more painful than 1) using an LLM to write quickly and not lying, or 2) doing what you claim to be doing and writing it with your feeble, non-metallic human hands.
(First version of this comment had an example from an HN thread of this SPECIFIC behaviour getting called out but ehh that’s not the right vibe. My point is it does happen.)
That's actually a good point. Plenty of people just want to let their beliefs out in public. And there's not much benefit to lie about AI use. Some people use it and some people don't, and it's two very different groups of people.
I mean the ability to fake all these metrics is relatively trivial, since I have written anti-bot fingerprinting scripts for automating various online services that want to keep you from automating them, you have the text you want to send from source A, you have random typing speed array you want to type them in, your chance of mistakes (put wrong character, backspace to remove, put in correct character).
And this doesn't even have to deal with all the stuff about mouse movements that you don't register?
Of course maybe I am just being typical programmer here, I guess lots of the people use generative AI would be defeated by copying pasting in the text and getting labeled AI, but that would also incorrectly label lots of people who have old texts in handwritten form they do not want to type all over again (of which I am one), and finally I assume that there is money in the field so producing something that allows bots to display "human heuristics" would probably get made and be profitable.
Funnily enough when I was automating things, generally twitter, I discovered that my real usage often got registered as bot, so I figured what's the use.
Also talking with someone who actually worked on bot-recognition by usage metrics said I was overly paranoid on some of the things I made my scripts do to appear human.
on the other hand - my automation was based on not wanting to spam services but provide the minimum level of content posting and regularity to benefit from algorithms that boost content based on the poster's engagement level, automation that wants to spam cannot benefit from this because slowing things down to show as if it was made in real time by a human (with fully non-headless browser etc.) does somewhat defeat the ability to spam like a machine.
i do actually think the general thrust of the idea is good. i also think it's sort of inherently cursed, and has some amount of a predeterminable fate.
if it doesnt take off, it dies.
if it does take off and becomes a relevant currency of some sort, it will need improving. if it needs improving how far are we willing to go? does some alternative system fork off to handle severity of provenance concern?
what happens when AI action becomes indistinguishable from human action? what happens when sticking computers in your head becomes vogue?
if it takes off and fills a small niche, maybe thats the best future.
Like I said, cool idea, but seemingly very cursed from the get go.
There's plenty of microbehavioral analysis we can do that is initially effective but will get bypassed (with GANs being the purest way, or something more domain-specific).
You could imagine a livestream that's permanently published somewhere. But the verification of the livestream takes longer than reading the piece itself (and is itself vulnerable to faking).
Personally I think it comes back to something like writing under your real name - staking your reputation - as the most trustworthy indicator. At least people in your circles can trust you.
You can read ahead as to how this would go by looking at the CAPTCHA world, which unbeknownst to a lot of people left behind "click this image" as the actual test a long time ago and does a lot of behavioral analysis of mouse motions and stuff. Which is itself a constant arms race. Which I would still characterize as "advantage attacker" with regard to that aspect of it, CAPTCHAs still have at least some utility more from the other streams they have access to that are harder to fake, e.g., "this IP address is from a rentable service space and thus more likely to be a bot" sorts of checks.
That means this basically works, as long as it never gets large enough to attack. Which may suit the author just fine. Not everything has to solve the world's problems. But it won't generalize very far, no.
I’ve been playing around with some similar ideas in a couple of pet projects. I truly believe that some type of “proof of work” system like this will be the only reliable way to have confidence in human writing. The various AI “detector” software that’s out there seems to be a dead end to me.
This is a good idea - and I might pick it up for my writing, since a lot of it still happens in Obsidian. If a document is worked on across multiple days / revisions / app sessions - how is that handled?
I'm choosing to read this as an option for people who want to put forward a serious attestation that the writing is not generated by AI, and are looking for ways to make that pledge tangible and give it a level of authority that feels more significant than "I promise" by getting that event-stream sent and signed. In that light, it seems like a useful offering!
I think it's important to consider the risks/rewards/benefits. There's definitely a sense in my circle of contacts that if a work is seen as purely human then it's somehow better and more authentic, and annecdotally, it's also possible to win points by taking something created with the help of AI and passing it off as your own independent work. Like social credit, there's a sense that you'll seem smarter than you feel yourself to be.
With that in mind, absolutely any technical solution to detect AI or attest to human authorship will be abused. The only context that a solution like this one supports is one where there isn't a risk or reward, the author just sincerely wants you to know that it's human-produced.
Plenty of people, after all, produce a draft with AI, then type it all out again editing and refining and updating as they go, and then ask an AI to look at the result and make suggestions, and then go back and make the changes they agree with. Such a workflow would be deemed human by semoi - so it's lucky that there's no point in lying about it.
i dunno why we dont do what the art world does and just blacklist you from everything if you're caught stealing/tracing and even revoke your degree(s) for it.
ai detectors always have so many false positives that im convinced they're only put in place by people too dumb to know any better and snakeoil salesmen.
im sorry but have you met the obsessive llm weirdos like
theyre already like that.
also im not talking about llm output? just... if you're caught passing it as your own in any field you get blacklisted which we already do with regular stealing
> One could work around Semoi by, for example, typing out a bunch of gibberish, leaving their editor open, and then pasting in an LLM generated texting and minting the proof. To which I would respond: why? That’s really pathetic.
I mean, I also think it's really pathetic to have an AI write something and then say "this text was written with no AI assistance"†. So if we've acknowledged that we're only going to stop non-pathetic people, why not skip the cryptographic hash signing and just go with the no-AI statement?
I understand that the goal is only to make lying hard, not impossible. However, I don't think this solution makes lying harder enough to meaningful.
I used to live with a ghostwriter for Simon Sinek, she was never attributed or acknowledged in "his" books. Since Ai authorship has become a common topic, I've wondered how the two means of producing a book relate and how it might inform this debate.
As a writer who publishes regularly on the web, I just keep the edit history of my longer articles on my github. Its not perfect and you can never really prove you wrote something unless you do it in front of an observer watching you write in real life imo. At some point though, you have to make a choice between not wanting to waste people's and your readers' time but also privacy and additional effort. One of my recent blogs: https://decodingvibes.com/blog/what-we-talk-about-when-we-ta...
The technical solution is interesting but I think the only way to truly go about it is web of trust. Everyone knows that my work doesn't use AI because I hate it so much and so many people know that I don't use AI. It's embedded into the core of my personality. These tools can always be gamed but a true belief against AI cannot.
A lot of people do, that's the point. And over time there will be zero evidence that I ever used an LLM because I haven't. Many people in real life also believe it and they can also vouch for me.
"No information about your text is ever sent to the server, and there’s no way to identify the author based on the minted certificate"
This means the certificate is independent of author and source text.
There's nothing stopping you from sending fake counts/duration to the semoi server. It's a certificate that only says "at this point in time, this is the information I was provided with".
You can then attach it to any piece of text you like.
At the very least, you'd need the ability to prove that there is an underlying event stream with these characteristics, and that this exact event stream creates the document in question. You still can fake that event stream, but it becomes enough work to distract at least casual abusers.
But really, it's the equivalent of saying "I wrote this without AI, honest" in-doc and signing that with your personal key. The value depends entirely on your willingess to be truthful. (IOW: I predict we'll see a resurgence of reputation systems, to some extent)
> IOW: I predict we'll see a resurgence of reputation systems, to some extent
I think you're right that reputation systems are the best solution to a low signal-to-noise ratio.
Consumers (and the agents under their control) will increasingly prioritize content and other products with reliable attestation that they come from a trustworthy publisher, organization, brand, or individual.
I'm not doing that. I'm not going to be coerced into a writing style that's not mine by LLMs. AI detectors think that 30% of the stuff I wrote 20 years ago is AI-generated. If someone wants to think, incorrectly, that my writing was done by an LLM, I can't stop them.
> “I didn’t think anyone would care” prevented me from writing, though
A few years ago I started writing Twitter threads [0]. A few weeks ago I passed 200 total threads.
When I started writing them, my thought process was: "Is anyone going to be interested in my stories/ideas??"
Dear HN comment reader, I can 100% assure you of two things:
1. If you write things, at least one person will read them.
2. It is VERY hard to predict what people will find interesting
e.g. some of the threads I thought people would find the least interesting got the most traction and vice versa. The only way to find out is to write it down.
I would also add that just writing, a LOT, helps you become a better thinker and writer. Twitter threads in particular are great as they force you to distill a story down into bite sized chunks.
One additional benefit: you meet amazing people when you write about what you are interested in. Why? Because if someone likes your writing, they would probably like talking to you and you to them.
0 - https://x.com/alexpotato/status/2012723178577985948?s=20
I recall a similar quote from Elton John that I'll paraphrase:
"I've had a lot of hits so you'd think I'd know in advance which ones will become hits. Songs that I was sure would become hits went nowhere and some songs that I didn't think anything of became my biggest hits"
It's been a long time since I heard this so I'm probably mangling it. A quick search shows that he probably did say something like this though https://www.birminghammail.co.uk/news/showbiz-tv/sir-elton-j...
The lesson is - just put it out there and see what happens
> how do you prove a human wrote something? Forget about the why or the value in it, just: how?
> Semoi is a plugin (currently only available for Obsidian) which tracks the length of time it took for a document to be written up
Trying to mechanistically prove that a human created some content as opposed to ai, in the age of LLMs and style transfer when you can just ask for something to be written in the style of Mark Twain or drawn in a style of van Gogh and get a great output, is a fool's errand.
All solutions to this end are going to be some form of attestation.
Even the proposed approach of tracking keystrokes and timing as a form of mechanical attestation, is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication. May not even need an ai for this, a stochastic program could conceivably reproduce this behavior.
Ok so what if I just rephrase the essay that the LLM gave me, while keeping all the overall message and ideas and arguments, just slightly rewritten in my own words, but without having any idea about what it all really means and whether it's correct. Is that something worth reading now, but the LLM output is not worth reading?
No, still not worth reading. Do you have any ideas? Tell me about them. If you don't have any idea what it all really means and whether it's correct, then don't waste my time, whether or not an AI is involved.
Right. But you'll never have a checker for that. How do you check for the overall idea being AI-shaped? It's a much deeper check than style, even typing patterns or anything like that. If I consult the chatbot, then memorize the overall arguments and ideas, and then record myself typing out an essay regurgitating it in my own words, how can you distinguish that from writing down my own ideas?
Fundamentally, I don't actually care about whether AI wrote it or not. I care about whether the ideas are well-thought-out or half-baked. I care whether the writing is slop (whether AI or corporate PR or politician-speak or whatever), or whether the writing is actually something a human would want to read.
But as you say, we don't know how to write a checker for that...
> is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication
You can remotely attest the input devices. You can do it anonymously (long story, but doable) and without requiring some kind of pre-signed image for the while OS. (The OS passes through recent-input attestations.)
I have a slightly different approach on my blog https://ezeugo.dev where my entire thought process (edits, original ideas, rewrites are a part of the actual essay). By making the event stream part of the product, it shows the process and output as one.
I think we are reaching a point where some kind of solid attestation that things are not AI generated would be very valuable, for all kinds of different media (printed word, photo, video, etc).
I'm not sure how you would actually do any of that attestation, if it's even possible. Text seems especially difficult. Maybe photo/video could be achieved with specialized hardware and cryptographic signing though I don't know much about either so I'm not sure how it would work. Maybe all that attestation would also tie in to some kind of universal personal identifier online, so that bad actors can be tracked or excluded and can't repeatedly spin up new accounts.
It might mean a big reduction in privacy for certain online spaces that opt in to such a system... but the alternative of all trust being eroded and voices drowned out by a sea of bots or generated content seems potentially worse.
Just like we trained on human language to create LLMs, we can train on human keystrokes with a similar algorithm and spit out believable (at least statistically) "human" keystrokes generated by machine.
Yeah I think the "how can you prove text was handwritten" question is a subset of the larger "how can you prove that a computer is being driven by a human" problem that all of the work around captcha, attestations, biometrics, and government-id auth has been aimed at. The fundamental issue, it seems to me, is that any signal that a human can provide to a computer (keystroke, camera frame, mouse click, etc) is inherently only parsable by code because a sensor has translated the analog signal into a digital one. That same requirement also ensures that the input can be digitally spoofed or automated. There's a similar problem on the output side: how can an analog user trust a digital certificate? What's stopping me from copying the certificate HTML or taking a screenshot and using it to trick people into thinking my AI content is handwritten?
I don't have any suggestions. I worry that the only strong solutions require a lot of power to be given to a centralized authority.
> One could work around Semoi by, for example, typing out a bunch of gibberish, leaving their editor open, and then pasting in an LLM generated texting and minting the proof. To which I would respond: why? That’s really pathetic.
Oh, right, the "it's pathetic" security mechanism.
Even before LLMs, proving that a specific human wrote some piece of text was difficult or imprecise due to coauthoring, editing, plagiarism, etc. The important question then, as now, is rather to determine whether a specific human approved some piece of text for publication under their name.
This is what puzzles me about the whole topic as well. Why do we need to treat AI texts as anything special? As long as there is a name next to the text, we should care about the message instead of the way it was written. It reminds me of the situation with political ads, when after the main part of a commercial we hear: "I'm John Doe and I approve this message." It doesn't matter who wrote the script of the commercial - John Doe confirms that it is his position, and that is what matters most.
For example, English is not my native language. I can speak, read, write - I have no issue using it for work or everyday life. My own kid only speaks English. But when we talk about writing an article, I would want to polish it. I would want to put my thoughts into a better form, so people may enjoy reading a well-written text which may have some fragments written or edited by AI so it will be simply better. I would use it for additional fact-check. Communication is not a competition in language skills.
I once heard a story told by a journalist. He used to write articles for the NYT from time to time, and the process was like this: he knows English, but the NYT asked him to write in his native language, a very experienced translator produced the English text, and then they polished it together with the editors. Once the article was published, it mentioned only his name — no mentions of editors or translators. Why? Because creating a text is not just writing or typing. Often it is a more complicated process which may or may not include other people or systems. What matters most is whether the author puts their signature at the end or not.
This means that we must hold humans responsible for every stupid thing they say, whether or not AI told it to them. And in fact, that's the road we should have been on from day one. "Computers do not lie" is an idea that should have died in the 1970s.
There is https://trulytyped.com/
https://news.ycombinator.com/item?id=48125186
I built and open sourced something similar between 2024 and 2025.
https://github.com/humthentic/itypedmypaper-v1
As others have pointed out, it's relatively a lot of effort to create an artifact that realistically current systems can pretty well forge.
I don't know that there is a scalable and comfortable solution to this problem (or at least one that is scalable and comfortable proportional to the demand for it).
AI detection has gotten a lot better recently, and I've found that Pangram is "good enough" for a lot of use cases. A lot of people are uncomfortable with opaque detection methods, but waiting for a perfect solution (if such a solution even exists) will cause communication channels to be filled with bad actors who turn your community/platform/etc... into something resembling LinkedIn.
Detecting Ai use based on "how many keystrokes did you type and when" doesn't solve the problem, because someone can just write a program to mimic human typing in the words.
So the issue isn't "did a human write something", it's what the actual content is
I'll be honest, I think this is a social issue that you're trying to apply a process-solution to solve.
If someone writes "this was completely hand-written with no AI assistance", I'll just believe them. I'm already committed to letting your words fill my brain for a bit, so I don't know why I would NEED a cryptographic signature to PROVE you aren't lying about WHO wrote it.
Being called out as a lier will be a LOT more painful than 1) using an LLM to write quickly and not lying, or 2) doing what you claim to be doing and writing it with your feeble, non-metallic human hands.
(First version of this comment had an example from an HN thread of this SPECIFIC behaviour getting called out but ehh that’s not the right vibe. My point is it does happen.)
That's actually a good point. Plenty of people just want to let their beliefs out in public. And there's not much benefit to lie about AI use. Some people use it and some people don't, and it's two very different groups of people.
I mean the ability to fake all these metrics is relatively trivial, since I have written anti-bot fingerprinting scripts for automating various online services that want to keep you from automating them, you have the text you want to send from source A, you have random typing speed array you want to type them in, your chance of mistakes (put wrong character, backspace to remove, put in correct character).
And this doesn't even have to deal with all the stuff about mouse movements that you don't register?
Of course maybe I am just being typical programmer here, I guess lots of the people use generative AI would be defeated by copying pasting in the text and getting labeled AI, but that would also incorrectly label lots of people who have old texts in handwritten form they do not want to type all over again (of which I am one), and finally I assume that there is money in the field so producing something that allows bots to display "human heuristics" would probably get made and be profitable.
Funnily enough when I was automating things, generally twitter, I discovered that my real usage often got registered as bot, so I figured what's the use.
Also talking with someone who actually worked on bot-recognition by usage metrics said I was overly paranoid on some of the things I made my scripts do to appear human.
on the other hand - my automation was based on not wanting to spam services but provide the minimum level of content posting and regularity to benefit from algorithms that boost content based on the poster's engagement level, automation that wants to spam cannot benefit from this because slowing things down to show as if it was made in real time by a human (with fully non-headless browser etc.) does somewhat defeat the ability to spam like a machine.
i do actually think the general thrust of the idea is good. i also think it's sort of inherently cursed, and has some amount of a predeterminable fate.
if it doesnt take off, it dies.
if it does take off and becomes a relevant currency of some sort, it will need improving. if it needs improving how far are we willing to go? does some alternative system fork off to handle severity of provenance concern?
what happens when AI action becomes indistinguishable from human action? what happens when sticking computers in your head becomes vogue?
if it takes off and fills a small niche, maybe thats the best future.
Like I said, cool idea, but seemingly very cursed from the get go.
Here's another similar project: https://writetrack.dev/
There's plenty of microbehavioral analysis we can do that is initially effective but will get bypassed (with GANs being the purest way, or something more domain-specific).
You could imagine a livestream that's permanently published somewhere. But the verification of the livestream takes longer than reading the piece itself (and is itself vulnerable to faking).
Personally I think it comes back to something like writing under your real name - staking your reputation - as the most trustworthy indicator. At least people in your circles can trust you.
You can read ahead as to how this would go by looking at the CAPTCHA world, which unbeknownst to a lot of people left behind "click this image" as the actual test a long time ago and does a lot of behavioral analysis of mouse motions and stuff. Which is itself a constant arms race. Which I would still characterize as "advantage attacker" with regard to that aspect of it, CAPTCHAs still have at least some utility more from the other streams they have access to that are harder to fake, e.g., "this IP address is from a rentable service space and thus more likely to be a bot" sorts of checks.
This is an arms race, advantage attacker.
That means this basically works, as long as it never gets large enough to attack. Which may suit the author just fine. Not everything has to solve the world's problems. But it won't generalize very far, no.
I’ve been playing around with some similar ideas in a couple of pet projects. I truly believe that some type of “proof of work” system like this will be the only reliable way to have confidence in human writing. The various AI “detector” software that’s out there seems to be a dead end to me.
I can’t conceive of any system that couldn’t be gamed. The better bet, in my opinion, is gauging the quality of the writing, as we already do.
What exactly do you think is wrong with Pangram?
I do wish we have a better way of doing this - the workaround given in the article just shows how fallible this is.
This is a good idea - and I might pick it up for my writing, since a lot of it still happens in Obsidian. If a document is worked on across multiple days / revisions / app sessions - how is that handled?
I'm choosing to read this as an option for people who want to put forward a serious attestation that the writing is not generated by AI, and are looking for ways to make that pledge tangible and give it a level of authority that feels more significant than "I promise" by getting that event-stream sent and signed. In that light, it seems like a useful offering!
I think it's important to consider the risks/rewards/benefits. There's definitely a sense in my circle of contacts that if a work is seen as purely human then it's somehow better and more authentic, and annecdotally, it's also possible to win points by taking something created with the help of AI and passing it off as your own independent work. Like social credit, there's a sense that you'll seem smarter than you feel yourself to be.
With that in mind, absolutely any technical solution to detect AI or attest to human authorship will be abused. The only context that a solution like this one supports is one where there isn't a risk or reward, the author just sincerely wants you to know that it's human-produced.
Plenty of people, after all, produce a draft with AI, then type it all out again editing and refining and updating as they go, and then ask an AI to look at the result and make suggestions, and then go back and make the changes they agree with. Such a workflow would be deemed human by semoi - so it's lucky that there's no point in lying about it.
i dunno why we dont do what the art world does and just blacklist you from everything if you're caught stealing/tracing and even revoke your degree(s) for it.
ai detectors always have so many false positives that im convinced they're only put in place by people too dumb to know any better and snakeoil salesmen.
(redacted) misunderstood oc
im sorry but have you met the obsessive llm weirdos like theyre already like that.
also im not talking about llm output? just... if you're caught passing it as your own in any field you get blacklisted which we already do with regular stealing
wtf is even your reply my dude.
> One could work around Semoi by, for example, typing out a bunch of gibberish, leaving their editor open, and then pasting in an LLM generated texting and minting the proof. To which I would respond: why? That’s really pathetic.
I mean, I also think it's really pathetic to have an AI write something and then say "this text was written with no AI assistance"†. So if we've acknowledged that we're only going to stop non-pathetic people, why not skip the cryptographic hash signing and just go with the no-AI statement?
I understand that the goal is only to make lying hard, not impossible. However, I don't think this solution makes lying harder enough to meaningful.
I guess I don't understand why the AI itself couldn't do the workaround.
I used to live with a ghostwriter for Simon Sinek, she was never attributed or acknowledged in "his" books. Since Ai authorship has become a common topic, I've wondered how the two means of producing a book relate and how it might inform this debate.
As a writer who publishes regularly on the web, I just keep the edit history of my longer articles on my github. Its not perfect and you can never really prove you wrote something unless you do it in front of an observer watching you write in real life imo. At some point though, you have to make a choice between not wanting to waste people's and your readers' time but also privacy and additional effort. One of my recent blogs: https://decodingvibes.com/blog/what-we-talk-about-when-we-ta...
The technical solution is interesting but I think the only way to truly go about it is web of trust. Everyone knows that my work doesn't use AI because I hate it so much and so many people know that I don't use AI. It's embedded into the core of my personality. These tools can always be gamed but a true belief against AI cannot.
And what's preventing you to LLM slop your next post?
We believe you on internet karma?
A lot of people do, that's the point. And over time there will be zero evidence that I ever used an LLM because I haven't. Many people in real life also believe it and they can also vouch for me.
This was already an issue before LLMs https://en.wikipedia.org/wiki/Keystroke_dynamics
"No information about your text is ever sent to the server, and there’s no way to identify the author based on the minted certificate"
This means the certificate is independent of author and source text.
There's nothing stopping you from sending fake counts/duration to the semoi server. It's a certificate that only says "at this point in time, this is the information I was provided with".
You can then attach it to any piece of text you like.
At the very least, you'd need the ability to prove that there is an underlying event stream with these characteristics, and that this exact event stream creates the document in question. You still can fake that event stream, but it becomes enough work to distract at least casual abusers.
But really, it's the equivalent of saying "I wrote this without AI, honest" in-doc and signing that with your personal key. The value depends entirely on your willingess to be truthful. (IOW: I predict we'll see a resurgence of reputation systems, to some extent)
> IOW: I predict we'll see a resurgence of reputation systems, to some extent
I think you're right that reputation systems are the best solution to a low signal-to-noise ratio.
Consumers (and the agents under their control) will increasingly prioritize content and other products with reliable attestation that they come from a trustworthy publisher, organization, brand, or individual.
I feel bad for actual humans named Claude. Cause that's this months —
I'm actively reviewing my writing and making it worse to avoid being accused of LLM slop.
It's so difficult to write, you review, edit, edit and you read it and you can't unsee the LLM criticism.
I'm not doing that. I'm not going to be coerced into a writing style that's not mine by LLMs. AI detectors think that 30% of the stuff I wrote 20 years ago is AI-generated. If someone wants to think, incorrectly, that my writing was done by an LLM, I can't stop them.
[dead]