I wrote something original in a language learning grammar online. I coined a term to describe something about how a certain language with negative language functions.
I tried talking to an LLM about it and sure and enough, it had scraped it and learned to use that terminology and explanation that I made up. I asked the LLM, "Where did that concept/term come from, who made it up?" It did a bunch of searching and tried to site a whole bunch of other sources which did NOT contain the terms of the explanation I had written. It simply would not cite or mention my source. It kept parroting my material semi-correctly and hiding the source. By any other actor that would be an egregious act of sloppiness, dishonesty, and plagiarism.
Generative AI is speed-running a widespread corruption of truth. I do not believe any of the gains are worth this.
Your text may well have been in the training corpus, but not searchable with whatever terms and search engine that LLM used after you prompted it. It doesn’t have recall of sources of documents comprising its training corpus, unless the source is widely cited in other of its training documents. Sourcing and provenance aren’t currently an intentional part of LLM training.
I would assume the parent understands that. The criticism is about the LLM behavior this results in. Being able to explain a behavior doesn’t necessarily excuse it. By some moral standards, you wouldn’t have trained an LLM that way, or wouldn’t have made it available, knowing the outcome.
Only to be topped by the massive amount of IP that Copilot is possibly secretly harvesting at present. We already believe OpenAI/Anthropic are doing it, what’s to stop Microsoft from attempting to improve their competitive advantage by surreptitiously using their role as a MiTM between end user Corporations and Model servers.
"Oh no did that 'Allow us to train on your org's data' toggle get turned on after the last update? We're so sorry. No we won;t delete the data, the toggle was -ON- what part of that don't you get??
Whether religious or not, the verse applies to anyone and everyone making such statements from the position of companies like Microsoft. I'll quite the KJV version 'cause that one lends itself well to situations like these. Read it with a booming thunder-loaded voice:
And why beholdest thou the mote that is in thy brother's eye, but considerest not the beam that is in thine own eye?
Or how wilt thou say to thy brother, Let me pull out the mote out of thine eye; and, behold, a beam is in thine own eye?
Thou hypocrite, first cast out the beam out of thine own eye; and then shalt thou see clearly to cast out the mote out of thy brother's eye.
It's a memo taken out of context from Jan 2023, literally 1 month after the ChatGPT public release. All the article cites:
> The brief cited an internal memo dated January 2023 by Microsoft director of Applied Science Brent Hecht, where he allegedly said, “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and also called it “the largest theft of labor in human history.”
No need to add religion into this. Let MS employees write their emails. If anything, what he said a month after ChatGPT's release was prophetic.
It is a translation of "dokós", a plank of wood that would serve as a joist or rafter. An ancient palestinian Phineas Gage stuck through with a classical age 2x10, if you will.
By stealing text from the Bible rather than using your own words, you've just committed labor theft!
Of course that accusation sounds outrageous as it should. All human knowledge and progress has come by building on the works of those who came before us, as it should. Turning this on its head and calling it "theft of labor" is insane, as is asking others not to build on today's knowledge.
Or to steal some more human labor: "If I have seen further than others, it is by standing on the shoulders of giants."
if he “had been made aware that OpenAI has scraped and trained on information that was behind a paywall,” the company would have required OpenAI “to retrain its models.”
An executive at this level is not going to risk their personal political capital to make such statements unless its part of plan and thought through. They might be right but in capitalistic system no company would given them a opportunity after such statements.
When we sing, draw, speak, play, we are not attacking the soul of someone else!
Culture has always had a kind of violence to it, but not FROM the THIEF: you HAVE to speak the language that was already made; you have to understand and participate in their song-structures, their plot devices, their ways of attesting.
Either way, a change did happen, financial impact aside this change is definitely going to disrupt the ecosystem as a whole which is a much bigger unknown at this point.
> A huge chunk of the content that was scraped was most likely provided voluntarily rather than being waged labour.
People committing work get rewards in different ways, waged labour could be one, attributions, acknowledgements, future job opportunities and more. Stripping those from the rightful owners can hardly be justified in the name of progress.
I wasn't implying that it was okay to "steal" things because they were offered voluntarily. My position on this is kind of complex and nuanced but this particular reply was meant to state the opposite.
I wrote something original in a language learning grammar online. I coined a term to describe something about how a certain language with negative language functions.
I tried talking to an LLM about it and sure and enough, it had scraped it and learned to use that terminology and explanation that I made up. I asked the LLM, "Where did that concept/term come from, who made it up?" It did a bunch of searching and tried to site a whole bunch of other sources which did NOT contain the terms of the explanation I had written. It simply would not cite or mention my source. It kept parroting my material semi-correctly and hiding the source. By any other actor that would be an egregious act of sloppiness, dishonesty, and plagiarism.
Generative AI is speed-running a widespread corruption of truth. I do not believe any of the gains are worth this.
Your text may well have been in the training corpus, but not searchable with whatever terms and search engine that LLM used after you prompted it. It doesn’t have recall of sources of documents comprising its training corpus, unless the source is widely cited in other of its training documents. Sourcing and provenance aren’t currently an intentional part of LLM training.
I would assume the parent understands that. The criticism is about the LLM behavior this results in. Being able to explain a behavior doesn’t necessarily excuse it. By some moral standards, you wouldn’t have trained an LLM that way, or wouldn’t have made it available, knowing the outcome.
Discussion from yesterday (811 comments): https://news.ycombinator.com/item?id=49752056
Only to be topped by the massive amount of IP that Copilot is possibly secretly harvesting at present. We already believe OpenAI/Anthropic are doing it, what’s to stop Microsoft from attempting to improve their competitive advantage by surreptitiously using their role as a MiTM between end user Corporations and Model servers.
"Oh no did that 'Allow us to train on your org's data' toggle get turned on after the last update? We're so sorry. No we won;t delete the data, the toggle was -ON- what part of that don't you get??
Who is this we?
FUD.
Here's a Bible verse this Microsoft individual should acquaint himself with:
https://www.biblegateway.com/verse/en/Matthew%207%3A5
Whether religious or not, the verse applies to anyone and everyone making such statements from the position of companies like Microsoft. I'll quite the KJV version 'cause that one lends itself well to situations like these. Read it with a booming thunder-loaded voice:
And why beholdest thou the mote that is in thy brother's eye, but considerest not the beam that is in thine own eye?
Or how wilt thou say to thy brother, Let me pull out the mote out of thine eye; and, behold, a beam is in thine own eye?
Thou hypocrite, first cast out the beam out of thine own eye; and then shalt thou see clearly to cast out the mote out of thy brother's eye.
It's a memo taken out of context from Jan 2023, literally 1 month after the ChatGPT public release. All the article cites:
> The brief cited an internal memo dated January 2023 by Microsoft director of Applied Science Brent Hecht, where he allegedly said, “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and also called it “the largest theft of labor in human history.”
No need to add religion into this. Let MS employees write their emails. If anything, what he said a month after ChatGPT's release was prophetic.
What is meant by "beam" in this quotation? A splinter of wood?
It is a translation of "dokós", a plank of wood that would serve as a joist or rafter. An ancient palestinian Phineas Gage stuck through with a classical age 2x10, if you will.
Yeah
By stealing text from the Bible rather than using your own words, you've just committed labor theft!
Of course that accusation sounds outrageous as it should. All human knowledge and progress has come by building on the works of those who came before us, as it should. Turning this on its head and calling it "theft of labor" is insane, as is asking others not to build on today's knowledge.
Or to steal some more human labor: "If I have seen further than others, it is by standing on the shoulders of giants."
if he “had been made aware that OpenAI has scraped and trained on information that was behind a paywall,” the company would have required OpenAI “to retrain its models.”
... yeah, sure.
Oh let's talk about sins of Microsoft since the 1980s, shall we?
He is correct, but why make the statement ? My guess, Microsoft is failing on their AI push.
Microsoft owns a quarter of OpenAI and 20% of it's revenue
An executive at this level is not going to risk their personal political capital to make such statements unless its part of plan and thought through. They might be right but in capitalistic system no company would given them a opportunity after such statements.
Presumably someone was paid when the content was created. So it's not really the labor that's being stolen, but the future value of the deliverable.
I'm pretty sure any of the eras of slavery throughout human history is more of an actual theft of labor.
> So it's not really the labor that's being stolen, but the future value of the deliverable.
the choice was stolen as well.
the agency was stolen as well.
When we sing, draw, speak, play, we are not attacking the soul of someone else!
Culture has always had a kind of violence to it, but not FROM the THIEF: you HAVE to speak the language that was already made; you have to understand and participate in their song-structures, their plot devices, their ways of attesting.
The "IP owners" should be paying ME!!!
A huge chunk of the content that was scraped was most likely provided voluntarily rather than being waged labour.
Either way, a change did happen, financial impact aside this change is definitely going to disrupt the ecosystem as a whole which is a much bigger unknown at this point.
> A huge chunk of the content that was scraped was most likely provided voluntarily rather than being waged labour.
People committing work get rewards in different ways, waged labour could be one, attributions, acknowledgements, future job opportunities and more. Stripping those from the rightful owners can hardly be justified in the name of progress.
I wasn't implying that it was okay to "steal" things because they were offered voluntarily. My position on this is kind of complex and nuanced but this particular reply was meant to state the opposite.
This is not how authors rights work