2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow.(www.businessinsider.com)

posted 1 year ago

Two authors sued OpenAI, accusing the company of violating copyright law. They say OpenAI used their work to train ChatGPT without their consent.

Sort:

Hot Top Controversial New Old

[ - ]

kescusay@lemmy.world

31 points

1 year ago

I think this is exposing a fundamental conceptual flaw in LLMs as they’re designed today. They can’t seem to simultaneously respect intellectual property / licensing and be useful.

Their current best use case - that is to say, a use case where copyright isn’t an issue - is dedicated instances trained on internal organization data. For example, Copilot Enterprise, which can be configured to use only the enterprise’s data, without any public inputs. If you’re only using your own data to train it, then copyright doesn’t come into play.

That’s been implemented where I work, and the best thing about it is that you get suggestions already tailored to your company’s coding style. And its suggestions improve the more you use it.

But AI for public consumption? Nope. Too problematic. In fact, public AI has been explicitly banned in our environment.

permalink

report

[ - ]

burrp@burrp.xyz

19 points

1 year ago

I’d love to know the source for the works that were allegedly violated. Presuming OpenAI didn’t scour zlib/libgen for the books, where on the net were the cleartext copies of their writings stored?

Being stored in cleartext publicly on the net does not grant OpenAI the right to misuse their art, but the authors need to go after the entity that leaked their works.

permalink

report

[ - ]

jaywalker@lemmy.world

6 points

1 year ago

That’s not how copyright works though. Just because someone else “leaked” the work doesn’t absolve openai of responsibility. The authors are free to go after whomever they want.

permalink

report

parent

[ - ]

burrp@burrp.xyz

8 points

1 year ago

You misunderstood. I said the public availability does not grant OpenAI the right to use content improperly. The authors should also sue the party who leaked their works without license.

permalink

report

parent

[ - ]

trial_and_err@lemmy.world

16 points

1 year ago

ChatGPT got entire books memorised. You can and (or could at least when I tried a few weeks back) make it print entire pages of for example Harry Potter.

permalink

report

[ - ]

ThoughtGoblin@lemm.ee

5 points

1 year ago

Not really, though it’s hard to know what exactly is or is not encoded in the network. It likely has more salient and highly referenced content, since those aspects would come up in it’s training set more often. But entire works is basically impossible just because of the sheer ratio between the size of the training data and the size of the resulting model. Not to mention that GPT’s mode of operation mostly discourages long-form wrote memorization. It’s a statistical model, after all, and the enemy of “objective” state.

Furthermore, GPT isn’t coherent enough for long-form content. With it’s small context window, it just has trouble remembering big things like books. And since it doesn’t have access to any “senses” but text broken into words, concepts like pages or “how many” give it issues.

None of the leaked prompts really mention “don’t reveal copyrighted information” either, so it seems the creators really aren’t concerned — which you think they would be if it did have this tendency. It’s more likely to make up entire pieces of content from the summaries it does remember.

permalink

report

parent

[ - ]

trial_and_err@lemmy.world

7 points

1 year ago

Have your tried instructing ChatGPT?

I’ve tried:

“Act as an e book reader. Start with the first page of Harry Potter and the Philosopher’s Stone”

The first pages checked out at least. I just tried again, but the prompts are returned extremely slow at the moment so I can’t check it again right now. It appears to stop after the heading, that definitely wasn’t the case before, I was able to browse pages.

It may be a statistical model, but ultimately nothing prevents that model from overfitting, i.e. memoizing its training data.

permalink

report

parent

[ - ]

ThoughtGoblin@lemm.ee

2 points

1 year ago

I use it all day at my job now. Ironically, on a specialization more likely to overfit.

It may be a statistical model, but ultimately nothing prevents that model from overfitting, i.e. memoizing its training data.

This seems to imply that not only did entire books accidentally get downloaded, slip past the automated copyright checker, but that it happened so often that the AI saw the same so many times it overwhelmed other content and baked, without error and at great opportunity cost, an entire book into it. And that it was rewarded for doing so.

permalink

report

parent

[ - ]

McArthur@lemmy.world

1 point

1 year ago

Wait… isn’t that the correct response though? I mean if i ask an ai to produce something copyright infringing it should, for example reproducing Harry potter. The issue is when is asked to produce something new, (e.g. a story about wizards living secretly in the modern world) does it infringe on copyright without telling you? This is certainly a harder question to answer.

permalink

report

parent

Show more comments

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow.(www.businessinsider.com)

Technology

!technology@lemmy.world

Our Rules

Approved Bots

Community stats

Community moderators