Google Researchers’ Attack Prompts ChatGPT to Reveal Its Training Data(www.404media.co)

posted 1 year ago

ChatGPT is full of sensitive private information and spits out verbatim text from CNN, Goodreads, WordPress blogs, fandom wikis, Terms of Service agreements, Stack Overflow source code, Wikipedia pages, news blogs, random internet comments, and much more.

Sort:

Hot Top Controversial New Old

You are viewing a single thread.

View all comments View context

[ - ]

tabarnaski@sh.itjust.works

4 points

1 year ago

You remember some dialogue from your favorite movie. Does this mean your neurons store copyrighted work?

permalink

report

parent

[ - ]

Fermion@mander.xyz

9 points

1 year ago

Shhh. Disney’s lawyers might get ideas.

permalink

report

parent

[ - ]

Excrubulent@slrpnk.net

5 points

1 year ago

Yes.

Just because they’re in a neural network and not ASCII or unicode doesn’t mean they’re not stored. It’s even more apt a concept since apparently those works can be retrieved fairly easily, even if the references to them are hard to isolate. It seems ChatGPT is storing eidetic copies of data, which would imply what other people have said in this thread, that it is overfitting itself to the data and not learning truly generalisable language.

permalink

report

parent

[ - ]

MxM111@kbin.social

2 points

1 year ago

The claim is that it contains entire copies of the book. It does not. AI memory is like our memory, we do not remember books word to word.

permalink

report

parent

[ - ]

Excrubulent@slrpnk.net

6 points

1 year ago

They are spitting out, as in the quote above, “verbatim text”, as in, word for word. That is copyrightable.

And that’s not what you said. You said it has no memory. That’s clearly wrong.

permalink

report

parent

[ - ]

anlumo@lemmy.world

1 point

1 year ago

It’s only under copyright if it’s a significant portion of the work. Single sentences are not enough, unless it’s a short poem.

permalink

report

parent

Show more comments

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

Community stats

17K
Monthly active users
12K
Posts
556K
Comments

Our Rules

Approved Bots

Community stats

Community moderators