I know there are other ways of accomplishing that, but this might be a convenient way of doing it. I’m wondering though if Reddit is still reverting these changes?

You are viewing a single thread.
View all comments View context
1 point
*

I’m saying reddit will not ship a trashed deliverable. Guaranteed.

Reddit will have already preprocessed for this type of data damage. This is basic data engineering and trivial to do to find events in the data and understanding timeseries of events.

Google will be receiving data that is uncorrupted, because they’ll get data properly versioned to before the damaging event.

If a high edit event happens on March 7th, they’ll ship march 7th - 1d. Guaranteed.

Edit to be clear: you’re ignoring/not accepting the practice of noting high volume of edits per user as an event, and using that timestamped event as a signal of data validity.

permalink
report
parent
reply
-1 points
*

I’m saying reddit will not ship a trashed deliverable. Guaranteed.

Nobody said anything about the database being trashed. What I’m saying is that the database is expected to have data unfit for LLM training, that Google will need to sort out, and Reddit won’t do it for Google.

Reddit will have already preprocessed for this type of data damage.

Do you know it, or are you assuming it?

If you know it, source it.

If you’re assuming, stop wasting my time with shit that you make up and your “huuuuh?” babble.

permalink
report
parent
reply
1 point
*

I know it because I’ve worked in corporate data engineering and large data migrations and it would be abnormal to do anything else. there’s a full review of test data, a scope of work, an acceptance period, etc.

You think reddit doesn’t know about these utilities? You think Google doesn’t?

You need to chill out and acknowledge how an industry works. I’m sure you are convinced but your idea of things isn’t how the industry works.

I don’t need to explain to you that the sky is blue. And I shouldn’t need to explain to you that Google isn’t going to accept a damaged product, and that reddit can or can’t do some basic querying and timeseries manipulations.

Edit like you literally asked for a textbook.

permalink
report
parent
reply
0 points
*

I know it because I’ve worked in corporate data migrations

In other words: “I dun have sauce, I’m assooming, but chruuuust me lol”

At this rate it’s safe to simply ignore your comments as noise. I’m not wasting further time with you.

permalink
report
parent
reply

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


Community stats

  • 18K

    Monthly active users

  • 12K

    Posts

  • 542K

    Comments