this post was submitted on 31 Jul 2026
734 points (98.9% liked)

Technology

88698 readers
3003 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] danc4498@lemmy.world 35 points 2 months ago (3 children)

Why would AI want a continuous deal with Reddit? Don’t they get all the data they need the first time? I doubt the new content is worth as much as the previous deal… maybe I don’t understand what these deals are for.

[–] UnderpantsWeevil@lemmy.world 18 points 2 months ago* (last edited 2 months ago)

Half the joke is that Reddit was ground zero for AI slop even before AI had gone mainstream.

The company got harvested back before the AI firms were overly worried with cross-contamination.

[–] XLE@piefed.social 5 points 2 months ago (2 children)

Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.

The returns on that value sharply diminish, of course, but I think they're still necessary. Which will leave everybody in a bind that is very funny.

[–] then_three_more@lemmy.world 4 points 2 months ago (2 children)

Why do they need a deal? Can't they just steal it like the rest of their training data?

[–] ryper@lemmy.ca 4 points 2 months ago (1 children)

The API restrictions and login requirements are meant to make scraping hard enough to make a deal worthwhile.

[–] then_three_more@lemmy.world 1 points 2 months ago

Ah that makes sense.

[–] XLE@piefed.social 2 points 2 months ago

Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit

[–] altkey@lemmy.dbzer0.com 3 points 2 months ago

Other than that, probably it's a licensing agreement that makes AI trainers keep paying if they still using that dataset.

[–] ctrl_alt_esc@lemmy.ml 1 points 2 months ago

Isn't it illegal for Reddit to sell user-created data without their consent?