Could Reddit's data be "poisoned" to prevent its use in training AI? : reddit

this post was submitted on 26 Feb 2024

166 points (93.7% liked)

17683 readers

88 users here now

News and Discussions about Reddit

Welcome to !reddit. This is a community for all news and discussions about Reddit.

The rules for posting and commenting, besides the rules defined here for lemmy.world, are as follows:

Rules

Rule 1- No brigading.

**You may not encourage brigading any communities or subreddits in any way. **

YSKs are about self-improvement on how to do things.

Rule 2- No illegal or NSFW or gore content.

**No illegal or NSFW or gore content. **

Rule 3- Do not seek mental, medical and professional help here.

Do not seek mental, medical and professional help here. Breaking this rule will not get you or your post removed, but it will put you at risk, and possibly in danger.

Rule 4- No self promotion or upvote-farming of any kind.

That's it.

Rule 5- No baiting or sealioning or promoting an agenda.

Posts and comments which, instead of being of an innocuous nature, are specifically intended (based on reports and in the opinion of our crack moderation team) to bait users into ideological wars on charged political topics will be removed and the authors warned - or banned - depending on severity.

Rule 6- Regarding META posts.

Provided it is about the community itself, you may post non-Reddit posts using the [META] tag on your post title.

Rule 7- You can't harass or disturb other members.

If you vocally harass or discriminate against any individual member, you will be removed.

Likewise, if you are a member, sympathiser or a resemblant of a movement that is known to largely hate, mock, discriminate against, and/or want to take lives of a group of people, and you were provably vocal about your hate, then you will be banned on sight.

Rule 8- All comments should try to stay relevant to their parent content.

Rule 9- Reposts from other platforms are not allowed.

Let everyone have their own content.

:::spoiler Rule 10- Majority of bots aren't allowed to participate here.

founded 1 year ago

MODERATORS

ja2@lemmy.world

_MoveSwiftly@lemmy.world

Thekingoflorda@lemmy.world

Rooki@lemmy.world

L3s@lemmy.world

166

Could Reddit's data be "poisoned" to prevent its use in training AI? (lemmy.world)

submitted 9 months ago* (last edited 9 months ago) by nodsocket@lemmy.world to c/reddit@lemmy.world

60 comments fedilink hide all child comments

In case you didn't know, you can't train an AI on content generated by another AI because it causes distortion that reduces the quality of the output. It is also very difficult to filter out AI text from human text in a database. This phenomenon is known as AI collapse.

So if you were to start using AI to generate comments and posts on Reddit, their database would be less useful for training AI and therefore the company wouldn't be able to sell it for that purpose.

you are viewing a single comment's thread
view the rest of the comments

[–] FaceDeer@kbin.social 35 points 9 months ago (2 children)

In case you didn’t know, you can’t train an AI on content generated by another AI because it causes distortion that reduces the quality of the output.

This is incorrect in the general case. You can run into problems if you do it incorrectly or in a naive manner. But this is stuff that the professionals have figured out months or years ago already. A lot of the better AIs these days are trained on "synthetic data", which is data that's been generated by other AIs.

I've seen a lot of people fall for wishful thinking on this subject. They don't like AI for whatever reason, they hear some news article that says something that sounds like "AI won't work because of problem X", and so they grab hold of that. "Model collapse" is one of those things, it's not really a problem that serious researchers consider insurmountable.

If you don't want Reddit to use your posts to train AI then don't post on Reddit. If you already did post on Reddit, it's too late, you already gave them your content. Bear this in mind next time you join a social media site, I guess.

[–] Windex007@lemmy.world 8 points 9 months ago (1 children)

Biased models are still absolutely a massive concern to serious researchers.

"AI collapse" isn't the only mechanism to throw a monkey wrench into someone's AI ambitions.

Intentionally introducing and reinforcing biases in an automated fashion adds an additional burden to those developing a model. I haven't actually looked into the economic asymmetry of those attacks, though.

[–] JeeBaiChow@lemmy.world 7 points 9 months ago

Absolutely this. Ai isn't some bastion of truth. I envision a future where AIS trained by different stakeholders, e.g. Dem vs repub, us vs Russia vs china. Etc... All fighting for eyeballs. It's just gonna get harder to tell what's real from fake because of the insane amount of content these bots are gonna churn out. It's already a huge problem with human monitored sources.

[–] Natanael 3 points 9 months ago (1 children)

Training on synthetic data is not a quality improvement, it's just an edge case reducer for a small set of edge cases by decreasing "overfitting", and it is only even able to achieve that if you're very very careful with what you add and how. If you're ONLY training on AI generated data repeatedly then it does start to degrade and loose coherence after a few generations of training

[–] FaceDeer@kbin.social 2 points 9 months ago (1 children)

Which is why nobody trains on ONLY AI generated data.

Really, experts have thought of this stuff already. Because they're experts. Synthetic data means that the amount of "real" data required is much less, so giant repositories like Reddit aren't so important.

[–] Natanael 1 points 9 months ago

No, "much less" training data isn't possible with synthetic data. That's not what it's there for. The experts would tell you as much if you asked them.