this post was submitted on 29 Aug 2026
58 points (92.6% liked)

Lemmy Today

328 readers
1 users here now

If you experience issues or problems with this instance (lemmy.today), this is the place to discuss them. Or if you just want to ask questions about how something works. Anything related to the instance or lemmy itself.

founded 2 years ago
MODERATORS
 

About using bots on lemmy.today

There has been quite a few reports lately about people using bots and llms to create and upvote content. I wanted to make a post to discuss this.

In my opinion, bots are a tool to quickly create or repost content. And its how users use that tool that is important for whether or not it will make lemmy a better or worse place.

Two examples:

User A

  • User A frequently visits reddit and would like to quickly share content that he thinks is actually good and would be liked on lemmy.
  • He uses a bot or some other tool to do this quickly and conveniently.
  • Most of the posts are generally liked by the community, sometimes ending up on the frontpage, with many comments from many different users.
  • A few users are complaining about bot content but the majority seems to appreciate the posts.

User B

  • User B also uses a similar repost bot or a llm to generate content, but they decide that lemmy needs a lot of content and it doesnt matter so much weather or not its actually good.
  • The main idea is to repost or generate as much as possible and lets users decide what they like.
  • They create 30 posts in a day, where almost none of them are actually interesting for the lemmy community. Posts have no comments, and most of them are only upvoted by other bot users.
  • Many users are complaining that they now see tons of low quality posts in their feeds because of this user.

How will we moderate this

In my mind, the lemmy community is fine with user A but not user B. So it requires moderators to make a decision about the content someone is creating. Does it make lemmy better or worse for the majority?

Thats how Im thinking about this, and if you want to share your own ideas or comments, I would love reading them. If you agree, easiest way is to upvote. If you dont, I would love a comment even more so I can understand some other point of view. :)

Thanks!

top 50 comments
sorted by: hot top controversial new old
[–] NewDark@lemmy.today 41 points 5 days ago (2 children)

Probably a good general rule that bots should be marked as a bot, either by name or the user setting. That's the only extra bit I can think of and I fully agree.

[–] nocturne@slrpnk.net 29 points 5 days ago (1 children)

Bots should also have their owner/maintainer's info in their bio.

load more comments (1 replies)
[–] blaggle42@lemmy.today 4 points 4 days ago

I'd like to piggy back - everything should be marked. AI marked AI, bot marked bot. Then the individual user can filter or weight as they choose. For this to work there shouldn't be intrinsic negative repercussions for marking bot, but there should be if it is not marked bot, when it is.

Basically users should be able to say, "that probably was bot", "that probably was AI" and then, if the AI/Bot is not marked, but the comment-tags signify it was there should be a down.

[–] Wren@lemmy.today 19 points 5 days ago* (last edited 5 days ago) (10 children)

I think all bots should be marked as such. The User B type bots are basically the only things I block on here, but I would like to know if I'm responding to someone who will engage or just an unmonitored bot posting whatever.

[–] mrmanager@lemmy.today 2 points 4 days ago (2 children)

I would like this concept also, but what do you call someone who uses an automated tool to post content and who reads the comments on his posts?

I mean, if someone makes a browser plugin that can one-click share content from other platforms (maybe it exists, I dont know), should it be banned? I think what we all want is quality content, no matter if its from an automated tool or manually created.

[–] Wren@lemmy.today 2 points 4 days ago (1 children)

It depends on the level of automation, and I want to preface this by saying I don't believe bot accounts are bad.

Someone who selects and schedules posts for later or has a one-click-posting tool shouldn't be considered a bot account. However, there's a community for Propublica where an account automatically posts all new articles (or at least it did a while ago) and that should be labelled. The difference is whether a human selected the content or not.

I appreciate a bot that posts everything from a specific news site, but I'd still like to see it tagged.

Maybe that would help ease the reports on accounts that post a high volume of content but are independently curated, if people know that admins are looking into and labeling accounts appropriately.

[–] mrmanager@lemmy.today 3 points 4 days ago (1 children)

Yeah I totally agree with you, I think the problem is mainly when someone reposts all new articles very quickly and doesnt screen them for quality or even read them. Thats when it becomes super annoying for users to have to see all the crap. Its the typical User B scenario above.

The problem with marking someone as a bot account is that they cant post their content anymore. Many communities have rules against all bots by default it seems. And if we put someone using a tool in that category, it will kill most of their posts.

[–] Wren@lemmy.today 3 points 4 days ago

In the case of the Propublica account, it seemed to have one purpose and lived in a community specifically for it. Since Propublica publishes sporadic investigative journalism, it worked well for that purpose without looking spammy. In that case it makes sense for the user to have another account for engagement, or rather that's my solution. I would still want to see the bot account labelled.

Otherwise, individual communities should be run however they see fit, even if it excludes bots or if their understanding of bots is incorrect. Anyone who really wants to post there can either follow their rules or DM the mods about changing their guidelines. In any case, everyone is free to make their own community here.

load more comments (1 replies)
load more comments (9 replies)
[–] arotrios@lemmy.world 6 points 4 days ago

Bot to repost human created content? Ok as long as it's not spam.

Bot to create and present AI content as human? That's just an invitation to fill your community with crap.

[–] CalypsoGirl@lemmy.today 3 points 4 days ago (1 children)
[–] mrmanager@lemmy.today 2 points 4 days ago

Thank you - fixed :)

[–] echo@lemmy.today 12 points 5 days ago (24 children)

Fuck both of them. If I wanted to read reddit then I'd be over there. If one can't manually create a post for something they found interesting them just don't. Ideally they would say what they found interesting, too.

load more comments (24 replies)
[–] SirEDCaLot@lemmy.today 12 points 5 days ago (3 children)

My problem isn't necessarily bots, it's automated or rapid manual posting of AI generated content. AI generated content is very rarely interesting or worth my time.
So yes, in your example I have little problem with Bot A but I want to get rid of Bot B.

I also think bot accounts should be required to be flagged, and bots that frequently accumulate negative scores on posts/comments should be removed.

[–] mrmanager@lemmy.today 3 points 5 days ago

Yes, I also think their content is the most important thing. If it's a net positive or a net negative for the community, that really matters.

load more comments (2 replies)
[–] sanitation@lemmy.today 1 points 3 days ago* (last edited 3 days ago) (1 children)

This is the numbers I want you guys to understand 90–9–1 rule:

Roughly 90% lurk, 9% comment, and 1% create.

That's why I wanted to post so much. If I don't post, most of these communities simply don't have enough people creating content to replace it. You need content before you can get lurkers to engage, commenters to participate, and eventually more people to start posting themselves.

I'm not posting for the sake of flooding Lemmy. I'm trying to help solve the exact problem Lemmy has: not enough content creators.

If you're counting the public moderation history as the fourth mechanism, call it quadruple moderation / quadruple jeopardy:

Quadruple moderation really doesn't help. Every active content creator on Lemmy has to deal with four layers:

  1. Community moderators can remove your posts or ban you from the community.
  2. Your home-instance admins can ban your entire account.
  3. Receiving-instance admins can block/ban you for everyone on their instance, even though you don't belong to it.
  4. Public moderation history means those actions follow you. One moderator's decision can influence the next and snowball into additional bans.

For prolific content creators, that's brutal. The more you post, the more chances you have for one of these independent moderators/admins to object to something. You effectively need all of them to remain aligned indefinitely.

It also creates an obvious attack surface: anyone who wants to disrupt Lemmy - trolls, bad actors, or even people deliberately trying to drive creators away - can exploit reports and visible moderation history to amplify a single moderation decision into others.

Reddit doesn't have this same federated quadruple-jeopardy problem. You're primarily dealing with the subreddit you're contributing to. Lemmy desperately needs content creators while simultaneously giving them more independent ways to get shut out the more they contribute.

I genuinely don't believe everyone complaining about prolific posters simply fails to understand this. At some point, I have to question whether some of this is deliberate.

If I were Reddit and wanted to undermine a growing federated competitor, this is exactly where I'd attack it: its content creators. Lemmy already has a tiny creator base, and its quadruple-moderation system makes those creators incredibly easy to target through reports, bans, and public moderation histories.

I suspect Reddit or people acting in Reddit's interests are exploiting this. You don't need to hack Lemmy or take servers offline. Just convince Lemmy moderators that its most prolific contributors are "spam," get them banned, and let the network starve itself of content.

Whether Reddit is actually doing that is something we'd need evidence to prove - but Lemmy's structure makes that kind of subversion remarkably easy.

[–] mrmanager@lemmy.today 2 points 3 days ago

I knew about those stats already, and I don't think the volume of your posts is a problem, at least not on this instance. But yes, moderators on other instances may act the way you describe if they feel you are a bot spammer.

But do you think it's the volume of your posts that are the problem or that some of them are low quality? I keep coming back to this in my mind. Let's say if you posted 30 posts per day and they all ended up on the frontpage. Would anyone be annoyed by that?

About the federated moderation, it's kind of necessary. I had someone from another instance start posting illegal content here and of course I needed to be able to block that user. But yes, there are some moderators on other instances that are too heavy handed in my opinion. I believe we should stay in the background until it's really needed.

I think you should continue posting like you did before. Many of your posts were liked and ended up on the frontpage. But yes, I know you want to post even more than you do. And I agree, Lemmy could use a lot more content! But it needs to be good content. It all boils down to that in the end. As long as it's actually good, it's great for everyone.

[–] Magnum@infosec.pub 6 points 5 days ago (1 children)

So this is only for lemmy and only today?

[–] mrmanager@lemmy.today 5 points 5 days ago (1 children)

It's for old.lemmy.today and all the other optional user interfaces as well. They are part of the same instance. :)

[–] Magnum@infosec.pub 5 points 5 days ago* (last edited 5 days ago)

The good ol' Lemmy getting some rules for today

[–] tachikoma@lemmy.today 3 points 5 days ago

Uhhh, without thinking too much about the topic...

Type A bots... eh. I guess I don't mind if it's meaningful engagement. But I also loath reddit now-a-days. I'm my perfect world; Everyone would stop interacting with reddit entirely and it would slowly die. Practically?... I still want reddit gone hehehe.

Type B bots? No thanks.

I can't really tell you if it will make lemmy better or worse. Evil is a point of view after all.

[–] goferking0@lemmy.sdf.org 3 points 5 days ago (4 children)

What about downvoting bots?

load more comments (4 replies)
[–] tal@lemmy.today 3 points 5 days ago* (last edited 5 days ago) (2 children)

Doesn't really answer the question, but I do have to say that I think that quite a number of posts and comments that are not generated by computers have been accused of such.

I remember one thread where some user posted an image. Some other user ran some "AI detector" on it, and it scored the image as having a 85% chance of being AI-generated. That second user became it was convinced that it was AI-generated and started a huge argument with User 1. I was pretty sure that the image probably was a real photograph, and pulled up several pieces of evidence supporting that, eventually digging up even an old TinEye result of the image predating the AI image generators. Didn't convince the suspicious user.

I've had people say that I must have used an LLM to write comments, as I write a lot of them (and often use a keyboard, and quote text). I'm the only person that I know the immediate source of comment text for on the Threadiverse, and so I know for a fact that I've never intentionally synthesized any comment text (outside of Google Translate and one or two comments specifically talking about what LLM text looks like and citing it as such). So I know that, in at least some cases, people aren't able to detect human-written text. I'm pretty sure that enforcement of any prohibition on LLM text would be a royal pain in the rear.

I think that in the context of images or video or audio being proof of something, it's important to know whether something is synthesized.

But I think that over time, it's going to become increasingly impractical to identify all content that has been synthesized. Like, okay. Go back before LLMs were a popular thing. I see all kinds of (human) quotes that are mis-attributed. Some sites, like Quote Investigator especially and Wikiquote do a pretty good job of doing the footwork and filtering out misattributions. Some sites do a really bad job. But point is, misattributed quotes are all over, get moved from site to site and into speeches and magazines and newspapers and so forth without people doing much validation of sources. "Winston Churchill said X." Like, they'll pull text from all kinda of sources.

Based on that, I'm pretty sure that the same thing is going to be true of LLM-generated text. Right now, yeah, you have these spam websites that are all generated by LLMs. But I'm pretty confident that we're going to increasingly have content on websites that aren't purely-aimed at spamming search engines that are going to incorporate LLM-generated text. It could be use of one to do machine translations---Google Translate's core is an LLM. That's been around for quite a while, and I don't think that most consider it to be particularly problematic. It could be people pulling responses from Web search engines, which often try to generate LLM-produced answers to searches. It could be people pulling text from other websites. But my bet is that there's just going to be an increasing amount of text that's passed through an LLM circulating on websites. And it's going to be really hard to avoid that when quoting text. People are simply not going to always try to go back to primary sources for every piece of text. Even if we could create a reliable LLM detector---and I'm confident that we could not do that, though we might be able to detect a particular model---it's going to wind up flagging text that people did not themselves generate with an LLM, because text on websites and in other forms of media is going to increasingly contain it.

Like, I just don't think that even if one wants to do so, it's going to be realistic to enforce an LLM prohibition.

I do think that it's more-practical to, as of 2026, identify many AI-generated images. I cannot do it myself with all images, and I've generated images that I cannot personally distinguish from a photograph, using the Flux diffusion model. But I am pretty sure that I could catch some. But I'm sure that it's going to get harder, and that people also mis-fire on that. There was a comic strip up the other day about how the comic's artist was exasperated about how many users kept accusing him of generating his comics with AI rather than doing them by hand.

I think that in the case of bots, it's reasonable to ask people to flag them. To some extent, it might be possible to maybe identify unflagged ones. I think that it's very likely that if someone wants to work at making a sophisticated human-impersonation bot, do something like post on a human schedule, stuff like that, maybe let a team of humans handle some responses (which is what I'd expect a commercial spam-bot to do) it's going to be hard to identify. Reddit had a number of bots that I was fine with, like doing imperial-metric conversions on units. I also really liked several bots in communities that dealt with certain products or fields that would identify and expand relevant acronyms (the !selfhosted@lemmy.world community here does this, and it looks like this), or link to wiki pages or similar.

The real concern I have with bots is organized influence campaigns. Trying to sell a product or affect opinions of a politician or country or whatever. I'm concerned about humans doing the same thing---bots just lower the costs and thus increase the viability.

I think that that has plenty of potential to be a problem, and I think that it is a real problem on some other social media platforms. I do not think that, as of 2026, that is a real thing (or a real thing at any scale, at least) on the Threadiverse. It just isn't worth the money. We have an active user population, MAU, of something like fifty thousand; it's not even a rounding error compared to far larger forms of social media. Facebook is over three billion. To some extent, someone could leverage work for another platform maybe, but I don't think that we provide enough of a return as things stand to warrant a lot of spammer time and effort. That may change down the line.

I think that most things that bots shouldn't do are also things that humans shouldn't do, and a lot of the prohibitions already are there for humans. If someone is spamming outside of some community that is explicitly okay with relevant products notifications, human or software, that should be dealt with the same way.

I do think that bot operators, even where they're beneficial and identify themselves as bots, like that acronym bot, should be asked to provide some way of getting in touch with a human operator, like in their user profile. Maybe some bots can maybe just route messages to them to a human, but others might rely on being able to process that message. An automated system can run haywire or something due to bugs, and it'd be preferable to have some solution other than just relying on instance or community bans and hoping that the operator notices.

I personally am not super worried about this at the moment, but I could imagine that maybe one day, in the future, we have so many (self-identifying) bots, some of which people want (like the acronym bot) and some of which they don't (I remember the Reddit haiku bot which identified unintentional comments that were haikus---I was okay with that, but I could see that being considered unnecessary noise) that it's hard to block just the ones you want without blocking all bots. I think that maybe, at that point, it might be necessary to also require bots to self-classify as things like novelty/entertainment or whatnot, to permit easier filtering. But I don't think that that's an immediate concern, as there are few self-identified bots active on the Threadiverse (at least in communities that I read). I think that if and when it becomes a problem, it'll be possible to have admins sit down and come up with some categories and for bot operators to retrofit existing bots.

I don't have any fundamental problem with bots that repost content from other sources, like RSS feeds or whatever (and in some cases, I think that they could be very useful). I do think that those should be expected to self-identify as bots, and it'd be up to a community moderator to make a call on one. Like, maybe a video games community likes having automated notifications of Steam sales show up, and another video games community doesn't want that. Unless they are producing load issues or chewing up instance resources at a high rate, which a bot could certainly do, I don't think that it really is an instance problem, but rather a community problem.

If a given user doesn't want to see content posted from a given bot, it's easy enough for them to block it unless we get some kind of massive explosion of bots. Maybe revisit it then.

I don't know if there's an existing popular "bot framework" library in Python or Rust or something for the Threadiverse, but I think that maybe if there is, it'd be nice to have some sane defaults, like having a built-in maximum rate limit unless the bot author goes out of their way to override it, to help avoid bugs from being problematic for other users. Not really an admin issue though. Something for devs to address.

[–] mrmanager@lemmy.today 3 points 4 days ago

Tal, you get the longest comment award for this one, congratulations. :)

But I think that over time, it’s going to become increasingly impractical to identify all content that has been synthesized.

For sure, 100%. Its only a matter of time before we cant tell anymore.

The real concern I have with bots is organized influence campaigns. Trying to sell a product or affect opinions of a politician or country or whatever. I’m concerned about humans doing the same thing—bots just lower the costs and thus increase the viability.

I think we have this problem on massive scale, in movie reviews as well as political compaigns. When nothing is actually real anymore, the trust will disappear. It already has happened for many people.

We have an active user population, MAU, of something like fifty thousand; it’s not even a rounding error compared to far larger forms of social media.

True but that is probably an advantage. Being small means there is not much interest in adding ads or other forms of monetization, and thats saving us from a lot of enshittification. Once money enters the picture, the product becomes worse, every time. Open AI recently added ads in Chat GPT by the way...

I don’t have any fundamental problem with bots that repost content from other sources, like RSS feeds or whatever (and in some cases, I think that they could be very useful). I do think that those should be expected to self-identify as bots, and it’d be up to a community moderator to make a call on one.

In this case, we are kind of in-between since reposting software is not like a bot and more like a convenience that saves time. As long as its high quality, I dont think its a problem. Actually I wonder how many of the memes on the popular All page today are generated by bots...

load more comments (1 replies)
load more comments
view more: next ›