this post was submitted on 29 Aug 2026
58 points (92.6% liked)

Lemmy Today

328 readers
1 users here now

If you experience issues or problems with this instance (lemmy.today), this is the place to discuss them. Or if you just want to ask questions about how something works. Anything related to the instance or lemmy itself.

founded 2 years ago
MODERATORS
 

About using bots on lemmy.today

There has been quite a few reports lately about people using bots and llms to create and upvote content. I wanted to make a post to discuss this.

In my opinion, bots are a tool to quickly create or repost content. And its how users use that tool that is important for whether or not it will make lemmy a better or worse place.

Two examples:

User A

  • User A frequently visits reddit and would like to quickly share content that he thinks is actually good and would be liked on lemmy.
  • He uses a bot or some other tool to do this quickly and conveniently.
  • Most of the posts are generally liked by the community, sometimes ending up on the frontpage, with many comments from many different users.
  • A few users are complaining about bot content but the majority seems to appreciate the posts.

User B

  • User B also uses a similar repost bot or a llm to generate content, but they decide that lemmy needs a lot of content and it doesnt matter so much weather or not its actually good.
  • The main idea is to repost or generate as much as possible and lets users decide what they like.
  • They create 30 posts in a day, where almost none of them are actually interesting for the lemmy community. Posts have no comments, and most of them are only upvoted by other bot users.
  • Many users are complaining that they now see tons of low quality posts in their feeds because of this user.

How will we moderate this

In my mind, the lemmy community is fine with user A but not user B. So it requires moderators to make a decision about the content someone is creating. Does it make lemmy better or worse for the majority?

Thats how Im thinking about this, and if you want to share your own ideas or comments, I would love reading them. If you agree, easiest way is to upvote. If you dont, I would love a comment even more so I can understand some other point of view. :)

Thanks!

you are viewing a single comment's thread
view the rest of the comments
[–] tal@lemmy.today 3 points 5 days ago* (last edited 5 days ago) (2 children)

Doesn't really answer the question, but I do have to say that I think that quite a number of posts and comments that are not generated by computers have been accused of such.

I remember one thread where some user posted an image. Some other user ran some "AI detector" on it, and it scored the image as having a 85% chance of being AI-generated. That second user became it was convinced that it was AI-generated and started a huge argument with User 1. I was pretty sure that the image probably was a real photograph, and pulled up several pieces of evidence supporting that, eventually digging up even an old TinEye result of the image predating the AI image generators. Didn't convince the suspicious user.

I've had people say that I must have used an LLM to write comments, as I write a lot of them (and often use a keyboard, and quote text). I'm the only person that I know the immediate source of comment text for on the Threadiverse, and so I know for a fact that I've never intentionally synthesized any comment text (outside of Google Translate and one or two comments specifically talking about what LLM text looks like and citing it as such). So I know that, in at least some cases, people aren't able to detect human-written text. I'm pretty sure that enforcement of any prohibition on LLM text would be a royal pain in the rear.

I think that in the context of images or video or audio being proof of something, it's important to know whether something is synthesized.

But I think that over time, it's going to become increasingly impractical to identify all content that has been synthesized. Like, okay. Go back before LLMs were a popular thing. I see all kinds of (human) quotes that are mis-attributed. Some sites, like Quote Investigator especially and Wikiquote do a pretty good job of doing the footwork and filtering out misattributions. Some sites do a really bad job. But point is, misattributed quotes are all over, get moved from site to site and into speeches and magazines and newspapers and so forth without people doing much validation of sources. "Winston Churchill said X." Like, they'll pull text from all kinda of sources.

Based on that, I'm pretty sure that the same thing is going to be true of LLM-generated text. Right now, yeah, you have these spam websites that are all generated by LLMs. But I'm pretty confident that we're going to increasingly have content on websites that aren't purely-aimed at spamming search engines that are going to incorporate LLM-generated text. It could be use of one to do machine translations---Google Translate's core is an LLM. That's been around for quite a while, and I don't think that most consider it to be particularly problematic. It could be people pulling responses from Web search engines, which often try to generate LLM-produced answers to searches. It could be people pulling text from other websites. But my bet is that there's just going to be an increasing amount of text that's passed through an LLM circulating on websites. And it's going to be really hard to avoid that when quoting text. People are simply not going to always try to go back to primary sources for every piece of text. Even if we could create a reliable LLM detector---and I'm confident that we could not do that, though we might be able to detect a particular model---it's going to wind up flagging text that people did not themselves generate with an LLM, because text on websites and in other forms of media is going to increasingly contain it.

Like, I just don't think that even if one wants to do so, it's going to be realistic to enforce an LLM prohibition.

I do think that it's more-practical to, as of 2026, identify many AI-generated images. I cannot do it myself with all images, and I've generated images that I cannot personally distinguish from a photograph, using the Flux diffusion model. But I am pretty sure that I could catch some. But I'm sure that it's going to get harder, and that people also mis-fire on that. There was a comic strip up the other day about how the comic's artist was exasperated about how many users kept accusing him of generating his comics with AI rather than doing them by hand.

I think that in the case of bots, it's reasonable to ask people to flag them. To some extent, it might be possible to maybe identify unflagged ones. I think that it's very likely that if someone wants to work at making a sophisticated human-impersonation bot, do something like post on a human schedule, stuff like that, maybe let a team of humans handle some responses (which is what I'd expect a commercial spam-bot to do) it's going to be hard to identify. Reddit had a number of bots that I was fine with, like doing imperial-metric conversions on units. I also really liked several bots in communities that dealt with certain products or fields that would identify and expand relevant acronyms (the !selfhosted@lemmy.world community here does this, and it looks like this), or link to wiki pages or similar.

The real concern I have with bots is organized influence campaigns. Trying to sell a product or affect opinions of a politician or country or whatever. I'm concerned about humans doing the same thing---bots just lower the costs and thus increase the viability.

I think that that has plenty of potential to be a problem, and I think that it is a real problem on some other social media platforms. I do not think that, as of 2026, that is a real thing (or a real thing at any scale, at least) on the Threadiverse. It just isn't worth the money. We have an active user population, MAU, of something like fifty thousand; it's not even a rounding error compared to far larger forms of social media. Facebook is over three billion. To some extent, someone could leverage work for another platform maybe, but I don't think that we provide enough of a return as things stand to warrant a lot of spammer time and effort. That may change down the line.

I think that most things that bots shouldn't do are also things that humans shouldn't do, and a lot of the prohibitions already are there for humans. If someone is spamming outside of some community that is explicitly okay with relevant products notifications, human or software, that should be dealt with the same way.

I do think that bot operators, even where they're beneficial and identify themselves as bots, like that acronym bot, should be asked to provide some way of getting in touch with a human operator, like in their user profile. Maybe some bots can maybe just route messages to them to a human, but others might rely on being able to process that message. An automated system can run haywire or something due to bugs, and it'd be preferable to have some solution other than just relying on instance or community bans and hoping that the operator notices.

I personally am not super worried about this at the moment, but I could imagine that maybe one day, in the future, we have so many (self-identifying) bots, some of which people want (like the acronym bot) and some of which they don't (I remember the Reddit haiku bot which identified unintentional comments that were haikus---I was okay with that, but I could see that being considered unnecessary noise) that it's hard to block just the ones you want without blocking all bots. I think that maybe, at that point, it might be necessary to also require bots to self-classify as things like novelty/entertainment or whatnot, to permit easier filtering. But I don't think that that's an immediate concern, as there are few self-identified bots active on the Threadiverse (at least in communities that I read). I think that if and when it becomes a problem, it'll be possible to have admins sit down and come up with some categories and for bot operators to retrofit existing bots.

I don't have any fundamental problem with bots that repost content from other sources, like RSS feeds or whatever (and in some cases, I think that they could be very useful). I do think that those should be expected to self-identify as bots, and it'd be up to a community moderator to make a call on one. Like, maybe a video games community likes having automated notifications of Steam sales show up, and another video games community doesn't want that. Unless they are producing load issues or chewing up instance resources at a high rate, which a bot could certainly do, I don't think that it really is an instance problem, but rather a community problem.

If a given user doesn't want to see content posted from a given bot, it's easy enough for them to block it unless we get some kind of massive explosion of bots. Maybe revisit it then.

I don't know if there's an existing popular "bot framework" library in Python or Rust or something for the Threadiverse, but I think that maybe if there is, it'd be nice to have some sane defaults, like having a built-in maximum rate limit unless the bot author goes out of their way to override it, to help avoid bugs from being problematic for other users. Not really an admin issue though. Something for devs to address.

[–] mrmanager@lemmy.today 3 points 4 days ago

Tal, you get the longest comment award for this one, congratulations. :)

But I think that over time, it’s going to become increasingly impractical to identify all content that has been synthesized.

For sure, 100%. Its only a matter of time before we cant tell anymore.

The real concern I have with bots is organized influence campaigns. Trying to sell a product or affect opinions of a politician or country or whatever. I’m concerned about humans doing the same thing—bots just lower the costs and thus increase the viability.

I think we have this problem on massive scale, in movie reviews as well as political compaigns. When nothing is actually real anymore, the trust will disappear. It already has happened for many people.

We have an active user population, MAU, of something like fifty thousand; it’s not even a rounding error compared to far larger forms of social media.

True but that is probably an advantage. Being small means there is not much interest in adding ads or other forms of monetization, and thats saving us from a lot of enshittification. Once money enters the picture, the product becomes worse, every time. Open AI recently added ads in Chat GPT by the way...

I don’t have any fundamental problem with bots that repost content from other sources, like RSS feeds or whatever (and in some cases, I think that they could be very useful). I do think that those should be expected to self-identify as bots, and it’d be up to a community moderator to make a call on one.

In this case, we are kind of in-between since reposting software is not like a bot and more like a convenience that saves time. As long as its high quality, I dont think its a problem. Actually I wonder how many of the memes on the popular All page today are generated by bots...

[–] PapaSkwat@lemmy.today -5 points 4 days ago

Great points. I agree with everything you said. I think the solution is pretty simple: rather than accusing someone of being a bot, using AI, or whatever else, people should just block users or communities they don't want to see.

To me, that solves most of the problem. If I thought you were a bot and everything you posted annoyed me, I'd just block you instead of deciding you should be banned from an instance or community or from the Fediverse.

The number of people on Lemmy who seem determined to decide what everyone else should or shouldn't see, often by calling for users or communities to be banned, is pretty shocking to me.

I use a keyboard to type too, so I feel you!