this post was submitted on 11 Aug 2026
29 points (96.8% liked)

Technology

7189 readers
234 users here now

News community around technology, social media platforms, information technology and governmental policy surrounding it.

What doesn't fit here?

The core of the story has to be technology focused.


Post guidelines

Title formatPost title should mirror the news source title. If you don't like the title of article, look for an alternative source instead of editorializing it.
URL formatPost URL should be the original link to the article (even if paywalled) and archived copies left in the body. It allows avoiding duplicate posts when cross-posting.
[Opinion] prefixOpinion (op-ed) articles must use [Opinion] prefix before the title. Opinion articles refer to articles that their publisher doesn't explictly endorse.
Country prefixCountry prefix can be added to the title with a separator (|, :, etc.) if the news is from a local publisher who doesn't clearly mention the country.


Rules

1. English onlyTitle and associated content has to be in English.
2. Use original linkPost URL should be the original link to the article (even if paywalled) and archived copies left in the body. It allows avoiding duplicate posts when cross-posting.
3. Respectful communicationAll communication has to be respectful of differing opinions, viewpoints, and experiences.
4. InclusivityEveryone is welcome here regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, education, socio-economic status, nationality, personal appearance, race, caste, color, religion, or sexual identity and orientation.
5. Ad hominem attacksAny kind of personal attacks are expressly forbidden. If you can't argue your position without attacking a person's character, you already lost the argument.
6. Off-topic tangentsStay on topic. Keep it relevant.
7. Instance rules may applyIf something is not covered by community rules, but are against lemmy.zip instance rules, they will be enforced.


Companion communities

!globalnews@lemmy.zip
!interestingshare@lemmy.zip


Icon attribution | Banner attribution


If someone is interested in moderating this community, message @brikox@lemmy.zip.

founded 2 years ago
MODERATORS
top 16 comments
sorted by: hot top controversial new old
[–] mrmaplebar@fedia.io 7 points 2 weeks ago (3 children)

Not enough. AI generated content should be required by law to contain both visible and invisible watermarks.

[–] floofloof@lemmy.ca 6 points 2 weeks ago (1 children)

You can't enforce that with text though, because people will just delete the words. They use a subtle statistical signal because people can't just spot it and delete it. Unfortunately it raises a problem: either they have to make the signal a secret, in which case only Anthropic and those contracted to keep the secret could detect it, or it's not, in which case people can use a detector and some other AI to discover how to eliminate the signal. I expect Anthropic would try to keep control over the detection secrets and provide a detection API for third-party software to use.

[–] Ek-Hou-Van-Braai@piefed.social 2 points 2 weeks ago (1 children)

Unless you're using AI to write books, I can't imagine how you'd implement some secret watermark

[–] floofloof@lemmy.ca 3 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

As I understand it, you're right that you need a good length of output to be able to detect the watermark, because only then can you see the statistical effect with confidence. And in code there are usually various options for how you get something done, and they could watermark generated code by adding a distinctive pattern to its preferences for certain constructs over others. But again, it would have to be subtle, so you'd need a large enough sample of its output before you could see the effect.

[–] Ek-Hou-Van-Braai@piefed.social 2 points 2 weeks ago

It should be fairly easy to stop this detention even in code.

You just need to set very strict "rules" the AI must follow when writing code.

Variable and method names etc. Must follow a specific formula.

Have a different AI write your docs etc.

If you're just vibe coding everything, the watermark would work. So maybe a good thing. But it'll be whack a mole, and if you're careful I'm sure you can stop the detection

[–] TrickDacy@lemmy.world 1 points 2 weeks ago

I'm so silly. I saw this and thought, oh cool a teeny tiny surprising win. Guess I should only be happy with the destruction of all LLMs.

[–] schnurrito@discuss.tchncs.de 1 points 2 weeks ago

You know, the evil bit used to be a joke, now people unironically suggest equivalent things...

[–] SnailMagnitude@mander.xyz 7 points 2 weeks ago

This seems a riot, trying to sneak in invisible copyright really does seem like the copyright system is well and truly fucked.

It arrived to try and help with the deluge of slop from Guttenberg nonsense but the whole idea, and copyleft, seems to be imploding at the moment.

[–] godsammitdam@lemmy.zip 6 points 2 weeks ago

So then they can lie and say real pictures have the watermark?

So they can lie and say fake images don't have the watermark?

So the AI looking for the watermark can hallucinate?

Yeah, I'm sure that's super reliable.

[–] undefinedTruth@lemmy.zip 5 points 2 weeks ago (1 children)

Can someone explain to me how is it possible to watermark text? Because it makes no sense to me.

[–] schnurrito@discuss.tchncs.de 3 points 2 weeks ago

Nor to me - but it doesn't make a lot of sense with images either.

Text is literally just an array of numbers (Unicode character points); where do you put "invisible" watermarks there? Images are, likewise, a three-dimensional array of pixel color values; while it's possible to very slightly vary those color values and call that "watermarking" (similar concept: steganography), how does that not get lost even unintentionally on the first lossy compression step?!

[–] MattR@feddit.org 1 points 1 week ago

There are multiple major flaws with watermarks for texts:

  1. Their method only works for longer texts as it relies on statistical effects that aren't clear enough to detect in a very short text.
  2. They are resting their approach on an incorrect assumption: their models prefer certain words and if they appear more frequently than in a random text, they assume it's generate by their model. But who says texts are random? Out of billions of people there will be many who have a similar preference for some words and their texts will always be wrongfully accused of AI-generated just because they happen to have a similar word choice preference. The shorter the texts the likelier that issue becomes.
  3. There will soon be tools that replace a random amount of words with synonyms automatically and therefore remove any chances of detecting the watermark.
[–] Double_A@discuss.tchncs.de -2 points 2 weeks ago (3 children)

And who will care about those? The slop consumers will still consume...

[–] BakedCatboy@lemmy.ml 8 points 2 weeks ago

Personally I would love a browser extension that detects these and adds a badge to AI content.

[–] Solumbran@lemmy.world 3 points 2 weeks ago

People who want to know if it is slop or not quickly?

[–] historicaldocuments@lemmy.world 1 points 2 weeks ago

If it's there and has been there, then there's always a chance it can be used later: