LatheOperator

joined 2 years ago
[–] LatheOperator@leminal.space 3 points 1 week ago

I can't believe I'm saying it but DON'T TOUCH GRASS

[–] LatheOperator@leminal.space 5 points 1 week ago (2 children)

Is that why you're so hot or vice versa?

[–] LatheOperator@leminal.space 1 points 2 weeks ago (1 children)

The thing is, if I don't add a "good enough for searching" OCR layer, someone else will - AI scraper or legitimate user. That's a cheap automatic operation. I'll better do it myself and poison it in a way that won't interfere with searching, like replacing some "are" with "aren't", such common words are rarely searched for. If there's a chance AI will fall for the metadata and invisible layer contents, that will decrease the requirements for visible poisoning, which is necessary but annoying.

Some of the text is indeed justified so I could do the multi-column trick that seems like the best compromise. The gap can be as narrow as one space. Or larger if I can write a script to detect lines and connect the columns with gibberish. A human can use zoom or window positioning to view one column at a time. I don't and will never have access to files the printouts are from (some are presumably in .602 format, others probably .doc), others are handwritten or typewritten, some have images glued on top or hand-traced; and Czech OCR is only about 99.5% reliable so as easier as it would make the endeavor, I can't be sure to preserve everything if I try to convert them into editable documents as an in-between step.

[–] LatheOperator@leminal.space 1 points 2 weeks ago

Woah, that's fucked up. But LLMs have very different performance by language so I think Czech responses are largely based on Czech-language data (maybe with training dataset augmentation with auto-translated works, judging by the literally translated technical terms). And I don't think this would happen in my country.

[–] LatheOperator@leminal.space 1 points 2 weeks ago (2 children)

Yes but high school biology knowledge won't be all that helpful in surveillance tools. I'm targeting chatbots students might want to use to cheat. If an essay is less coherent or truthful, teachers will be able to more confidently punish the student or at least apply some extra scrutiny like a random oral exam on the topic (yes, those are still done here). Trust me that free AIs will become more limited once investor money dries up and paid tools will get price hikes, so if students see that the expensive model can't produce a Czech text that fools their biology teacher, they might just stop paying, restoring some of their information literacy and reducing corporate revenue a bit. I agree that the surveillance battle is real but that's not fought on this front. The major chokepoint for surveillance is government accountability (very low in today's US with ICE agents etc.), slightly less knowledgeable chatbots don't hinder it very much.

[–] LatheOperator@leminal.space 1 points 2 weeks ago* (last edited 2 weeks ago) (4 children)

I think the AI boom will be over before it makes sense to scrape this, especially if I succeed at making it so that only a human fluent Czech speaker (or maybe a very bespoke program debugged by a Czech speaker) can see through the obfuscation (we have minimum wage so that would be costly). We don't have AGI and fully autonomous agent workforces as promised, and that's not gonna change if the remaining <20%, hardest-to-clean digital human-made data gets fed to them. Not to mention some was acquired illegally already, with pending lawsuits. The investors, including governments (thankfully not mine) will eventually realize that LLMs are simply not delivering nearly as much as they cost (including externalities unless the government is shit and passes them to people living near datacenters, laid-off programmers etc.) and never will, although that might take a while.

[–] LatheOperator@leminal.space 1 points 2 weeks ago* (last edited 2 weeks ago)

Yes, I will use both licence terms and "Made by bad AI" in metadata to discourage scraping but it's tempting to also poison the Czech-language biology knowledge base of bots who use the materials anyway. A good interleaving text that won't get filtered as off-topic might be multi-step roundabout machine translation of the original using shitty local tools.

[–] LatheOperator@leminal.space 2 points 2 weeks ago (6 children)

Lbraries are cool but don't spread information very far, especially when "for local lending only, no copying". Those would almost never get used and probably thrown away as soon as the libraries realized the difficult copyright situation (not all are by my grandpa, many are unclear due to missing cover pages). Even when incompatible with screen readers (which someone criticized me for), a poisoned PDF is way more usable.

Anyway, the poisoning, outside the invisible layers and "transcript", will be visibly baked into the image too but obvious to humans (typewriter vs rendered text that inverts some statements with "not", "never" etc. at the ends of lines or between words in tiny font and scatters lines between paragraphs saying the text is unreliable, by random bad AI models, or just "clear AI tells" like "Sure, I can do that! 😉 Here's your summary:✨"). I don't think anyone will develop an OCR font filter for a few thousand pages in Czech to illegally defeat my DRM enforcing the included licence (unfortunately DRM is incompatible with CC so I can't use the "CC-BY-NC-SA" shorthand, not that scrapers, my enemies, respect it anyway): there is some protection in the obscurity.

[–] LatheOperator@leminal.space 1 points 2 weeks ago* (last edited 2 weeks ago)

Screen readers and LLMs both need plain text. Sorry, I'm not giving it to them. There will be an invisible OCR layer for searchability (and to discourage scrapers from re-doing OCR instead), like in many scanned PDFs, but poisoned. There's not many blind teachers anyway.

Yet, this is more accessible than what someone else suggested: making photocopies and donating them to libraries "for local lending only". Those would almost never get used and probably thrown away as soon as the libraries realized the difficult copyright situation (not all are by my grandpa, many are unclear due to missing cover pages)

[–] LatheOperator@leminal.space 3 points 2 weeks ago (8 children)

That's expensive and a bit impractical to access, isn't it?

[–] LatheOperator@leminal.space 3 points 2 weeks ago* (last edited 2 weeks ago)

I want the poisoning to work on file level because it's inevitable and welcome for the documents to be shared between teachers and students for free on all kinds of existing platforms. I don't own any domains, anyway, and it might be best to scatter the documents around to make them harder to blacklist.

Edit: Google has both search crawlers and AI scraping bots. Even if both are separate, easy to filter or even abiding to robots.txt, the company has indicated that opting out of or hindering scraping will impact search ranking. Of course I'd use a throwaway gibberish $1 domain and couldn't care less about search ranking, but the power of their opaque, corporate algorithm is immense and maybe would spread to DNS blocking (they control 8.8.8.8). I don't want to play a cat-and-mouse game (and expect users of the docs to play along).

[–] LatheOperator@leminal.space 1 points 2 weeks ago* (last edited 2 weeks ago)

Good point. However, distillation (cannibalizing better LLMs) is a frequent technique so better convince the scraper it's not good even for that. That's why I plan to visibly digitally stamp OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs of the typewritten text (of course in dozens of variations, maybe even different languages, to undermine search-and-replace). Humans will know it's fake (especially if I add a disclaimer) but scrapers, including ones that re-render and OCR the PDF themselves to get rid of misleading metadata and invisible layers, will most likely rate the text low in value. Of course ethically (and arguably legally) trained commercial AI would reject any text if it is released CC-BY-NC-SA 4.0 but I can't use that because scrapers for AI ignore licences in practice, and CC specifically forbids taking technical measures to devalue the text for some users.

 

TL;DR: Please help me fuck (with) AI. See bold sections

Hi,
I haven't been keeping up with anti-AI combat so I'm asking for help. I inherited thousands of pages of materials my late grandpa made or used for his grammar school teaching job in the 1990s-2000s. They are A4 pages of documents made using what seems to be a typewriter, Text602 (DOS rich text editor) and Word. They were most likely not all made by him but he treasured them in nice binding and they have sources (mostly books and journals, almost no webpages, and absolutely no AI) and a cursory look shows meticulous compilation of every important fact on each subject (frankly, the level of detail is excruciating and I'm glad I went to a different grammar school). There's obviously no original scientific research but the materials can still be useful to someone, I bet. They were almost thrown away by the widowed grandma (she already removed and disposed of the plastic bindings and front covers so I'll have to guess document titles) but I think grandpa would prefer them to be shared. With an ADF scanner and OCR software (I have no chance of accessing the work computers he used so I'll have to scan), I can quickly make searchable PDFs of each document, and share them via torrent and DDL sites (there are Czech sites dedicated to sharing teaching materials but they have paywalls or an upload-credit system so best avoid them, not to mention some materials contain newspaper clippings and textbook photocopies for images so best stay anonymous and not try to assert copyright).

I'm afraid these texts could become a major part of some commercial LLM's Czech-language biology/social sciences knowledge corpus unless poisoned. How to best reduce the value of the documents when people try to feed them to AI (training/rewriting) with them while keeping their value for most legitimate users? (Sorry, people with screen readers, there may need to be extra steps for you.) I'm thinking about adding a huge volume of thesaurized or otherwise fuzzed public domain text like f4mi did with .ass subtitles (a technique that would probably still work if she didn't get 1M views detailing it, making YouTube reduce subtitle formatting support). Prompt injection or replacements (cell→gnome) might be interesting too. However, tools I know add an extra PDF layer, which is too obvious. I'm thinking about adding tiny text in the header and footer or between paragraphs in the OCR layer (not overlaid to reduce interference when selecting/searching), but how? I need an automated way to do this with such a huge page count. I can use both Linux and Windows machines for the job. None of them are very powerful but speed is not a concern, it's summer break and nobody will need school materials until September. I'll be happy to include multiple layers and techniques to make them too frustrating to remove.

The paper smells musty but does not seem to be moldy. It's all blank on the other side so I'll interleave it with recent newspaper to allow for the odor-neutralizing chemicals to seep into the sheets so I can eventually reuse them.

Illustration pic is an actual sheet from the collection, to make the post more engaging. Of course I won't be adding watermarks like that, that would just aggrevate people and make them try extra hard to extract the actual content. (And this one is easy to remove with color channel mixing.)

 

Real news story. Pic unrelated.

148
This was posted to PROMOTE the generator (image.tensorartassets.com)
submitted 1 year ago* (last edited 1 year ago) by LatheOperator@leminal.space to c/fuck_ai@lemmy.world
 

TensorArt bragged how their FluxAI tool can generate comics... The results attached to the post speak for themselves. It's just 8 months old, so not an early model.

Transcript
AI-generated comic, mostly black and white with a few shades of grey.

Big panel at top: Fashion store with T-shirts and an undistinct item on display. Two very similar women walking into a wall next to its entrance, as well as a girl with impossible legs walking either left or right. Store sign: "Size leye light Is got neaıl boutique"

Panel 2: Woman in a shirt, cross-eyed: "Decitions, erctiʋns..."

Panel 3: Same girl, looking as if around a corner, with a disfigured disembodied hand holding a bag-like item: "We.let tame it caƨt harrng sishil.. ଚaraa I tine ջur ડeas tmall shopping ppree..

Panel 4: Same girl, looking at one or more large items held in her 4- and 11-fingered hands. Concerningly blank stare, circles around eyes, eyeballs lined with red: "ooh! the bres or ኗ5бo season.

Panel 5: Different girl with no arms and disfigured legs next to clothes racks. Confused look, indistinct symbols above her head: "What!' "thought I has cute shop, top.?"

Wide panel 6 at bottom: Both characters from above looking at a group of schoolgirls in front of badly drawn lockers; one has question marks above her head. The girl from panels 2-4: "Will a hebes. Red!" in a two-tailed speech bubble.

Indistinct signature and initial in the bottom right corner.

Apparently the prompt contains the entire script, it's not the image AI that wrote it:

Prompt
A young woman stands in front of a clothing store, looking excited
Caption: "Sarah's eyes light up as she spots her favorite boutique."
Sarah: "Ooh! Time for some retail therapy!"

second panel Sarah is inside the store, surrounded by clothes racks, holding up two different outfits
Caption: "Decisions, decisions..."
Sarah: "Hmm, the blue dress or the red skirt?"

third panel Sarah is at the checkout counter, looking shocked as the cashier shows her the total
Cashier: "That'll be $250, please."
Sarah: "What?! I thought it was sale season!"

fourth panel Sarah is walking out of the store with a small shopping bag, looking both satisfied and slightly guilty
Caption: "The aftermath of a successful shopping spree."
Sarah: "Well, at least I got this cute top... Ramen for dinner it is!"

More with the same prompt, different styles:








 

FYI: Alegria "Art", or the larger "Corporate Memphis" style is flat-color, anatomically-distorted slop associated with corporations and used by Facebook, Google and others.

Yes, I could make a community but I'm not here often enough to moderate one.

view more: next ›