this post was submitted on 28 Sep 2026
26 points (84.2% liked)

Selfhosted

62672 readers
277 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

Detailed Rules Post

  1. Be civil.

  2. No spam.

  3. Posts are to be related to self-hosting.

  4. Don't duplicate the full text of your blog or readme if you're providing a link.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.

  8. AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

I've tried giving screenshots of phishing emails to a local Qwen instance and so far it always correctly detected it as scam, even points out the exact elements that it based its judgement on. Sending screenshots to it ad-hoc isn't too scalable for family and friends. I'd like to be able to either forward emails for screening, or perhaps have it screen everything from a mailbox.

Has anyone done anything like this? Is there anything self-hostable that does this?

you are viewing a single comment's thread
view the rest of the comments
[–] avidamoeba@lemmy.ca 10 points 5 days ago (2 children)

Good point. It'll have to have no access to the internet or anything local outside of its container. Just text in, text out.

[–] Dran_Arcana@lemmy.world 10 points 5 days ago (1 children)

You could (and probably should) use a system-one style inference system for spam classification. Much cheaper and the structured output means it's impossible to go rogue and curl some malware or whatever. It can absolutely misclassify but its output is programmatically structured and just ranks a pre-selected set of output tokens.

In your case that's

Spam

Not_spam

[–] tigerhawkvok@startrek.website 1 points 1 day ago (1 children)

I only read about that today, so I hadn't even considered that angle. They're doing some cool stuff in that space. jeff is the self-hostable one that does best as far as I know.

[–] Dran_Arcana@lemmy.world 1 points 4 hours ago

You actually have a lot of options in this space. You can take the prefill engine of just about any LLM and turn it's transformer into a classifier by lobotomizing out the decoder. You can also use a diffusion model on single pass to surprisingly competent result.

If Jeff looks easy enough to deploy by all means start there, but don't discount the idea if Jeff sucks; you have a lot of options.

[–] Zikeji@programming.dev 8 points 5 days ago

Yeah if it's just a basic input with a function call for spam or not spam the risk is low. What's the worst case outcome, it tricks it into saying no it isn't a scam and you have to delete it manually? Hahaha