this post was submitted on 28 Sep 2026
26 points (84.2% liked)

Selfhosted

62755 readers
584 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

Detailed Rules Post

  1. Be civil.

  2. No spam.

  3. Posts are to be related to self-hosting.

  4. Don't duplicate the full text of your blog or readme if you're providing a link.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.

  8. AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

I've tried giving screenshots of phishing emails to a local Qwen instance and so far it always correctly detected it as scam, even points out the exact elements that it based its judgement on. Sending screenshots to it ad-hoc isn't too scalable for family and friends. I'd like to be able to either forward emails for screening, or perhaps have it screen everything from a mailbox.

Has anyone done anything like this? Is there anything self-hostable that does this?

all 16 comments
sorted by: hot top controversial new old
[–] KairuByte@lemmy.dbzer0.com 3 points 4 days ago

The ability to do it aside, if feel it should be brought up that handing off things like this can lead to giving anything that gets through an “unofficial seal of approval” so to speak.

“This email looks phishy, but AI says it’s safe so I’ll ignore my initial concerns” is going to inevitably happen.

There’s also the fact that using a model to detect scams, means you can use that same model to train a different model on how to build scams that bypass the detection.

This also feels a little like a hammer looking for a nail. There have been tools to handle this without the insane computational requirements of an LLM for many years.

[–] Shadow@lemmy.ca 30 points 1 week ago (1 children)

I'd be pretty concerned about prompt injection risks with feeding a LLM unsanitized data. You definitely need a good harness around it..

[–] avidamoeba@lemmy.ca 10 points 1 week ago (2 children)

Good point. It'll have to have no access to the internet or anything local outside of its container. Just text in, text out.

[–] Dran_Arcana@lemmy.world 10 points 1 week ago (1 children)

You could (and probably should) use a system-one style inference system for spam classification. Much cheaper and the structured output means it's impossible to go rogue and curl some malware or whatever. It can absolutely misclassify but its output is programmatically structured and just ranks a pre-selected set of output tokens.

In your case that's

Spam

Not_spam

[–] tigerhawkvok@startrek.website 1 points 3 days ago (1 children)

I only read about that today, so I hadn't even considered that angle. They're doing some cool stuff in that space. jeff is the self-hostable one that does best as far as I know.

[–] Dran_Arcana@lemmy.world 2 points 2 days ago

You actually have a lot of options in this space. You can take the prefill engine of just about any LLM and turn it's transformer into a classifier by lobotomizing out the decoder. You can also use a diffusion model on single pass to surprisingly competent result.

If Jeff looks easy enough to deploy by all means start there, but don't discount the idea if Jeff sucks; you have a lot of options.

[–] Zikeji@programming.dev 8 points 1 week ago

Yeah if it's just a basic input with a function call for spam or not spam the risk is low. What's the worst case outcome, it tricks it into saying no it isn't a scam and you have to delete it manually? Hahaha

[–] BartyDeCanter@piefed.social 16 points 1 week ago

You definitely don’t have to go full LLM for this. There are a wide variety of local spam filtering tools available, depending on what email application you use. If you really want to use modern ML, I’d go with something like Laya as it is tunable and runs locally.

[–] TheMightyCat@ani.social 14 points 1 week ago

I would recommend using a classifier model instead of a text generation one.

As for the way to do it a simple program that connects to you mailbox via imap and calls the vllm or whatever api you are using and then moves it to spam could work?

[–] hendrik@palaver.p3x.de 10 points 1 week ago* (last edited 1 week ago)

LLM plugins should be available in standard spamfilters like Rspamd. Some mail servers like Stalwart have it as well. You pick a prompt and provide it with an OpenAI-compatible endpoint and it'll ask the AI for every mail. Can be a local model.

Not sure if it's in the Email clients as well. I found a few Thunderbird addons, but that was just a quick Google search.

[–] MagicShel@lemmy.zip 9 points 1 week ago

If your email service offers an MCP connection (Google does), you can have your AI connect and examine emails and manipulate them. You can hook that up to qwen using maybe lmstudio. I haven't tried it self-hosted, but it sounds pretty straightforward.

[–] Talaraine@fedia.io 6 points 1 week ago (1 children)

Just be aware that those phishers are now also scanning their phishing emails to see how they can get better too! The wheels on the bus go round and round! Round and Round! Round and Round!

[–] avidamoeba@lemmy.ca 5 points 1 week ago

Kinda what prompted me to think about this. I started seeing some too good phishing emails recently.

[–] cybervseas@lemmy.world 5 points 1 week ago (1 children)

Probably better to ask this in an ai community. However fwiw there are a few skills for generic imap checking. I've set up a local openclaw gateway and model with an automated task to check my email in the morning and basically prod me into replying. The same concept should be applicable to spam detection too.

[–] bretton.dev@coves.social 1 points 1 week ago

OP if you don't mind running this through a Hermes / OpenClaw harness this is probably the easiest way with pre built email connectors. Also if you don't mind spending a few bucks a month DeepSeek v4 will blow out anything you run locally for pennies (or gpt luna as of last wed.)

It can also set up its own custom connector if needed and run as needed via CRON.