this post was submitted on 14 Aug 2026
233 points (97.9% liked)
Technology
87184 readers
3713 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
But it doesn't work. It looks like only the owner of the text generator is able to check if some text is written with this concrete generator (with some probability). I see no use of this technique.
It works for them not scraping their own slop back into training data. I assume that is actually the real purpose of the system. They don't want to share the key with the public. But they probably will with other llm companies in exchange for theirs.
Hmm, yes. I just thought about outside usage, somehow haven't thought about it as an inner LLM maintenance tool.
The article seemed to claim that some models allow outsiders to submit text for detection if I read it right. That seems like a decent way of doing it. If you "open source" the raw data, it means people can do things to try and get around it - same reason most websites don't reveal their anti-spam techniques.
That's still not very useful unless the test is for all AI models otherwise you have to manually submit the text to every AI companies detection system.
I see an opportunity for an AI test aggregator company that checks all of them for you at once.
Vibe-coding one seems more appropriate to me ;-)
Yeah, but you'll have to submit to every known provider and hope the user used one of the ones that provide this checking service, and didn't do something like ask a local model to just randomize synonyms in a text.
That would only work for their own slop though. Anthropic cannot recognize Google's watermark, only theirs.
I assumed the goal might be so they could check whether other models have been trained on their output. Like anthropic using that as a "proof" when they start whining again about Chinese "distillation attacks".
Of course, since they're the only ones able to check their watermark, it would be rather shit as evidence anyway. "We've run the numbers, and we know you can't, but trust us, this chatbot is totally copying ours!"
It's practically not of any use to end-users. It's a tool made by the AI company for themselves, to be able to claim a specific text was generated from their model.
Probability becomes a non-problem the longer the scanned text is.