Suppose I gave you a series of 1000 coin flips. I tell you they were generated by a fair coin. You’re suspicious, you think I hid a watermark in them. But you look at the sequence and 485 are heads, close enough to 50%. You look at the correlation between a coin flip and the next and it’s 0.0433. Close to zero. You do a bunch more stats and everything checks out. So you’re convinced.
Then I tell you: calculate the average of every 3rd flip minus the average of every second. It should come out close to zero, and unlikely to be higher than +/- 30 but for my sequence it comes out to +89. Ok, that’s very unlikely.
So from now on we share this secret. Whenever I generate sequences of coin flips they all came out with this particular statistic out of whack, but unless you’re looking for it, you can’t detect it. In fact, in the real world I use cryptography such that unless you know my secret key, the specific statistic in question is mathematically undetectable.
Now use this sequence of coin flips to pick amongst next tokens in an LLM. It doesn’t change the distribution of words used in the LLM. It doesn’t change its writing style. In fact, unless you know the secret key you cannot detect the watermark. So it’s not that some words are used more often. That would be detectable. It’s just that if you translate back to a series of 0’s and 1’s the average of every 3rd bit minus the average of every 2nd is out of whack.
But that signal wouldnt carry through if i used your coin for 30% of flips and the results aren’t contiguous right? And for that specific case even 50% and contiguous might not be detected.
It just seems like once you try and detect a signal against something that wasnt one shot it breaks down fast. Unless the signal is constrained to a very small space (pairs of words) but then quality and stealthiness must suffer massively.
Im curious about situations like this.
Paper submitted with some headings, title, and 3 paragraphs. 1 and 3 mostly generated. generated ones have a small percentage of sentences rewritten or deleted. A few find and replaces on terms like load bearing and provenance and glue words around them. The teacher runs the entire thing through a checker. What happens?
Or you have 100 paragraphs and 15 are generated, whole things scanned, what happens?
I guess the implementation and edits matter. And the “resolution “ (ie every sentence-ish bears the mark vs every paragraph). But it seems likely the signal would be lost pretty badly. And i think this is the typical sort of way people actually use AI for important things you might want to validate against.
Absolutely you can hide a watermark in a big chunk of text. But what happens when it’s inside a larger body work that gets checked or split up even a little?
Edit: im reading about synthid and see a splicing doesn’t hurt it much.
I don't have the math for it, but I wonder how much text/words/tokens are needed before it become easy to embed an ID for authorship as well. There's got to be some number of bytes where it becomes both reliable and hard to find.
"I'm sure glad the LLM was able to fix up my grammar and imagery in the anonymous treatise I made criticizing the authoritarian regime. Hold up, someone's knocking at my door..."
Then I tell you: calculate the average of every 3rd flip minus the average of every second. It should come out close to zero, and unlikely to be higher than +/- 30 but for my sequence it comes out to +89. Ok, that’s very unlikely.
So from now on we share this secret. Whenever I generate sequences of coin flips they all came out with this particular statistic out of whack, but unless you’re looking for it, you can’t detect it. In fact, in the real world I use cryptography such that unless you know my secret key, the specific statistic in question is mathematically undetectable.
Now use this sequence of coin flips to pick amongst next tokens in an LLM. It doesn’t change the distribution of words used in the LLM. It doesn’t change its writing style. In fact, unless you know the secret key you cannot detect the watermark. So it’s not that some words are used more often. That would be detectable. It’s just that if you translate back to a series of 0’s and 1’s the average of every 3rd bit minus the average of every 2nd is out of whack.