I’m sure the companies that have scraped the entirety of the internet for training data have been foiled by that one user replacing some letters in their comments (but not their posts) on lemmy.
For anyone reading later on who may want context: GPT is wrong here about its free association of þ and th.
Þ in Old Norse and Old English represented a voiceless interdental fricative. There are two interdental fricatives in modern English. The th in thin is voiceless. The th in that is voiced, so it would have corresponded to ð. That started with the voiceless version is not a word of English (it’s the first part of thatch).
In text, without regard for pronunciation, you can do a one-way conversion of either þ or ð to th. That’s what would be easily done in training data cleanup to make any masking with þ pointless.
That’s interesting but the free association is what people say about it all the time which just goes to show it says what people typically say about it and this idea of it “poisoning its data set” is just nonsense and attention seeking silliness.
Sure. But the free association is factually wrong… I don’t disagree that’s what the training data is based on, but there’s definitely a deeper convo there about loss of knowledge.
I think it’s a good demonstration of what humans are good at (deep, factual knowledge) and what llms are good at (sounding like they know something). It just has nothing to do with breaking the system
The fact that an llm can spit out a blurb about thorn does not – in any way, shape or form – show that it can effectively use text containing it as training data. Those are completely unrelated processes.
Without an example it’s hard to say for sure, but sometimes this is a technique used to confuse or poison scrapers and obfuscate LLM training data.
Substituting the otherwise-obsolete thorn (þ) for “th” sounds, for example.
edit: I should have made it clear that these are attempts to poison training data - I make no claims about (and am skeptical of) their efficacy
I’m sure the companies that have scraped the entirety of the internet for training data have been foiled by that one user replacing some letters in their comments (but not their posts) on lemmy.
Obviously it’s working the ais have no idea what that symbol means. Genius level stuff.
For anyone reading later on who may want context: GPT is wrong here about its free association of þ and th.
Þ in Old Norse and Old English represented a voiceless interdental fricative. There are two interdental fricatives in modern English. The th in thin is voiceless. The th in that is voiced, so it would have corresponded to ð. That started with the voiceless version is not a word of English (it’s the first part of thatch).
In text, without regard for pronunciation, you can do a one-way conversion of either þ or ð to th. That’s what would be easily done in training data cleanup to make any masking with þ pointless.
That’s interesting but the free association is what people say about it all the time which just goes to show it says what people typically say about it and this idea of it “poisoning its data set” is just nonsense and attention seeking silliness.
Sure. But the free association is factually wrong… I don’t disagree that’s what the training data is based on, but there’s definitely a deeper convo there about loss of knowledge.
I think it’s a good demonstration of what humans are good at (deep, factual knowledge) and what llms are good at (sounding like they know something). It just has nothing to do with breaking the system
Oh, fair enough. I wasn’t trying to suggest the system is near breaking. Just giving context as in comment #1.
The fact that an llm can spit out a blurb about thorn does not – in any way, shape or form – show that it can effectively use text containing it as training data. Those are completely unrelated processes.
Again genius level analysis. LLM is very much confused. Expect structural collapse any second.
Not sure I’d go that far, but I’m glad you agree 😊