In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.
We have all become Milhouse now.
"multiple providers disclose sensitive conversation-derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation."
Not good at all.
No, they don't "leak data", data is sold. Leaking data requires a mistake. This is intentional.
This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).
This is why I built my own chat interface. https://ai.ivx.run/
Access through: https://ai.ivx.run/chat/
A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you. They don't care about the agreements you've signed. You're just a pile of cash to them
I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.
First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.
You thought... they didn't?
All sells but I saw Chinese models are upfront about that most of the time.
The word is "sell" not "leak". This title takes away all agency from the thieves selling private data to advertisers.
I am very surprised.
Is it still called a leak if it was the whole point and purpose of the deal?
Someone should have to investigate, but I suppose it's all "legal"?
The paper doesn't say when the app sends the conversion artifact.
my first guess is always Gboard.
Insert surprised pikachu meme
Hanlon's Razor given that these tools are almost certainly vibe coded at this point?
data is sold i think
Entire conversations via exposed permalinks. For Grok: trackers receiving the conversation URL could access the full chat because the link lacked access controls.
Screenshots of conversations. TikTok received screenshots of Grok chats during sharing, exposing the actual visible conversation content.
Conversation-derived content tied to persistent identifiers, including prompts and automatically generated chat titles revealing sensitive facts. "Salary 85k NYC: mortgage 280–350k".
[dead]
Tangential:
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.