I would like an unambiguously clear statement from OpenAI as to what they do with data collected from non-business accounts when:
(a) The data controls setting to train on the data is unchecked.
(b) The privacy controls opt-out has been submitted.
(c) Both.
Both things can be true:
1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.
2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.
The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
[1]: https://openai.com/index/chatgpt-for-academic-researchers/
[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m
"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."
This is the third day of total hysteria that is based on nothing of substance. Move on folks.
It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
Most people here are missing the forest for the trees.
We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights.
We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line.
This is double plus ungood.
When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1]https://mathoverflow.net/a/513885
I think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.
This is the second wake up call.
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
I am not using any of these AI tools. I thought it was basically given that any single thing you write in these systems also used by the companies. On the other hand now I understand why there are so much projects about hosting AI systems locally.
Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.
My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
I wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery
If they didn't care about the artists, why would they care about academia?
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".
They don't even claim to have had a proof, only to have been working on it.
The oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
Most scientific breakthroughs are simply a continuation of previous work.
I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.
Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.
I think it seems sensible to _assume_ anything the LLM reads (if you aren't inferencing it) has a chance of ending up in some database somewhere. Regardless of whether you trust the other party its a sensible thing to plan around.
Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).
How was I not aware of this before?
How to steal ideas with AI.
step 1, identify high value users by net worth, citation count, or number of followers
step 2, select all prompts by high value users
step 3, invest 10 billion thinking tokens in modeling an objective for each user
step 4, build an RL environment for each user
step 5, rollout 10 billion tokens per environment
step 6, train on resulting traces
I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
ClosedAI has every incentive to scoop academics to juice their valuation. Their public statements are worthless, only the incentive 'alignment' matters and theirs will never be on the side of the user.
Recently, one of the services (not OpenAI) has felt need to mention that something I've tried to discuss with it would 'be a new way to approach the problem, no other company ...'
But it wasn't asked about the novel side of it, twice in last two weeks I've seen that, not seen it before ...
OpenAI trying their best to put the Navier-Stokes episode behind them by making the GPUs go brrrr. NYT:
> In its Wednesday night statement, OpenAI said: “In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”
https://x.com/markchen90/status/2097400166554993041?s=20
that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
If you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources.
It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.
This article explains the controversy and the mathematical problem much better than the tweet and toots: https://www.science.org/content/article/how-ai-math-breakthr...
This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
Thing build from stolen data continues stealing data - for some reason, I am not surprised. ;-)
Never ever trust OpenAI, they are evil.
It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value.
And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.
That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.
Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.
Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.
Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!