logoalt Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

575 pointsby pred_today at 6:49 AM578 commentsview on HN

https://mathstodon.xyz/@andreasthom/117240536885387540

https://mathstodon.xyz/@andreasthom/117240537520615623

https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/209779183626997799...

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...


Comments

SwellJoetoday at 4:25 PM

It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value.

And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.

That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.

Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.

show 1 reply
foogazitoday at 2:10 PM

Even when you pay you are the product

MetaverseClubtoday at 6:35 PM

Never ever trust OpenAI, they are evil.

AyanamiKainetoday at 8:50 AM

I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.

There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.

Why would you need to train a model on certain specific near prove chat if you just query it?

Besides that, its hard to believe that its the case for every "company stole my prove".

throwaway63467today at 7:34 PM

Isn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.

maxglutetoday at 2:00 PM

300 billion tokens is like.. $5-25 million giving range of OpenAI ouput prices, I"m sure they pay less at cost so, I wonder if more $$$ in wage hours have been spend by humans on the problem. My feeling is yes?

gps372today at 9:08 AM

If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.

square_usualtoday at 1:56 PM

I think this is stupid, for three reasons:

1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.

2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.

3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)

show 2 replies
gentleraintoday at 1:42 PM

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?

How do people become that trusting?

The phrasing itself is guilt tripping

show 6 replies
BatchJobtoday at 8:57 PM

I have a better question? Why would you trust OpenAI or any AI company, at all? Or you crazy?

pred_today at 6:49 AM

See https://openai.com/index/ten-advances-in-mathematics/ for the announcement this refers to.

Davidzhengtoday at 4:45 PM

Tbh it won't really matter soon.

semiquavertoday at 11:11 AM

In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.

int32_64today at 2:17 PM

Doesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?

show 1 reply
spindump8930today at 1:37 PM

Reminder that there are degrees of "trained on conversations". From John Schulman:

> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper

> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this

> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"

source: https://x.com/johnschulman2/status/2097440545853637108

show 2 replies
jijjitoday at 10:47 PM

The oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.

blactuarytoday at 10:49 PM

I wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery

galkktoday at 5:26 AM

I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.

I would like to see chat logs etc and understand how much of a progress was done by human.

foogazitoday at 2:13 PM

What’s the limit ?

Will Microsoft Word publish your novel on Amazon behind your back ?

Will VS Code setup a website with your app idea ?

segmondytoday at 5:02 PM

Question: Can you trust the cloud?

No.

Henchman21today at 9:25 PM

Why is anyone expecting decency from people who have already proven to have none?

gnfargbltoday at 10:41 AM

In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.

The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.

In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.

bambaxtoday at 7:23 AM

All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?

The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.

That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.

show 2 replies
xbartoday at 1:59 PM

How can OpenAI figure out how to be trustworthy?

overfeedtoday at 8:24 AM

I can't wait for OpenAI to do this to companies firing people to free up AI budgets

simianwordstoday at 8:07 PM

> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

> The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

https://archive.is/lWzkk

avereveardtoday at 6:05 PM

Eh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.

uoaeitoday at 7:38 PM

I'm confused by a lot of this discourse...

What have they done to show they can be trusted?

sdcfgytoday at 8:04 AM

Theft machines be thieving.

keedatoday at 4:03 PM

It would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced.

I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.

oergiRtoday at 10:13 AM

One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.

The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.

show 1 reply
_DeadFred_today at 6:47 PM

Forget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.

737mintoday at 9:28 PM

Imagine what happens when you use a Chinese model. Seriously, just think about how much more control and visibility you have w US companies compared to CCP-controlled ones.

show 1 reply
dyauspitrtoday at 6:44 PM

Astra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.

mrbluecoattoday at 1:37 PM

"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"

Welcome to the party, with the rest of humanity.

ur-whaletoday at 6:35 PM

Its the "with unpublished math" that I have a problem with.

lf88today at 4:47 PM

short answer seems to be "no"

qg127today at 1:52 PM

There are so many naive academics. They still believe an "opt-out" button.

Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.

Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.

hn1rig3raktoday at 1:18 PM

the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.

show 1 reply
bakugotoday at 9:17 AM

Interesting that this is already off the front page after just 4 hours.

buellerbuellertoday at 3:19 PM

Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.

You will not be able to opt out unless you completely isolate yourself from society, tough shit.

wslhtoday at 2:45 PM

Worth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.

protocolturetoday at 10:10 PM

>Trust

No you cant do that lmao.

esafaktoday at 2:05 PM

What happens if you use a different harness?? Does opting out online suffice?

stego-techtoday at 5:19 PM

I hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences.

Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop.

Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again.

If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.

nisegamitoday at 11:35 AM

One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?

Grimblewaldtoday at 5:24 AM

people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.

show 1 reply
vrganjtoday at 8:43 AM

OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.

If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?

They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.

I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.

This is American AI companies committing suicide.

🔗 View 23 more comments