logoalt Hacker News

tristanjyesterday at 10:28 PM5 repliesview on HN

The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.


Replies

Pulcinellayesterday at 10:47 PM

Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.

show 1 reply
whimsicalismyesterday at 11:10 PM

I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).

I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.

I also don't really believe that whether or not this model was trained on these conversations is unknowable information.

show 1 reply
DetroitThrowyesterday at 10:41 PM

>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.

Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.

Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.

show 1 reply
hellohello2yesterday at 11:46 PM

Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.

A model being trained on lots of irrelevant information does not mean relevant information was not used.

fwipyesterday at 11:02 PM

The IP laundering machine strikes again.