logoalt Hacker News

tristanjyesterday at 10:46 PM1 replyview on HN

Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.

It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.

We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.


Replies

biophysboyyesterday at 11:58 PM

I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math