logoalt Hacker News

AceyManyesterday at 8:20 PM4 repliesview on HN

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training".

I would expect academic institutions to require equivalent contractual terms.


Replies

buzeryesterday at 9:14 PM

Some of the recent statements have caused at least me to look those claims in a bit more nuanced light. In particular what does OpenAI consider to be "your data"? I would assume input (prompt) to be it at least. However it becomes more murky when you consider other aspects. Is output "your data"? Is the chain of thought that you are not even allowed to see? Can they use these and possibly even inputs to generate synthetic data that is then used?

All of these would seem to be "your data", but when they are carefully only including certain aspects (like prompts) in their statements it starts to sound they want to hide something.

show 1 reply
tecleandortoday at 12:25 AM

Thing is... if OpenAI cannot even confidently say if some data was used for training or not, as their models and weights and stuff are mostly black boxes, how could you enforce or demonstrate in court that case?

About researchers, lots of them are probably using personal plans that aren't even reimbursed by their institutions. I could ask Cordova's research institution (I MAY) but I wouldn't be surprised at all if that was the case.

rainprincessyesterday at 9:01 PM

Sure but they could also rewrite your data to create synthetic reconstructions and many academics, sign up for their own accounts.

For example, at school they can have an agreement with Gemini, but the student / academic could have bought an individual pro subscription to any other model provider.

stefan_yesterday at 9:21 PM

The open internet is now a cesspit, with very little new good data. Expect everyone to train on user data always. They just got clever about whitening it.