If they could then it wouldn't be de-identified data...
The researcher could share a string from one of their conversations and OpenAI can confirm whether it exists in their training data.
Or OpenAI could just look at their code and say what it does (maybe have their AI do it if they're having so much trouble with this?)
Well you could search for elements similar to the proof / problem in the training data, even if it's de-identified, right? OpenAI can probably do better than Ctrl-f "Navier Stokes".