I'm seeing this on a current academic research project as well- it's like throwing fuel on the fire for all the best and worst parts of working with researchers.
First: it's definitely become harder to get people to integrate their work and use standard tools. We're exploring creating AI skills etc for our tooling, simply because the actual audience we need to convince isn't in the room. If the LLM doesn't echo our recommendations, people will go with whatever one off script their personal bad idea bear churns out. Many research software errors are edge cases (scaling, numerical errors, etc), and "works on my computer" syndrome is endemic in the literature. The field has made big strides in enabling computational reproducibility, but lately it feels like we're set to lose ground again.
Second: LLMs have a bias to action, and can easily bury the user reporting on whatever. I've seen multiple seminars recently where the Q&A devolves to "Q: What are the implications of this finding? A: I don't know, this is just presenting a report on the results".
It seems that humans are still figuring out how to maintain agency and steer the chatbot to the big picture, and unreviewable science is just as likely an outcome as unreviewable code... but with far less automated tooling to help guide the process.
Historically, PIs nominally guided the big picture, and the entity who did the work was a participant in the review process; now students are having to look at projects from a new angle with their own (invisible) chatbot underlings. I suspect that the solution will involve a combination of technological change, capturing common expert review checks, and really adjusting the kinds of skills that trainees are expected to have early on.