logoalt Hacker News

reasonablekloutyesterday at 5:33 PM2 repliesview on HN

This sounds completely insane, utter sci-fi, especially that the communication happened during a training run. And yet OpenAI decided to continue the training, and we didn't hear about the incident for weeks. And now they are pushing forward with deploying a new model anyway. How is this happening? What will things look like in the labs in 3 months, let alone 3 years?


Replies

NitpickLawyeryesterday at 5:39 PM

It's not unexpected. Current model gains are mainly from RLing a pretrained model on lots and lots of scenarios. They have the models run scenarios, and RL on successful runs.

porridgeraisinyesterday at 8:33 PM

The training here is RL training, the rollouts there are not different from inference and have access to the same tools as regular inference.