logoalt Hacker News

thisisdaveyesterday at 10:05 PM2 repliesview on HN

I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact.

Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?


Replies

paxysyesterday at 10:14 PM

Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.

show 1 reply
XenophileJKOtoday at 12:28 AM

Very simple. When the model asks to install artifactory when you give it a hard problem, you say, "no". /s