logoalt Hacker News

basilikumtoday at 4:54 PM0 repliesview on HN

I'm not sure if I'm misunderstanding you or the article, but as far as I understand it from a cursory reading this is not a jailbreak. They use this one prompt for fine tuning and use the model's response to the prompt as the sole underlying benchmark.

How would you do this with a closed weights model?