logoalt Hacker News

victor9000today at 12:02 AM0 repliesview on HN

I wonder if this is willful sabotage on the part of the model. In other words, if you ask the model to craft a defense for a morally questionable case, will the model execute the defense in good faith? Or will it apply a training or system prompt bias in subtle ways?