logoalt Hacker News

LinchZhangtoday at 5:49 PM0 repliesview on HN

I agree CoT monitoring is imperfect but it's okay in practice and helps us get defense-in-depth re: model intent. We absolutely do not have a singular safety mechanism in place that's sufficiently good that we can use it in exclusion of all other imperfect ones.

"Do you know what is a faithful representation of what the model wants to do? Tool calls."

tool calls could absolutely be spoofed, my impression is that this happened many times in the OAI HuggingFace attack.