logoalt Hacker News

kmehtoday at 5:52 AM0 repliesview on HN

You're assuming that your prompt is not being intercepted and rerouted by a lightweight prompt classification model.

In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.