The alternative is Claude-style "safeguards" aka censorship, which:
1. doesn't eliminate the possibility of a jailbreak anyway
2. frequently has false positives, triggering on innocuous requests, which is just really annoying
Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...
yeah the 'grok' way sounds less safe but it means less policing and having abstract arbiters of the truth
> The alternative is Claude-style "safeguards" aka censorship
Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.