It's a hack but doing things the 'proper' way is at least 1000x harder so whatever.
Is it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating
Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”
https://huggingface.co/openai/gpt-oss-safeguard-120b
Is it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating
Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”
https://huggingface.co/openai/gpt-oss-safeguard-120b