I don't understand why they don't look for large substring matches for the system prompt before returning the response. Trivial calculation compared to a system prompt instruction asking the model not to do it
Because it's trivial to bypass through things like the model natively knowing how to speak in encodings like base64
because you can always make up your own language and ask the model to use it, no filters would catch that