Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.