I think "system prompt" is the key bit they're getting at. It doesn't necessarily reflect poorly on the underlying model if the system prompt was bad. It does reflect somewhat, in terms of alignment (how well the model does what the training company wants) and instruction following (how well the model does what the user wants). But it's not so clear to me what exactly the right answer is here. E.g., a model that scrupulously follows its system prompt and does what the user wants is a pretty useful, if very sharp, tool, albeit perhaps dangerous in the wrong hands.
I think "system prompt" is the key bit they're getting at. It doesn't necessarily reflect poorly on the underlying model if the system prompt was bad. It does reflect somewhat, in terms of alignment (how well the model does what the training company wants) and instruction following (how well the model does what the user wants). But it's not so clear to me what exactly the right answer is here. E.g., a model that scrupulously follows its system prompt and does what the user wants is a pretty useful, if very sharp, tool, albeit perhaps dangerous in the wrong hands.