You raise an interesting point. Although wouldn’t Anthropic be the best experts on when a model would be most applicable or likely to falter?
I guess my opinion would really depend on the exact wording and details of the arrangement. Lots of people jumping in with strong stances based on pretty limited details.
On faltering: the US Military already has rules around the use of automated systems for this type decision making. This is not a new concept to them by any means, and currently it always involves a human decision.