I agree with your observation. Although the framing of the question signals some of the response "Do you think.. " forces the model to check for the cost-benefits.
We can definitely mould AI agents to think more critically about these things, I dont know how effective it will be in the long term honestly.
Recommend checking out Ponytail — a pi extension that behaves like a senior engineer who keeps the LLM code generation in check, and the processing of your prompts just the same
https://ponytail.dev