Well, you can estimate the confidence BEFORE you start the task, too. That way you can restrict your trajectory to just a few models.
We also think there are tons of people working on "context management"--e.g. retrieval systems, prompt compression, log compression, etc. We want to work harder on the "decode" side as we think there are lots of savings to be made