I agree with the author if we are trying to use this as some sort of absolute scale of confidence. It only develops meaning when we control for many other variables. Looking at confidence scores across two different models or prompts is probably not a good idea.