I agree that there are differences between Jev and LLM-as-a-judge (e.g. Almeida's assertion that LLM probabilities have been irrevocably biased by RLHF), but I am not sure about your description of 'confidence'. Perhaps I misinterpret you or the docs, but I understand it as simply being computed from the probability distribution: https://docs.typesafe.ai/confidence#how-confidence-is-calcul...