I missed the hypewave so can't say a lot about Jev, but the double standards are entertaining:
About Jev:
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities (even though they are not always correct).
Only 99% correctness! Borderline unusable!
About their model:
> It classifies: it gets a prompt with choices and outputs probabilities.
You want numbers, it gives you numbers! What more could you want?