logoalt Hacker News

beng-nltoday at 8:41 AM0 repliesview on HN

I agree: it feels to me that judging this implies knowing whether there is a solution at all (or a solution available per model, example: whether either model will answer it given known guardrails), which is as powerful as answering the question in the first place (the router can answer the decision problem, which can polynomially be transformed into getting a specific solution).

It reminds me of someone I met at a poster session who had an incredible project: his ai could include a confidence score with its answer, and he had evidence that x% confident answers were in fact correct x% of the time. That also gave me the gut feeling of that being impossible (in the general case) as it implies more powerful capability than the ai answering the question, in an oracle like way.