logoalt Hacker News

wmfyesterday at 11:25 PM1 replyview on HN

There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience


Replies

daishi55today at 1:26 AM

I wanted some examples they actually experienced. Because I use these things daily and haven’t seen a hallucination in a long long time.

show 1 reply