logoalt Hacker News

smackeyackytoday at 10:18 PM3 repliesview on HN

How can these models do anything close to RSI when they can’t even self check their output? Gemini for example is so self confidently wrong about 30% of the time for me on certain tasks. I tell it that its answer is wrong and it issues a mea culpa but goes back to being wrong in short order. I feel like the AI industry is still massively overstating their projections.


Replies

StevenWatermantoday at 10:36 PM

As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like

show 2 replies
drodgerstoday at 10:42 PM

> Gemini

That's definitely part of your problem.

In my recent experience, error rates for astra/fable are at or below human level. Just like when directing humans, it pays to ask probing questions ('Are you sure about X?', 'Did you check for Y?', 'Please run Z just to double check.') if you really care about the result being correct.

show 1 reply
kakugawatoday at 10:35 PM

They can only do it in the (narrow) domains that are verifiable.