logoalt Hacker News

ricardobeat • yesterday at 10:36 PM • 6 replies • view on HN

Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.


Replies

famouswaffles • today at 12:01 AM

It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization.

Byte Latent Transformer - https://arxiv.org/abs/2412.09871

1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.

LikesPwsh • yesterday at 10:54 PM

One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.

hbcdbff • yesterday at 10:39 PM

Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.

➕ show 1 reply
mitxela • yesterday at 11:58 PM

They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.

vanuatu • yesterday at 11:30 PM

we already fixed it with reasoning