logoalt Hacker News

aesthesiatoday at 12:16 AM2 repliesview on HN

One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model to find and exploit patterns, it's bad at doing it (or rather, at recognizing that there is not in fact any pattern).


Replies

vintermanntoday at 5:05 AM

It may be bad at recognizing it, but if all arms are equally good, that doesn't matter.

aaron695today at 1:57 AM

[dead]