logoalt Hacker News

MichaelGlasstoday at 12:50 PM0 repliesview on HN

I see a lot of claims in this article without ... any proof?

Both can be true: - It's useful to anthropomorphize agents when predicting behavior and - we have to use specific language to specify what we mean.

What does the author mean by "confuse the models" ? Are they talking about not picking right information? Picking the wrong information? Losing their previous context / task?

Part of setting up a proper eval is also deciding what we actually mean ourself. What are we actually optimizing for? It's not, e.g. % confusion, %rubbish, etc.

The article does point to it: retrieval latency, accuracy, etc.