logoalt Hacker News

sebastostoday at 4:19 PM1 replyview on HN

But if the model is an LLM, you actually COULD ask it why it drove under the semi, and it would give you an answer. Now, you may argue that it will just be generating a whole new, backwards-rationalized post-hoc explanation of its own behavior given the logs that it managed to take before the crash. But then I ask you: how do you think a person explains why they did what they did after a crash? I direct you to all of the unsettling split-brain neuroscience literature demonstrating that humans are incorrigible backwards rationalizers who make for unreliable witnesses.


Replies

JoshTripletttoday at 6:08 PM

That's a bug, and it would be an awful mistake to replicate that bug rather than fix it.