LLMs fail in such bizzare, obscure methods (to the average observer) at times. Sometimes even simple questions ("who was that x person who was super famous I'm thinking of") type questions fail terribly.
The more vague and non committal and hand-wavey and subjective the field for AI to answer, the better the results (imo).
I daresay it's because the assistant training corpus is heavily biased towards one-shot solution answers.
Because the correct response to that query is "I have no idea -- you will need to provide more information"
and LLM Agents suck at that.