So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers.
I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.
I repeated it with another one - again correct answer. I then selected the question they said Grok specifically answered incorrectly and it again answered it correctly.
I am sticking with my first intuition: people are terrible at testing tools and probably wanted them to answer incorrectly/not fully (the questions are constructed in a way to make it difficult as well). They also have vested interest in the conclusion (they are financial advisory firm) so there is that to consider.
People reading ft will now think chat boxes are bad at answering financial questions while they are pretty good at it. Zero consequences for spreading fake news for Financial Times there but good for financial advisors I guess.
> each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.
Okay but why do you feel 3 trials say as much as 10,000?
LLMs use random numbers, so a single test won't necessarily match someone else's experience.