Glad to see a benchmark outside of programming tasks. Presumably this one was not part of the training and shows the models performance (or lack thereof) on knowledge tasks.