Another interesting question is why the frontier labs are piling on pure maths, which has little direct economic value compared to something like law or improving the efficiency of their own models? How much OpenAI and Anthropic are paying to serve these models for ordinary users is the elephant in the room. A cynical take is that the frontier labs are trying their best to pump up their pre-IPO valuation through flashy headlines.
Because it's a tool in search of a use case (or many use cases) and mathematics is the most natural use case for it. Mathematics is by definition the art of putting words on a page in a rigorously defined "correct manner" (i.e. in the form of a valid logical argument, a proof) and all LLMs do is put words on pages and evaluating if they're good words is by far easiest when there is a strict definition of right and wrong.
It's one of the few areas where you can verify results. That fits nicely into training models. They aren't just making judgement calls on what would be nice, it's "what can we do?".