The first thing I'd do if working with an LLM on tabular data is to ask what the best tool would be to work with that data and build up a proper harness to work with the data sensibly. Rawdogging LLM isn't the tool for forecasting like this, as they found.
Unsure if it's LLMs that fail at tabular data or its just that tree boosting are spectacular at that task.
Nowhere in the paper do they mention the reasoning level or budget used for the experiments?
You’ve got to be kidding me. That one variable could make a huge difference in the results. I can’t understand why they would leave that out.
>We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning
Sigh. So this is somewhat interesting niche academic research but utterly irrelevant to real-world use cases.
Just have 2 LLMs debate whether tabs or spaces are the superior choice
Look at the white text on white background in Appendix F. Pretty funny.