logoalt Hacker News

lyelibi • today at 5:49 AM • 1 reply • view on HN

I have never seen tabular transformer models beat xgboost/catboost in industrial context where datasets is gigantic. They most produce these results on relatively small datasets, clearly not in the tens of millions of rows.


Replies

icfly2 • today at 6:09 AM

I work with what generally qualifies as big data (a large fraction of European e commerce payments). The vast majority of this data holds no new insights. So even for xgboost the training data is trimmed down. To get these models to work one can trim the data down further, so split out by some known characteristics. For titanic (obviously a way to small dataset) split by gender and/or class.