Please, do not use LLMs to do actual statistics. They suck at doing actual statistics (like they suck at everything else).
Use proven traditional statistical methods like supervised learning for this. You can use traditional (not large) language models to tokenize the content and then a supervised learning trained on the most popular models to detect if those models generated the content.