Its not an LLM if there's no pretraining. AR transformers were around before LLMs and will be there after LLMs.
When I made this, the point was to show that you dont need pretraining (which is what makes an LLM) to perform well on complex tasks
And yes it is not a language model either. I did not train it on any language data. Only ARC puzzles
Out of interest, would you call BERT an LLM? It’s pre trained but not particularly large.