The problem i have with this argument is they still struggle to generate images of so many basic things that should be explained in their "world model" (text, fingers, reflections, faces) - so clearly they dont have a "world model" in the way that we have one.
And if you use image generation capabilities, you clearly see that LLM's suck as understanding "next" to a thing.