Because the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside", "outside", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.
The problem i have with this argument is they still struggle to generate images of so many basic things that should be explained in their "world model" (text, fingers, reflections, faces) - so clearly they dont have a "world model" in the way that we have one.
And if you use image generation capabilities, you clearly see that LLM's suck as understanding "next" to a thing.
I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., "beside the rock" or "in the array", it's doing the same kind of inferential activity.