If you think about it there's overlap in what the VAE encodes and main model encodes. Objects at a distance resemble texture and textures zoomed in gain structure. The VAE makes textures more efficiently representable at the cost of reducing the representable space of pixels. So things like tiny text become nonsense scribbles. Working in pixel space, especially with something with recursive or cascaded structure, opens the possibility of using the learnt structure of real writing at a higher level to perfect tiny details that actually cannot be approximated without being obviously wrong.
Time is "just" another dimension. There's temporal continuity between frames, a video VAE would be learning and representing those temporal shifts, but there's nothing to say that e.g. a recursively applied generative model at the pixel level also doesn't learn and represent those things.
As ever, figuring out how to train the thing is the hard bit I expect.
(Handwaving over "textures" here, VAEs encode more like somewhat macro blocks of image whose content is also conditioned on surrounding blocks, rather than tiny patches of patterned pixels.)
(And yes I'm a total imposter layman here, I just see VAEs as seeming to be a crutch that reduce data size - super super helpful of course - but being strictly speaking redundant and inhibiting correct fine detail.)