logoalt Hacker News

schopra909 • today at 7:56 PM • 1 reply • view on HN

This is a totally fair point and definitely worth exploring!

The jumping point for this no-VAE work was three-fold:

1) Our goal here is to get 32x32 token reduction to make video training and inference downstream cheaper. To date, the best open-weight Image VAEs like Flux-2 seem to cap out at 16x16 token reduction (8x8 VAE + 2x2 linear patchification). Others like H3 have pushed to 32x32 reduction but requires them swapping out the small VAE decoder with a 2B parameter decoder. So, this is a foray to get 32x32 compression without compromising quality.

2) We believe that end-to-end trained networks will tend to perform better than modularly trained networks (e.g. VAE + DiT). This hypothesis comes from work like REPA-E, where authors are able to get much better results by backpropagating through the pretrained VAE. The latent space for perception / reconstruction seems to have a different "optimal" configuration than a latent space specifically built for generation. That's why we liked the idea of trained E2E here.

3) There's work with VAEs that show that providing additional modalities (e.g. text captions) can help the VAEs improve as well. That's natural to this construction, so we thought language might actually help with the compression, not hurt. To be honest, this this is the most hypothetical of three ideas; and definitely warrants specific ablations.


Replies

schopra909 • today at 8:01 PM

Also i failed to clarify that the VAE for Linum v2 was the Wan2 VAE