logoalt Hacker News

keunhongyesterday at 9:40 PM2 repliesview on HN

Atlas project lead here.

Atlas is an auto-regressive diffusion model, so context length limitations apply similar to LLMs and video models.

Where Atlas has an edge is that its context comprised of an arbitrary sequence of images with camera poses, which lends itself to managing the context in creative ways (we called this "context juggling" in our RTFM blog, https://www.worldlabs.ai/blog/rtfm). So yes through clever context management you could potentially build an entire 3D model of the world.


Replies

cman1444today at 12:48 AM

Would it be more reasonable to take images from movies and create worlds of various IPs?

My first thought is a detailed Hogwarts that is fully explorable using scenes from the movies (or even descriptions from the books?)

stranded-manyesterday at 9:42 PM

can atlas also generate 3D without pose information attached to the input images?

show 1 reply