logoalt Hacker News

strangecasts • today at 10:18 AM • 1 reply • view on HN

The technical report suggests it does not use the depth buffer: https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf

> The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.

> Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields such as depth, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.


Replies

mikepurvis • today at 3:00 PM

Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn't be of benefit to bring as much of that kind of metadata to the model as possible. You'd think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.

I'd also be interested in how post-processing fits in with this. Like if you've got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.

➕ show 2 replies