Jensen : Nobody needs to code anymore...
Programmer: OpenDLSS...
Jensen : Wait. Not like that! (╯°□°)╯︵┻━┻
I'm rather surprised that it doesn't take the z-buffer as an input. I would have thought that would have provided useful information, it's one of the more useful forms of contolnet.
> bit-exact against the original
What kind of sorcery is this ? Very impressive work !
I think it's interesting that Nvidia is so interested in producing the hardware that fuels the future of software development, given that their primary business advantage is their software moat. This is an interesting project for sure, but turning a bunch of GPUs at Zluda[0] (an open implementation of Cuda) could be far more destructive for them, right?
Almost 8ms on 1080p resolution seems extremely expensive, does the original also eat into the rendering budget as much?
How useful is this without weights?
Isn't the mote that Nvidia has is they work with studios to generate the training data from the game, then they ship a model per game?
Or is my knowledge outdated here and they're just using a single generalised model?
Am I the only one who feels a sense of disinterest in a project where the main README is LLM-generated? Does the author not have time to write what they did and how it's used?
At what point does this neural rendering take away the human touch on the art styles?
So that means AMD implementation is on the horizon?
Something will this would historically guarantee a Senior Staff+ position at Nvidia. Wondering why Jensen doesn't put money where his mouth is ("were seeking exceptional engineers blabla") and offer him a job?
> It takes one rendered frame (a low dynamic range proxy of it, three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars) and produces four f32 channels per pixel: an RGB residual and one temporal-blend logit.
> The temporal path is implemented, but in the demo: the network's history input lanes and its per-pixel blend logit drive a reprojected feedback loop (docs/frame.md). The dlss5vk tool runs single frames with no history, which is what the reference captures were made with.
From this I assume the network uses the (via motion vectors) reprojected previous frame in order to increase temporal stability, i.e. similarity over adjacent frames. But this isn't strictly necessary, and apart from it, DLSS 5 is a pure post-process filter. So you could apply it to an old animated CGI movie like Final Fantasy (2001) [1]. Which should make it look significantly more realistic, at the cost of some flicker or other temporal instability.
One could also apply it to still images, like old renders from Tomb Raider [2], where temporal stability is not a factor. The difference to conventional text-to-image models with a "make it photorealistic" prompt would be that DLSS 5 strongly adheres to the underlying geometry.
1: https://www.imdb.com/title/tt0173840/
2: https://www.tombraiderchronicles.com/images/artwork-high-res...
The big IP laundering engines doing IP laundering again?
Seems quite similar to this repository?
https://github.com/aloshdenny/open-dlss