Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn't be of benefit to bring as much of that kind of metadata to the model as possible. You'd think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.
I'd also be interested in how post-processing fits in with this. Like if you've got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.
I think it has more to do with what kind of data they have access to at runtime - IIRC DLSS upscaling has only required the previous frames and motion vectors, so requiring depth buffers would mean it was no longer a "drop-in" replacement
> I'd also be interested in how post-processing fits in with this.
I think screenspace effects like film grain and tonemapping are excluded in the same way UI elements are rendered separately from the game.
Hmm well the depth buffer only has accurate depth for opaque objects (and even that's not really true). Things like hair wouldn't be in there. Its not ground truth depth.