I'm rather surprised that it doesn't take the z-buffer as an input. I would have thought that would have provided useful information, it's one of the more useful forms of contolnet.
I think this is mainly so it can use the existing hooks for DLSS upscaling without requiring changes to the renderer, AMD is working on a comparable method which uses adapter networks to slot normals and material properties from the renderer into the diffusion model: https://gpuopen.com/learn/temporally-stable-generative-illum...
Even relatively small RGB -> depth models are pretty good. Which kind of implies depth is well encoded in the RGB, and adding depth would not really reduce entropy, while costing bandwidth.
The official one seems to do, as well as other info from the engine (I think remember their mentioning LOD/UV map hints in one of the public demos, or articles, a few months back--or it might have been an Unreal Engine podcast)