Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving.
RAM was probably the bottleneck for the amount of context they were offering.
I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful