Basically it's: skip decode, just do prefill then do some measurements. That's the crude description anyways.
And prefill is way faster on GPU type hardware.