Fair callout. Strongly discounting (b) and (a) is underspecified for me to respond reasonably.
Your premise of full token commoditization has a major wrinkle though. GPU-centric data centers with never allow tokens to go to zero since you have a 1.5 year technical deprecation and 5 year life on the GPUs. Maybe Cerebras or similar chip builders, after writing model weights to silicon, will improve those returns on capital and enabling the scaling you are envisioning.
From my perch, AI at the edge is severely underserved, and the GTM for the hyperscalers mostly ignore it. So if someone figures out how to scale AI at the edge (nvidia + HF perhaps) then I think your premise is on point.
Maybe they still want to serve the feee tokens if that gets them good training data.