> And we can also implement llm inference deterministically if we want, it’s just that it’s not worth the loss in performance to do it.
My understanding is thats not possible (different from being practical), wonder if you have any literature, research to back up that claim?