We had processor level branch predictors. Now do we not only pre fill the next prompt, why not just start generating the response as well?
Interesting thought at least.
[dead]
[dead]