I’d be curious to play with this with some different open weights models. I’d think that a single next token would end up having similar probability distributions so given that this a probabilistic method, you’d be able to get at least some signal.
I suppose if that did work someone would have been able to work backwards and crack their key already, so I must be missing something.