Perhaps it might be interesting: a latent thinking version is here https://huggingface.co/nmitchko/DeepSeek-V4-Flash-0731-Laten...
Does no thinking emissions for context saving.
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?