Wouldn't just putting tokens in a ring buffer work?
Not unless you want to cheat the attention mechanism or do extra computations running prefill in a front-truncated version of the conversation.
Also, to the extent that the model reasons and thus learns something, if you blindly truncate the front, you will lose that knowledge.
Not unless you want to cheat the attention mechanism or do extra computations running prefill in a front-truncated version of the conversation.
Also, to the extent that the model reasons and thus learns something, if you blindly truncate the front, you will lose that knowledge.