logoalt Hacker News

visarga • today at 6:35 PM • 1 reply • view on HN

Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.


Replies

nsingh2 • today at 6:37 PM

One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.

Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.

I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.

➕ show 3 replies