Maybe have A second model do the management?
Looks like they tested that in the paper, and the code allows for it as well. https://github.com/facebookresearch/context-language-models/... Of course, this is for suggesting context management strategies rather than the actual management afaict.
I was thinking this, not so dissimilar from co-training dflash drafters
Yeah, that would be the canonical solution. Modularity is better
Looks like they tested that in the paper, and the code allows for it as well. https://github.com/facebookresearch/context-language-models/... Of course, this is for suggesting context management strategies rather than the actual management afaict.