Bigger context windows help but they don't remove the need to chunk. Embedding 8K tokens into one vector smears everything, retrieval quality drops even though nothing got truncated.
Are there any good practices on multi-resolution embedding? Like, a strategy where you embed an entire document, and multiple levels of smearing, perhaps going all the way down to 512-token chunks?
Are there any good practices on multi-resolution embedding? Like, a strategy where you embed an entire document, and multiple levels of smearing, perhaps going all the way down to 512-token chunks?