logoalt Hacker News

robrorcroptrertoday at 11:35 AM2 repliesview on HN

What about splitting bigger content into chunks before embedding?


Replies

freakynittoday at 12:25 PM

How are you gonna handle the relations that span across individual chunks... if a later chunk refers something from 2 chunks before using `it`, rather than proper name, how will you handle that? Because at query time, that later chunk would not match.

show 2 replies
mdp2021today at 1:20 PM

What member freakynit said nearby about chunks and relations between chunks, plus the storage and information efficiency problem: make some calculations about storing vectors - for paragraphs and for collections of paragraphs -, then compare the needed space with the original data...

Because you could have clever ideas about vectors related to more paragraphs related in the document structure - but that would multiply the vectors. The index can become much bigger than the corpus.