The 1M-document numbers are a great reality check. I’d happily start with the NumPy version and only reach for a vector DB once it actually hurts.
Nothing new under the sun, the same thing has been happening since the "big data" keyword became popular in the 2010s.
384 is a tad small. How does it perform on 1k/2k/4k..?
I 'member this is what Karpathy used in his restaurant menu app, and John Carmack said this is the way. I also agree that embeddings are overkill for the majority of tasks.
Yes! I'm working on an MCP server called Liveclip (not released yet) and we just store embeddings as text in SQLite.. it works over ~40MB of text, 1450 SCOTUS opinions
start=$(date +%s); curl -s -X POST "https://[urlredacted].workers.dev/mcp" -H "Content-Type: application/json" -H "Accept: application/json, text/event-stream" -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"table_search_similar","arguments":{"key":"scotus_2010s","col":"M","query":"endangered fish"}}}' 2>&1 | grep '^data:' | sed 's/^data: //' | jq; end=$(date +%s); echo "elapsed: $((end-start))s" { "result": { "content": [], "structuredContent": { "has_more": true, "headers": { "A": "author_name", "B": "category", "C": "per_curiam", "D": "case_name", "E": "date_filed", "F": "federal_cite_one", "G": "absolute_url", "H": "year_filed", "I": "scdb_id", "J": "scdb_decision_direction", "K": "scdb_votes_majority", "L": "scdb_votes_minority", "M": "text" }, "offset": 0, "ok": true, "rows": [ { "A": "Justice Thomas", "B": "dissenting", "C": "False", "D": "Florida v. Georgia", "E": "2018-06-27", "F": "", "G": "https://www.courtlistener.com/opinion/4511641/florida-v-geor...", "H": "2018", "I": "", "J": "", "K": "", "L": "", "M": "Florida asks this Court to cap Georgia’s use of water in\nthe Apalachicola-Chattahoochee-Flint River Basin (Basin).\nFlorida claims that such a cap would allow additional\nwater to flow into the Apalachicola River and Bay, which\nwould benefit Florida by alleviating certain ecological\nharms. To prevail under our precedents, Florida must\npresent clear and convincing evidence that its proposed\ncap will benefit Florida more than it harms Georgia. See\nColorado v. New Mexico, 459 U.S. 176, 187 (1982) (Colo-\nrado I ). The Special Master applied this balance-of-harms\nstandard and, after presiding over a 1-month trial involv\ning 40 witnesses and more than 2,000 exhibits, found that\nFlorida had not met its burden. Because that finding\nis well supported by the evidence, I would have over-\nruled Florida’s objections to the Special Master’s Report\n(Report) and denied Florida’s request for relief. I respectfully\ndissent.\n I\n The Court’s recitation of the facts focuses o", "_match": true, "_row": 1189 },
elapsed: 12s (on my laptop)
Umm, is pgvector relevant to this usecase?
This works especially well if your embedding model was trained to perform well with quantized embeddings. Binary + hamming distance = incredibly fast.
This post is from 2024 but I wrote about using this technique in https://emschwartz.me/binary-vector-embeddings-are-so-cool/