logoalt Hacker News

trimethylpurine • today at 2:37 PM • 1 reply • view on HN

I'm not an expert, but I think that:

1. Storing Shakespeare's work costs almost no $ in regards to disk space.

2. If the prompt doesn't include "Shakespeare" or relevant terms then no regression is performed for that topic and therefore there is no effective token cost.

Someone may correct me, but I think it's not a big $ win to exclude relevant topics from the models' overall capabilities. Instead you'd tune weights so that #2 better identifies what is or isn't among the relevant terms on which to run regressions.


Replies

smaudet • today at 3:02 PM

I think you're probably missing the forest for the trees here... broadly speaking, these multibillion "parameter" (whatever that actually means) models store a lot more than shakespeare, and have storage costs in the 100s of GB/TB (which translates to $$$$$$ in SSD/RAM costs), nevermind the (kilos/mega/giga)watts involved, all the pollution, etc...

Meanwhile, a template (maybe a couple KB) costs less than a couple cents to store and run. Large Languages Models are not really interesting, (smaller) LLMs that only contain "what you need" are.

➕ show 1 reply