logoalt Hacker News

ShinTakuya • yesterday at 8:32 AM • 3 replies • view on HN

Yes. Not as fast/cheap as Typesafe claims, but lots of benchmarks suggest around 3-4 times cheaper, and 7-14 times faster.

- https://www.ml6.eu/en/blog/jev-vs-gpt-6-luna-vs-bert-text-cl... - https://tessl.io/blog/jev-is-136x-faster-and-27x-cheaper-tha... - https://x.com/fazxes/status/2100300097695232164 (this last one is Luna 5.6 but that isn't too different from 6 besides accuracy and cost)


Replies

CharlieDigital • yesterday at 11:08 AM

If you can do it offline, batching solves this. We used a test dataset from CFPB and at n=20, it was 1.6x faster and 1.2x more expensive with no statistically meaningful accuracy dropoff. Did not tune for batch size, but it's possible that we could get to 30 and see better perf.

➕ show 1 reply
saberience • yesterday at 1:48 PM

No shit its cheaper and faster, thats the case with every small model.

It's a trade off between wanting something dumber but fast and cheap, or something slower and more expensive that is vastly more capable and smarter.

➕ show 2 replies