logoalt Hacker News

anana_yesterday at 6:22 PM1 replyview on HN

And to read the tea leaves a little:

3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas.

It also produces nearly twice as many tokens per task as 3.6 (and by extension, time), which may be a tradeoff required to achieve correctness at this parameter size.


Replies

skohanyesterday at 6:40 PM

Imo it makes sense for things to move in the direction of small, focused models that excel in one area. I use LLMs for technical work 99% of the time, I could care less about general world knowledge, or if the model is good at creative writing.

With good orchestration and delegation you can get surprisingly far with small models running on consumer hardware.

show 3 replies