logoalt Hacker News

IsTomtoday at 11:43 AM2 repliesview on HN

How is 30B smaller than 27B?


Replies

LeBittoday at 11:48 AM

It uses fractal compression

lostmsutoday at 12:27 PM

They say it is trained with quantization awareness, so it should only be 15GB or so. Qwen was only trained in FP8 with QAT.

UPD, NVM, got misled by comments here. It is actually almost 60 GB so much larger

show 2 replies