How is 30B smaller than 27B?
It uses fractal compression
They say it is trained with quantization awareness, so it should only be 15GB or so. Qwen was only trained in FP8 with QAT.
UPD, NVM, got misled by comments here. It is actually almost 60 GB so much larger
It uses fractal compression