À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.