I have never seen anyone report "this produced really great results" from intentionally quantizing their context vs. leaving it at full precision which is the ordinary default.
Gemma's QAT is surprisingly good (although Gemma isn't that great to begin with).
Gemma's QAT is surprisingly good (although Gemma isn't that great to begin with).