I have a 3 year old gaming system. RTX 4080 w/128GB of DDR5. It runs Qwen 38 Flash around 44-40 t/s with 128K context. It is on a specialized build that caches MoE experts and uses an optimized 3bit quant that basically is within a few points of the full 8 bit quant. In general, in casual benchmarking with Alibaba's endpoint I could not tell much of a difference. Overall this model is very good on long horizon agentic work. The main pain point for it is that its input processing speed is slow. Regardless, it gets meaningful work done.
I paid $500 for the RAM in Nov 2023 :)
> "I paid $500 for the RAM in Nov 2023 :)"
No wonder Warren Buffet gave up and resigned.