logoalt Hacker News

walrus01today at 12:49 PM1 replyview on HN

3.8-flash-next quantized in a "large" Q4 that just fits in 128GB RAM even more so, in how close it can get to state of the art in a number of benchmarks. Or a large Q8 version of it that fits in under 190GB. Competing against things that are closed weights/opaque information about the model and might very well be 600B+ in size.


Replies

nicman23today at 1:00 PM

it "fits" in 64 ram with mmap. granted it runs at 15 tk/s with a 9070xt but it runs

show 1 reply