logoalt Hacker News

ComputerGuruyesterday at 9:05 PM3 repliesview on HN

What quantization level is that? Because official endpoints are slow.


Replies

ak_tyesterday at 9:25 PM

It doesn't need extra quantization. The official weights are natively mixed precision FP4/FP8, so it fits in ~160GB. The API slowness is probably from being batched with other concurrent user requests. The provider's aggregate throughput gets higher but per-stream speed slows down.

zargonyesterday at 10:14 PM

V4 Flash fits entirely in two RTX Pro 6000s without any quantization at all.

bel8yesterday at 9:23 PM

From opencode go $10/mo plan I get between 60 t/s and 100 token/s even with large contexts of 150k+ tokens.

I wouldn't call 80 t/s slow.

show 1 reply