logoalt Hacker News

throwa356262 • today at 5:53 PM • 1 reply • view on HN

The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.

In practice, you will be able to run models a bit bigger than 35B.

https://unsloth.ai/docs/models/qwen3.8-next


Replies

cyanydeez • today at 7:12 PM

https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.

38GB of vram resident. more tk/s, more prefill.