logoalt Hacker News

Tade0 • today at 4:49 PM • 0 replies • view on HN

> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?

Turns out that actually - no. Researchers have managed to prune half the Experts in a MoE model that had a low probability of getting activated during coding tasks, resulting in a more focused model:

https://arxiv.org/abs/2607.16721

Main benefit is that it greatly reduces the amount of RAM required to run these models. Of course you could just cache those unused experts on disk instead, but the main point here is that you know which ones matter.

But aside from that recent models, like Qwen3.8-27b are reportedly more durable under heavy quantisation, e.g. 3bits or even ternary. With additional techniques like TurboQuant, you can feasibly run these models on consumer hardware - even if at 1/4th the speed you'd get from rented infrastructure.

VS Code has extensions such as Kilo Code or llama-vscode which let you work with local models much like you would with cloud based solutions.