logoalt Hacker News

Neutrino-1 8B

67 pointsby handfuloflighttoday at 4:49 AM12 commentsview on HN

Comments

kamranjontoday at 5:53 AM

There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format.

PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https://prismml.com/news/ternary-bonsai

I wonder how proprietary it all is though, since the BitNet b1.58 paper has been out for a couple years now: https://arxiv.org/abs/2402.17764

From the wikipedia on 1.58 bit llms: "BitNet derives its performance from being trained natively in 1.58 bit instead of being quantized from a full-precision model after training. Still, training is an expensive process, and it would be desirable to be able to somehow convert an existing model to 1.58 bits. In 2024, HuggingFace reported a way to gradually ramp up the 1.58-bit quantization in fine-tuning an existing model down to 1.58 bits."

The section from huggingface is here: https://huggingface.co/blog/1_58_llm_extreme_quantization#fi...

I just wonder how many of these labs are basically following the huggingface recipe here and possibly tweaking it and releasing models without huge training costs.

show 2 replies
moinismtoday at 6:40 AM

The content on that page is too AI-generated to make sense to me; I don't understand what the model is for.

show 1 reply
Havoctoday at 7:50 AM

Can’t say I’m a fan of containers for this. A big chunk of local LLM gains come (imo) from the open modular nature of llama.cpp and friends. Easy to modify. Easy to experiment.

Containers are the proprietary binary blob in hardware world equivalent

yborgtoday at 5:46 AM

Largely outperformed by Ternary-Bonsai-8B by their own chart, doesn't seem clear what their special sauce is here.

NetOpWibbytoday at 7:38 AM

There’s a new announcement every other day wrt models. How do y’all keep track of them all same know what’s decent? Good grief!

And if it’s decent today, it’s shit in eight months! I tool hop as much as the next dev but this is a bit much.

show 1 reply
madhu_ghalametoday at 7:09 AM

The real strength of an 8B model is efficiency. It will be interesting to see the balance between performance and inference cost.

Alien1Beingtoday at 8:04 AM

AI slop site with AI slop research...

Blog populated with incoherent PR material generated by Yet Another AI.

Sigh...

runtime_lenstoday at 7:40 AM

[flagged]