logoalt Hacker News

mikert89 • yesterday at 6:26 PM • 1 reply • view on HN

you probably aren’t using the models to their full capacity if you don’t notice the difference


Replies

verdverm • yesterday at 6:29 PM

vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering

here's an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models

https://github.com/verdverm/pge-jax#note-from-author

are open weights lagging, yes, are they way behind, no

if open weights were so inferior, they would not be >50% of all token processing

➕ show 1 reply