logoalt Hacker News

verdverm • yesterday at 6:29 PM • 1 reply • view on HN

vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering

here's an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models

https://github.com/verdverm/pge-jax#note-from-author

are open weights lagging, yes, are they way behind, no

if open weights were so inferior, they would not be >50% of all token processing


Replies

mikert89 • yesterday at 7:55 PM

these models are trash

➕ show 1 reply