logoalt Hacker News

Nvidia Nemotron 3.5 Lightning and NeMo Switchyard

110 pointsby droidjjtoday at 7:35 PM41 commentsview on HN

Comments

jmward01today at 8:48 PM

One major consequence of the ramapocalypse, I think, is an even higher focus on small efficient models. I personally believe that the multi-trillion parameter models are fundamentally missing things and the push to smaller, more efficient will drive evolutionary structural changes that will lead to future gains

show 3 replies
thehamkercattoday at 7:50 PM

> NeMo Switchyard, an open source library for smart routing

> When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job

How do routers like this handle prompt caching when you send the second request?

Sticky models per session? but then the second message of that session won't be sent to a suitable model, and will only be sent to the same model as previous one.

show 3 replies
docheinestagestoday at 9:28 PM

They conveniently decided not to include the Qwen range of models in the Artificial Analysis graph, except the out-of-league Max variant. At least be brave and honest.

average_bloketoday at 8:16 PM

I would like to propose something:

- problem: massive deluge of information because of AI

- solution: human beings should adopt a minimalist style of communicating in writing.

- e.g. this entire website page can be ten bullet points.

show 5 replies
WalterGRtoday at 7:57 PM

24 comments so far about Nemotron on this earlier submission: https://news.ycombinator.com/item?id=49257947

CurbStompertoday at 8:44 PM

[dead]

XCSmetoday at 8:03 PM

The new Meta 30B models seems A LOT better:

https://aibenchy.com/compare/meta-muse-glimmer-30b-xhigh/nvi...

show 5 replies