logoalt Hacker News

I built non-autoregressive decision models with RL a year ago

898 pointsby nandakishor_mltoday at 10:46 AM222 commentsview on HN

Comments

m3kw9today at 6:48 PM

in two weeks, a chinese lab will have a Pev-2.7-flash-qwen for 0.00004cents/million

yogthostoday at 2:57 PM

I just built a server based on Jev API to run Laya here https://github.com/jlt-commons/laya-jolt

avaertoday at 2:26 PM

> Seeing the hype online feels both validating and deeply frustrating.

The post is conflating hype and money with technical innovation, they are not really correlated. Kurzweil is known for saying most innovations succeed based not on technology but on timing. Today, who talks about it might matter even more than timing.

Superior research often gets overlooked in favor of someone raising millions, sometimes people who have produced literally nothing manage to sell it. Not saying that's happening here, but I've seen this pattern a lot over my career.

Someone riding (or manufacturing) a hype wave is playing a completely different game from a researcher. If you're a researcher you can't really feel dejected when someone is making a business on the back of what seems like your research; legal protections are decades out of date, even ignoring vibe coding. If you want to make money/hype/whatever off of your work, do that. But realize that it's a path that's often orthogonal to research.

moinismtoday at 2:49 PM

I'm just glad to see focus being shifted (albeit slowly) to conventional ML. Enough with LLM guys

legions-lovetoday at 2:19 PM

[flagged]

reso_codestoday at 3:29 PM

[flagged]

Kuyawatoday at 1:43 PM

[dead]

rexthonyytoday at 3:04 PM

[dead]

zurfertoday at 12:14 PM

I've been deeply impressed with Jev as it made a bunch of workloads we had on Luna or Gemini 10x cheaper and 2x faster (previously used non reasoning version for latency reasons).

Now Laya promises another speed up and it's open source. Tbh if it can't run on a CPU I anyway want to buy it from an inference provider. Managing gpus in production is a non trivial problem.

What I also wondered about Jev is how different it is from something like tabular foundation models. They seem to overlap in use cases. Which then leads to the question, what is actually learned? A lot of people in machine learning spend time to making things explainable and always struggled to move beyond data induced biases.

Having it open source is awesome as fine tuning might give additional performance on the task we care about.

show 1 reply
prodigycorptoday at 2:22 PM

This post is either majorly astroturfed or nobody upvoting this has read past a couple sentences.

Thee author of this post has a specialized model that does routing for his sales pipeline.

After jev came out, the author adapted his model to produce jev-shaped output: choice, noul, and score primitives.

Then, author writes a post about how jev copied him. (even though this was released in the past couple days, after jev was released)

If I were to download this model, I would get nothing like jev. I would get a model that is able to beat jev in a very small domain that this model has been fine tuned for.