logoalt Hacker News

spaintechtoday at 4:21 PM1 replyview on HN

You can like the character or not, but there is a trend I’m following ( heavily vested in NVIDIA, so tongue in cheek when I say this ) that might be highly align with Zitron. Looking at the moves from NVIDIA ( Groq )and AMD ( Taalas ) which are pure inference plays. I believe this shows that the impetus to train a better-bigger model might be coming to a level of maturity that might merit a serious threat to the frontier labs.

For frontier labs, their fund-train-new model play might not be as effective, and a shift of spent of compute cost moving away from training to inference might be a tell-tell sign of the LLM as we know it plateauing out as scale is just not as effective. Open models might also be placing a major pressure on meeting then revenue targets need to sustain the model, lots of customer hosting their own inference to mitigate costs.

If you only move the needle just slightly in the direction of inference, frontier labs will soon loose their alphas. Becoming just another SaaS for inference might not be as attractive unless you are Google/MSF ( IMHO ).

Should this pan out, it could be a scenario where the NeoClouds could soon loose their biggest customers, so I tend to agree with that aspect of Zitron’s view.

Thoughts?


Replies

drivebyhootingtoday at 4:28 PM

Inference time compute is the new scaling.