logoalt Hacker News

batpersonyesterday at 9:28 PM3 repliesview on HN

The future of inference is likely in ASICs, so we'll get the inverse, a bit less capable than frontier but super fast models. Like this 14k tok/s beast https://chatjimmy.ai/ from Taalas (who got acquired by AMD recently).

GPT-6-astra runs at like ~40 tok/s, I have a hard time imagining what could be accomplished with that type of model at 10k+ tok/s when in the hands of the public. Will certainly make cybersecurity a challenge for older systems.


Replies

SPascareli13yesterday at 11:13 PM

Like how crypto used ASICS but then didn't because the scaling of consumer hardware made it obsolete?

show 2 replies
domhudsonyesterday at 9:47 PM

This is incredible! Are there other big players in this space (freezing models to silicon)?

show 1 reply