logoalt Hacker News

jrfloyesterday at 6:30 PM6 repliesview on HN

Burning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real


Replies

jaggederestyesterday at 6:36 PM

https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second.

https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.

show 4 replies
5555watchtoday at 12:14 AM

I'm curious, how hard/expensive it is to burn a really large model into silicon, and why aren't we doing this already?

Or, when we will start doing this, who's going to be able to do that in scale?

I'm seeing the TAALAS example, but it's only an 8B model, suggesting some real limitations parameter wise. And for 2.5kW?

Yopoloyesterday at 6:48 PM

And don't underestimate how much money Google, Microsoft, Amazon and Meta still have to spend on this tech.

Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.

and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.

It seems China is already able to do DUV a lot sooner than others expected.

show 1 reply
FuriouslyAdriftyesterday at 7:33 PM

Yep it will be ASICs and DSPs all over again. Orders of magnitude changes.

show 1 reply
vjvjvjvjghvyesterday at 8:19 PM

Not an expert on this but wouldn’t this be possible with something similar to an FPGA?

show 2 replies
coffeebeqnyesterday at 7:02 PM

What does that mean though? Like some kind of a ROM memory ?

show 1 reply