logoalt Hacker News

harrouettoday at 6:30 AM0 repliesview on HN

I could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens.

Who needs memory when your model is set in silicon ?