logoalt Hacker News

guyomes • yesterday at 6:36 PM • 1 reply • view on HN

If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document


Replies

mdp2021 • yesterday at 8:33 PM

> hardware dedicated to a specific LLM

That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...