impressive, i wish someone takes a stab at using this technique on mobile gpu's even if it does...

pdyc • today at 12:36 PM • 0 replies • view on HN

impressive, i wish someone takes a stab at using this technique on mobile gpu's even if it does not use storage it would still be a win. I am running llama.cpp on adreno 830 with oepncl and i am getting pathetic 2-3t/s for output tokens

alt Hacker News