logoalt Hacker News

benjiro29today at 12:32 PM2 repliesview on HN

Ironically, we are also moving to more capable / faster models that use less power. DeepSeek V4 Flash 0731 is so extreme capable and comparability to a lot of models cheap to run.

We are seeing stuff like AMD buying Taalas, with their Llama 3 8B on a chip, being able to push 17.000 tokens / second.

https://chatjimmy.ai/

Generated in 0,033s • 14.212 tok/s ...

People need to understand, that even AI companies do not like to build datacenters and the high energy needs. It cost them money and they are also looking at ways to develop better hardware / solutions that gives them more inference for cheaper (what means less power draw / heat generated).


Replies

smoldertoday at 12:53 PM

I just tested out chatjimmy and a bit and it did some misinterpretation of my prompts, E.g. "write me a python script to generate insults for a name given by command line" had multiple of the same insult in the insult list, and it had a list of names like John and Jane instead of doing the obvious thing and making the name a variable. Not good code.

show 1 reply
niek_pastoday at 12:43 PM

Unfortunately, "more capable / faster models that use less power" will likely mean those models will be used in more places than ever before. LLMs on laptops, on tablets, on phones, on fridges, in your web browser. See the Jevons paradox.