logoalt Hacker News

freehorsetoday at 7:06 PM2 repliesview on HN

I have used their gemma 4 31b model through kagi and getting real instantaneous answers is absolutely crazy. A very different feeling and UX. Even if the model is smaller, there is definitely a use case for these. I was wondering if they would put the qwen 27b model, it sounds very interesting to try.


Replies

codazodatoday at 9:25 PM

Really an aside, but yesterday I got the Gemma-4-12b (128k context) to build it's first web app in the minimal Dark Software Factory I've been building for myself.

https://joeldare.com/a-local-open-weight-model-builds-its-fi...

bitexplodertoday at 7:29 PM

The thing I didn’t realize for a while is 27B is rather smart. As many (or more) activated parameters as the flash models of the universe that we know about. It reasons very well. It just doesn’t have a lot of knowledge.

show 1 reply