logoalt Hacker News

Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

48 pointsby edwardbzhangyesterday at 7:44 PM11 commentsview on HN

Comments

beautiful_appletoday at 12:25 AM

A benchmark table comparing to Qwen 3.5 35B-A3B seems strange when Qwen 3.6 35B-A3B has been out for some time and is significantly better.

I didn't notice the version difference when first reading the article! So this is a heads up to people like me.

walrus01yesterday at 11:56 PM

I wish that "small" LLMs would stop being confidently very incorrect. Admittedly this is a bit of an intentionally esoteric test, but the confident way in which it presents a totally incorrect answer is a bit concerning.

"please write 250 words on the etymology and history of the word schlong"

https://pastes.io/uhshFgn4

The actual origin of the word is from middle high German and Yiddish-speaking Ashkenazi Jewish communities.

For comparison qwen 3.6 35B A3B does perfect on this and will give a solid description of the word's real origins and how it has made it into casual profanity/vulgarity as used in US English, and even mentions specific stand-up comedians and famous public figures of specific ethnic/religious origin in the US NE who introduced it into wider use.

Ask it for something that's not a narrow niche scientific or technical field, but something that would be less common to make it into a 20B size model, and see just how it does.

chat test link: https://chat.deepgrove.ai/

arjietoday at 12:31 AM

For small models like this, it’s super important that it works well at tool calling etc. imho because it can’t memorize facts and isn’t big enough to tell when it doesn’t know. I could use it for high quality tool routing or a backup fast model for smaller task set. E.g. I use GPT-5.6 for voice channels at home. I’d prefer to be able to have this do basic tool calls and stuff because of the local speed.

Will give it a crack as a quick model in my clawlike.

show 1 reply
getcrunktoday at 1:27 AM

Anyone compare this to ternary bonsai vs the 1 bit

Havocyesterday at 11:28 PM

Looks promising though much like the bonsai tenary one it hallucinates knowledge quite aggressively. The online chat having search tools covers this up somewhat, but it's still there.

...from very unscientific casual vibes it does seem pretty good though considering the speed

Would also be curious what their search tool backend looks like - that too is very fast for very rapid multiple searches

show 1 reply
jsphweidyesterday at 11:36 PM

As of now, 3 of the 5 comments on this page are just 0-1 karma accounts high-fiving the article. Suspicious.

show 2 replies
zooloo99yesterday at 11:34 PM

Edge is edging closer!

Super cool and a taste of what's to come with local AI becoming more accessible to low-end hardware.

HenryNdubuakuyesterday at 9:02 PM

This is cool!

Baghramian11yesterday at 11:49 PM

[dead]

aatmoyesterday at 7:55 PM

[flagged]

praneel_patelyesterday at 8:31 PM

[flagged]

tokenmunchingyesterday at 8:01 PM

i didnt think this was possible