logoalt Hacker News

Kimi-K3 Releases on HuggingFace 7/27

452 pointsby nateb2022today at 6:18 AM207 commentsview on HN

Comments

NitpickLawyertoday at 6:37 AM

This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".

Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.

Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)

show 13 replies
KronisLVtoday at 8:14 AM

I feel like most hardware to run LLMs on is shaped wrong for individuals.

It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful (albeit NVLink or equivalent would need to be commonplace).

Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.

How unfortunate.

show 10 replies
gorgmahtoday at 7:56 AM

We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers

I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong

show 3 replies
elo02today at 10:07 AM

I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models

show 4 replies
maelitotoday at 7:36 AM

Did someone run censorship and political bias tests on this ? Must be interesting.

show 5 replies
bertilitoday at 10:43 AM

Is there any (near future) technology that would permit burning this terrabyte into some kind of ROM chip?

show 2 replies
khanhnguyen8386today at 10:24 AM

Can't wait to run this at 0.02 tokens/sec on my CPU so I can get a response just in time for next month.

2001zhaozhaotoday at 10:28 AM

I wonder how long it will take for this to get fully decensored and for bad, BAD things to happen

show 1 reply
docheinestagestoday at 8:43 AM

Given the frontier-level capabilities of Kimi K3, I'm wondering if it's possible to extract the core capabilities (fundamental reasoning and tool calling) of the model into a smaller one that consumer devices could run? Not sure exactly how, but either by heavy distillation or some other surgical method since Kimi has a Mixture of Experts architecture.

I think it's very valuable to have a smaller model that doesn't have any domain knowledge or facts built into its weights, but given the right context, could accurately reason about what to do and use the right tools.

I'm aware of colibri [1], but so far I've only seen extremely slow performance.

[1] https://github.com/JustVugg/colibri

show 3 replies
wwwhizztoday at 7:27 AM

That would be 7/27.

show 4 replies
apexalphatoday at 10:59 AM

This has to be one of the craziest uploads on the internet up until now.

Raw fucking intelligence at your disposal, free to download.

If you'd describe what's happening now to someone from five years ago they'd think you're hallucinating or mad.

davidkunztoday at 7:33 AM

This is historic. For the first time, an open-weights LLM is right at the top.

We won't be able to run this ourselves, but many providers can.

show 1 reply
martindelophytoday at 9:49 AM

The VRAM, power, networking, and operational requirements put it beyond the reach of many enterprises.

rs38today at 8:07 AM

is there a realistic way to distill 2 consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?

show 1 reply
hneqy2wqlstoday at 10:24 AM

Good enough is often the right call

show 1 reply
pmg1991today at 7:56 AM

Hoping no issues on Huggingface due to download rush.

show 2 replies
padolseytoday at 9:04 AM

There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.

show 1 reply
colortilestoday at 8:07 AM

This looks really promising. Excited to see where this goes. Looking forward to trying it out!

dsrtslnd23today at 9:36 AM

maybe a quantized version on a GB300 would work? unsloth hopefully working on it.

CodeComposttoday at 7:32 AM

Why is there a countdown?

show 3 replies
marvinLucktoday at 7:37 AM

The weightings should be released on July 27.

sreekanth850today at 7:49 AM

how feasible its will be to run on modal or deepinfra? anyone here tried and tested such large models running?

minimaxirtoday at 8:08 AM

...does Hugging Face have enough bandwidth to let people download en masse however much file size a 2 trillion parameters model is?

seydortoday at 9:47 AM

It's thankful that openAI or anthropic haven't IPOed yet. Or terrible for some

dxxvitoday at 10:25 AM

Now I hope that nvidia will host it for free :-)

m00dytoday at 7:27 AM

There’s going to be a lot of competition around this model. Let’s see how low AI providers are willing to push prices.

show 2 replies
rvztoday at 9:16 AM

This is (actually) AGI, that truly benefits all of humanity with zero gatekeeping.

Now the US government has 5 hours left to (attempt to) stop the release. (and save Anthropic)

Let competition run its course and the market (not government) determine the winners and losers.

hellajack3dtoday at 8:31 AM

So... Now we give huggingface the hug of death - right? ;)

wayknowtoday at 9:15 AM

[dead]

extra-AItoday at 8:11 AM

[flagged]

luciana1utoday at 9:47 AM

[flagged]

VladoIvankovictoday at 10:53 AM

[dead]

youre-wrong3today at 7:51 AM

[dead]

threeroutertoday at 6:18 AM

[dead]

throwaw12today at 7:50 AM

Strange communists, giving away such an expensive model to the public.

On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"

TurdF3rgusontoday at 10:48 AM

FYI huggingface refers to the alien from the Aliens movies that we need to prevent from reaching earth at any cost because it means the end of civilization.

Just checking in because y'all sound good with that.

show 3 replies