logoalt Hacker News

Who's afraid of Chinese models?

592 pointsby mfiguiereyesterday at 11:05 AM400 commentsview on HN

Comments

tristanjyesterday at 10:22 PM

The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models for free. If the frontier labs are forced to cut prices and join the race to the bottom in token prices, these valuations are unjustified, and VCs will face enormous (paper) losses.

show 16 replies
benruttertoday at 8:19 AM

I loved this article! Regardless of how you feel about AI as an industry or tool, the economics of AI is fascinating. It's awesome to see something like this that gets into the business side a bit more.

I don't know if I agreed totally with the assessment of the risk Chinese labs pose to US labs though, in particular I think the main part I wasn't sure about was this:

> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.

How true is this? My understanding from Deepseek's original paper was that they focused heavily on optimising training and inference costs, in particular so that they can operate on cheaper (and more accessible to China) hardware.

It's possible I'm just not in the loop, but nobody seems to talk about US models innovating in this way (I'm just talking about cost-to-serve/train, not saying US AI companies don't innovate in other ways).

It seems to me at least, like there's a fair bit of evidence that AI shifting to a price based commodity market (vs a "best-model takes all" type market) would put China at a significant advantage? And even more significantly, require a pretty hefty correction of company valuations in the US?

show 1 reply
wxwyesterday at 10:10 PM

> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users.

My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.

[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]

show 9 replies
jadamczyktoday at 8:03 AM

I think releasing models for free is some 4d chess move by the chinese. Big chunk of the us stock market is fueled by ai mania, if the frontier labs turn out to be drastically less valuable than first believed, the downturm may be very bad. Think of all the big tech companies that have a ton of debt that they took to pour money into AI. It seems like a similar tactic to what Chinese car manufacturers are doing in Europe but the result may be more dramatic.

faangguyindiayesterday at 10:53 PM

I operate an analytics site (pretty big one B2B where client's backend feeds data into our system), and we see tons of traffic originating from northwestern China (Xinjiang) from Shenzhen Tencent Computer Systems Company Limited.

There are also half a dozen other companies from China continuously hammering our clients’ websites.

I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.

Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region

credit:

'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopoc

Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.

Assuming that China only distills is a huge mistake.

It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.

Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.

The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.

Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.

This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....

There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.

ballon_monkeytoday at 3:18 AM

The 2 things people need to remember:

1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.

2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.

show 8 replies
aaronrobinsontoday at 7:22 AM

“ Let the frontier labs win by being better; don’t let them define safety or security, or pull up the ladder of humanity’s collective knowledge”

Love this.

kinj28today at 3:37 AM

I am afraid — if Chinese models go mainstream it has a clear way of pushing its narrative way beyond its otherwise borders. More like a Trojan horse it is for the Chinese.

Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question

https://x.com/jinen83/status/2079406993979383902?s=46&t=D7hQ...

show 4 replies
_aavaa_yesterday at 2:23 PM

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation

Sounds great to me; live by the sword, die by the sword.

show 11 replies
jke_kangyesterday at 10:33 PM

People seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."

show 2 replies
OleksandrCyesterday at 9:57 PM

The article makes a point about agent harnesses being sticky (the supposed moat). I have been building my own agent harness for a while, and I can tell with confidence that the harness almost does not matter, the entirety of the AI magic is the model itself. The harness can be almost barebones (like, for example, mini-swe-agent used for benchmarks), and yet the model still does the task just fine.

So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).

show 5 replies
bg24today at 12:32 AM

I think in general rest of the world needs to take notice (not saying afraid), starting with the US. It cannot be taken for granted that China's frontier labs will be a few months behind. They might be at par or exceed.

The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.

At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.

Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.

show 3 replies
abhinaitoday at 7:20 AM

The author doesn't seem to realize that a healthy margin has been built into the inference pricing. Once low cost open source inference providers get their hands on powerful frontier level models, there would be a severe margin compression for OpenAI and Anthropic.

Source: https://martinalderson.com/posts/the-upcoming-ai-margin-coll...

show 1 reply
credit_guytoday at 2:01 AM

People who claim that the Chinese open weight models have some type of manifest advantage don't realize that the close weight models have a huge advantage as well: the researchers from OpenAI, Anthropic, Google, xAI, Meta are not dumb, they can read the white papers written by DeepSeek, Moonshot, etc, and they can inspect all those architectures and they can pick and choose the best tricks there are out there, and of course, they have access to their own in-house secret sauces.

Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.

But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:

1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.

2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.

3. The frontier labs are also investing more and more in building an ecosystem around their models.

I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.

[1] https://www.anthropic.com/careers/jobs

show 4 replies
oezitoday at 6:32 AM

The article makes a great point that the token industry is going to be commoditized as time goes on.

Following this argument the key for each player will be the underlying cost structure and serving capacity to offset the upfront R&D cost.

The cost infrastructure will be driven by access to cheap electricity and cheap chips. The capacity will be driven primarily by depth of pockets now to buy all available supply in chips/mem/data center building capacity. While China is certainly in the lead on cheap energy, I am wondering if they can/want to beat the > 1tn USD being spent on data centers right now. Following the example in the article:

If company C from China sells 10 units for 20 USD produced for 10 USD they pocket 100 USD.

If company A from America can sell 100 units for 20 USD produced for 15 units, they pocket 500 USD or 5/6th of the market's profits.

0x38Btoday at 2:57 AM

Excellent article; the argument towards the end for allowing distillation for US companies is compelling:

> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?

> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.

show 1 reply
throwa356262yesterday at 10:03 PM

According to openAI's own @deanwball: Even OpenAI isn't buying this distillation talk:

https://xcancel.com/deanwball/status/2078133895766114412#m

show 6 replies
atmosxtoday at 6:27 AM

No one should be afraid of anything. Fear is a terrible advisor. Keep your eyes open, try to read the context as careful as you can and adapt as best as you can. Don’t spent too much time trying to be an oracle, never works out…

deauxtoday at 2:33 AM

> This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?

This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.

And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".

spenvotoday at 1:05 AM

"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence"

That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.

His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".

show 1 reply
simonreifftoday at 12:19 AM

I fully agree with everything in this essay. Make distillation fair use. And let us use Mythos/Fable and Sol and successor or future models for all cybersecurity purposes.

zkmontoday at 2:39 AM

What's wrong if the roles of USA and China are reversed in technology? Why does the rest of the world care? It's not as if USA has done a great good for the world, and China has evil intentions towards the world. Infact it is the opposite in the case of AI so far.

ArtRichardstoday at 5:56 AM

New, smaller models can outperform the previous generation's foundation models.

What if there's a way to extract the commodity of intelligence from smaller models?

I've seen for many use cases it's well enough. :)

mattastoday at 3:18 AM

"Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute."

It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.

The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.

nottorptoday at 7:15 AM

It's say Anthropic, Allegedly OpenAI...

softwaredougtoday at 2:12 AM

Haven’t we been in this “China is 3-6 months behind” for a while now (maybe up to a year? Longer?)

The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI story

Then Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.

It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.

golly_nedtoday at 1:41 AM

> I expect the inference market to grow much faster than training costs

This was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.

But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.

overfeedtoday at 1:56 AM

> [Anthropic/OpenAI] are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs. Second, intelligence isn’t in fact a perfect commodity, in part because applied intelligence makes itself smarter

Is he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.

softwaredougyesterday at 11:23 PM

> By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage.

I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].

If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.

But that's a big if we just don't know for sure.

1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...

holoduketoday at 7:40 AM

So happy that we have finally 2 countries playing the competitive game. No more secret deals between competitors. No real competition. A race to the bottom is always a good thing for consumers.

Jyaiftoday at 7:40 AM

In case it's useful for other non-native speakers:

I didn't know about the word "undergirding", and it turns out it's misused here. The much more common word "underlying" would have been more appropriate.

minrawsyesterday at 9:54 PM

Me I am, so very afraid of actually decently priced inference.

ggmtoday at 12:00 AM

A reminder any comment about risk FROM china, invites a "Tu Qoque" facing the other way. The paranoia here is probably fully symmetrical.

I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.

A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"

coretxtoday at 1:57 AM

The best model is the model that runs best on your hardware.

throwitaway222today at 4:36 AM

A company making a decision to allow use of chinese models is a company also choosing to send tons of various credentials to chinese model companies. These will just get scooped up, OpenAI and Anthropic can probably hack into anything at this point if they wanted to.

show 1 reply
hexatoryesterday at 11:00 PM

I'm worried that any ban on Chinese AI models might be an excuse to get mass surveillance.

show 1 reply
nuneztoday at 5:16 AM

I really enjoyed reading this.

This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.

Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)

I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.

Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.

ab_wahab01today at 3:59 AM

Honestly, as someone from a developing country, this shift is good for us. US frontier models are too expensive for us to use regularly. Chinese open-source models/subscriptions are really good to use.

nltoday at 12:05 AM

> because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?

Is this an assertion that is backed by evidence?

From the Elon/OpenAI trial:

> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”

https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...

ilamontyesterday at 9:58 PM

But it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum.

I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.

The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.

IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.

alizakitoday at 12:16 AM

There is no “Chinese LLM”. Each “lab” is distinct and their models behavior is as unique as those from OpenAI and Anthropic

show 1 reply
purplepatricktoday at 1:50 AM

Commenting wholesale on some folks who are asking for hard evidence. I cannot provide that either but can contribute some empirical data.

I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.

As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.

Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.

I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.

So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.

Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.

jdla1otoday at 3:37 AM

The future is SLMs and China is going there..

jmclnxyesterday at 10:15 PM

One thing I have not seen mentioned between Chinese AI vs US, population.

China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.

Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.

So I believe, China will end up owing AI.

show 2 replies
fellowniusmonkyesterday at 10:12 PM

The U.S. "executive" class is so obsessed with the "exploit" part of the explore/exploit cycle that it's very clear they are prematurely closing advancement. Better a little money and power for them now than a lot of money and power for their country/humanity.

This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.

You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.

show 1 reply
joshttoday at 1:36 AM

Someone (anyone!) get David Sacks on the horn and tell him to read this.

Havocyesterday at 11:37 PM

oh wow - hadn't realized they decided to opensource Qwen 3.8 Max. That's pretty big news.

NooneAtAll3yesterday at 10:06 PM

I don't understand the premise in the beginning

how is running servers supposed to be 0 cost, while running ai inferrence isn't?

show 2 replies
ChrisArchitecttoday at 6:10 AM

Related:

Ben Thompson is wrong: US frontier labs are right to be panicking

https://news.ycombinator.com/item?id=48982061

pupskippertoday at 12:55 AM

The fact that Anthropic has a model like Mythos means that counterpart countries like Russia and China are not far behind, if they haven't already developed something similar or better.

show 3 replies

🔗 View 24 more comments