logoalt Hacker News

ma2kxtoday at 6:11 PM5 repliesview on HN

It would be interesting to know how many optimizations of the Chinese models were incorporated back into Claude and Codex.


Replies

janalsncmtoday at 10:17 PM

Many of the Chinese optimizations are public because they are published by the Chinese labs themselves. Hard to say what the labs are doing of course.

ekiddtoday at 9:57 PM

DeepSeek has published some really good papers. Lately they're pushing really hard for dramatically cheaper serving costs.

largbaetoday at 9:13 PM

I may be remembering wrong, but reasoning was first demonstrated by DeepSeek. Edit: I am indeed remembering wrong, seems o1 was first.

show 1 reply
Frost1xtoday at 7:39 PM

To some degree even unintentionally this is all inevitable. I know this isn’t what you’re taking about, but it’s a similar line of thought. These models are consuming public free information, and they’re all also producing free public information, so they’re all pissing and drinking into same pool.

If we go down the line of dead internet theory which I’m becoming more convinced of these days, the volume of information that’s not necessarily original or extracted from reality and just interpolated and extrapolated from existing information in different ways by LLMs will greatly outnumber human information coming up.

In which case these models should.. start to converge on the same data I imagine, with slightly different behaviors within that. One big generative orgy feedback loop.

lenerdenatortoday at 8:35 PM

Let me put it this way: if they're not ripping off all that they can, their investors need to shitcan the leadership.