logoalt Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

278 pointsby bmulhollandtoday at 2:06 PM190 commentsview on HN

https://www.bloomberg.com/news/articles/2026-08-25/openai-cl..., https://archive.ph/yCTrr


Comments

mchusmatoday at 7:44 PM

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves.

For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.

While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.

I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.

show 7 replies
corfordtoday at 8:53 PM

These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be

show 4 replies
epistasistoday at 6:47 PM

It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.

One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)

show 1 reply
frabonifacetoday at 7:11 PM

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

show 8 replies
anthonypasqtoday at 6:38 PM

Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.

show 8 replies
tecoholictoday at 11:08 PM

The reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.

jimmySixDOFtoday at 6:52 PM

I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al

show 9 replies
lelanthrantoday at 7:36 PM

This means that they're going to want to IPO soon - this is good news for investors + they need the capital.

show 1 reply
luciana1utoday at 10:43 PM

everyone's silicon beats everyone else's benchmarks until it has to run someone's actual production workload. the real test is six months of your own inference traffic, not a vendor's chart.

m4rtinktoday at 11:17 PM

So this will make GPUs and associated affordable for people, rigjt ?

show 1 reply
bjournetoday at 11:38 PM

The article is a bit naive:

> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.

How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?

ChoosesBarbecuetoday at 5:54 PM

This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.

show 1 reply
thebeardisredtoday at 7:16 PM

All of these words spilled and no mention of the ISA.

show 2 replies
a2ff6eeb0today at 10:09 PM

Sounds like a great way to get deals out of Nvidia.

throwaw12today at 7:20 PM

Competition is good for all of us, we will get better and faster chips.

Or at least Nvidia GPUs will become slightly cheaper for regular consumers again

show 4 replies
danielovichdktoday at 8:04 PM

I guess special hardware is the new moat in AI.

Maybe the money will still flow into this industry after all

empath75today at 7:04 PM

When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.

Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.

show 2 replies
einpoklumtoday at 8:59 PM

I hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.

show 1 reply
calldacopsidgaftoday at 10:57 PM

Any article that features Sam's fucking creepy face should be marked with a jumpscare warning

mkw5053today at 9:56 PM

Warning, this is a long comment! (I’m trying to stick to sourced facts here and not overstate what they mean)

I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw

I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.

So, I started digging while waiting for various day-job inference calls to return, ha.

Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]

In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]

Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.

The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]

And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]

I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.

The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.

[1] https://www.dwarkesh.com/p/dylan-jon

[2] https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-w...

[3] https://www.theinformation.com/articles/dylan-patel-semianal...

[4] https://www.latent.space/p/dylanpatel-cooking

[5] https://www.dwarkesh.com/p/dylan-patel-3

[6] https://www.theinformation.com/briefings/exclusive-semianaly...

[7] https://news.ycombinator.com/item?id=31065646

show 2 replies
simianwordstoday at 7:14 PM

How can OpenAI mass produce this chip at scale more economically than Nvidia which has experience in the supply chain and scale efficiencies to do it efficiently?

show 4 replies
0xbadcafebeetoday at 7:18 PM

Story says they're power limited. That's half-true. Actually they're water-limited. To generate power, you need water. To cool chips, you need water. If you try to use less water on one side, you need more water on the other side (it's physics ya'll, making and using energy generates heat which requires dissipation). The world's freshwater is diminishing while also being consumed at an alarming rate. The future AI oligarchs are whoever controls the most water.

The other side of the conversation is the idea that large models in DCs on custom silicon is the future. Maybe for enterprise? But consumers will eventually (10 yrs) have affordable hardware designed to run crazy-good local models (more RAM + higher bandwidth). That will take pressure off of datacenters, but also reduce AI profits, and move that money to consumer chip/device makers. Apple is once again the biggest winner. Nvidia consumer chips might get cheaper, but nerfed, to encourage datacenter use where they make more money. I'm hoping AMD can stop being terrible at software so that when we finally have their better hardware we can actually use it.

show 2 replies
Alien1Beingtoday at 8:38 PM

WARNING AI HYPE

varispeedtoday at 6:42 PM

Why they don't research how to make their own RAM and they have to buy it from the common market?

They should GTFO with this crap.

Create barriers to computing for ordinary people while milking businesses for tokens.

show 2 replies
LarsDu88today at 7:12 PM

Well Sam Altman finally has built a moat against Chinese open weight AI. Well done. But what will this mean for Cerebras?

I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq

show 5 replies