logoalt Hacker News

Getting 50 GB/S Back from the Apple Neural Engine

76 pointsby eilnlast Thursday at 12:16 AM13 commentsview on HN

Comments

VladVladikoffyesterday at 11:05 PM

This website hijacked my back button during a simple page load. You should fix that, it’s not an acceptable way to behave.

show 3 replies
bee_rideryesterday at 11:28 PM

Nice investigation.

It is always surprising to me when a nice round number like 1MiB results in the “bad performance” configuration (although it happens).

Are you sure erratum is the right word in this context? I usually see it used to describe the notice that a document has an error in it.

eilnlast Thursday at 12:16 AM

RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.

Neywinyyesterday at 10:41 PM

Just checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL?

thenewwazooyesterday at 11:41 PM

"Apparently the memory controller's throughput has a dominant harmonic with wavelength 2048 in tensor-dimension space."

That got a laugh out of me.

RantyDaveyesterday at 11:27 PM

Ummm, wow. That's really bad.