Getting 50 GB/S Back from the Apple Neural Engine

eiln.github.io

83 points by eiln 3 days ago


nelsonfigueroa - 4 minutes ago

This goes way over my head and I don't understand most of it lol. I noticed you're still in the middle of getting your B.S. degree and you're already writing things like this...amazing.

VladVladikoff - 3 hours ago

This website hijacked my back button during a simple page load. You should fix that, it’s not an acceptable way to behave.

anemll - 12 minutes ago

Not all systems affected M1 and M5MAX are OK

https://x.com/anemll/status/2098454204478366132?s=20

bee_rider - 3 hours ago

Nice investigation.

It is always surprising to me when a nice round number like 1MiB results in the “bad performance” configuration (although it happens).

Are you sure erratum is the right word in this context? I usually see it used to describe the notice that a document has an error in it.

eiln - 3 days ago

RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.

Neywiny - 4 hours ago

Just checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL?

thenewwazoo - 3 hours ago

"Apparently the memory controller's throughput has a dominant harmonic with wavelength 2048 in tensor-dimension space."

That got a laugh out of me.

RantyDave - 3 hours ago

Ummm, wow. That's really bad.