M5 Ultra Mac Studio Review
macstories.net205 points by piotrgrabowski 7 hours ago
205 points by piotrgrabowski 7 hours ago
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:
Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
M3 Ultra 31 23.5 20 15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
can confirm.
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
How are you deciding which work to send to the 5090 vs a frontier model, or making the two work together nicely?
Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.
Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
> it's like the worst ADHD team member and requires constant supervision
Perhaps consider some non-offensive language for your comparison?
Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
Personally I think the price is way too high right now. It’s a power hungry gaming GPU. The efficient single card equivalent would be a 4500 Blackwell which launched at about $3500. Or you could get a 9700 32GB or an Arc B70 for well under $2k, today. You only buy a 5090 if you want absolute speed.
32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.
People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.
Exactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1.
My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."
This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
Once I get some kind of settlement after getting beaten up by a cop my first purchase will be some RTX Pro 6000s.
A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
> for me at least a GPU is completely useless for anything but being a token generator.
No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.
> No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.
Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.
A GPU would be better-off attached to your Mac in an eGPU enclosure. There is not a single Apple Silicon GPU on the market that leads the industry in prefill, decode or power efficiency.
But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.
Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
same reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).
If you seriously think apple products are nothing but a status item, you're deluding yourself and probably have been for decades.
This 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
So if it isn't a comparative ability, now it's a power cost issue? This reads like goal post moving.
Oh, it's absolutely both. The power you waste waiting for TFTT on prefill will absolutely compound at the "medium sized labs and businesses" scale.
If you seriously think apple cares about anything other than cell phones, you're deluding yourself and probably have been for decades.
...did you mean profit? I don't think they're manufacturing iphones just on the hope they delight you. This is also true of Google et al.
I don't get these weird parasocial emotional attachments/beefs people have with brands. Talk to a therapist.
brother my point is they don't care about their product offerings outside of their phones. this post/thread is about one of their product offerings which is not a phone which is inferior to their competitors'. simple.
My M1 Pro MBP is 6 years old and continues to be the best computer I own, so if that’s Apple not trying, god help everybody else once they do.