It's impressive and yet it surely depends what you're doing with it? 50% better performance isn't going to make a local LLM feel quick and yet so many other aspects of computing are quite fast anyhow. I can browse the web fairly comfortably on a Raspberry Pi and it's a bit slow but manageable.
If your goal is web browsing you will want faster single core performance. Unfortunately that means Apple M series. If your goal is local LLM, I’m afraid a several-year-old GPU will smoke the fastest CPU available today.
IIRC, compared to GPUs, M-series is currently still stuck in memory bandwidths from around 2016. M7 might catch up to 2019 or so. So GPUs will be better for LLMs for a pretty decent while.
How so? I don't know of any other mainstream platform with >1TB/s memory bandwidth. Personally I don't want to deal with macOS but between the memory bandwidth and out of box Thunderbolt networking it's hard to argue that Apple doesn't have a couple significant advantages over the current alternatives.
> I don't know of any other mainstream platform with >1TB/s memory bandwidth.
I mean, RTX 5080 has nearly 1TB/s, 5090 has nearly 2TB/s. Maybe you are talking about CPUs / unified memory platforms? I agree nobody else does it better. But for LLMs, GPUs can still be significantly faster than even the most advanced Apple silicon on the planet. TTFT in particular is super inferior with Apple, for now.
That's probably also the reason Apple had to reluctantly give into Nvidia servers for the initial rollout of Siri AI, though they claim to use trusted computing extensions to reach an acceptable level of privacy. (I do not trust that nearly as much as the Apple Silicon nodes)
It'll be amazing five years or whatever down the line to see Apple reaching those figures. They seem to be heading in that direction lately.
I suppose it depends what models, but even heavily quantized the mid tier local models need >100GB of memory so unified memory platforms are the only option I would consider a mainstream option (doesn't require special order etc). If you want to do it in two slots that's two Blackwell RTX 6000s which is around $50k list just for the cards, and if you want to do it in four+ slots that's outside the realm of normal desktop PCs. Right now 2x DGX Spark, 2x Ryzen 395, or 1x Mac Studio are the three configurations that get to 256GB of reasonably fast memory that you can order online and plug in, for around $8k to $12k depending on the details.