I did the math the other day, and the cost of the mac versus the token rate that you get off of it for a decent model would take you over 70 years of 100% 24/7/365 usage to pay for itself versus just using hosted inference.
They are wildly poor options for local AI models, at least on a cost basis. If you don’t care about spending $10,000 to run a halfway okay model at 100x the cost, then go for it.
Does that include energy usage (which I would think Apple being in the top 5 of being energy efficient) and just for AI?
I’m sure most would be doing more than just hosting and running a single model.
You could include energy usage in the equation, but if you’re buying a $10,000 Mac Studio to edit your PDFs and watch YouTube, then that’s your own decision. That’s also not what we’re talking about here.
The Mac is still going to be incredibly energy efficient.
Let’s break it down for Qwen 3.5 MoE.
M4 Ultra Mac Studio:
22t/s
~ 210Wh
9.55 watts per token/s
B300 GPU:
~ 2160t/s per GPU
~ 1750Wh (14Kwh 8 GPU System)
0.88 watts per tokens/s
The Mac Studio comes out using ~ 11-12x more energy per token. Making it extremely energy inefficient at this task in comparison.
Not saying it wasn’t implied but didn’t offer it in my original comment but yeah, that’s what I’m talking about. It’s a computer, PC, desktop, workstation…… mostly like someone would be doing other things instead of just AI. You might be talking about just AI on the Mac Studio. To be fair, neither of us clarified our standings, till I ask further questions based on on my original comment. So you’re not wrong, just not what I was getting at. Thanks for sharing though.
Well local models are getting better. Open weight models are a real thread to their business. OpenAI’s prices certainly wouldn’t stay that low if they were the only player. Their prices are not profitable at the moment anyway.
Have you used open weight models for serious jobs? Ones that you can actually run effectively on a Mac studio? (Not deepseek v4 pro, not Kimi K3, not GLM 5.3)
They work alright, at best, sometimes.
And the ones you can’t run effectively, locally, like Kimi K3 are considerably better at demanding tasks like software engineering. But even then they still suck at that job compared to frontier Anthropic and OpenAI models. I’ve been building with all of the above, and the open weight models just overall suck at serious high demand software workloads.
No, no, they have not.
I did the math the other day, and the cost of the mac versus the token rate that you get off of it for a decent model would take you over 70 years of 100% 24/7/365 usage to pay for itself versus just using hosted inference.
They are wildly poor options for local AI models, at least on a cost basis. If you don’t care about spending $10,000 to run a halfway okay model at 100x the cost, then go for it.
Does that include energy usage (which I would think Apple being in the top 5 of being energy efficient) and just for AI? I’m sure most would be doing more than just hosting and running a single model.
You could include energy usage in the equation, but if you’re buying a $10,000 Mac Studio to edit your PDFs and watch YouTube, then that’s your own decision. That’s also not what we’re talking about here.
The Mac is still going to be incredibly energy efficient.
Let’s break it down for Qwen 3.5 MoE.
The Mac Studio comes out using ~ 11-12x more energy per token. Making it extremely energy inefficient at this task in comparison.
Not saying it wasn’t implied but didn’t offer it in my original comment but yeah, that’s what I’m talking about. It’s a computer, PC, desktop, workstation…… mostly like someone would be doing other things instead of just AI. You might be talking about just AI on the Mac Studio. To be fair, neither of us clarified our standings, till I ask further questions based on on my original comment. So you’re not wrong, just not what I was getting at. Thanks for sharing though.
Well local models are getting better. Open weight models are a real thread to their business. OpenAI’s prices certainly wouldn’t stay that low if they were the only player. Their prices are not profitable at the moment anyway.
They really aren’t.
Have you used open weight models for serious jobs? Ones that you can actually run effectively on a Mac studio? (Not deepseek v4 pro, not Kimi K3, not GLM 5.3)
They work alright, at best, sometimes.
And the ones you can’t run effectively, locally, like Kimi K3 are considerably better at demanding tasks like software engineering. But even then they still suck at that job compared to frontier Anthropic and OpenAI models. I’ve been building with all of the above, and the open weight models just overall suck at serious high demand software workloads.
I just tried some local llm on my gaming rig, sufficient for chat and images.
But buying the rig just for this is most likely not worth it.