I know. It hurts. But I have the feeling that you get what you pay for.
The purchase cost of H100 or B200 systems with comparable VRAM is a one order of magnitude higher. Although I can only guess how much lower the token/sec output of the Mac Studio will be. Probably 2-3 magnitudes lower?
While a cluster has to work with many users simultanously, and is a good investment for a company, perhaps the Mac Studio will be a good use case for a personal larger LLM deployment configuration.
Perhaps someone has the token/sec numbers for larger models running on older Mac Studios?
Back of the envelope is that compute doesn't matter for inference, only memory access speed matters. Tokens/sec is going to be in the ballpark of how much time it takes for the compute element to read the entire model. So you can give GBtok/sec (sec = seconds, GB gigabyte size for the chosen quantization)
Q8 is a nice quant for this calculation, since 1 byte = 1 weight, so Qwen 27B Q8 will be the tok/secGB value divided by 27, and Deepseek V4 Flash 162B, by 162.
NVIDIA B200: 8,000 GBtok/sec
NVIDIA H100: 3,350 GBtok/sec
NVIDIA A100: 2,039 GBtok/sec
NVIDIA RTX 5090: 1,792 GBtok/sec (deepseek only fits at Q2 or less)
Apple M5 Ultra: 1,200 GBtok/sec
NVIDIA RTX 5060 Ti: 448 GBtok/sec (deepseek doesn't even fit at Q1)
Apple M5 Pro: 307 GBtok/sec (can do deepseek only at Q4 or less, at maxed out sped)
Apple M1 Pro: 200 GBtok/sec (can do deepseek only at Q2 or less, at maxed out sped)
Apple M6: 170 GBtok/sec (can do deepseek only at Q2 or less, at maxed out sped)
So a base model apple M6, Qwen3.8 Q8 = 170/26 tok/s = 6 tok/s (or 12 tok/s at Q4), and a B200 will do ~40 times that, or 240 tok/s. Which is kind of sad as a base model M1 pro will beat it comfortably despite 6 years of chip advancements. To add insult to injury an M1 pro ... is cheaper secondhand.
512GB memory available in October. Stellar move, Apple.
$18,299 ... sigh
I know. It hurts. But I have the feeling that you get what you pay for.
The purchase cost of H100 or B200 systems with comparable VRAM is a one order of magnitude higher. Although I can only guess how much lower the token/sec output of the Mac Studio will be. Probably 2-3 magnitudes lower?
While a cluster has to work with many users simultanously, and is a good investment for a company, perhaps the Mac Studio will be a good use case for a personal larger LLM deployment configuration.
Perhaps someone has the token/sec numbers for larger models running on older Mac Studios?
Back of the envelope is that compute doesn't matter for inference, only memory access speed matters. Tokens/sec is going to be in the ballpark of how much time it takes for the compute element to read the entire model. So you can give GBtok/sec (sec = seconds, GB gigabyte size for the chosen quantization)
Q8 is a nice quant for this calculation, since 1 byte = 1 weight, so Qwen 27B Q8 will be the tok/secGB value divided by 27, and Deepseek V4 Flash 162B, by 162.
NVIDIA B200: 8,000 GBtok/sec
NVIDIA H100: 3,350 GBtok/sec
NVIDIA A100: 2,039 GBtok/sec
NVIDIA RTX 5090: 1,792 GBtok/sec (deepseek only fits at Q2 or less)
Apple M5 Ultra: 1,200 GBtok/sec
NVIDIA RTX 5060 Ti: 448 GBtok/sec (deepseek doesn't even fit at Q1)
Apple M5 Pro: 307 GBtok/sec (can do deepseek only at Q4 or less, at maxed out sped)
Apple M1 Pro: 200 GBtok/sec (can do deepseek only at Q2 or less, at maxed out sped)
Apple M6: 170 GBtok/sec (can do deepseek only at Q2 or less, at maxed out sped)
So a base model apple M6, Qwen3.8 Q8 = 170/26 tok/s = 6 tok/s (or 12 tok/s at Q4), and a B200 will do ~40 times that, or 240 tok/s. Which is kind of sad as a base model M1 pro will beat it comfortably despite 6 years of chip advancements. To add insult to injury an M1 pro ... is cheaper secondhand.