← All Mac Studio models

Mac Studio M5 Max

The M5 Max Mac Studio runs local models at 460 to 614 GB/s of memory bandwidth with 36 to 128 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 128 GB you can hold roughly a 172B dense model at Q4. The Max is where the bus gets wide enough that model size, not bandwidth, becomes the thing you plan around.

By , Founder & Lead Engineer— Updated

Specifications

ChipApple M5 Max
CPU cores18
GPU cores32 or 40
Unified memory36, 48, 64, or 128 GB
Memory bandwidth460 to 614 GB/s
Neural Engine38 TOPS
Released2026
AvailabilitySold new by Apple

Memory bandwidth is faster than 75% of the Apple Silicon chips shipped in a Mac, against a 1228 GB/s peak.

Its memory ceiling is above 65% of them, against a 512 GB peak.

Pick your configuration

Every option Apple sells with this chip. The model list below recomputes against the one you pick.

GPU cores

Unified memory

Apple couples memory to the core count on this chip, so the options change with the bin above.

What each memory option runs

Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.

36 GB unified memory

31 GB usable for weights · 460 GB/s

6,120 of 6,563 models fit, and 5,827 of them run with headroom rather than as a squeeze.

Largest model at Q4
Chinese Mixtral 8x7B · 46.91B
Best all-round pick
Kimi K3 DSpark · Q8_0 · ~107 tok/s

48 GB unified memory

42 GB usable for weights · 614 GB/s

6,273 of 6,563 models fit, and 6,033 of them run with headroom rather than as a squeeze.

Largest model at Q4
Yi 34Bx2 MoE 60B DPO · 60.81B
Best all-round pick
Kimi K3 DSpark · Q8_0 · ~143 tok/s

64 GB unified memory

56 GB usable for weights · 614 GB/s

6,357 of 6,563 models fit, and 6,086 of them run with headroom rather than as a squeeze.

Largest model at Q4
Hy3 REAP 48e · 84.02B
Best all-round pick
Kimi K3 DSpark · Q8_0 · ~143 tok/s

128 GB unified memory

112 GB usable for weights · 614 GB/s

6,420 of 6,563 models fit, and 6,334 of them run with headroom rather than as a squeeze.

Largest model at Q4
Kimi K2 Thinking converted · 170.27B
Best all-round pick
Ornith 1.0 35B · Q8_0 · ~118 tok/s

What a 36 GB M5 Max Mac Studio can run

Every model in the database against this exact configuration, at 460 GB/s. Ratings and speeds are the same numbers the model pages show.

Showing 6563 of 6563 models

General · radixark · 2026-07-27

Q8_0Excellent
3.0 GB8% of RAM~107 tok/sEstimated2.25B params
Run with ToolPiper

General · openbmb · 2026-05-21

Q8_0Excellent
1.7 GB5% of RAM~223 tok/sEstimated1.08B params
Run with ToolPiper

General · farbodtavakkoli · 2026-06-17

Q8_0Excellent
10.8 GB30% of RAMBenchmark needed9.23B params
Run with ToolPiper

General · Liquid AI · 2026-05-28

Q8_0Excellent
9.9 GB28% of RAMBenchmark needed8.47B params
Run with ToolPiper

General · weiboai · 2026-06-12

Q8_0Excellent
3.9 GB11% of RAM~78 tok/sEstimated3.09B params
Run with ToolPiper

General · Liquid AI · 2026-07-28

Q8_0Excellent
3.5 GB10% of RAM~89 tok/sEstimated2.7B params
Run with ToolPiper

General · Liquid AI · 2026-06-24

Q8_0Excellent
0.8 GB2% of RAM~1,048 tok/sEstimated0.23B params
Run with ToolPiper

General · nanbeige · 2026-07-21

Q8_0Excellent
5.2 GB14% of RAM~58 tok/sEstimated4.17B params
Run with ToolPiper

General · ma7ee7 · 2026-07-30

Q8_0Excellent
4.0 GB11% of RAM~77 tok/sEstimated3.13B params
Run with ToolPiper

General · goekdeniz-guelmez · 2026-07-31

Q8_0Excellent
3.0 GB8% of RAM~106 tok/sEstimated2.27B params
Run with ToolPiper

General · ktruestory · 2026-05-28

Q8_0Excellent
1.7 GB5% of RAM~223 tok/sEstimated1.08B params
Run with ToolPiper

General · petrouil · 2026-07-23

Q8_0Excellent
3.4 GB9% of RAMBenchmark needed2.61B params
Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0Excellent
4.3 GB12% of RAM~71 tok/sEstimated3.4B params
Run with ToolPiper

General · internscience · 2026-07-13

Q8_0Excellent
5.6 GB15% of RAM~53 tok/sEstimated4.54B params
Run with ToolPiper

Multimodal · ibm-granite · 2026-04-16

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4B params
Run with ToolPiper

General · Liquid AI · 2026-03-31

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
1.5 GB4% of RAM~277 tok/sEstimated0.87B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
3.0 GB8% of RAM~106 tok/sEstimated2.27B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
1.5 GB4% of RAM~277 tok/sEstimated0.87B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
3.0 GB8% of RAM~106 tok/sEstimated2.27B params
Run with ToolPiper

General · nvidia · 2026-03-02

Q8_0Excellent
4.8 GB13% of RAM~63 tok/sEstimated3.83B params
Run with ToolPiper

General · cagrigungor · 2026-08-06

Q8_0Excellent
0.8 GB2% of RAM~892 tok/sEstimated0.27B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
5.7 GB16% of RAM~52 tok/sEstimated4.66B params
Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0Excellent
6.2 GB17% of RAM~47 tok/sEstimated5.12B params
Run with ToolPiper

General · deepreinforce-ai · 2026-06-21

Q8_0Excellent
9.7 GB27% of RAM~29 tok/sEstimated8.21B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
5.7 GB16% of RAM~52 tok/sEstimated4.66B params
Run with ToolPiper

Multimodal · Liquid AI · 2026-01-05

Q8_0Excellent
2.3 GB6% of RAM~151 tok/sEstimated1.6B params
Run with ToolPiper

General · lukebailey181pub · 2026-04-21

Q8_0Excellent
8.2 GB23% of RAM~35 tok/sEstimated6.91B params
Run with ToolPiper

General · Liquid AI · 2026-01-20

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · ai21labs · 2026-01-06

Q8_0Excellent
3.9 GB11% of RAMBenchmark needed3.03B params
Run with ToolPiper

Reasoning · openonerec · 2026-06-09

Q8_0Excellent
1.4 GB4% of RAM~301 tok/sEstimated0.8B params
Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2026-01-04

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-12-25

Q8_0Excellent
3.4 GB9% of RAM~94 tok/sEstimated2.57B params
Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0Excellent
3.4 GB9% of RAM~94 tok/sEstimated2.57B params
Run with ToolPiper

Multimodal · davidau · 2026-02-02

Q8_0Excellent
5.3 GB15% of RAM~56 tok/sEstimated4.3B params
Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0Excellent
9.4 GB26% of RAM~30 tok/sEstimated8B params
Run with ToolPiper

Chat · Liquid AI · 2026-01-06

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · inclusionai · 2026-02-09

Q8_0Excellent
18.6 GB52% of RAMBenchmark needed16.26B params
Run with ToolPiper

General · ravichandranj · 2026-02-13

Q8_0Excellent
7.7 GB21% of RAM~37 tok/sEstimated6.43B params
Run with ToolPiper

General · Liquid AI · 2025-10-28

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-10-07

Q8_0Excellent
9.8 GB27% of RAMBenchmark needed8.34B params
Run with ToolPiper

General · obliteratus · 2026-04-15

Q8_0Excellent
9.4 GB26% of RAM~30 tok/sEstimated8B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-10-22

Q8_0Excellent
3.8 GB11% of RAM~80 tok/sEstimated3B params
Run with ToolPiper

Reasoning · jackrong · 2026-03-16

Q8_0Excellent
11.3 GB31% of RAM~25 tok/sEstimated9.65B params
Run with ToolPiper

General · frontiersmind · 2026-08-03

Q8_0Excellent
1.2 GB3% of RAM~371 tok/sEstimated0.65B params
Run with ToolPiper

General · typhoon-ai · 2025-09-23

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

General · Alibaba · 2025-08-05

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

General · hmellor · 2025-07-22

Q8_0Excellent
1.9 GB5% of RAM~194 tok/sEstimated1.24B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
8.2 GB23% of RAMBenchmark needed6.94B params
Run with ToolPiper

General · nvidia · 2026-03-18

Q8_0Excellent
10.0 GB28% of RAM~28 tok/sEstimated8.49B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
4.3 GB12% of RAM~71 tok/sEstimated3.4B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0Excellent
1.0 GB3% of RAM~535 tok/sEstimated0.45B params
Run with ToolPiper

General · ibm-granite · 2026-04-16

Q8_0Excellent
9.8 GB27% of RAM~29 tok/sEstimated8.38B params
Run with ToolPiper

General · openonerec · 2025-12-30

Q8_0Excellent
2.9 GB8% of RAM~113 tok/sEstimated2.13B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
4.1 GB11% of RAM~76 tok/sEstimated3.19B params
Run with ToolPiper

General · bytedance · 2025-10-28

Q8_0Excellent
2.1 GB6% of RAM~168 tok/sEstimated1.43B params
Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0Excellent
2.3 GB6% of RAM~153 tok/sEstimated1.58B params
Run with ToolPiper

General · Liquid AI · 2025-09-22

Q8_0Excellent
3.4 GB9% of RAM~94 tok/sEstimated2.57B params
Run with ToolPiper

General · Liquid AI · 2025-09-30

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-08-22

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-08-25

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

General · NCAI · 2025-12-29

Q8_0Excellent
8.6 GB24% of RAMBenchmark needed7.25B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
11.3 GB31% of RAM~25 tok/sEstimated9.65B params
Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0Excellent
10.3 GB29% of RAM~27 tok/sEstimated8.79B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-26

Q8_0Excellent
11.3 GB31% of RAM~25 tok/sEstimated9.65B params
Run with ToolPiper

General · inclusionai · 2025-11-25

Q8_0Excellent
18.6 GB52% of RAMBenchmark needed16.26B params
Run with ToolPiper

General · ibm-granite · 2025-04-30

Q8_0Excellent
7.9 GB22% of RAMBenchmark needed6.67B params
Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0Excellent
1.8 GB5% of RAM~206 tok/sEstimated1.17B params
Run with ToolPiper

Reasoning · Microsoft · 2025-04-29

Q8_0Excellent
4.8 GB13% of RAM~63 tok/sEstimated3.84B params
Run with ToolPiper

General · allenai · 2025-11-18

Q8_0Excellent
8.6 GB24% of RAM~33 tok/sEstimated7.3B params
Run with ToolPiper

General · amd · 2025-05-17

Q8_0Excellent
2.2 GB6% of RAM~161 tok/sEstimated1.5B params
Run with ToolPiper

General · stefanruseti · 2025-06-04

Q8_0Excellent
1.9 GB5% of RAM~194 tok/sEstimated1.24B params
Run with ToolPiper

General · empero-ai · 2026-06-19

Q8_0Excellent
11.0 GB31% of RAM~26 tok/sEstimated9.41B params
Run with ToolPiper

General · allenai · 2026-01-28

Q8_0Excellent
8.8 GB24% of RAM~32 tok/sEstimated7.43B params
Run with ToolPiper

Chat · pearl-ai · 2026-02-26

Q8_0Excellent
9.5 GB26% of RAM~30 tok/sEstimated8.03B params
Run with ToolPiper

General · ibm-granite · 2025-10-07

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · lgai-exaone · 2025-07-11

Q8_0Excellent
1.9 GB5% of RAM~188 tok/sEstimated1.28B params
Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

General · menlo · 2025-06-25

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0Excellent
1.3 GB4% of RAM~326 tok/sEstimated0.74B params
Run with ToolPiper

General · goekdeniz-guelmez · 2026-07-31

Q8_0Excellent
11.3 GB31% of RAM~25 tok/sEstimated9.65B params
Run with ToolPiper

Multimodal · NCAI · 2025-12-29

Q8_0Excellent
9.0 GB25% of RAMBenchmark needed7.58B params
Run with ToolPiper

General · alibaba-nlp · 2026-03-31

Q8_0Excellent
9.6 GB27% of RAM~29 tok/sEstimated8.19B params
Run with ToolPiper

Multimodal · Alibaba · 2025-01-26

Q8_0Excellent
4.7 GB13% of RAM~64 tok/sEstimated3.75B params
Run with ToolPiper

Multimodal · google · 2026-05

Q8_0Excellent
13.8 GB38% of RAM~20 tok/sEstimated11.96B params
Run with ToolPiper

Reasoning · DeepSeek · 2025-05-29

Q8_0Excellent
9.6 GB27% of RAM~29 tok/sEstimated8.19B params
Run with ToolPiper

General · huggingfacetb · 2025-07-08

Q8_0Excellent
3.9 GB11% of RAM~78 tok/sEstimated3.08B params
Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0Excellent
2.5 GB7% of RAM~135 tok/sEstimated1.78B params
Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0Excellent
5.4 GB15% of RAM~55 tok/sEstimated4.41B params
Run with ToolPiper

General · huggingfacetb · 2025-06-19

Q8_0Excellent
3.9 GB11% of RAM~78 tok/sEstimated3.08B params
Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0Excellent
1.3 GB4% of RAM~321 tok/sEstimated0.75B params
Run with ToolPiper

General · allenai · 2025-09-12

Q8_0Excellent
8.6 GB24% of RAM~33 tok/sEstimated7.3B params
Run with ToolPiper

General · anton-hugging · 2026-02-06

Q8_0Excellent
9.0 GB25% of RAM~32 tok/sEstimated7.62B params
Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0Excellent
9.4 GB26% of RAM~30 tok/sEstimated8B params
Run with ToolPiper

Chat · xcuros · 2026-02-28

Q8_0Excellent
9.0 GB25% of RAM~32 tok/sEstimated7.62B params
Run with ToolPiper

General · pfnet · 2025-02-05

Q8_0Excellent
1.9 GB5% of RAM~187 tok/sEstimated1.29B params
Run with ToolPiper

General · bytedance-seed · 2025-04-09

Q8_0Excellent
11.0 GB30% of RAMBenchmark needed9.37B params
Run with ToolPiper

Chat · baseten · 2025-09-12

Q8_0Excellent
4.1 GB11% of RAM~75 tok/sEstimated3.21B params
Run with ToolPiper

General · sapientinc · 2026-05-17

Q8_0Excellent
1.8 GB5% of RAM~204 tok/sEstimated1.18B params
Run with ToolPiper

General · fableforge-ai · 2026-07-05

Q8_0Excellent
2.2 GB6% of RAM~156 tok/sEstimated1.54B params
Run with ToolPiper

General · viorikaai-org · 2026-07-05

Q8_0Excellent
0.6 GB2% of RAM~4,819 tok/sEstimated0.05B params
Run with ToolPiper

General · bananamind · 2026-07-17

Q8_0Excellent
0.5 GB1% of RAM~24,095 tok/sEstimated0.01B params
Run with ToolPiper

General · lgai-exaone · 2025-03-12

Q8_0Excellent
3.2 GB9% of RAM~100 tok/sEstimated2.41B params
Run with ToolPiper

General · maliosdark · 2026-07-09

Q8_0Excellent
0.6 GB2% of RAM~4,819 tok/sEstimated0.05B params
Run with ToolPiper

General · raidium · 2026-06-15

Q8_0Excellent
0.5 GB1% of RAM~12,048 tok/sEstimated0.02B params
Run with ToolPiper

General · shaungves · 2026-06-19

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

General · preparebuddy · 2026-06-02

Q8_0Excellent
3.9 GB11% of RAM~78 tok/sEstimated3.08B params
Run with ToolPiper

Embedding · taide · 2026-06-12

Q8_0Excellent
0.8 GB2% of RAM~803 tok/sEstimated0.3B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
1.3 GB4% of RAM~321 tok/sEstimated0.75B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
2.8 GB8% of RAM~119 tok/sEstimated2.03B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

Multimodal · Alibaba

Q8_0Excellent
5.5 GB15% of RAM~54 tok/sEstimated4.44B params
Run with ToolPiper

Multimodal · zai-org

Q8_0Excellent
2.0 GB6% of RAM~181 tok/sEstimated1.33B params
Run with ToolPiper

Chat · Alibaba · 2025-08-05

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

Multimodal · Alibaba

Q8_0Excellent
2.9 GB8% of RAM~113 tok/sEstimated2.13B params
Run with ToolPiper

General · prefeitura-rio

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

Multimodal · datalab-to

Q8_0Excellent
1.3 GB4% of RAM~349 tok/sEstimated0.69B params
Run with ToolPiper

Multimodal · Microsoft

Q8_0Excellent
5.1 GB14% of RAM~58 tok/sEstimated4.15B params
Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0Excellent
2.4 GB7% of RAM~140 tok/sEstimated1.72B params
Run with ToolPiper

Multimodal · openbmb

Q8_0Excellent
2.0 GB5% of RAM~185 tok/sEstimated1.3B params
Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0Excellent
1.2 GB3% of RAM~402 tok/sEstimated0.6B params
Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4.02B params
Run with ToolPiper

Multimodal · nanonets

Q8_0Excellent
4.7 GB13% of RAM~64 tok/sEstimated3.75B params
Run with ToolPiper

Multimodal · rednote-hilab

Q8_0Excellent
3.9 GB11% of RAM~79 tok/sEstimated3.04B params
Run with ToolPiper

Multimodal · Microsoft

Q8_0Excellent
5.8 GB16% of RAM~51 tok/sEstimated4.74B params
Run with ToolPiper

Chat · allenai · 2025-11-19

Q8_0Excellent
8.6 GB24% of RAM~33 tok/sEstimated7.3B params
Run with ToolPiper

Multimodal · moonshotai

Q8_0Excellent
18.8 GB52% of RAMBenchmark needed16.41B params
Run with ToolPiper

General · t-tech · 2025-12-22

Q8_0Excellent
9.6 GB27% of RAM~29 tok/sEstimated8.19B params
Run with ToolPiper

Multimodal · typhoon-ai

Q8_0Excellent
4.7 GB13% of RAM~64 tok/sEstimated3.75B params
Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0Excellent
9.5 GB26% of RAM~30 tok/sEstimated8.03B params
Run with ToolPiper

Multimodal · tencent

Q8_0Excellent
1.7 GB5% of RAMBenchmark needed1.12B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
5.3 GB15% of RAM~56 tok/sEstimated4.3B params
Run with ToolPiper

Chat · allenai · 2025-11-17

Q8_0Excellent
8.6 GB24% of RAM~33 tok/sEstimated7.3B params
Run with ToolPiper

Multimodal · dots-studio

Q8_0Excellent
3.9 GB11% of RAM~79 tok/sEstimated3.04B params
Run with ToolPiper

General · distil-labs

Q8_0Excellent
0.9 GB2% of RAM~688 tok/sEstimated0.35B params
Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0Excellent
9.0 GB25% of RAM~32 tok/sEstimated7.62B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
1.8 GB5% of RAM~199 tok/sEstimated1.21B params
Run with ToolPiper

Multimodal · raxcore-dev

Q8_0Excellent
3.0 GB8% of RAM~106 tok/sEstimated2.27B params
Run with ToolPiper

Multimodal · moonshotai

Q8_0Excellent
18.8 GB52% of RAMBenchmark needed16.41B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
5.2 GB15% of RAM~57 tok/sEstimated4.25B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
5.0 GB14% of RAM~60 tok/sEstimated4B params
Run with ToolPiper

Multimodal · phanviethoang1512

Q8_0Excellent
5.7 GB16% of RAM~52 tok/sEstimated4.66B params
Run with ToolPiper

General · jinaai

Q8_0Excellent
2.2 GB6% of RAM~156 tok/sEstimated1.54B params
Run with ToolPiper

Multimodal · ath-maas

Q8_0Excellent
1.4 GB4% of RAM~283 tok/sEstimated0.85B params
Run with ToolPiper

32-core vs 40-core GPU

Apple ties the memory bus to the bin on this chip: 460 GB/s at 32 cores and 614 GB/s at 40. That moves tokens per second. It does not move which models fit, because that is memory, not cores.

ConfigurationMemory bandwidthMemory optionsModels that fit
18-core CPU, 32-core GPU460 GB/s36 GBIdentical
18-core CPU, 40-core GPU614 GB/s48, 64, 128 GBIdentical

On openbuddy zero 56b v21.2 32k the 32-core generates about 12 tok/s and the 40-core about 16 tok/s, a 33% difference. Both hold the model at the same quantization.

Measured on the M5 Max

Nobody has submitted a benchmark on the M5 Max yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.

ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.

Why unified memory is the number that matters

On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M5 Max's 614 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 128 GB machine can hand almost all of that to a model with no copy across a bus.

The Studio exists for this workload. It carries the widest memory buses and the highest capacities Apple sells, and it runs at full clocks indefinitely. Buy the memory, not the cores: every extra GB raises what you can load, while the core count only moves throughput on models that already fit.

Common questions

Can the M5 Max Mac Studio run a 70B model?

Yes, at 128 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 128 GB of unified memory leaves about 112 GB for weights once macOS takes its share. At 36 GB it does not fit at any quantization worth running.

How much unified memory should I get with the M5 Max Mac Studio?

Memory is the only spec that changes what you can run at all. 36 GB holds about a 47B model at Q4; 128 GB holds about 172B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.

How fast are local LLMs on the M5 Max?

Token generation is bandwidth-bound, so M5 Max throughput scales with its 614 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M5 Max lands in the tens of tokens per second and a 70B model lands in the single digits.

Is the 40-core GPU worth it over the 32-core on the M5 Max?

For throughput, yes: Apple ties bandwidth to the bin here, so the 32-core runs at 460 GB/s and the 40-core at 614 GB/s, about 33% more. For fit, no: both bins hold exactly the same models, because that is set by memory rather than by cores.

Should I buy the M5 Max Mac Studio now or wait for the next one?

Buy on the memory you need today. Apple raises memory ceilings slowly and bandwidth in steps, and the M5 Max already holds about a 172B model at Q4. If your target model fits in 128 GB, waiting buys throughput rather than capability.

Run these models on your Mac Studio

ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.