The M5 Max Mac Studio runs local models at 460 to 614 GB/s of memory bandwidth with 36 to 128 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 128 GB you can hold roughly a 172B dense model at Q4. The Max is where the bus gets wide enough that model size, not bandwidth, becomes the thing you plan around.
Memory bandwidth is faster than 75% of the Apple Silicon chips shipped in a Mac, against a 1228 GB/s peak.
Its memory ceiling is above 65% of them, against a 512 GB peak.
Every option Apple sells with this chip. The model list below recomputes against the one you pick.
GPU cores
Unified memory
Apple couples memory to the core count on this chip, so the options change with the bin above.
Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.
6,120 of 6,563 models fit, and 5,827 of them run with headroom rather than as a squeeze.
6,273 of 6,563 models fit, and 6,033 of them run with headroom rather than as a squeeze.
6,357 of 6,563 models fit, and 6,086 of them run with headroom rather than as a squeeze.
6,420 of 6,563 models fit, and 6,334 of them run with headroom rather than as a squeeze.
Every model in the database against this exact configuration, at 460 GB/s. Ratings and speeds are the same numbers the model pages show.
Showing 6563 of 6563 models
General · radixark · 2026-07-27
General · openbmb · 2026-05-21
General · farbodtavakkoli · 2026-06-17
General · Liquid AI · 2026-05-28
General · weiboai · 2026-06-12
General · Liquid AI · 2026-07-28
General · Liquid AI · 2026-06-24
General · nanbeige · 2026-07-21
General · ma7ee7 · 2026-07-30
General · goekdeniz-guelmez · 2026-07-31
General · ktruestory · 2026-05-28
General · petrouil · 2026-07-23
General · ibm-granite · 2026-04-06
General · internscience · 2026-07-13
Multimodal · ibm-granite · 2026-04-16
General · Liquid AI · 2026-03-31
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
General · nvidia · 2026-03-02
General · cagrigungor · 2026-08-06
Multimodal · Alibaba · 2026-02-27
Multimodal · Google · 2026-03-02
General · deepreinforce-ai · 2026-06-21
Multimodal · Alibaba · 2026-02-27
Multimodal · Liquid AI · 2026-01-05
General · lukebailey181pub · 2026-04-21
General · Liquid AI · 2026-01-20
General · ai21labs · 2026-01-06
Reasoning · openonerec · 2026-06-09
General · Liquid AI · 2026-01-05
General · Liquid AI · 2026-01-04
General · Liquid AI · 2025-12-25
General · Liquid AI · 2026-01-05
Multimodal · davidau · 2026-02-02
Multimodal · Google · 2026-03-02
Chat · Liquid AI · 2026-01-06
General · inclusionai · 2026-02-09
General · ravichandranj · 2026-02-13
General · Liquid AI · 2025-10-28
General · Liquid AI · 2025-10-07
General · obliteratus · 2026-04-15
Multimodal · Liquid AI · 2025-10-22
Reasoning · jackrong · 2026-03-16
General · frontiersmind · 2026-08-03
General · typhoon-ai · 2025-09-23
General · Alibaba · 2025-08-05
General · hmellor · 2025-07-22
General · ibm-granite · 2025-09-16
General · nvidia · 2026-03-18
General · ibm-granite · 2025-09-16
Multimodal · Liquid AI · 2025-08-12
General · ibm-granite · 2026-04-16
General · openonerec · 2025-12-30
General · ibm-granite · 2025-09-16
General · bytedance · 2025-10-28
General · z-lab · 2026-01-04
Multimodal · Liquid AI · 2025-08-12
General · Liquid AI · 2025-09-22
General · Liquid AI · 2025-09-30
General · Liquid AI · 2025-08-22
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-08-25
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · NCAI · 2025-12-29
Multimodal · Alibaba · 2026-02-27
General · ibm-granite · 2026-04-06
Multimodal · Alibaba · 2026-02-26
General · inclusionai · 2025-11-25
General · ibm-granite · 2025-04-30
General · Liquid AI · 2025-07-10
Reasoning · Microsoft · 2025-04-29
General · allenai · 2025-11-18
General · amd · 2025-05-17
General · stefanruseti · 2025-06-04
General · empero-ai · 2026-06-19
General · allenai · 2026-01-28
Chat · pearl-ai · 2026-02-26
General · ibm-granite · 2025-10-07
General · lgai-exaone · 2025-07-11
General · Liquid AI · 2025-07-10
General · menlo · 2025-06-25
General · Liquid AI · 2025-07-10
General · goekdeniz-guelmez · 2026-07-31
Multimodal · NCAI · 2025-12-29
General · alibaba-nlp · 2026-03-31
Multimodal · Alibaba · 2025-01-26
Multimodal · google · 2026-05
Reasoning · DeepSeek · 2025-05-29
General · huggingfacetb · 2025-07-08
Reasoning · DeepSeek · 2025-01-20
General · Alibaba · 2025-09-23
General · huggingfacetb · 2025-06-19
General · Alibaba · 2025-09-23
General · allenai · 2025-09-12
General · anton-hugging · 2026-02-06
General · z-lab · 2026-01-04
Chat · xcuros · 2026-02-28
General · pfnet · 2025-02-05
General · bytedance-seed · 2025-04-09
Chat · baseten · 2025-09-12
General · sapientinc · 2026-05-17
General · fableforge-ai · 2026-07-05
General · viorikaai-org · 2026-07-05
General · bananamind · 2026-07-17
General · lgai-exaone · 2025-03-12
General · maliosdark · 2026-07-09
General · raidium · 2026-06-15
General · shaungves · 2026-06-19
General · preparebuddy · 2026-06-02
Embedding · taide · 2026-06-12
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
Multimodal · Alibaba
Multimodal · zai-org
Chat · Alibaba · 2025-08-05
Multimodal · Alibaba
General · prefeitura-rio
Multimodal · datalab-to
Multimodal · Microsoft
General · Alibaba · 2025-04-28
Multimodal · openbmb
General · Alibaba · 2025-04-28
General · Alibaba · 2025-04-28
Multimodal · nanonets
Multimodal · rednote-hilab
Multimodal · Microsoft
Chat · allenai · 2025-11-19
Multimodal · moonshotai
General · t-tech · 2025-12-22
Multimodal · typhoon-ai
Reasoning · DeepSeek · 2025-01-20
Multimodal · tencent
General · farbodtavakkoli
Chat · allenai · 2025-11-17
Multimodal · dots-studio
General · distil-labs
Reasoning · DeepSeek · 2025-01-20
General · farbodtavakkoli
Multimodal · raxcore-dev
Multimodal · moonshotai
General · farbodtavakkoli
General · ibm-granite
Multimodal · phanviethoang1512
General · jinaai
Multimodal · ath-maas
Apple ties the memory bus to the bin on this chip: 460 GB/s at 32 cores and 614 GB/s at 40. That moves tokens per second. It does not move which models fit, because that is memory, not cores.
| Configuration | Memory bandwidth | Memory options | Models that fit |
|---|---|---|---|
| 18-core CPU, 32-core GPU | 460 GB/s | 36 GB | Identical |
| 18-core CPU, 40-core GPU | 614 GB/s | 48, 64, 128 GB | Identical |
On openbuddy zero 56b v21.2 32k the 32-core generates about 12 tok/s and the 40-core about 16 tok/s, a 33% difference. Both hold the model at the same quantization.
Nobody has submitted a benchmark on the M5 Max yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.
ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.
What each step actually changes for local models, rather than which one is newer.
On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M5 Max's 614 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 128 GB machine can hand almost all of that to a model with no copy across a bus.
The Studio exists for this workload. It carries the widest memory buses and the highest capacities Apple sells, and it runs at full clocks indefinitely. Buy the memory, not the cores: every extra GB raises what you can load, while the core count only moves throughput on models that already fit.
Yes, at 128 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 128 GB of unified memory leaves about 112 GB for weights once macOS takes its share. At 36 GB it does not fit at any quantization worth running.
Memory is the only spec that changes what you can run at all. 36 GB holds about a 47B model at Q4; 128 GB holds about 172B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.
Token generation is bandwidth-bound, so M5 Max throughput scales with its 614 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M5 Max lands in the tens of tokens per second and a 70B model lands in the single digits.
For throughput, yes: Apple ties bandwidth to the bin here, so the 32-core runs at 460 GB/s and the 40-core at 614 GB/s, about 33% more. For fit, no: both bins hold exactly the same models, because that is set by memory rather than by cores.
Buy on the memory you need today. Apple raises memory ceilings slowly and bandwidth in steps, and the M5 Max already holds about a 172B model at Q4. If your target model fits in 128 GB, waiting buys throughput rather than capability.
ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.