The M4 Max MacBook Pro 16" runs local models at 410 to 546 GB/s of memory bandwidth with 36 to 128 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 128 GB you can hold roughly a 172B dense model at Q4. The Max is where the bus gets wide enough that model size, not bandwidth, becomes the thing you plan around.
Apple no longer sells this configuration new. It stays fully evaluated here because the used market is where most of its local AI value now sits.
Memory bandwidth is faster than 72% of the Apple Silicon chips shipped in a Mac, against a 819 GB/s peak.
Its memory ceiling is above 67% of them, against a 512 GB peak.
Every option Apple sells with this chip. The model list below recomputes against the one you pick.
GPU cores
Unified memory
Apple couples memory to the core count on this chip, so the options change with the bin above.
Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.
3,417 of 3,641 models fit, and 3,252 of them run with headroom rather than as a squeeze.
3,503 of 3,641 models fit, and 3,351 of them run with headroom rather than as a squeeze.
3,541 of 3,641 models fit, and 3,385 of them run with headroom rather than as a squeeze.
3,582 of 3,641 models fit, and 3,532 of them run with headroom rather than as a squeeze.
Every model in the database against this exact configuration, at 410 GB/s. Ratings and speeds are the same numbers the model pages show.
Showing 3641 of 3641 models
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-27
Multimodal · Alibaba · 2026-02-27
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Reasoning · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Reasoning · Liquid AI · 2025-11-28
General · ibm-granite · 2025-09-16
Chat · Liquid AI · 2025-11-28
Chat · Liquid AI · 2025-11-28
General · ibm-granite · 2025-09-16
Reasoning · jackrong · 2026-03-16
Chat · Liquid AI · 2025-11-28
General · NCAI · 2025-12-29
Reasoning · HuggingFace · 2025-07-08
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Multimodal · NCAI · 2025-12-29
Multimodal · Alibaba · 2026-02-27
Multimodal · Google · 2025-07-30
Multimodal · Google · 2025-06-25
Multimodal · Alibaba · 2026-02-26
Multimodal · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
General · lgai-exaone · 2025-03-12
Multimodal · google · 2026-05
General · LG AI · 2025-07-15
General · raidium · 2026-06-15
General · NCAI · 2025-12-29
Embedding · taide · 2026-06-12
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
Multimodal · Google · 2025-07-30
General · Alibaba · 2025-04-27
Multimodal · zai-org
Multimodal · Alibaba
Multimodal · Microsoft
General · Alibaba
General · Alibaba
Reasoning · NVIDIA · 2025-06-01
Multimodal · openbmb
Reasoning · DeepSeek
Multimodal · rednote-hilab
General · DeepSeek
Multimodal · datalab-to
General · ibm-granite
Multimodal · moonshotai
General · openbmb
Multimodal · rednote-hilab
General · hmellor
General · distil-labs
Multimodal · nanonets
General · farbodtavakkoli
General · Liquid AI
General · nanonets
General · nanbeige
Multimodal · ibm-granite
General · ibm-granite
General · Google
Multimodal · moonshotai
Multimodal · lkhl
Multimodal · typhoon-ai
General · jinaai
General · Alibaba
General · kristaller486
Reasoning · typhoon-ai
General · Upstage
General · zstanjj
General · ibm-granite
General · nvidia
General · weiboai
General · x-izhang
General · openai
General · paddlepaddle
General · twinkle-ai
General · Liquid AI
General · pfnet
General · openbmb
General · ibm-granite
General · baidu
General · jetbrains
General · amd
General · bytedance-seed
General · arcee-ai
General · adamlucek
Coding · shahriarferdoush
General · ahczhg
General · abaryan
General · etherll
General · stanford-oval
General · jetbrains
General · flowaicom
General · bllossom
General · farbodtavakkoli
General · z-lab
Reasoning · khazarai
General · kamilamila
General · radheneev
General · getonit
General · paddlepaddle
General · moonshotai
General · jangq-ai
Reasoning · jackrong
General · agentica-org
General · NousResearch
Reasoning · nvidia
General · inference-net
General · novaciano
General · dmusingu
Reasoning · nvidia
General · kgrabko
General · ordenwills
General · smcleish
General · carsenk
General · dealignai
General · ibm-granite
Coding · ibm-granite
Reasoning · iffyuan
Reasoning · nvidia
General · zero-point-ai
General · stanfordaimi
General · trillionlabs
General · Microsoft
General · dealignai
General · huihui-ai
General · ibm-granite
General · openbmb
General · ibm-granite
General · osaurusai
General · pyoakum
General · ibm-granite
General · anakin87
General · redix
Apple ties the memory bus to the bin on this chip: 410 GB/s at 32 cores and 546 GB/s at 40. That moves tokens per second. It does not move which models fit, because that is memory, not cores.
| Configuration | Memory bandwidth | Memory options | Models that fit |
|---|---|---|---|
| 14-core CPU, 32-core GPU | 410 GB/s | 36 GB | Identical |
| 16-core CPU, 40-core GPU | 546 GB/s | 48, 64, 128 GB | Identical |
On openbuddy zero 56b v21.2 32k the 32-core generates about 11 tok/s and the 40-core about 14 tok/s, a 33% difference. Both hold the model at the same quantization.
Nobody has submitted a benchmark on the M4 Max yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.
ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.
What each step actually changes for local models, rather than which one is newer.
On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M4 Max's 546 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 128 GB machine can hand almost all of that to a model with no copy across a bus.
Apple stopped selling this one, which is exactly why it is interesting. The 16-inch chassis has the most thermal headroom Apple ships in a laptop, so sustained token throughput stays close to the burst figure. A used M4 Max at 128 GB still gives you 546 GB/s and a hard 172B ceiling, and neither number degrades with age the way a battery does.
Yes, at 128 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 128 GB of unified memory leaves about 112 GB for weights once macOS takes its share. At 36 GB it does not fit at any quantization worth running.
Memory is the only spec that changes what you can run at all. 36 GB holds about a 47B model at Q4; 128 GB holds about 172B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.
Token generation is bandwidth-bound, so M4 Max throughput scales with its 546 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M4 Max lands in the tens of tokens per second and a 70B model lands in the single digits.
For throughput, yes: Apple ties bandwidth to the bin here, so the 32-core runs at 410 GB/s and the 40-core at 546 GB/s, about 33% more. For fit, no: both bins hold exactly the same models, because that is set by memory rather than by cores.
For inference, the specs that matter do not age: 546 GB/s and up to 128 GB of unified memory are the same numbers today as they were in 2024. A used M4 Max at the top memory option usually beats a new base-tier machine at the same price on both. Check the battery and the display, not the silicon.
ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.