The M5 Pro MacBook Pro 14" runs local models at 307 GB/s of memory bandwidth with 24 to 64 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 64 GB you can hold roughly a 85B dense model at Q4. The Pro roughly doubles the base chip's memory bus. That moves mid-size models from usable to comfortable without moving you into desktop money.
Memory bandwidth is faster than 50% of the Apple Silicon chips shipped in a Mac, against a 1228 GB/s peak.
Its memory ceiling is above 45% of them, against a 512 GB peak.
Every option Apple sells with this chip. The model list below recomputes against the one you pick.
GPU cores
Unified memory
Apple couples memory to the core count on this chip, so the options change with the bin above.
Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.
6,032 of 6,563 models fit, and 5,271 of them run with headroom rather than as a squeeze.
6,273 of 6,563 models fit, and 6,033 of them run with headroom rather than as a squeeze.
6,357 of 6,563 models fit, and 6,086 of them run with headroom rather than as a squeeze.
Every model in the database against this exact configuration, at 307 GB/s. Ratings and speeds are the same numbers the model pages show.
Showing 6563 of 6563 models
General · radixark · 2026-07-27
General · openbmb · 2026-05-21
General · farbodtavakkoli · 2026-06-17
General · Liquid AI · 2026-05-28
General · Liquid AI · 2026-07-28
General · Liquid AI · 2026-06-24
General · goekdeniz-guelmez · 2026-07-31
General · ktruestory · 2026-05-28
General · petrouil · 2026-07-23
General · weiboai · 2026-06-12
General · ma7ee7 · 2026-07-30
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
General · ibm-granite · 2026-04-06
General · internscience · 2026-07-13
Multimodal · ibm-granite · 2026-04-16
General · Liquid AI · 2026-03-31
General · nanbeige · 2026-07-21
General · cagrigungor · 2026-08-06
Multimodal · Alibaba · 2026-02-27
Multimodal · Alibaba · 2026-02-27
Multimodal · Liquid AI · 2026-01-05
General · Liquid AI · 2026-01-20
Reasoning · openonerec · 2026-06-09
General · Liquid AI · 2026-01-05
General · Liquid AI · 2026-01-04
General · Liquid AI · 2025-12-25
General · Liquid AI · 2026-01-05
Multimodal · Google · 2026-03-02
General · Liquid AI · 2025-10-28
General · ai21labs · 2026-01-06
General · nvidia · 2026-03-02
General · frontiersmind · 2026-08-03
Multimodal · davidau · 2026-02-02
Chat · Liquid AI · 2026-01-06
General · ibm-granite · 2025-09-16
Multimodal · Liquid AI · 2025-08-12
General · openonerec · 2025-12-30
General · bytedance · 2025-10-28
General · Liquid AI · 2025-10-07
Multimodal · Liquid AI · 2025-08-12
Multimodal · Liquid AI · 2025-10-22
General · Liquid AI · 2025-09-22
General · Liquid AI · 2025-09-30
General · Liquid AI · 2025-08-22
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-08-25
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · NCAI · 2025-12-29
General · typhoon-ai · 2025-09-23
General · hmellor · 2025-07-22
General · ibm-granite · 2025-04-30
General · Liquid AI · 2025-07-10
General · ibm-granite · 2025-09-16
General · amd · 2025-05-17
General · stefanruseti · 2025-06-04
General · ibm-granite · 2025-09-16
General · lgai-exaone · 2025-07-11
General · Liquid AI · 2025-07-10
General · Liquid AI · 2025-07-10
Multimodal · NCAI · 2025-12-29
General · Alibaba · 2025-08-05
General · Alibaba · 2025-09-23
Reasoning · Microsoft · 2025-04-29
General · pfnet · 2025-02-05
General · bytedance-seed · 2025-04-09
General · sapientinc · 2026-05-17
General · z-lab · 2026-01-04
General · ibm-granite · 2025-10-07
General · menlo · 2025-06-25
General · fableforge-ai · 2026-07-05
General · viorikaai-org · 2026-07-05
Reasoning · jackrong · 2026-03-16
General · bananamind · 2026-07-17
General · lgai-exaone · 2025-03-12
General · maliosdark · 2026-07-09
General · raidium · 2026-06-15
Embedding · taide · 2026-06-12
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
Multimodal · Alibaba · 2025-01-26
Multimodal · Google · 2026-03-02
Multimodal · zai-org
Multimodal · Alibaba
Multimodal · datalab-to
General · Alibaba · 2025-04-28
General · huggingfacetb · 2025-07-08
Multimodal · openbmb
General · Alibaba · 2025-04-28
Multimodal · rednote-hilab
Reasoning · DeepSeek · 2025-01-20
Multimodal · tencent
Multimodal · dots-studio
General · distil-labs
General · farbodtavakkoli
Multimodal · raxcore-dev
General · huggingfacetb · 2025-06-19
General · jinaai
Multimodal · ath-maas
Multimodal · Alibaba
Multimodal · Liquid AI
Reasoning · typhoon-ai
Multimodal · paddlepaddle
General · Liquid AI
Chat · uzlm · 2025-09-03
Multimodal · paddlepaddle
Chat · baseten · 2025-09-12
General · adamlucek
Coding · shahriarferdoush
General · ahczhg
General · onnx-community · 2025-04-28
General · etherll
General · baidu
General · Liquid AI
General · openbmb
Multimodal · lkhl
Multimodal · infly
General · benjamin
General · lemonelabs
General · openbmb · 2025-06-05
General · farbodtavakkoli
General · novachronoai
Multimodal · paddlepaddle
Reasoning · khazarai
General · kamilamila
General · launch
General · arcee-ai
General · getonit
General · Liquid AI
General · saidutta69
General · agentica-org
General · novaciano
General · dmusingu
General · kgrabko
General · ordenwills
General · smcleish
General · carsenk
Reasoning · nvidia
Multimodal · zero-point-ai
General · openbmb
General · ibm-granite
General · osaurusai
General · pyoakum
General · ibm-granite
General · treadon
General · tencent
Both bins run the same 307 GB/s memory bus, so token generation is the same on either one. The extra cores show up in image and video work, not in tokens per second.
| Configuration | Memory bandwidth | Memory options | Models that fit |
|---|---|---|---|
| 15-core CPU, 16-core GPU | 307 GB/s | 24, 48 GB | Identical |
| 18-core CPU, 20-core GPU | 307 GB/s | 24, 48, 64 GB | Identical |
Nobody has submitted a benchmark on the M5 Pro yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.
ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.
What each step actually changes for local models, rather than which one is newer.
MacBook Pro 14" M5 Max
2x the memory bandwidth, up to 128 GB instead of 64 GB
Used market alternativeMacBook Pro 14" M4 Pro
11% less memory bandwidth, 48 GB ceiling instead of 64 GB
Same chip, other MacMacBook Pro 16" M5 Pro
The same chip in a different Mac
Same chip, other MacMac mini M5 Pro
The same chip in a different Mac
Apple has not announced any of this. The M7 Max and Ultra are the first parts Apple designed after cancelling a generation to reach them, so extrapolating from the M5 under-represents them. These rows assume LPDDR6, whose wider channels grow every bus by half, at its top speed bin by the time the Max and Ultra ship. The 1.5 TB Ultra ceiling is Bloomberg's reported design target, and whether that configuration ships depends on the memory market. Stacked memory or a new package fabric would land above these numbers; nobody outside Apple can price that yet.
Projected chip · expected 2027
691 GB/s · 24 to 96 GB unified memory
18-core CPU · 20 or 24-core GPU · 76 TOPS Neural Engine
Would hold about a 129B model at Q4
On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M5 Pro's 307 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 64 GB machine can hand almost all of that to a model with no copy across a bus.
The 14-inch chassis cools well enough to hold its clocks through a long generation run, and it is the smallest machine Apple puts a Max chip in. Buy the memory, not the cores: every extra GB raises what you can load, while the core count only moves throughput on models that already fit.
Yes, at 64 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 64 GB of unified memory leaves about 56 GB for weights once macOS takes its share. At 24 GB it does not fit at any quantization worth running.
Memory is the only spec that changes what you can run at all. 24 GB holds about a 31B model at Q4; 64 GB holds about 85B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.
Token generation is bandwidth-bound, so M5 Pro throughput scales with its 307 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M5 Pro lands in the tens of tokens per second and a 70B model lands in the single digits.
Not for LLMs. Both bins run the same 307 GB/s memory bus and take the same memory options, and token generation is bound by bandwidth rather than GPU cores. The extra cores show up in image generation and video work, not in tokens per second.
Buy on the memory you need today. Apple raises memory ceilings slowly and bandwidth in steps, and the M5 Pro already holds about a 85B model at Q4. If your target model fits in 64 GB, waiting buys throughput rather than capability.
ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.