← All MacBook Pro 16" models

MacBook Pro 16" M1 Pro

The M1 Pro MacBook Pro 16" runs local models at 200 GB/s of memory bandwidth with 16 to 32 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 32 GB you can hold roughly a 42B dense model at Q4. The Pro roughly doubles the base chip's memory bus. That moves mid-size models from usable to comfortable without moving you into desktop money.

Apple no longer sells this configuration new. It stays fully evaluated here because the used market is where most of its local AI value now sits.

Specifications

ChipApple M1 Pro
CPU cores10
GPU cores16
Unified memory16 or 32 GB
Memory bandwidth200 GB/s
Neural Engine11 TOPS
Released2021
AvailabilityUsed market

Memory bandwidth is faster than 33% of the Apple Silicon chips shipped in a Mac, against a 819 GB/s peak.

Its memory ceiling is above 17% of them, against a 512 GB peak.

Pick your configuration

Every option Apple sells with this chip. The model list below recomputes against the one you pick.

Unified memory

What each memory option runs

Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.

16 GB unified memory

14 GB usable for weights · 200 GB/s

3,128 of 3,641 models fit, and 2,907 of them run with headroom rather than as a squeeze.

Largest model at Q4
DeepSeek V4 Flash JANGTQ2 · 20.2B
Best all-round pick
Qwen3.5 0.8B · Q8_0 · ~120 tok/s

32 GB unified memory

28 GB usable for weights · 200 GB/s

3,388 of 3,641 models fit, and 3,127 of them run with headroom rather than as a squeeze.

Largest model at Q4
Phi 3.5 MoE instruct · 41.87B
Best all-round pick
Qwen3.5 0.8B · Q8_0 · ~120 tok/s

What a 16 GB M1 Pro MacBook Pro 16" can run

Every model in the database against this exact configuration, at 200 GB/s. Ratings and speeds are the same numbers the model pages show.

Showing 3641 of 3641 models

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
1.5 GB9% of RAM~120 tok/sEstimated0.87B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
1.5 GB9% of RAM~120 tok/sEstimated0.87B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
3.0 GB19% of RAM~46 tok/sEstimated2.27B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
3.0 GB19% of RAM~46 tok/sEstimated2.27B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

Reasoning · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.3 GB8% of RAM~142 tok/sEstimated0.74B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

Reasoning · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
8.2 GB52% of RAMBenchmark needed6.94B params
Run with ToolPiper

Chat · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

Chat · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
3.4 GB21% of RAM~41 tok/sEstimated2.57B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.5B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
3.4 GB21% of RAM~41 tok/sEstimated2.57B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.5B params
Run with ToolPiper

General · NCAI · 2025-12-29

Q8_0Excellent
8.6 GB54% of RAMBenchmark needed7.25B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
5.7 GB36% of RAM~22 tok/sEstimated4.66B params
Run with ToolPiper

Reasoning · HuggingFace · 2025-07-08

Q8_0Excellent
3.8 GB24% of RAM~35 tok/sEstimated3B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
5.7 GB36% of RAM~22 tok/sEstimated4.66B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
1.0 GB6% of RAM~233 tok/sEstimated0.45B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
2.3 GB14% of RAM~65 tok/sEstimated1.6B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
9.8 GB61% of RAMBenchmark needed8.3B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
2.3 GB14% of RAM~66 tok/sEstimated1.58B params
Run with ToolPiper

Chat · Liquid AI · 2025-11-28

Q8_0Excellent
3.4 GB21% of RAM~41 tok/sEstimated2.57B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
4.1 GB25% of RAM~33 tok/sEstimated3.19B params
Run with ToolPiper

General · LG AI · 2025-07-15

Q8_0Excellent
1.8 GB11% of RAM~87 tok/sEstimated1.2B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
3.8 GB24% of RAM~35 tok/sEstimated3B params
Run with ToolPiper

General · raidium · 2026-06-15

Q8_0Excellent
0.5 GB3% of RAM~5,238 tok/sEstimated0.02B params
Run with ToolPiper

Multimodal · NCAI · 2025-12-29

Q8_0Excellent
9.0 GB56% of RAMBenchmark needed7.58B params
Run with ToolPiper

Embedding · taide · 2026-06-12

Q8_0Excellent
0.8 GB5% of RAM~349 tok/sEstimated0.3B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
1.3 GB8% of RAM~140 tok/sEstimated0.75B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
2.8 GB17% of RAM~52 tok/sEstimated2.03B params
Run with ToolPiper

Multimodal · zai-org

Q8_0Excellent
2.0 GB12% of RAM~79 tok/sEstimated1.33B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.54B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
1.0 GB7% of RAM~214 tok/sEstimated0.49B params
Run with ToolPiper

Multimodal · openbmb

Q8_0Excellent
2.0 GB12% of RAM~81 tok/sEstimated1.3B params
Run with ToolPiper

Reasoning · DeepSeek

Q8_0Excellent
2.5 GB16% of RAM~59 tok/sEstimated1.78B params
Run with ToolPiper

Multimodal · datalab-to

Q8_0Excellent
1.3 GB8% of RAM~152 tok/sEstimated0.69B params
Run with ToolPiper

General · openbmb

Q8_0Excellent
1.7 GB11% of RAM~97 tok/sEstimated1.08B params
Run with ToolPiper

General · hmellor

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

General · distil-labs

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
1.8 GB12% of RAM~87 tok/sEstimated1.21B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
7.9 GB50% of RAMBenchmark needed6.67B params
Run with ToolPiper
Q8_0Excellent
0.5 GB3% of RAM~10,476 tok/sEstimated0.01B params
Run with ToolPiper

Multimodal · lkhl

Q8_0Excellent
2.7 GB17% of RAM~53 tok/sEstimated1.96B params
Run with ToolPiper

General · jinaai

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.54B params
Run with ToolPiper

Reasoning · typhoon-ai

Q8_0Excellent
2.9 GB18% of RAM~49 tok/sEstimated2.13B params
Run with ToolPiper

General · openai

Q8_0Excellent
3.1 GB20% of RAMBenchmark needed2.37B params
Run with ToolPiper

General · paddlepaddle

Q8_0Excellent
1.6 GB10% of RAM~109 tok/sEstimated0.96B params
Run with ToolPiper

General · Liquid AI

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · pfnet

Q8_0Excellent
1.9 GB12% of RAM~81 tok/sEstimated1.29B params
Run with ToolPiper

General · openbmb

Q8_0Excellent
1.7 GB11% of RAM~97 tok/sEstimated1.08B params
Run with ToolPiper

General · baidu

Q8_0Excellent
0.9 GB6% of RAM~291 tok/sEstimated0.36B params
Run with ToolPiper

General · amd

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.5B params
Run with ToolPiper

General · arcee-ai

Q8_0Excellent
7.3 GB46% of RAMBenchmark needed6.12B params
Run with ToolPiper

General · adamlucek

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

Coding · shahriarferdoush

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

General · ahczhg

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

General · abaryan

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

General · etherll

Q8_0Excellent
1.3 GB8% of RAM~142 tok/sEstimated0.74B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
2.5 GB15% of RAMBenchmark needed1.77B params
Run with ToolPiper

General · kamilamila

Q8_0Excellent
1.2 GB7% of RAM~169 tok/sEstimated0.62B params
Run with ToolPiper
Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.5B params
Run with ToolPiper

General · paddlepaddle

Q8_0Excellent
1.6 GB10% of RAM~109 tok/sEstimated0.96B params
Run with ToolPiper

General · agentica-org

Q8_0Excellent
2.5 GB16% of RAM~59 tok/sEstimated1.78B params
Run with ToolPiper

General · novaciano

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.5B params
Run with ToolPiper

General · kgrabko

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.5B params
Run with ToolPiper

General · ordenwills

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper
Q8_0Excellent
2.1 GB13% of RAM~75 tok/sEstimated1.39B params
Run with ToolPiper

General · carsenk

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

Reasoning · nvidia

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.54B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
2.0 GB12% of RAMBenchmark needed1.33B params
Run with ToolPiper

General · openbmb

Q8_0Excellent
2.0 GB12% of RAM~81 tok/sEstimated1.3B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
2.1 GB13% of RAM~72 tok/sEstimated1.46B params
Run with ToolPiper

General · pyoakum

Q8_0Excellent
6.5 GB40% of RAMBenchmark needed5.35B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
4.2 GB26% of RAMBenchmark needed3.3B params
Run with ToolPiper
Q8_0Excellent
2.0 GB12% of RAM~81 tok/sEstimated1.3B params
Run with ToolPiper

General · skis-ai-research

Q8_0Excellent
1.8 GB11% of RAM~90 tok/sEstimated1.17B params
Run with ToolPiper

General · menlo

Q8_0Excellent
2.4 GB15% of RAM~61 tok/sEstimated1.72B params
Run with ToolPiper

General · thkim0305

Q8_0Excellent
1.6 GB10% of RAM~105 tok/sEstimated1B params
Run with ToolPiper
Q8_0Excellent
1.5 GB9% of RAM~120 tok/sEstimated0.87B params
Run with ToolPiper

Coding · rahul7star

Q8_0Excellent
1.5 GB9% of RAM~120 tok/sEstimated0.87B params
Run with ToolPiper

General · TII

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.55B params
Run with ToolPiper

Coding · z-lab

Q8_0Excellent
1.7 GB11% of RAM~97 tok/sEstimated1.08B params
Run with ToolPiper
Q8_0Excellent
6.9 GB43% of RAMBenchmark needed5.75B params
Run with ToolPiper

General · artificialguybr

Q8_0Excellent
1.9 GB12% of RAM~84 tok/sEstimated1.24B params
Run with ToolPiper

General · Microsoft

Q8_0Excellent
0.7 GB4% of RAMBenchmark needed0.17B params
Run with ToolPiper

General · roystar

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.54B params
Run with ToolPiper

General · weiboai

Q8_0Excellent
2.5 GB16% of RAM~59 tok/sEstimated1.78B params
Run with ToolPiper

General · Microsoft

Q8_0Excellent
0.7 GB4% of RAMBenchmark needed0.17B params
Run with ToolPiper

General · TII

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.55B params
Run with ToolPiper

General · tencent

Q8_0Excellent
1.1 GB7% of RAM~194 tok/sEstimated0.54B params
Run with ToolPiper

General · primeintellect

Q8_0Excellent
1.1 GB7% of RAMBenchmark needed0.54B params
Run with ToolPiper

General · lgai-exaone · 2025-03-12

Q8_0Excellent
3.2 GB20% of RAM~43 tok/sEstimated2.41B params
Run with ToolPiper

Multimodal · Alibaba

Q8_0Excellent
2.9 GB18% of RAM~49 tok/sEstimated2.13B params
Run with ToolPiper

Multimodal · rednote-hilab

Q8_0Excellent
3.9 GB24% of RAM~34 tok/sEstimated3.04B params
Run with ToolPiper

Multimodal · Google · 2025-06-25

Q8_0Excellent
5.0 GB31% of RAM~26 tok/sEstimated4B params
Run with ToolPiper

Multimodal · rednote-hilab

Q8_0Excellent
3.9 GB24% of RAM~34 tok/sEstimated3.04B params
Run with ToolPiper

General · internlm

Q8_0Excellent
2.1 GB13% of RAM~75 tok/sEstimated1.4B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
2.9 GB18% of RAM~49 tok/sEstimated2.13B params
Run with ToolPiper
Q8_0Excellent
4.0 GB25% of RAM~33 tok/sEstimated3.13B params
Run with ToolPiper

General · bytedance

Q8_0Excellent
2.1 GB13% of RAM~73 tok/sEstimated1.43B params
Run with ToolPiper

General · dmusingu

Q8_0Excellent
2.9 GB18% of RAM~49 tok/sEstimated2.13B params
Run with ToolPiper

Coding · ibm-granite

Q8_0Excellent
4.4 GB27% of RAM~30 tok/sEstimated3.48B params
Run with ToolPiper

General · zero-point-ai

Q8_0Excellent
3.0 GB19% of RAM~46 tok/sEstimated2.27B params
Run with ToolPiper

General · bezzam

Q8_0Excellent
1.4 GB9% of RAM~134 tok/sEstimated0.78B params
Run with ToolPiper

General · tencent

Q8_0Excellent
2.7 GB17% of RAM~53 tok/sEstimated1.96B params
Run with ToolPiper

General · huihui-ai

Q8_0Excellent
3.0 GB19% of RAM~46 tok/sEstimated2.27B params
Run with ToolPiper

General · inclusionai

Q8_0Excellent
3.2 GB20% of RAM~43 tok/sEstimated2.44B params
Run with ToolPiper

Reasoning · ai21labs

Q8_0Excellent
4.1 GB25% of RAMBenchmark needed3.2B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
1.0 GB7% of RAM~214 tok/sEstimated0.49B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
2.4 GB15% of RAM~61 tok/sEstimated1.72B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
1.2 GB7% of RAM~175 tok/sEstimated0.6B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
4.3 GB27% of RAM~31 tok/sEstimated3.4B params
Run with ToolPiper

General · Microsoft

Q8_0Excellent
1.2 GB7% of RAM~175 tok/sEstimated0.6B params
Run with ToolPiper

Multimodal · opengvlab

Q8_0Excellent
1.5 GB10% of RAM~111 tok/sEstimated0.94B params
Run with ToolPiper

Multimodal · nanonets

Q8_0Excellent
4.7 GB29% of RAM~28 tok/sEstimated3.75B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
1.3 GB8% of RAM~140 tok/sEstimated0.75B params
Run with ToolPiper

General · voyageai

Q8_0Excellent
0.9 GB6% of RAM~299 tok/sEstimated0.35B params
Run with ToolPiper

General · Microsoft

Q8_0Excellent
0.8 GB5% of RAM~388 tok/sEstimated0.27B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
1.0 GB6% of RAM~233 tok/sEstimated0.45B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
0.8 GB5% of RAM~388 tok/sEstimated0.27B params
Run with ToolPiper

Multimodal · typhoon-ai

Q8_0Excellent
4.7 GB29% of RAM~28 tok/sEstimated3.75B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
1.6 GB10% of RAM~105 tok/sEstimated1B params
Run with ToolPiper

General · kristaller486

Q8_0Excellent
3.9 GB24% of RAM~34 tok/sEstimated3.04B params
Run with ToolPiper

General · laap-ai

Q8_0Excellent
2.2 GB14% of RAM~68 tok/sEstimated1.54B params
Run with ToolPiper

General · jakobhuss

Q8_0Excellent
0.8 GB5% of RAM~388 tok/sEstimated0.27B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.49B params
Run with ToolPiper

General · Upstage

Q8_0Excellent
9.5 GB59% of RAMBenchmark needed8.05B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
4.3 GB27% of RAM~31 tok/sEstimated3.4B params
Run with ToolPiper

Coding · DeepSeek

Q8_0Excellent
2.0 GB13% of RAM~78 tok/sEstimated1.35B params
Run with ToolPiper

General · weiboai

Q8_0Excellent
3.9 GB25% of RAM~34 tok/sEstimated3.09B params
Run with ToolPiper

General · llava-hf

Q8_0Excellent
1.1 GB7% of RAM~210 tok/sEstimated0.5B params
Run with ToolPiper

General · tomg-group-umd

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.49B params
Run with ToolPiper

General · x-izhang

Q8_0Excellent
4.1 GB26% of RAM~32 tok/sEstimated3.25B params
Run with ToolPiper

General · bytedance

Q8_0Excellent
2.7 GB17% of RAM~52 tok/sEstimated2.01B params
Run with ToolPiper

General · contextboxai

Q8_0Excellent
2.4 GB15% of RAM~61 tok/sEstimated1.72B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
4.3 GB27% of RAM~31 tok/sEstimated3.4B params
Run with ToolPiper

General · openbmb

Q8_0Excellent
1.0 GB6% of RAM~244 tok/sEstimated0.43B params
Run with ToolPiper
Q8_0Excellent
1.2 GB8% of RAM~156 tok/sEstimated0.67B params
Run with ToolPiper

General · onnx-community

Q8_0Excellent
2.2 GB14% of RAM~70 tok/sEstimated1.49B params
Run with ToolPiper

Measured on the M1 Pro

Nobody has submitted a benchmark on the M1 Pro yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.

ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.

Why unified memory is the number that matters

On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M1 Pro's 200 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 32 GB machine can hand almost all of that to a model with no copy across a bus.

Apple stopped selling this one, which is exactly why it is interesting. The 16-inch chassis has the most thermal headroom Apple ships in a laptop, so sustained token throughput stays close to the burst figure. A used M1 Pro at 32 GB still gives you 200 GB/s and a hard 42B ceiling, and neither number degrades with age the way a battery does.

Common questions

Can the M1 Pro MacBook Pro 16" run a 70B model?

No. A 70B model at Q4_K_M needs about 46 GB, and the largest M1 Pro MacBook Pro 16" tops out at 32 GB, which leaves about 28 GB for weights. The practical ceiling on this machine is around 42B parameters at Q4.

How much unified memory should I get with the M1 Pro MacBook Pro 16"?

Memory is the only spec that changes what you can run at all. 16 GB holds about a 20B model at Q4; 32 GB holds about 42B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.

How fast are local LLMs on the M1 Pro?

Token generation is bandwidth-bound, so M1 Pro throughput scales with its 200 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M1 Pro lands in the tens of tokens per second and a 70B model lands in the single digits.

Is a used M1 Pro MacBook Pro 16" still worth buying for local AI?

For inference, the specs that matter do not age: 200 GB/s and up to 32 GB of unified memory are the same numbers today as they were in 2021. A used M1 Pro at the top memory option usually beats a new base-tier machine at the same price on both. Check the battery and the display, not the silicon.

Run these models on your MacBook Pro 16"

ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.