← All Mac Studio models

Mac Studio M3 Ultra

The M3 Ultra Mac Studio runs local models at 819 GB/s of memory bandwidth with 96 to 512 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 512 GB you can hold roughly a 696B dense model at Q4. Two Max dies fused together: the widest memory bus and the highest capacity Apple sells. This is the tier that runs frontier-size open weights locally.

Specifications

ChipApple M3 Ultra
CPU cores28 or 32
GPU cores60 or 80
Unified memory96, 256, or 512 GB
Memory bandwidth819 GB/s
Neural Engine36 TOPS
Released2025
AvailabilitySold new by Apple

Memory bandwidth is faster than 94% of the Apple Silicon chips shipped in a Mac, against a 819 GB/s peak.

Its memory ceiling is above 94% of them, against a 512 GB peak.

Pick your configuration

Every option Apple sells with this chip. The model list below recomputes against the one you pick.

GPU cores

Unified memory

Apple couples memory to the core count on this chip, so the options change with the bin above.

What each memory option runs

Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.

96 GB unified memory

84 GB usable for weights · 819 GB/s

3,556 of 3,641 models fit, and 3,503 of them run with headroom rather than as a squeeze.

Largest model at Q4
Qwen3.5 122B A10B · 125.09B
Best all-round pick
Qwen3.6 35B A3B · Q8_0 · ~143 tok/s

256 GB unified memory

225 GB usable for weights · 819 GB/s

3,608 of 3,641 models fit, and 3,581 of them run with headroom rather than as a squeeze.

Largest model at Q4
DeepSeek V3.2 REAP 345B A37B · 344.89B
Best all-round pick
Qwen3.6 35B A3B · Q8_0 · ~143 tok/s

512 GB unified memory

450 GB usable for weights · 819 GB/s

3,637 of 3,641 models fit, and 3,607 of them run with headroom rather than as a squeeze.

Largest model at Q4
Kimi K2.6 REAP Solidity · 688.12B
Best all-round pick
Qwen3.6 35B A3B · Q8_0 · ~143 tok/s

What a 96 GB M3 Ultra Mac Studio can run

Every model in the database against this exact configuration, at 819 GB/s. Ratings and speeds are the same numbers the model pages show.

Showing 3641 of 3641 models

Multimodal · Alibaba · 2026-04-15

Q8_0Excellent
40.6 GB42% of RAMBenchmark needed35.95B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
5.7 GB6% of RAM~92 tok/sEstimated4.66B params
Run with ToolPiper

Multimodal · google · 2026-05

Q8_0Excellent
13.8 GB14% of RAM~36 tok/sEstimated11.96B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
1.5 GB2% of RAM~493 tok/sEstimated0.87B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-24

Q8_0Excellent
40.6 GB42% of RAMBenchmark needed35.95B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
3.0 GB3% of RAM~189 tok/sEstimated2.27B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
1.5 GB2% of RAM~493 tok/sEstimated0.87B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
5.7 GB6% of RAM~92 tok/sEstimated4.66B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0Excellent
3.0 GB3% of RAM~189 tok/sEstimated2.27B params
Run with ToolPiper

Reasoning · jackrong · 2026-03-16

Q8_0Excellent
11.3 GB12% of RAM~44 tok/sEstimated9.65B params
Run with ToolPiper

Reasoning · jackrong · 2026-03-07

Q8_0Excellent
40.6 GB42% of RAMBenchmark needed35.95B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0Excellent
11.3 GB12% of RAM~44 tok/sEstimated9.65B params
Run with ToolPiper

Multimodal · Alibaba · 2026-02-26

Q8_0Excellent
11.3 GB12% of RAM~44 tok/sEstimated9.65B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

Reasoning · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
9.8 GB10% of RAMBenchmark needed8.3B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
27.1 GB28% of RAMBenchmark needed23.84B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
3.4 GB4% of RAM~167 tok/sEstimated2.57B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.3 GB1% of RAM~580 tok/sEstimated0.74B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
3.4 GB4% of RAM~167 tok/sEstimated2.57B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

Reasoning · Liquid AI · 2025-11-28

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

General · alibaba-nlp · 2026-03-31

Q8_0Excellent
9.6 GB10% of RAM~52 tok/sEstimated8.19B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
36.4 GB38% of RAMBenchmark needed32.21B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
8.2 GB9% of RAMBenchmark needed6.94B params
Run with ToolPiper

Chat · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

Chat · Liquid AI · 2025-11-28

Q8_0Excellent
3.4 GB4% of RAM~167 tok/sEstimated2.57B params
Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0Excellent
4.1 GB4% of RAM~134 tok/sEstimated3.19B params
Run with ToolPiper

Chat · Liquid AI · 2025-11-28

Q8_0Excellent
1.8 GB2% of RAM~367 tok/sEstimated1.17B params
Run with ToolPiper

General · NCAI · 2025-12-29

Q8_0Excellent
8.6 GB9% of RAMBenchmark needed7.25B params
Run with ToolPiper

General · NCAI · 2025-12-29

Q8_0Excellent
22.4 GB23% of RAMBenchmark needed19.6B params
Run with ToolPiper

Multimodal · Google · 2025-07-30

Q8_0Excellent
29.5 GB31% of RAMBenchmark needed26B params
Run with ToolPiper

Multimodal · Google · 2025-07-30

Q8_0Excellent
9.4 GB10% of RAM~54 tok/sEstimated8B params
Run with ToolPiper

Multimodal · Google · 2025-07-30

Q8_0Excellent
6.2 GB6% of RAM~84 tok/sEstimated5.1B params
Run with ToolPiper

Reasoning · HuggingFace · 2025-07-08

Q8_0Excellent
3.8 GB4% of RAM~143 tok/sEstimated3B params
Run with ToolPiper

Multimodal · Google · 2025-06-25

Q8_0Excellent
5.0 GB5% of RAM~107 tok/sEstimated4B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
2.2 GB2% of RAM~286 tok/sEstimated1.5B params
Run with ToolPiper

General · Liquid AI · 2025-11-28

Q8_0Excellent
2.2 GB2% of RAM~286 tok/sEstimated1.5B params
Run with ToolPiper

Multimodal · NCAI · 2025-12-29

Q8_0Excellent
9.0 GB9% of RAMBenchmark needed7.58B params
Run with ToolPiper

Reasoning · NVIDIA · 2025-06-01

Q8_0Excellent
10.5 GB11% of RAM~48 tok/sEstimated9B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
1.0 GB1% of RAM~953 tok/sEstimated0.45B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
2.3 GB2% of RAM~268 tok/sEstimated1.6B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
3.8 GB4% of RAM~143 tok/sEstimated3B params
Run with ToolPiper

Multimodal · Liquid AI · 2025-11-28

Q8_0Excellent
2.3 GB2% of RAM~272 tok/sEstimated1.58B params
Run with ToolPiper

General · lgai-exaone · 2025-03-12

Q8_0Excellent
3.2 GB3% of RAM~178 tok/sEstimated2.41B params
Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0Excellent
9.0 GB9% of RAM~56 tok/sEstimated7.62B params
Run with ToolPiper

Reasoning · jackrong · 2026-02-27

Q8_0Excellent
31.5 GB33% of RAM~15 tok/sEstimated27.78B params
Run with ToolPiper

General · LG AI · 2025-07-15

Q8_0Excellent
1.8 GB2% of RAM~358 tok/sEstimated1.2B params
Run with ToolPiper

General · raidium · 2026-06-15

Q8_0Excellent
0.5 GB1% of RAM~21,450 tok/sEstimated0.02B params
Run with ToolPiper

Embedding · taide · 2026-06-12

Q8_0Excellent
0.8 GB1% of RAM~1,430 tok/sEstimated0.3B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
1.3 GB1% of RAM~572 tok/sEstimated0.75B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
5.0 GB5% of RAM~107 tok/sEstimated4.02B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
9.6 GB10% of RAM~52 tok/sEstimated8.19B params
Run with ToolPiper

General · openai

Q8_0Excellent
24.5 GB26% of RAMBenchmark needed21.51B params
Run with ToolPiper

Multimodal · Alibaba · 2026-04-21

Q8_0Excellent
31.5 GB33% of RAM~15 tok/sEstimated27.78B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
2.8 GB3% of RAM~211 tok/sEstimated2.03B params
Run with ToolPiper

Multimodal · Alibaba

Q8_0Excellent
5.5 GB6% of RAM~97 tok/sEstimated4.44B params
Run with ToolPiper

Multimodal · zai-org

Q8_0Excellent
2.0 GB2% of RAM~323 tok/sEstimated1.33B params
Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0Excellent
34.6 GB36% of RAMBenchmark needed30.53B params
Run with ToolPiper

Multimodal · Google · 2025-03-01

Q8_0Excellent
13.9 GB14% of RAM~36 tok/sEstimated12B params
Run with ToolPiper

Multimodal · Alibaba

Q8_0Excellent
2.9 GB3% of RAM~201 tok/sEstimated2.13B params
Run with ToolPiper

Coding · Alibaba

Q8_0Excellent
34.6 GB36% of RAMBenchmark needed30.53B params
Run with ToolPiper

General · zai-org

Q8_0Excellent
35.3 GB37% of RAMBenchmark needed31.22B params
Run with ToolPiper

Multimodal · datalab-to

Q8_0Excellent
6.4 GB7% of RAM~81 tok/sEstimated5.3B params
Run with ToolPiper

General · prefeitura-rio

Q8_0Excellent
5.0 GB5% of RAM~107 tok/sEstimated4.02B params
Run with ToolPiper

Multimodal · Microsoft

Q8_0Excellent
5.1 GB5% of RAM~103 tok/sEstimated4.15B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
5.5 GB6% of RAM~97 tok/sEstimated4.44B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
2.2 GB2% of RAM~279 tok/sEstimated1.54B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
1.0 GB1% of RAM~876 tok/sEstimated0.49B params
Run with ToolPiper

Multimodal · Google

Q8_0Excellent
29.3 GB31% of RAMBenchmark needed25.82B params
Run with ToolPiper

Multimodal · openbmb

Q8_0Excellent
2.0 GB2% of RAM~330 tok/sEstimated1.3B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
5.0 GB5% of RAM~107 tok/sEstimated4.02B params
Run with ToolPiper

Reasoning · DeepSeek

Q8_0Excellent
2.5 GB3% of RAM~241 tok/sEstimated1.78B params
Run with ToolPiper

Multimodal · rednote-hilab

Q8_0Excellent
3.9 GB4% of RAM~141 tok/sEstimated3.04B params
Run with ToolPiper

Multimodal · Alibaba

Q8_0Excellent
35.2 GB37% of RAMBenchmark needed31.07B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
9.0 GB9% of RAM~56 tok/sEstimated7.62B params
Run with ToolPiper

Multimodal · Microsoft · 2025-04-01

Q8_0Excellent
16.1 GB17% of RAM~31 tok/sEstimated14B params
Run with ToolPiper

Reasoning · DeepSeek

Q8_0Excellent
9.6 GB10% of RAM~52 tok/sEstimated8.19B params
Run with ToolPiper

Multimodal · bytedance-seed

Q8_0Excellent
9.7 GB10% of RAM~52 tok/sEstimated8.29B params
Run with ToolPiper

Multimodal · huggingfacem4

Q8_0Excellent
9.9 GB10% of RAM~51 tok/sEstimated8.46B params
Run with ToolPiper

General · tiger-lab

Q8_0Excellent
5.1 GB5% of RAM~103 tok/sEstimated4.15B params
Run with ToolPiper

General · DeepSeek

Q8_0Excellent
18.0 GB19% of RAMBenchmark needed15.71B params
Run with ToolPiper

Multimodal · Google

Q8_0Excellent
30.1 GB31% of RAMBenchmark needed26.54B params
Run with ToolPiper

Multimodal · datalab-to

Q8_0Excellent
1.3 GB1% of RAM~622 tok/sEstimated0.69B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
4.3 GB4% of RAM~126 tok/sEstimated3.4B params
Run with ToolPiper

Reasoning · DeepSeek

Q8_0Excellent
9.5 GB10% of RAM~53 tok/sEstimated8.03B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
5.3 GB6% of RAM~100 tok/sEstimated4.3B params
Run with ToolPiper

Multimodal · moonshotai

Q8_0Excellent
18.8 GB20% of RAMBenchmark needed16.41B params
Run with ToolPiper

General · openbmb

Q8_0Excellent
1.7 GB2% of RAM~397 tok/sEstimated1.08B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
8.3 GB9% of RAM~62 tok/sEstimated6.97B params
Run with ToolPiper

Multimodal · rednote-hilab

Q8_0Excellent
3.9 GB4% of RAM~141 tok/sEstimated3.04B params
Run with ToolPiper

General · hmellor

Q8_0Excellent
1.9 GB2% of RAM~346 tok/sEstimated1.24B params
Run with ToolPiper

General · distil-labs

Q8_0Excellent
0.9 GB1% of RAM~1,226 tok/sEstimated0.35B params
Run with ToolPiper

Multimodal · nanonets

Q8_0Excellent
4.7 GB5% of RAM~114 tok/sEstimated3.75B params
Run with ToolPiper

General · trevorjs

Q8_0Excellent
29.3 GB31% of RAMBenchmark needed25.81B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
1.8 GB2% of RAM~355 tok/sEstimated1.21B params
Run with ToolPiper

General · poolside

Q8_0Excellent
37.8 GB39% of RAMBenchmark needed33.44B params
Run with ToolPiper

General · applied-innovation-center

Q8_0Excellent
45.9 GB48% of RAMBenchmark needed40.67B params
Run with ToolPiper

Multimodal · reducto

Q8_0Excellent
9.7 GB10% of RAM~52 tok/sEstimated8.29B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
34.6 GB36% of RAMBenchmark needed30.53B params
Run with ToolPiper

General · lgai-exaone

Q8_0Excellent
9.2 GB10% of RAM~55 tok/sEstimated7.82B params
Run with ToolPiper

General · Liquid AI

Q8_0Excellent
9.9 GB10% of RAMBenchmark needed8.47B params
Run with ToolPiper

General · nanonets

Q8_0Excellent
4.7 GB5% of RAM~114 tok/sEstimated3.75B params
Run with ToolPiper

General · nanbeige

Q8_0Excellent
4.9 GB5% of RAM~109 tok/sEstimated3.93B params
Run with ToolPiper
Q8_0Excellent
40.6 GB42% of RAMBenchmark needed35.95B params
Run with ToolPiper

General · lmms-lab

Q8_0Excellent
5.8 GB6% of RAM~91 tok/sEstimated4.74B params
Run with ToolPiper

Multimodal · ibm-granite

Q8_0Excellent
5.0 GB5% of RAM~107 tok/sEstimated4B params
Run with ToolPiper

Multimodal · allenai

Q8_0Excellent
9.2 GB10% of RAM~55 tok/sEstimated7.76B params
Run with ToolPiper

General · farbodtavakkoli

Q8_0Excellent
5.2 GB5% of RAM~101 tok/sEstimated4.25B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
5.0 GB5% of RAM~107 tok/sEstimated4B params
Run with ToolPiper

General · arliai

Q8_0Excellent
23.8 GB25% of RAMBenchmark needed20.91B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
7.9 GB8% of RAMBenchmark needed6.67B params
Run with ToolPiper
Q8_0Excellent
0.5 GB1% of RAM~42,900 tok/sEstimated0.01B params
Run with ToolPiper

Multimodal · moonshotai

Q8_0Excellent
18.8 GB20% of RAMBenchmark needed16.41B params
Run with ToolPiper

Multimodal · lkhl

Q8_0Excellent
2.7 GB3% of RAM~219 tok/sEstimated1.96B params
Run with ToolPiper

Multimodal · typhoon-ai

Q8_0Excellent
4.7 GB5% of RAM~114 tok/sEstimated3.75B params
Run with ToolPiper

General · jinaai

Q8_0Excellent
2.2 GB2% of RAM~279 tok/sEstimated1.54B params
Run with ToolPiper
Q8_0Excellent
40.6 GB42% of RAMBenchmark needed35.95B params
Run with ToolPiper

General · aoxo

Q8_0Excellent
36.4 GB38% of RAMBenchmark needed32.15B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
2.9 GB3% of RAM~201 tok/sEstimated2.13B params
Run with ToolPiper
Q8_0Excellent
5.5 GB6% of RAM~97 tok/sEstimated4.44B params
Run with ToolPiper

Coding · coder3101

Q8_0Excellent
29.3 GB31% of RAMBenchmark needed25.81B params
Run with ToolPiper

General · kristaller486

Q8_0Excellent
3.9 GB4% of RAM~141 tok/sEstimated3.04B params
Run with ToolPiper
Q8_0Excellent
39.7 GB41% of RAMBenchmark needed35.11B params
Run with ToolPiper

Reasoning · typhoon-ai

Q8_0Excellent
2.9 GB3% of RAM~201 tok/sEstimated2.13B params
Run with ToolPiper

General · hcompany

Q8_0Excellent
39.7 GB41% of RAMBenchmark needed35.11B params
Run with ToolPiper

General · sarvamai

Q8_0Excellent
36.4 GB38% of RAMBenchmark needed32.15B params
Run with ToolPiper

General · Upstage

Q8_0Excellent
9.5 GB10% of RAMBenchmark needed8.05B params
Run with ToolPiper

General · 01.ai

Q8_0Excellent
7.3 GB8% of RAM~71 tok/sEstimated6.06B params
Run with ToolPiper

General · zstanjj

Q8_0Excellent
4.8 GB5% of RAM~112 tok/sEstimated3.82B params
Run with ToolPiper

General · ibm-granite

Q8_0Excellent
4.3 GB4% of RAM~126 tok/sEstimated3.4B params
Run with ToolPiper

General · nvidia

Q8_0Excellent
4.8 GB5% of RAM~112 tok/sEstimated3.83B params
Run with ToolPiper
Q8_0Excellent
39.7 GB41% of RAMBenchmark needed35.11B params
Run with ToolPiper

General · weiboai

Q8_0Excellent
3.9 GB4% of RAM~139 tok/sEstimated3.09B params
Run with ToolPiper

General · nvidia

Q8_0Excellent
35.7 GB37% of RAMBenchmark needed31.58B params
Run with ToolPiper

Coding · ibm-granite

Q8_0Excellent
9.5 GB10% of RAM~53 tok/sEstimated8.05B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
40.6 GB42% of RAMBenchmark needed35.95B params
Run with ToolPiper

General · Alibaba

Q8_0Excellent
9.0 GB9% of RAM~56 tok/sEstimated7.62B params
Run with ToolPiper

General · x-izhang

Q8_0Excellent
4.1 GB4% of RAM~132 tok/sEstimated3.25B params
Run with ToolPiper

General · idea-research

Q8_0Excellent
5.0 GB5% of RAM~105 tok/sEstimated4.07B params
Run with ToolPiper

General · openai

Q8_0Excellent
3.1 GB3% of RAMBenchmark needed2.37B params
Run with ToolPiper

General · paddlepaddle

Q8_0Excellent
1.6 GB2% of RAM~447 tok/sEstimated0.96B params
Run with ToolPiper

60-core vs 80-core GPU

Both bins run the same 819 GB/s memory bus, so token generation is the same on either one. The extra cores show up in image and video work, not in tokens per second.

ConfigurationMemory bandwidthMemory optionsModels that fit
28-core CPU, 60-core GPU819 GB/s96 GBIdentical
32-core CPU, 80-core GPU819 GB/s96, 256, 512 GBIdentical

Measured on the M3 Ultra

Nobody has submitted a benchmark on the M3 Ultra yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.

ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.

Where to go from here

What each step actually changes for local models, rather than which one is newer.

Why unified memory is the number that matters

On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M3 Ultra's 819 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 512 GB machine can hand almost all of that to a model with no copy across a bus.

The Studio exists for this workload. It carries the widest memory buses and the highest capacities Apple sells, and it runs at full clocks indefinitely. Buy the memory, not the cores: every extra GB raises what you can load, while the core count only moves throughput on models that already fit.

Common questions

Can the M3 Ultra Mac Studio run a 70B model?

Yes, at 512 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 512 GB of unified memory leaves about 450 GB for weights once macOS takes its share. At 96 GB it does not fit at any quantization worth running.

How much unified memory should I get with the M3 Ultra Mac Studio?

Memory is the only spec that changes what you can run at all. 96 GB holds about a 129B model at Q4; 512 GB holds about 696B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.

How fast are local LLMs on the M3 Ultra?

Token generation is bandwidth-bound, so M3 Ultra throughput scales with its 819 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M3 Ultra lands in the tens of tokens per second and a 70B model lands in the single digits.

Is the 80-core GPU worth it over the 60-core on the M3 Ultra?

Not for LLMs. Both bins run the same 819 GB/s memory bus and take the same memory options, and token generation is bound by bandwidth rather than GPU cores. The extra cores show up in image generation and video work, not in tokens per second.

Should I buy the M3 Ultra Mac Studio now or wait for the next one?

Buy on the memory you need today. Apple raises memory ceilings slowly and bandwidth in steps, and the M3 Ultra already holds about a 696B model at Q4. If your target model fits in 512 GB, waiting buys throughput rather than capability.

Run these models on your Mac Studio

ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.