The M3 Ultra Mac Studio runs local models at 819 GB/s of memory bandwidth with 96 to 512 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 512 GB you can hold roughly a 696B dense model at Q4. Two Max dies fused together: the widest memory bus and the highest capacity Apple sells. This is the tier that runs frontier-size open weights locally.
Memory bandwidth is faster than 94% of the Apple Silicon chips shipped in a Mac, against a 819 GB/s peak.
Its memory ceiling is above 94% of them, against a 512 GB peak.
Every option Apple sells with this chip. The model list below recomputes against the one you pick.
GPU cores
Unified memory
Apple couples memory to the core count on this chip, so the options change with the bin above.
Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.
3,556 of 3,641 models fit, and 3,503 of them run with headroom rather than as a squeeze.
3,608 of 3,641 models fit, and 3,581 of them run with headroom rather than as a squeeze.
3,637 of 3,641 models fit, and 3,607 of them run with headroom rather than as a squeeze.
Every model in the database against this exact configuration, at 819 GB/s. Ratings and speeds are the same numbers the model pages show.
Showing 3641 of 3641 models
Multimodal · Alibaba · 2026-04-15
Multimodal · Alibaba · 2026-02-27
Multimodal · google · 2026-05
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-24
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-27
Multimodal · Alibaba · 2026-02-28
Reasoning · jackrong · 2026-03-16
Reasoning · jackrong · 2026-03-07
Multimodal · Alibaba · 2026-02-27
Multimodal · Alibaba · 2026-02-26
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Reasoning · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Reasoning · Liquid AI · 2025-11-28
General · alibaba-nlp · 2026-03-31
General · ibm-granite · 2025-09-16
General · ibm-granite · 2025-09-16
Chat · Liquid AI · 2025-11-28
Chat · Liquid AI · 2025-11-28
General · ibm-granite · 2025-09-16
Chat · Liquid AI · 2025-11-28
General · NCAI · 2025-12-29
General · NCAI · 2025-12-29
Multimodal · Google · 2025-07-30
Multimodal · Google · 2025-07-30
Multimodal · Google · 2025-07-30
Reasoning · HuggingFace · 2025-07-08
Multimodal · Google · 2025-06-25
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Multimodal · NCAI · 2025-12-29
Reasoning · NVIDIA · 2025-06-01
Multimodal · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
General · lgai-exaone · 2025-03-12
Reasoning · DeepSeek · 2025-01-20
Reasoning · jackrong · 2026-02-27
General · LG AI · 2025-07-15
General · raidium · 2026-06-15
Embedding · taide · 2026-06-12
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
General · openai
Multimodal · Alibaba · 2026-04-21
General · Alibaba · 2025-04-27
Multimodal · Alibaba
Multimodal · zai-org
General · Alibaba · 2025-04-27
Multimodal · Google · 2025-03-01
Multimodal · Alibaba
Coding · Alibaba
General · zai-org
Multimodal · datalab-to
General · prefeitura-rio
Multimodal · Microsoft
General · Alibaba
General · Alibaba
General · Alibaba
Multimodal · Google
Multimodal · openbmb
General · Alibaba
Reasoning · DeepSeek
Multimodal · rednote-hilab
Multimodal · Alibaba
General · Alibaba
Multimodal · Microsoft · 2025-04-01
Reasoning · DeepSeek
Multimodal · bytedance-seed
Multimodal · huggingfacem4
General · tiger-lab
General · DeepSeek
Multimodal · Google
Multimodal · datalab-to
General · ibm-granite
Reasoning · DeepSeek
General · farbodtavakkoli
Multimodal · moonshotai
General · openbmb
General · farbodtavakkoli
Multimodal · rednote-hilab
General · hmellor
General · distil-labs
Multimodal · nanonets
General · trevorjs
General · farbodtavakkoli
General · poolside
General · applied-innovation-center
Multimodal · reducto
General · Alibaba
General · lgai-exaone
General · Liquid AI
General · nanonets
General · nanbeige
Reasoning · lordx64
General · lmms-lab
Multimodal · ibm-granite
Multimodal · allenai
General · farbodtavakkoli
General · ibm-granite
General · arliai
General · ibm-granite
General · Google
Multimodal · moonshotai
Multimodal · lkhl
Multimodal · typhoon-ai
General · jinaai
General · huihui-ai
General · aoxo
General · Alibaba
Multimodal · goekdeniz-guelmez
Coding · coder3101
General · kristaller486
General · llmfan46
Reasoning · typhoon-ai
General · hcompany
General · sarvamai
General · Upstage
General · 01.ai
General · zstanjj
General · ibm-granite
General · nvidia
General · llmfan46
General · weiboai
General · nvidia
Coding · ibm-granite
General · Alibaba
General · Alibaba
General · x-izhang
General · idea-research
General · openai
General · paddlepaddle
Both bins run the same 819 GB/s memory bus, so token generation is the same on either one. The extra cores show up in image and video work, not in tokens per second.
| Configuration | Memory bandwidth | Memory options | Models that fit |
|---|---|---|---|
| 28-core CPU, 60-core GPU | 819 GB/s | 96 GB | Identical |
| 32-core CPU, 80-core GPU | 819 GB/s | 96, 256, 512 GB | Identical |
Nobody has submitted a benchmark on the M3 Ultra yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.
ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.
What each step actually changes for local models, rather than which one is newer.
On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M3 Ultra's 819 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 512 GB machine can hand almost all of that to a model with no copy across a bus.
The Studio exists for this workload. It carries the widest memory buses and the highest capacities Apple sells, and it runs at full clocks indefinitely. Buy the memory, not the cores: every extra GB raises what you can load, while the core count only moves throughput on models that already fit.
Yes, at 512 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 512 GB of unified memory leaves about 450 GB for weights once macOS takes its share. At 96 GB it does not fit at any quantization worth running.
Memory is the only spec that changes what you can run at all. 96 GB holds about a 129B model at Q4; 512 GB holds about 696B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.
Token generation is bandwidth-bound, so M3 Ultra throughput scales with its 819 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M3 Ultra lands in the tens of tokens per second and a 70B model lands in the single digits.
Not for LLMs. Both bins run the same 819 GB/s memory bus and take the same memory options, and token generation is bound by bandwidth rather than GPU cores. The extra cores show up in image generation and video work, not in tokens per second.
Buy on the memory you need today. Apple raises memory ceilings slowly and bandwidth in steps, and the M3 Ultra already holds about a 696B model at Q4. If your target model fits in 512 GB, waiting buys throughput rather than capability.
ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.