The M1 Ultra Mac Studio runs local models at 800 GB/s of memory bandwidth with 64 to 128 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 128 GB you can hold roughly a 172B dense model at Q4. Two Max dies fused together: the widest memory bus and the highest capacity Apple sells. This is the tier that runs frontier-size open weights locally.
Apple no longer sells this configuration new. It stays fully evaluated here because the used market is where most of its local AI value now sits.
Memory bandwidth is faster than 80% of the Apple Silicon chips shipped in a Mac, against a 1228 GB/s peak.
Its memory ceiling is above 65% of them, against a 512 GB peak.
Every option Apple sells with this chip. The model list below recomputes against the one you pick.
GPU cores
Unified memory
Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.
6,357 of 6,563 models fit, and 6,086 of them run with headroom rather than as a squeeze.
6,420 of 6,563 models fit, and 6,334 of them run with headroom rather than as a squeeze.
Every model in the database against this exact configuration, at 800 GB/s. Ratings and speeds are the same numbers the model pages show.
Showing 6563 of 6563 models
General · radixark · 2026-07-27
General · openbmb · 2026-05-21
General · farbodtavakkoli · 2026-06-17
General · internscience · 2026-07-13
General · Liquid AI · 2026-05-28
General · weiboai · 2026-06-12
General · Liquid AI · 2026-07-28
General · Liquid AI · 2026-06-24
General · nanbeige · 2026-07-21
General · ma7ee7 · 2026-07-30
General · goekdeniz-guelmez · 2026-07-31
General · ktruestory · 2026-05-28
General · petrouil · 2026-07-23
General · deepreinforce-ai · 2026-06-21
General · ibm-granite · 2026-04-06
Multimodal · ibm-granite · 2026-04-16
General · lukebailey181pub · 2026-04-21
General · trevorjs · 2026-04-03
General · empero-ai · 2026-06-19
Coding · coherelabs · 2026-06-05
General · goekdeniz-guelmez · 2026-07-31
Multimodal · Google · 2026-03-11
Multimodal · Alibaba · 2026-02-27
General · ibm-granite · 2026-04-06
Multimodal · Google · 2026-03-02
Multimodal · google · 2026-05
Multimodal · Alibaba · 2026-02-28
General · deepreinforce-ai · 2026-06-21
Multimodal · Alibaba · 2026-02-28
General · bahushruth · 2026-06-11
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
General · yuxinlu1 · 2026-06-28
Multimodal · Alibaba · 2026-02-27
General · flywheel-ai · 2026-06-21
General · Liquid AI · 2026-03-31
General · Alibaba · 2026-06-22
General · ravichandranj · 2026-02-13
General · poolside · 2026-06-20
General · ibm-granite · 2026-04-16
General · obliteratus · 2026-06-05
General · huihui-ai · 2026-07-11
General · obliteratus · 2026-04-15
General · internscience · 2026-06-22
General · nvidia · 2026-03-02
General · Liquid AI · 2026-02-24
Reasoning · jackrong · 2026-03-16
General · cagrigungor · 2026-08-06
Multimodal · Alibaba · 2026-02-27
Multimodal · Google · 2026-03-02
Multimodal · Alibaba · 2026-02-26
General · nvidia · 2026-03-18
Multimodal · Liquid AI · 2026-01-05
General · nvidia · 2026-03-18
General · poolside · 2026-04-23
General · Liquid AI · 2026-01-20
General · apodex · 2026-06-07
General · sarvamai · 2026-03-03
General · ai21labs · 2026-01-06
Reasoning · openonerec · 2026-06-09
General · Liquid AI · 2026-01-05
General · Liquid AI · 2026-01-04
General · Liquid AI · 2025-12-25
General · Liquid AI · 2026-01-05
Multimodal · davidau · 2026-02-02
Multimodal · Alibaba · 2026-04-15
General · zai-org · 2026-01-19
General · inclusionai · 2026-02-09
General · Liquid AI · 2025-10-28
General · allenai · 2025-11-18
General · allenai · 2026-01-28
Chat · pearl-ai · 2026-02-26
General · huihui-ai · 2026-04-21
Multimodal · huihui-ai · 2026-04-18
Multimodal · Liquid AI · 2025-10-22
General · frontiersmind · 2026-08-03
General · alibaba-nlp · 2026-03-31
General · openai · 2025-08-04
Multimodal · Alibaba · 2026-02-24
Chat · Liquid AI · 2026-01-06
General · typhoon-ai · 2025-09-23
General · Alibaba · 2025-08-05
General · ibm-granite · 2025-09-16
Chat · bineric · 2026-01-12
General · ibm-granite · 2025-09-16
Multimodal · Liquid AI · 2025-08-12
General · allenai · 2025-09-12
General · openai · 2025-09-18
General · anton-hugging · 2026-02-06
General · openonerec · 2025-12-30
General · ibm-granite · 2025-09-16
General · bytedance · 2025-10-28
General · dreamfast · 2026-01-11
General · Liquid AI · 2025-10-07
General · z-lab · 2026-01-04
Multimodal · Liquid AI · 2025-08-12
General · Liquid AI · 2025-09-22
General · Liquid AI · 2025-09-30
General · Liquid AI · 2025-08-22
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-08-25
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
Reasoning · jackrong · 2026-03-07
General · NCAI · 2025-12-29
General · NCAI · 2025-12-29
Reasoning · DeepSeek · 2025-05-29
Coding · Alibaba · 2025-07-31
Chat · allenai · 2025-11-19
General · t-tech · 2025-12-22
General · hmellor · 2025-07-22
Chat · allenai · 2025-11-17
General · nvidia · 2025-08-12
General · inclusionai · 2025-11-25
General · ibm-granite · 2025-04-30
General · Liquid AI · 2025-07-10
General · z-lab · 2026-01-04
Reasoning · Microsoft · 2025-04-29
General · Alibaba · 2025-07-29
General · ibm-granite · 2025-09-16
Chat · xcuros · 2026-02-28
General · amd · 2025-05-17
General · stefanruseti · 2025-06-04
General · lgai-exaone · 2025-07-11
General · Liquid AI · 2025-07-10
General · menlo · 2025-06-25
General · Liquid AI · 2025-07-10
Multimodal · NCAI · 2025-12-29
Multimodal · Alibaba · 2025-01-26
General · huggingfacetb · 2025-07-08
General · Alibaba · 2025-09-23
General · nvidia · 2025-08-21
General · avitotech · 2025-10-20
General · huggingfacetb · 2025-06-19
General · Alibaba · 2025-09-23
General · dream-org · 2025-04-03
General · openbmb · 2025-09-02
General · Alibaba · 2025-09-23
General · pfnet · 2025-02-05
General · bytedance-seed · 2025-04-09
Chat · baseten · 2025-09-12
General · sapientinc · 2026-05-17
General · ibm-granite · 2025-10-07
Chat · utter-project · 2026-01-26
General · fableforge-ai · 2026-07-05
General · viorikaai-org · 2026-07-05
General · bananamind · 2026-07-17
General · lgai-exaone · 2025-03-12
General · maliosdark · 2026-07-09
Both bins run the same 800 GB/s memory bus, so token generation is the same on either one. The extra cores show up in image and video work, not in tokens per second.
| Configuration | Memory bandwidth | Memory options | Models that fit |
|---|---|---|---|
| 20-core CPU, 48-core GPU | 800 GB/s | 64, 128 GB | Identical |
| 20-core CPU, 64-core GPU | 800 GB/s | 64, 128 GB | Identical |
Nobody has submitted a benchmark on the M1 Ultra yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.
ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.
What each step actually changes for local models, rather than which one is newer.
Apple has not announced any of this. The M7 Max and Ultra are the first parts Apple designed after cancelling a generation to reach them, so extrapolating from the M5 under-represents them. These rows assume LPDDR6, whose wider channels grow every bus by half, at its top speed bin by the time the Max and Ultra ship. The 1.5 TB Ultra ceiling is Bloomberg's reported design target, and whether that configuration ships depends on the memory market. Stacked memory or a new package fabric would land above these numbers; nobody outside Apple can price that yet.
Projected chip · expected 2028
2765 GB/s · 256 to 1536 GB unified memory
40-core CPU · 112 or 128-core GPU · 152 TOPS Neural Engine
Would hold about a 2092B model at Q4
On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M1 Ultra's 800 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 128 GB machine can hand almost all of that to a model with no copy across a bus.
Apple stopped selling this one, which is exactly why it is interesting. The Studio exists for this workload. It carries the widest memory buses and the highest capacities Apple sells, and it runs at full clocks indefinitely. A used M1 Ultra at 128 GB still gives you 800 GB/s and a hard 172B ceiling, and neither number degrades with age the way a battery does.
Yes, at 128 GB. A 70B model at Q4_K_M needs about 46 GB including an 8K context, and 128 GB of unified memory leaves about 112 GB for weights once macOS takes its share. At 64 GB it does not fit at any quantization worth running.
Memory is the only spec that changes what you can run at all. 64 GB holds about a 85B model at Q4; 128 GB holds about 172B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.
Token generation is bandwidth-bound, so M1 Ultra throughput scales with its 800 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M1 Ultra lands in the tens of tokens per second and a 70B model lands in the single digits.
Not for LLMs. Both bins run the same 800 GB/s memory bus and take the same memory options, and token generation is bound by bandwidth rather than GPU cores. The extra cores show up in image generation and video work, not in tokens per second.
For inference, the specs that matter do not age: 800 GB/s and up to 128 GB of unified memory are the same numbers today as they were in 2022. A used M1 Ultra at the top memory option usually beats a new base-tier machine at the same price on both. Check the battery and the display, not the silicon.
ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.