The M4 MacBook Air runs local models at 120 GB/s of memory bandwidth with 16 to 32 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 32 GB you can hold roughly a 42B dense model at Q4. A base M chip is the narrow end of the memory bus. It runs small models pleasantly and stops hard at the memory ceiling, which is the constraint you will hit first.
Apple no longer sells this configuration new. It stays fully evaluated here because the used market is where most of its local AI value now sits.
Memory bandwidth is faster than 17% of the Apple Silicon chips shipped in a Mac, against a 819 GB/s peak.
Its memory ceiling is above 17% of them, against a 512 GB peak.
Every option Apple sells with this chip. The model list below recomputes against the one you pick.
GPU cores
Unified memory
Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.
3,128 of 3,641 models fit, and 2,907 of them run with headroom rather than as a squeeze.
3,352 of 3,641 models fit, and 3,007 of them run with headroom rather than as a squeeze.
3,388 of 3,641 models fit, and 3,127 of them run with headroom rather than as a squeeze.
Every model in the database against this exact configuration, at 120 GB/s. Ratings and speeds are the same numbers the model pages show.
Showing 3641 of 3641 models
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Reasoning · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Reasoning · Liquid AI · 2025-11-28
Multimodal · Alibaba · 2026-02-28
General · ibm-granite · 2025-09-16
Multimodal · Alibaba · 2026-02-28
Chat · Liquid AI · 2025-11-28
Chat · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
General · NCAI · 2025-12-29
Multimodal · Liquid AI · 2025-11-28
General · Liquid AI · 2025-11-28
Multimodal · Liquid AI · 2025-11-28
General · raidium · 2026-06-15
Multimodal · NCAI · 2025-12-29
Embedding · taide · 2026-06-12
General · Alibaba · 2025-04-27
General · Alibaba
Multimodal · datalab-to
General · openbmb
General · distil-labs
General · ibm-granite
General · Google
General · openai
General · paddlepaddle
General · Liquid AI
General · openbmb
General · baidu
General · arcee-ai
Coding · shahriarferdoush
General · etherll
General · LG AI · 2025-07-15
General · farbodtavakkoli
General · kamilamila
General · paddlepaddle
General · ordenwills
General · ibm-granite
General · pyoakum
General · ibm-granite
General · thkim0305
Reasoning · jackrong
Coding · rahul7star
Coding · z-lab
General · osaurusai
General · Microsoft
General · Microsoft
General · tencent
General · primeintellect
Multimodal · zai-org
Multimodal · openbmb
Reasoning · HuggingFace · 2025-07-08
Reasoning · DeepSeek
General · hmellor
General · farbodtavakkoli
Multimodal · lkhl
Reasoning · typhoon-ai
General · pfnet
General · adamlucek
General · ahczhg
General · abaryan
General · smcleish
General · carsenk
Reasoning · nvidia
General · bezzam
General · openbmb
General · ibm-granite
General · treadon
General · skis-ai-research
General · artificialguybr
Multimodal · Alibaba · 2026-02-27
General · Alibaba · 2025-04-27
General · Alibaba
Multimodal · Alibaba
General · Alibaba
General · Alibaba
General · Microsoft
Multimodal · opengvlab
General · Alibaba
Multimodal · Alibaba · 2026-02-27
General · voyageai
General · Microsoft
General · farbodtavakkoli
General · farbodtavakkoli
General · jinaai
General · farbodtavakkoli
General · jakobhuss
General · Upstage
General · llava-hf
General · amd
General · openbmb
General · mrs83
Coding · DeepSeek
General · ibm-granite
General · numind
General · toxicityprompts
General · turing-motors
General · ibm-granite
General · qnguyen3
General · getonit
General · ysmao
General · opengvlab
General · agentica-org
General · Liquid AI · 2025-11-28
General · novaciano
General · lmms-lab
General · tabularisai
General · numind
General · kgrabko
General · isotonic
General · thisisiron
General · nvidia
General · manycore-research
General · dllm-hub
General · huihui-ai
General · tencent
General · amd
General · menlo
General · lunahr
General · TII
General · ibm-granite
General · lunahr
General · roystar
General · dllm-hub
General · weiboai
General · TII
General · pjmixers-dev
General · Liquid AI · 2025-11-28
General · Google
Multimodal · opendatalab
Multimodal · tencent
Both bins run the same 120 GB/s memory bus, so token generation is the same on either one. The extra cores show up in image and video work, not in tokens per second.
| Configuration | Memory bandwidth | Memory options | Models that fit |
|---|---|---|---|
| 10-core CPU, 8-core GPU | 120 GB/s | 16, 24, 32 GB | Identical |
| 10-core CPU, 10-core GPU | 120 GB/s | 16, 24, 32 GB | Identical |
Nobody has submitted a benchmark on the M4 yet, so every speed on this page is the formula estimate rather than a measured run. The estimate is bandwidth-driven and calibrated against chips that do have data, which makes it a good guide and not a promise.
ToolPiper contributes a result anonymously when you run the benchmark, and the leaderboard shows every chip that already has one.
What each step actually changes for local models, rather than which one is newer.
MacBook Air M5
1.3x the memory bandwidth
Used market alternativeMacBook Air M3
17% less memory bandwidth, 24 GB ceiling instead of 32 GB
Same chip, other MacMacBook Pro 14" M4
The same chip in a different Mac
Same chip, other MacMac mini M4
The same chip in a different Mac
Same chip, other MaciMac M4
The same chip in a different Mac
On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M4's 120 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 32 GB machine can hand almost all of that to a model with no copy across a bus.
Apple stopped selling this one, which is exactly why it is interesting. The Air has no fan, so a long generation run settles into a lower sustained clock than the same chip in a Pro. Model loading and short chats are unaffected; a multi-hour batch job is where you would notice. A used M4 at 32 GB still gives you 120 GB/s and a hard 42B ceiling, and neither number degrades with age the way a battery does.
No. A 70B model at Q4_K_M needs about 46 GB, and the largest M4 MacBook Air tops out at 32 GB, which leaves about 28 GB for weights. The practical ceiling on this machine is around 42B parameters at Q4.
Memory is the only spec that changes what you can run at all. 16 GB holds about a 20B model at Q4; 32 GB holds about 42B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.
Token generation is bandwidth-bound, so M4 throughput scales with its 120 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M4 lands in the tens of tokens per second and a 70B model lands in the single digits.
Not for LLMs. Both bins run the same 120 GB/s memory bus and take the same memory options, and token generation is bound by bandwidth rather than GPU cores. The extra cores show up in image generation and video work, not in tokens per second.
For inference, the specs that matter do not age: 120 GB/s and up to 32 GB of unified memory are the same numbers today as they were in 2024. A used M4 at the top memory option usually beats a new base-tier machine at the same price on both. Check the battery and the display, not the silicon.
ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.