Everything this chip runs, machine by machine: MacBook Pro 14" M5 Pro, MacBook Pro 16" M5 Pro
Showing 6563 of 6563 models
General · radixark · 2026-07-27
General · openbmb · 2026-05-21
General · farbodtavakkoli · 2026-06-17
General · Liquid AI · 2026-05-28
General · Liquid AI · 2026-07-28
General · Liquid AI · 2026-06-24
General · goekdeniz-guelmez · 2026-07-31
General · ktruestory · 2026-05-28
General · petrouil · 2026-07-23
General · weiboai · 2026-06-12
General · Liquid AI · 2026-03-31
General · ma7ee7 · 2026-07-30
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
Multimodal · Alibaba · 2026-02-28
General · ibm-granite · 2026-04-06
General · internscience · 2026-07-13
Multimodal · ibm-granite · 2026-04-16
General · nanbeige · 2026-07-21
General · cagrigungor · 2026-08-06
Multimodal · Alibaba · 2026-02-27
Multimodal · Google · 2026-03-02
Multimodal · Alibaba · 2026-02-27
Multimodal · Liquid AI · 2026-01-05
General · Liquid AI · 2026-01-20
General · ai21labs · 2026-01-06
Reasoning · openonerec · 2026-06-09
General · Liquid AI · 2026-01-05
General · nvidia · 2026-03-02
General · Liquid AI · 2026-01-04
General · Liquid AI · 2025-12-25
General · Liquid AI · 2026-01-05
Chat · Liquid AI · 2026-01-06
General · Liquid AI · 2025-10-28
General · Liquid AI · 2025-10-07
Multimodal · Liquid AI · 2025-10-22
General · Liquid AI · 2025-09-30
General · frontiersmind · 2026-08-03
Multimodal · davidau · 2026-02-02
General · hmellor · 2025-07-22
General · ibm-granite · 2025-09-16
Multimodal · Liquid AI · 2025-08-12
General · openonerec · 2025-12-30
General · ibm-granite · 2025-09-16
General · bytedance · 2025-10-28
Multimodal · Liquid AI · 2025-08-12
General · Liquid AI · 2025-09-22
General · Liquid AI · 2025-08-22
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-08-25
General · Liquid AI · 2025-09-03
General · Liquid AI · 2025-09-03
General · NCAI · 2025-12-29
General · typhoon-ai · 2025-09-23
General · ibm-granite · 2025-04-30
General · Liquid AI · 2025-07-10
General · ibm-granite · 2025-09-16
General · amd · 2025-05-17
General · stefanruseti · 2025-06-04
General · z-lab · 2026-01-04
General · ibm-granite · 2025-10-07
General · lgai-exaone · 2025-07-11
General · Liquid AI · 2025-07-10
General · Liquid AI · 2025-07-10
Multimodal · NCAI · 2025-12-29
Reasoning · DeepSeek · 2025-01-20
General · Alibaba · 2025-08-05
General · Alibaba · 2025-09-23
Reasoning · Microsoft · 2025-04-29
General · pfnet · 2025-02-05
Chat · uzlm · 2025-09-03
General · bytedance-seed · 2025-04-09
General · sapientinc · 2026-05-17
General · menlo · 2025-06-25
General · fableforge-ai · 2026-07-05
General · viorikaai-org · 2026-07-05
Reasoning · jackrong · 2026-03-16
General · bananamind · 2026-07-17
General · lgai-exaone · 2025-03-12
General · maliosdark · 2026-07-09
General · raidium · 2026-06-15
Embedding · taide · 2026-06-12
General · darthcrawl · 2026-05-07
General · Alibaba · 2025-04-27
General · Alibaba · 2025-04-27
Multimodal · Alibaba · 2025-01-26
Multimodal · Google · 2026-03-02
Multimodal · zai-org
Multimodal · Alibaba
Multimodal · datalab-to
General · Alibaba · 2025-04-28
General · huggingfacetb · 2025-07-08
Multimodal · openbmb
General · Alibaba · 2025-04-28
Multimodal · rednote-hilab
Multimodal · tencent
Multimodal · dots-studio
General · distil-labs
General · farbodtavakkoli
Multimodal · raxcore-dev
General · huggingfacetb · 2025-06-19
General · jinaai
Multimodal · ath-maas
Multimodal · Alibaba
Multimodal · Liquid AI
General · ravichandranj · 2026-02-13
Reasoning · typhoon-ai
Multimodal · paddlepaddle
General · lukebailey181pub · 2026-04-21
General · Liquid AI
Multimodal · paddlepaddle
Chat · baseten · 2025-09-12
General · adamlucek
Coding · shahriarferdoush
General · ahczhg
General · onnx-community · 2025-04-28
General · etherll
General · baidu
General · Liquid AI
General · openbmb
Multimodal · lkhl
Multimodal · infly
General · benjamin
General · lemonelabs
General · openbmb · 2025-06-05
General · farbodtavakkoli
General · novachronoai
Multimodal · paddlepaddle
Reasoning · khazarai
General · kamilamila
General · launch
General · arcee-ai
General · getonit
General · Liquid AI
General · saidutta69
General · agentica-org
General · novaciano
General · dmusingu
General · kgrabko
General · ordenwills
General · smcleish
General · carsenk
Reasoning · nvidia
Multimodal · zero-point-ai
General · openbmb
General · ibm-granite
General · osaurusai
General · pyoakum
Model not in the list? Paste a HuggingFace URL or ID for an instant fit check.
It depends on the model size and quantization. A 7B parameter model at Q4 quantization needs about 5 GB of RAM, while a 70B model needs 40+ GB. Apple Silicon Macs use unified memory, so your entire RAM pool is available for model weights — no separate VRAM required.
Not comfortably. A 70B model at Q4 quantization needs about 40 GB of RAM. The MacBook Air maxes out at 24-32 GB depending on the generation. You'd need a Mac Studio or MacBook Pro with 48+ GB for a 70B model to run well.
Speed depends on your chip's memory bandwidth and the model size. Smaller models (3-7B) run fastest — expect 40-70+ tokens per second on M2 Pro or better. Use the calculator above to see estimated speeds for your specific Mac.
Quantization reduces model precision to use less memory. Q8 (8-bit) is nearly lossless. Q4 (4-bit) reduces memory by ~75% with minor quality loss — it's the sweet spot for most users. Q2 (2-bit) saves the most memory but noticeably degrades output quality.
Apple Silicon uses unified memory — CPU and GPU share the same RAM pool. A Mac with 32 GB can load a 28 GB model directly. On NVIDIA systems, you're limited by GPU VRAM (typically 8-24 GB on consumer cards), even if the PC has 64 GB of system RAM.
ToolPiper uses Metal GPU acceleration via llama.cpp for LLM inference on Apple Silicon. The GPU and CPU share unified memory, so there's no data transfer overhead. The Neural Engine (ANE) is used for specific tasks like super-resolution and pose detection.
Yes, if you have enough RAM. ToolPiper manages model loading and can keep multiple models in memory simultaneously. When memory gets tight, it automatically evicts the least recently used model to make room for a new one.
GGUF is the standard format for running quantized models with llama.cpp (and ToolPiper). It supports all quantization levels and runs on CPU+GPU. MLX is Apple's format optimized for Apple Silicon. AWQ and GPTQ are NVIDIA-focused formats that don't run natively on Mac.
A cross-section of all 6563 models. Open any one to see which Macs run it, then follow its related models to work through the rest of the catalog.
Model database updated: 2026-08-21 · 6563 models