← All MacBook Pro 14" models

MacBook Pro 14" M4 Max

The M4 Max MacBook Pro 14" runs local models at 410 to 546 GB/s of memory bandwidth with 36 to 128 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 128 GB you can hold roughly a 207B dense model at Q4. The Max is where the bus gets wide enough that model size, not bandwidth, becomes the thing you plan around.

Apple no longer sells this configuration new. It stays fully evaluated here because the used market is where most of its local AI value now sits.

By , Founder & Lead Engineer— Updated

Specifications

ChipApple M4 Max
CPU cores14 or 16
GPU cores32 or 40
Unified memory36, 48, 64, or 128 GB
Memory bandwidth410 to 546 GB/s
Neural Engine38 TOPS
Released2024
AvailabilityUsed market

Memory bandwidth is faster than 70% of the Apple Silicon chips shipped in a Mac, against a 1228 GB/s peak.

Its memory ceiling is above 65% of them, against a 512 GB peak.

Pick your configuration

Every option Apple sells with this chip. The model list below recomputes against the one you pick.

GPU cores

Unified memory

Apple couples memory to the core count on this chip, so the options change with the bin above.

What each memory option runs

Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.

36 GB unified memory

31 GiB forecast budget for models · 410 GB/s

6,270 of 6,563 models fit by their known memory.

A browser cannot measure a runtime's working memory, so no fit here is confirmed. ToolPiper measures the file on your Mac and confirms it.

Largest model at Q4
Qwen2 57B A14B Instruct · 57.41B
Best all-round pick
JOSIE 2 2B OSS · Q8_0 · ~94 tok/s

48 GB unified memory

42 GiB forecast budget for models · 546 GB/s

6,339 of 6,563 models fit by their known memory.

A browser cannot measure a runtime's working memory, so no fit here is confirmed. ToolPiper measures the file on your Mac and confirms it.

Largest model at Q4
GLM 5.1 JANG_1L · 74.39B
Best all-round pick
JOSIE 2 2B OSS · Q8_0 · ~126 tok/s

64 GB unified memory

56 GiB forecast budget for models · 546 GB/s

6,373 of 6,563 models fit by their known memory.

A browser cannot measure a runtime's working memory, so no fit here is confirmed. ToolPiper measures the file on your Mac and confirms it.

Largest model at Q4
Midnight Miqu 103B v1.5 · 103.2B
Best all-round pick
JOSIE 2 2B OSS · Q8_0 · ~126 tok/s

128 GB unified memory

112 GiB forecast budget for models · 546 GB/s

6,446 of 6,563 models fit by their known memory.

A browser cannot measure a runtime's working memory, so no fit here is confirmed. ToolPiper measures the file on your Mac and confirms it.

Largest model at Q4
Step 3.7 Flash · 201.37B
Best all-round pick
JOSIE 2 2B OSS · Q8_0 · ~126 tok/s

What a 36 GB M4 Max MacBook Pro 14" can run

Every model in the database against this exact configuration, at 410 GB/s. Ratings and speeds are the same numbers the model pages show.

Showing 6563 of 6563 models

General · goekdeniz-guelmez · 2026-07-31

Q8_0 Fit unconfirmed
at least 2.2 GiB6% of RAM~94 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-07-28

Q8_0 Fit unconfirmed
at least 2.6 GiB7% of RAM~80 tok/sEstimated2.7B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-06-24

Q8_0 Fit unconfirmed
at least 0.2 GiB1% of RAM~935 tok/sEstimated0.23B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ma7ee7 · 2026-07-30

Q8_0 Fit unconfirmed
at least 3.1 GiB8% of RAM~69 tok/sEstimated3.13B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · radixark · 2026-07-27

Q8_0 Fit unconfirmed
at least 2.2 GiB6% of RAM~95 tok/sEstimated2.25B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · weiboai · 2026-06-12

Q8_0 Fit unconfirmed
at least 3.0 GiB8% of RAM~70 tok/sEstimated3.09B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ktruestory · 2026-05-28

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~199 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · internscience · 2026-07-13

Q8_0 Fit unconfirmed
at least 4.4 GiB12% of RAM~47 tok/sEstimated4.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nanbeige · 2026-07-21

Q8_0 Fit unconfirmed
at least 4.1 GiB11% of RAM~52 tok/sEstimated4.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openbmb · 2026-05-21

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~199 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-03-31

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~606 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0 Fit unconfirmed
at least 3.3 GiB9% of RAM~63 tok/sEstimated3.4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · ibm-granite · 2026-04-16

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~54 tok/sEstimated4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · openonerec · 2026-06-09

Q8_0 Fit unconfirmed
at least 0.8 GiB2% of RAM~268 tok/sEstimated0.8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 0.9 GiB2% of RAM~246 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 0.9 GiB2% of RAM~246 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 2.2 GiB6% of RAM~94 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 2.2 GiB6% of RAM~94 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · cagrigungor · 2026-08-06

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~801 tok/sEstimated0.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-20

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nvidia · 2026-03-02

Q8_0 Fit unconfirmed
at least 3.7 GiB10% of RAM~56 tok/sEstimated3.83B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · davidau · 2026-02-02

Q8_0 Fit unconfirmed
at least 4.2 GiB12% of RAM~50 tok/sEstimated4.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · frontiersmind · 2026-08-03

Q8_0 Fit unconfirmed
at least 0.6 GiB2% of RAM~331 tok/sEstimated0.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-12-25

Q8_0 Fit unconfirmed
at least 2.5 GiB7% of RAM~84 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 2.5 GiB7% of RAM~84 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Liquid AI · 2026-01-06

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-04

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 1.6 GiB4% of RAM~135 tok/sEstimated1.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lukebailey181pub · 2026-04-21

Q8_0 Fit unconfirmed
at least 6.8 GiB19% of RAM~31 tok/sEstimated6.91B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 4.6 GiB13% of RAM~46 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 4.6 GiB13% of RAM~46 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0 Fit unconfirmed
at least 5.0 GiB14% of RAM~42 tok/sEstimated5.12B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-10-28

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~608 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-10-22

Q8_0 Fit unconfirmed
at least 2.9 GiB8% of RAM~72 tok/sEstimated3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ravichandranj · 2026-02-13

Q8_0 Fit unconfirmed
at least 6.3 GiB17% of RAM~33 tok/sEstimated6.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · deepreinforce-ai · 2026-06-21

Q8_0 Fit unconfirmed
at least 8.0 GiB22% of RAM~26 tok/sEstimated8.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-22

Q8_0 Fit unconfirmed
at least 2.5 GiB7% of RAM~84 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-30

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~606 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openonerec · 2025-12-30

Q8_0 Fit unconfirmed
at least 2.1 GiB6% of RAM~101 tok/sEstimated2.13B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · baseten · 2025-09-12

Q8_0 Fit unconfirmed
at least 3.1 GiB9% of RAM~67 tok/sEstimated3.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0 Fit unconfirmed
at least 3.1 GiB9% of RAM~67 tok/sEstimated3.19B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0 Fit unconfirmed
at least 3.3 GiB9% of RAM~63 tok/sEstimated3.4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · bytedance · 2025-10-28

Q8_0 Fit unconfirmed
at least 1.4 GiB4% of RAM~150 tok/sEstimated1.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lgai-exaone · 2025-07-11

Q8_0 Fit unconfirmed
at least 1.3 GiB3% of RAM~168 tok/sEstimated1.28B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-08-22

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~606 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~606 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~606 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-08-25

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~606 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~289 tok/sEstimated0.74B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0 Fit unconfirmed
at least 1.5 GiB4% of RAM~136 tok/sEstimated1.58B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0 Fit unconfirmed
at least 0.4 GiB1% of RAM~476 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · obliteratus · 2026-04-15

Q8_0 Fit unconfirmed
at least 7.8 GiB22% of RAM~27 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · hmellor · 2025-07-22

Q8_0 Fit unconfirmed
at least 1.2 GiB3% of RAM~174 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · typhoon-ai · 2025-09-23

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · uzlm · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.2 GiB3% of RAM~174 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~54 tok/sEstimated4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · menlo · 2025-06-25

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Alibaba · 2025-08-05

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-08-05

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · allenai · 2025-11-19

Q8_0 Fit unconfirmed
at least 7.1 GiB20% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · allenai · 2025-11-17

Q8_0 Fit unconfirmed
at least 7.1 GiB20% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · allenai · 2025-11-18

Q8_0 Fit unconfirmed
at least 7.1 GiB20% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · amd · 2025-05-17

Q8_0 Fit unconfirmed
at least 1.5 GiB4% of RAM~143 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · empero-ai · 2026-06-19

Q8_0 Fit unconfirmed
at least 9.2 GiB26% of RAM~23 tok/sEstimated9.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0 Fit unconfirmed
at least 7.8 GiB22% of RAM~27 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-16

Q8_0 Fit unconfirmed
at least 8.2 GiB23% of RAM~26 tok/sEstimated8.38B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · pearl-ai · 2026-02-26

Q8_0 Fit unconfirmed
at least 7.9 GiB22% of RAM~27 tok/sEstimated8.03B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · stefanruseti · 2025-06-04

Q8_0 Fit unconfirmed
at least 1.2 GiB3% of RAM~174 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · bananamind · 2026-07-17

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~17,800 tok/sEstimated0.01B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · goekdeniz-guelmez · 2026-07-31

Q8_0 Fit unconfirmed
at least 9.4 GiB26% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huggingfacetb · 2025-07-08

Q8_0 Fit unconfirmed
at least 3.0 GiB8% of RAM~70 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lgai-exaone · 2025-03-12

Q8_0 Fit unconfirmed
at least 2.4 GiB7% of RAM~89 tok/sEstimated2.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · maliosdark · 2026-07-09

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~3,967 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~286 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · viorikaai-org · 2026-07-05

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~4,728 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · allenai · 2026-01-28

Q8_0 Fit unconfirmed
at least 7.3 GiB20% of RAM~29 tok/sEstimated7.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · fableforge-ai · 2026-07-05

Q8_0 Fit unconfirmed
at least 1.5 GiB4% of RAM~139 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · ibm-granite · 2025-04-09

Q8_0 Fit unconfirmed
at least 2.5 GiB7% of RAM~85 tok/sEstimated2.53B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-10-07

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~609 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0 Fit unconfirmed
at least 8.6 GiB24% of RAM~24 tok/sEstimated8.79B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · Microsoft · 2025-04-29

Q8_0 Fit unconfirmed
at least 3.8 GiB10% of RAM~56 tok/sEstimated3.84B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · naver-hyperclovax · 2025-04-22

Q8_0 Fit unconfirmed
at least 3.6 GiB10% of RAM~58 tok/sEstimated3.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nvidia · 2026-03-18

Q8_0 Fit unconfirmed
at least 8.3 GiB23% of RAM~25 tok/sEstimated8.49B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · raidium · 2026-06-15

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~9,594 tok/sEstimated0.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Embedding · taide · 2026-06-12

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~709 tok/sEstimated0.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · alibaba-nlp · 2026-03-31

Q8_0 Fit unconfirmed
at least 8.0 GiB22% of RAM~26 tok/sEstimated8.19B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huggingfacetb · 2025-06-19

Q8_0 Fit unconfirmed
at least 3.0 GiB8% of RAM~70 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2025-01-26

Q8_0 Fit unconfirmed
at least 3.7 GiB10% of RAM~57 tok/sEstimated3.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0 Fit unconfirmed
at least 4.3 GiB12% of RAM~49 tok/sEstimated4.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · shaungves · 2026-06-19

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · xcuros · 2026-02-28

Q8_0 Fit unconfirmed
at least 7.4 GiB21% of RAM~28 tok/sEstimated7.62B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · allenai · 2025-09-12

Q8_0 Fit unconfirmed
at least 7.1 GiB20% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · anton-hugging · 2026-02-06

Q8_0 Fit unconfirmed
at least 7.4 GiB21% of RAM~28 tok/sEstimated7.62B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · darthcrawl · 2026-05-07

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~5,688 tok/sEstimated0.04B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0 Fit unconfirmed
at least 1.7 GiB5% of RAM~121 tok/sEstimated1.78B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2025-02-19

Q8_0 Fit unconfirmed
at least 3.8 GiB10% of RAM~56 tok/sEstimated3.84B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · pfnet · 2025-02-05

Q8_0 Fit unconfirmed
at least 1.3 GiB4% of RAM~166 tok/sEstimated1.29B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · preparebuddy · 2026-06-02

Q8_0 Fit unconfirmed
at least 3.0 GiB8% of RAM~70 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · sapientinc · 2026-05-17

Q8_0 Fit unconfirmed
at least 1.2 GiB3% of RAM~182 tok/sEstimated1.18B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · k-intelligence · 2025-07-03

Q8_0 Fit unconfirmed
at least 2.3 GiB6% of RAM~93 tok/sEstimated2.31B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · lgai-exaone · 2024-12-01

Q8_0 Fit unconfirmed
at least 2.4 GiB7% of RAM~89 tok/sEstimated2.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~286 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 2.0 GiB6% of RAM~106 tok/sEstimated2.03B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · sunbird · 2026-07-19

Q8_0 Fit unconfirmed
at least 5.0 GiB14% of RAM~42 tok/sEstimated5.1B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Coding · mahiatlinux · 2026-06-30

Q8_0 Fit unconfirmed
at least 4.6 GiB13% of RAM~46 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · onnx-community · 2025-04-28

Q8_0 Fit unconfirmed
at least 0.4 GiB1% of RAM~478 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openbmb · 2025-06-05

Q8_0 Fit unconfirmed
at least 0.4 GiB1% of RAM~495 tok/sEstimated0.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · jackrong · 2026-03-16

Q8_0 Fit unconfirmed
at least 9.4 GiB26% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-05-31

Q8_0 Fit unconfirmed
at least 0.5 GiB1% of RAM~435 tok/sEstimated0.49B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-05-31

Q8_0 Fit unconfirmed
at least 1.5 GiB4% of RAM~139 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-09-15

Q8_0 Fit unconfirmed
at least 1.5 GiB4% of RAM~139 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 0.6 GiB2% of RAM~360 tok/sEstimated0.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 1.7 GiB5% of RAM~125 tok/sEstimated1.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 3.9 GiB11% of RAM~53 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 9.4 GiB26% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-26

Q8_0 Fit unconfirmed
at least 9.4 GiB26% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Google · 2022-03-02

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~19,362 tok/sEstimated0.01B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · huihui-ai · 2024-10-01

Q8_0 Fit unconfirmed
at least 1.5 GiB4% of RAM~143 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2025-02-24

Q8_0 Fit unconfirmed
at least 5.5 GiB15% of RAM~39 tok/sEstimated5.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · neuralcrew · 2025-09-14

Q8_0 Fit unconfirmed
at least 5.9 GiB16% of RAM~36 tok/sEstimated6.04B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · swiss-ai · 2026-04-12

Q8_0 Fit unconfirmed
at least 3.7 GiB10% of RAM~56 tok/sEstimated3.83B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0 Fit unconfirmed
at least 7.8 GiB22% of RAM~27 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · obliteratus · 2026-06-05

Q8_0 Fit unconfirmed
at least 11.7 GiB32% of RAM~18 tok/sEstimated11.96B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · supralabs · 2026-06-03

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~4,147 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · tiger-lab · 2024-10-08

Q8_0 Fit unconfirmed
at least 4.1 GiB11% of RAM~52 tok/sEstimated4.15B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huihui-ai · 2026-07-11

Q8_0 Fit unconfirmed
at least 11.7 GiB32% of RAM~18 tok/sEstimated11.96B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2024-04-22

Q8_0 Fit unconfirmed
at least 3.7 GiB10% of RAM~56 tok/sEstimated3.82B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2024-05-19

Q8_0 Fit unconfirmed
at least 4.1 GiB11% of RAM~52 tok/sEstimated4.15B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2024-08-16

Q8_0 Fit unconfirmed
at least 3.7 GiB10% of RAM~56 tok/sEstimated3.82B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · swiss-ai · 2025-08-13

Q8_0 Fit unconfirmed
at least 7.9 GiB22% of RAM~27 tok/sEstimated8.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · t-tech · 2025-12-22

Q8_0 Fit unconfirmed
at least 8.0 GiB22% of RAM~26 tok/sEstimated8.19B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · yuxinlu1 · 2026-06-28

Q8_0 Fit unconfirmed
at least 11.7 GiB32% of RAM~18 tok/sEstimated11.96B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · 8f-ai

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~285 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · abdallalswaiti

Q8_0 Fit unconfirmed
at least 3.1 GiB9% of RAM~67 tok/sEstimated3.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · ath-maas

Q8_0 Fit unconfirmed
at least 0.8 GiB2% of RAM~252 tok/sEstimated0.85B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · adamlucek

Q8_0 Fit unconfirmed
at least 1.2 GiB3% of RAM~174 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · alignmentresearch

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~2,859 tok/sEstimated0.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.2 GiB3% of RAM~174 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · andrew0425

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~199 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · bytedance

Q8_0 Fit unconfirmed
at least 3.7 GiB10% of RAM~57 tok/sEstimated3.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~3,356 tok/sEstimated0.06B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · daremodels

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~199 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB3% of RAM~184 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

32-core vs 40-core GPU

Apple ties the memory bus to the bin on this chip: 410 GB/s at 32 cores and 546 GB/s at 40. That moves tokens per second. It does not move which models fit, because that is memory, not cores.

ConfigurationMemory bandwidthMemory optionsModels that fit
14-core CPU, 32-core GPU410 GB/s36 GBIdentical
16-core CPU, 40-core GPU546 GB/s48, 64, 128 GBIdentical

On GLM 5.1 JANG_1L the 32-core generates about 8 tok/s and the 40-core about 11 tok/s, a 33% difference. Both hold the model at the same quantization.

Published observations for M4 Max

No compatible observations are published here. Formula estimates remain labelled, and unsupported forecasts remain unavailable.

Benchmark contribution is disclosed before a run, and the leaderboard shows every chip that already has one.

What comes next

Projection dated 2026-09-08

Apple has not announced any of this. The M7 Max and Ultra are the first parts Apple designed after cancelling a generation to reach them, so extrapolating from the M5 under-represents them. These rows assume LPDDR6, whose wider channels grow every bus by half, at its top speed bin by the time the Max and Ultra ship. The 1.5 TB Ultra ceiling is Bloomberg's reported design target, and whether that configuration ships depends on the memory market. Stacked memory or a new package fabric would land above these numbers; nobody outside Apple can price that yet.

M7 MaxProjected

Projected chip · expected 2027

1382 GB/s · 48 to 384 GB unified memory

20-core CPU · 56 or 64-core GPU · 76 TOPS Neural Engine

Would hold about a 624B model at Q4

Why unified memory is the number that matters

On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M4 Max's 546 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 128 GB machine can hand almost all of that to a model with no copy across a bus.

Apple stopped selling this one, which is exactly why it is interesting. The 14-inch chassis cools well enough to hold its clocks through a long generation run, and it is the smallest machine Apple puts a Max chip in. A used M4 Max at 128 GB still gives you 546 GB/s and a hard 207B ceiling, and neither number degrades with age the way a battery does.

Common questions

Can the M4 Max MacBook Pro 14" run a 70B model?

Yes, at 128 GB. A 70B model at Q4_K_M needs about 38 GB including an 8K context, and 128 GB of unified memory leaves about 112 GB for weights once macOS takes its share. At 36 GB it does not fit at any quantization worth running.

How much unified memory should I get with the M4 Max MacBook Pro 14"?

Memory is the only spec that changes what you can run at all. 36 GB holds about a 57B model at Q4; 128 GB holds about 207B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.

How fast are local LLMs on the M4 Max?

Token generation is bandwidth-bound, so M4 Max throughput scales with its 546 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M4 Max lands in the tens of tokens per second and a 70B model lands in the single digits.

Is the 40-core GPU worth it over the 32-core on the M4 Max?

For throughput, yes: Apple ties bandwidth to the bin here, so the 32-core runs at 410 GB/s and the 40-core at 546 GB/s, about 33% more. For fit, no: both bins hold exactly the same models, because that is set by memory rather than by cores.

Is a used M4 Max MacBook Pro 14" still worth buying for local AI?

For inference, the specs that matter do not age: 546 GB/s and up to 128 GB of unified memory are the same numbers today as they were in 2024. A used M4 Max at the top memory option usually beats a new base-tier machine at the same price on both. Check the battery and the display, not the silicon.

Run these models on your MacBook Pro 14"

ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.