← All Mac Studio models

Mac Studio M1 Max

The M1 Max Mac Studio runs local models at 400 GB/s of memory bandwidth with 32 to 64 GB of unified memory. On Apple Silicon that memory is shared with the GPU, so the whole pool is available for weights: at 64 GB you can hold roughly a 103B dense model at Q4. The Max is where the bus gets wide enough that model size, not bandwidth, becomes the thing you plan around.

Apple no longer sells this configuration new. It stays fully evaluated here because the used market is where most of its local AI value now sits.

By , Founder & Lead Engineer— Updated

Specifications

ChipApple M1 Max
CPU cores10
GPU cores24 or 32
Unified memory32 or 64 GB
Memory bandwidth400 GB/s
Neural Engine11 TOPS
Released2021
AvailabilityUsed market

Memory bandwidth is faster than 55% of the Apple Silicon chips shipped in a Mac, against a 1228 GB/s peak.

Its memory ceiling is above 45% of them, against a 512 GB peak.

Pick your configuration

Every option Apple sells with this chip. The model list below recomputes against the one you pick.

GPU cores

Unified memory

What each memory option runs

Unified memory is the ceiling and it is soldered, so this is the decision you cannot revisit.

32 GB unified memory

28 GiB forecast budget for models · 400 GB/s

6,260 of 6,563 models fit by their known memory.

A browser cannot measure a runtime's working memory, so no fit here is confirmed. ToolPiper measures the file on your Mac and confirms it.

Largest model at Q4
ODM_1B_params_50B_tokens · 50B
Best all-round pick
JOSIE 2 2B OSS · Q8_0 · ~92 tok/s

64 GB unified memory

56 GiB forecast budget for models · 400 GB/s

6,373 of 6,563 models fit by their known memory.

A browser cannot measure a runtime's working memory, so no fit here is confirmed. ToolPiper measures the file on your Mac and confirms it.

Largest model at Q4
Midnight Miqu 103B v1.5 · 103.2B
Best all-round pick
JOSIE 2 2B OSS · Q8_0 · ~92 tok/s

What a 32 GB M1 Max Mac Studio can run

Every model in the database against this exact configuration, at 400 GB/s. Ratings and speeds are the same numbers the model pages show.

Showing 6563 of 6563 models

General · goekdeniz-guelmez · 2026-07-31

Q8_0 Fit unconfirmed
at least 2.2 GiB7% of RAM~92 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-07-28

Q8_0 Fit unconfirmed
at least 2.6 GiB8% of RAM~78 tok/sEstimated2.7B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-06-24

Q8_0 Fit unconfirmed
at least 0.2 GiB1% of RAM~912 tok/sEstimated0.23B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ma7ee7 · 2026-07-30

Q8_0 Fit unconfirmed
at least 3.1 GiB10% of RAM~67 tok/sEstimated3.13B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · radixark · 2026-07-27

Q8_0 Fit unconfirmed
at least 2.2 GiB7% of RAM~93 tok/sEstimated2.25B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · weiboai · 2026-06-12

Q8_0 Fit unconfirmed
at least 3.0 GiB9% of RAM~68 tok/sEstimated3.09B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ktruestory · 2026-05-28

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~194 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · internscience · 2026-07-13

Q8_0 Fit unconfirmed
at least 4.4 GiB14% of RAM~46 tok/sEstimated4.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nanbeige · 2026-07-21

Q8_0 Fit unconfirmed
at least 4.1 GiB13% of RAM~50 tok/sEstimated4.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openbmb · 2026-05-21

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~194 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-03-31

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~591 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0 Fit unconfirmed
at least 3.3 GiB10% of RAM~62 tok/sEstimated3.4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · openonerec · 2026-06-09

Q8_0 Fit unconfirmed
at least 0.8 GiB2% of RAM~261 tok/sEstimated0.8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 0.9 GiB3% of RAM~240 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 0.9 GiB3% of RAM~240 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 2.2 GiB7% of RAM~92 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 2.2 GiB7% of RAM~92 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · cagrigungor · 2026-08-06

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~782 tok/sEstimated0.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · ibm-granite · 2026-04-16

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-20

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nvidia · 2026-03-02

Q8_0 Fit unconfirmed
at least 3.7 GiB12% of RAM~55 tok/sEstimated3.83B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · davidau · 2026-02-02

Q8_0 Fit unconfirmed
at least 4.2 GiB13% of RAM~49 tok/sEstimated4.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · frontiersmind · 2026-08-03

Q8_0 Fit unconfirmed
at least 0.6 GiB2% of RAM~323 tok/sEstimated0.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-12-25

Q8_0 Fit unconfirmed
at least 2.5 GiB8% of RAM~82 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 2.5 GiB8% of RAM~82 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Liquid AI · 2026-01-06

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-04

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 1.6 GiB5% of RAM~131 tok/sEstimated1.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lukebailey181pub · 2026-04-21

Q8_0 Fit unconfirmed
at least 6.8 GiB21% of RAM~30 tok/sEstimated6.91B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 4.6 GiB14% of RAM~45 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 4.6 GiB14% of RAM~45 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0 Fit unconfirmed
at least 5.0 GiB16% of RAM~41 tok/sEstimated5.12B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-10-28

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~593 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-10-22

Q8_0 Fit unconfirmed
at least 2.9 GiB9% of RAM~70 tok/sEstimated3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ravichandranj · 2026-02-13

Q8_0 Fit unconfirmed
at least 6.3 GiB20% of RAM~33 tok/sEstimated6.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-22

Q8_0 Fit unconfirmed
at least 2.5 GiB8% of RAM~82 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-30

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~591 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openonerec · 2025-12-30

Q8_0 Fit unconfirmed
at least 2.1 GiB7% of RAM~98 tok/sEstimated2.13B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · baseten · 2025-09-12

Q8_0 Fit unconfirmed
at least 3.1 GiB10% of RAM~65 tok/sEstimated3.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · deepreinforce-ai · 2026-06-21

Q8_0 Fit unconfirmed
at least 8.0 GiB25% of RAM~26 tok/sEstimated8.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0 Fit unconfirmed
at least 3.1 GiB10% of RAM~66 tok/sEstimated3.19B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0 Fit unconfirmed
at least 3.3 GiB10% of RAM~62 tok/sEstimated3.4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · bytedance · 2025-10-28

Q8_0 Fit unconfirmed
at least 1.4 GiB4% of RAM~146 tok/sEstimated1.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lgai-exaone · 2025-07-11

Q8_0 Fit unconfirmed
at least 1.3 GiB4% of RAM~164 tok/sEstimated1.28B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-08-22

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~591 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~591 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~591 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-08-25

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~591 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~282 tok/sEstimated0.74B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0 Fit unconfirmed
at least 1.5 GiB5% of RAM~132 tok/sEstimated1.58B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0 Fit unconfirmed
at least 0.4 GiB1% of RAM~465 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · obliteratus · 2026-04-15

Q8_0 Fit unconfirmed
at least 7.8 GiB24% of RAM~26 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · hmellor · 2025-07-22

Q8_0 Fit unconfirmed
at least 1.2 GiB4% of RAM~170 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · typhoon-ai · 2025-09-23

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · uzlm · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.2 GiB4% of RAM~170 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Alibaba · 2025-08-05

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-08-05

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · amd · 2025-05-17

Q8_0 Fit unconfirmed
at least 1.5 GiB5% of RAM~140 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-16

Q8_0 Fit unconfirmed
at least 8.2 GiB26% of RAM~25 tok/sEstimated8.38B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · stefanruseti · 2025-06-04

Q8_0 Fit unconfirmed
at least 1.2 GiB4% of RAM~170 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · bananamind · 2026-07-17

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~17,366 tok/sEstimated0.01B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · goekdeniz-guelmez · 2026-07-31

Q8_0 Fit unconfirmed
at least 9.4 GiB29% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huggingfacetb · 2025-07-08

Q8_0 Fit unconfirmed
at least 3.0 GiB9% of RAM~68 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lgai-exaone · 2025-03-12

Q8_0 Fit unconfirmed
at least 2.4 GiB7% of RAM~87 tok/sEstimated2.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · maliosdark · 2026-07-09

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~3,870 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · menlo · 2025-06-25

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~279 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · viorikaai-org · 2026-07-05

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~4,613 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · allenai · 2025-11-19

Q8_0 Fit unconfirmed
at least 7.1 GiB22% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · allenai · 2025-11-17

Q8_0 Fit unconfirmed
at least 7.1 GiB22% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · allenai · 2025-11-18

Q8_0 Fit unconfirmed
at least 7.1 GiB22% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · empero-ai · 2026-06-19

Q8_0 Fit unconfirmed
at least 9.2 GiB29% of RAM~22 tok/sEstimated9.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · fableforge-ai · 2026-07-05

Q8_0 Fit unconfirmed
at least 1.5 GiB5% of RAM~136 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0 Fit unconfirmed
at least 7.8 GiB24% of RAM~26 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · ibm-granite · 2025-04-09

Q8_0 Fit unconfirmed
at least 2.5 GiB8% of RAM~83 tok/sEstimated2.53B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-10-07

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~595 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · Microsoft · 2025-04-29

Q8_0 Fit unconfirmed
at least 3.8 GiB12% of RAM~55 tok/sEstimated3.84B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · naver-hyperclovax · 2025-04-22

Q8_0 Fit unconfirmed
at least 3.6 GiB11% of RAM~56 tok/sEstimated3.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · pearl-ai · 2026-02-26

Q8_0 Fit unconfirmed
at least 7.9 GiB25% of RAM~26 tok/sEstimated8.03B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · raidium · 2026-06-15

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~9,360 tok/sEstimated0.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Embedding · taide · 2026-06-12

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~692 tok/sEstimated0.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huggingfacetb · 2025-06-19

Q8_0 Fit unconfirmed
at least 3.0 GiB9% of RAM~68 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0 Fit unconfirmed
at least 4.3 GiB13% of RAM~47 tok/sEstimated4.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · shaungves · 2026-06-19

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · xcuros · 2026-02-28

Q8_0 Fit unconfirmed
at least 7.4 GiB23% of RAM~28 tok/sEstimated7.62B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · allenai · 2026-01-28

Q8_0 Fit unconfirmed
at least 7.3 GiB23% of RAM~28 tok/sEstimated7.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · darthcrawl · 2026-05-07

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~5,549 tok/sEstimated0.04B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0 Fit unconfirmed
at least 1.7 GiB5% of RAM~118 tok/sEstimated1.78B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0 Fit unconfirmed
at least 8.6 GiB27% of RAM~24 tok/sEstimated8.79B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2025-02-19

Q8_0 Fit unconfirmed
at least 3.8 GiB12% of RAM~55 tok/sEstimated3.84B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nvidia · 2026-03-18

Q8_0 Fit unconfirmed
at least 8.3 GiB26% of RAM~25 tok/sEstimated8.49B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · pfnet · 2025-02-05

Q8_0 Fit unconfirmed
at least 1.3 GiB4% of RAM~162 tok/sEstimated1.29B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · preparebuddy · 2026-06-02

Q8_0 Fit unconfirmed
at least 3.0 GiB9% of RAM~68 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · sapientinc · 2026-05-17

Q8_0 Fit unconfirmed
at least 1.2 GiB4% of RAM~177 tok/sEstimated1.18B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · alibaba-nlp · 2026-03-31

Q8_0 Fit unconfirmed
at least 8.0 GiB25% of RAM~26 tok/sEstimated8.19B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · k-intelligence · 2025-07-03

Q8_0 Fit unconfirmed
at least 2.3 GiB7% of RAM~91 tok/sEstimated2.31B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · lgai-exaone · 2024-12-01

Q8_0 Fit unconfirmed
at least 2.4 GiB7% of RAM~87 tok/sEstimated2.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2025-01-26

Q8_0 Fit unconfirmed
at least 3.7 GiB11% of RAM~56 tok/sEstimated3.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~279 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 2.0 GiB6% of RAM~103 tok/sEstimated2.03B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · sunbird · 2026-07-19

Q8_0 Fit unconfirmed
at least 5.0 GiB16% of RAM~41 tok/sEstimated5.1B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · allenai · 2025-09-12

Q8_0 Fit unconfirmed
at least 7.1 GiB22% of RAM~29 tok/sEstimated7.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · anton-hugging · 2026-02-06

Q8_0 Fit unconfirmed
at least 7.4 GiB23% of RAM~28 tok/sEstimated7.62B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Coding · mahiatlinux · 2026-06-30

Q8_0 Fit unconfirmed
at least 4.6 GiB14% of RAM~45 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · onnx-community · 2025-04-28

Q8_0 Fit unconfirmed
at least 0.4 GiB1% of RAM~466 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openbmb · 2025-06-05

Q8_0 Fit unconfirmed
at least 0.4 GiB1% of RAM~483 tok/sEstimated0.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · jackrong · 2026-03-16

Q8_0 Fit unconfirmed
at least 9.4 GiB29% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-05-31

Q8_0 Fit unconfirmed
at least 0.5 GiB2% of RAM~424 tok/sEstimated0.49B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-05-31

Q8_0 Fit unconfirmed
at least 1.5 GiB5% of RAM~136 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-09-15

Q8_0 Fit unconfirmed
at least 1.5 GiB5% of RAM~136 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 0.6 GiB2% of RAM~352 tok/sEstimated0.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 1.7 GiB5% of RAM~122 tok/sEstimated1.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 3.9 GiB12% of RAM~52 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Google · 2022-03-02

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~18,890 tok/sEstimated0.01B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · huihui-ai · 2024-10-01

Q8_0 Fit unconfirmed
at least 1.5 GiB5% of RAM~140 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2025-02-24

Q8_0 Fit unconfirmed
at least 5.5 GiB17% of RAM~38 tok/sEstimated5.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · neuralcrew · 2025-09-14

Q8_0 Fit unconfirmed
at least 5.9 GiB18% of RAM~35 tok/sEstimated6.04B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · swiss-ai · 2026-04-12

Q8_0 Fit unconfirmed
at least 3.7 GiB12% of RAM~55 tok/sEstimated3.83B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · obliteratus · 2026-06-05

Q8_0 Fit unconfirmed
at least 11.7 GiB37% of RAM~18 tok/sEstimated11.96B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 9.4 GiB29% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-26

Q8_0 Fit unconfirmed
at least 9.4 GiB29% of RAM~22 tok/sEstimated9.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · supralabs · 2026-06-03

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~4,046 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · tiger-lab · 2024-10-08

Q8_0 Fit unconfirmed
at least 4.1 GiB13% of RAM~51 tok/sEstimated4.15B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huihui-ai · 2026-07-11

Q8_0 Fit unconfirmed
at least 11.7 GiB37% of RAM~18 tok/sEstimated11.96B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2024-04-22

Q8_0 Fit unconfirmed
at least 3.7 GiB12% of RAM~55 tok/sEstimated3.82B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2024-05-19

Q8_0 Fit unconfirmed
at least 4.1 GiB13% of RAM~51 tok/sEstimated4.15B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2024-08-16

Q8_0 Fit unconfirmed
at least 3.7 GiB12% of RAM~55 tok/sEstimated3.82B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · yuxinlu1 · 2026-06-28

Q8_0 Fit unconfirmed
at least 11.7 GiB37% of RAM~18 tok/sEstimated11.96B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0 Fit unconfirmed
at least 7.8 GiB24% of RAM~26 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · 8f-ai

Q8_0 Fit unconfirmed
at least 0.7 GiB2% of RAM~278 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · abdallalswaiti

Q8_0 Fit unconfirmed
at least 3.1 GiB10% of RAM~65 tok/sEstimated3.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · ath-maas

Q8_0 Fit unconfirmed
at least 0.8 GiB3% of RAM~246 tok/sEstimated0.85B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · adamlucek

Q8_0 Fit unconfirmed
at least 1.2 GiB4% of RAM~170 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · alignmentresearch

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~2,790 tok/sEstimated0.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.2 GiB4% of RAM~170 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · andrew0425

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~194 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · bytedance

Q8_0 Fit unconfirmed
at least 3.7 GiB11% of RAM~56 tok/sEstimated3.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~3,274 tok/sEstimated0.06B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · daremodels

Q8_0 Fit unconfirmed
at least 1.1 GiB3% of RAM~194 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB4% of RAM~179 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

24-core vs 32-core GPU

Both bins run the same 400 GB/s memory bus, so token generation is the same on either one. The extra cores show up in image and video work, not in tokens per second.

ConfigurationMemory bandwidthMemory optionsModels that fit
10-core CPU, 24-core GPU400 GB/s32, 64 GBIdentical
10-core CPU, 32-core GPU400 GB/s32, 64 GBIdentical

Published observations for M1 Max

No compatible observations are published here. Formula estimates remain labelled, and unsupported forecasts remain unavailable.

Benchmark contribution is disclosed before a run, and the leaderboard shows every chip that already has one.

What comes next

Projection dated 2026-09-08

Apple has not announced any of this. The M7 Max and Ultra are the first parts Apple designed after cancelling a generation to reach them, so extrapolating from the M5 under-represents them. These rows assume LPDDR6, whose wider channels grow every bus by half, at its top speed bin by the time the Max and Ultra ship. The 1.5 TB Ultra ceiling is Bloomberg's reported design target, and whether that configuration ships depends on the memory market. Stacked memory or a new package fabric would land above these numbers; nobody outside Apple can price that yet.

M7 MaxProjected

Projected chip · expected 2027

1382 GB/s · 48 to 384 GB unified memory

20-core CPU · 56 or 64-core GPU · 76 TOPS Neural Engine

Would hold about a 624B model at Q4

Why unified memory is the number that matters

On a PC the model has to fit in GPU VRAM, which is a separate pool from system RAM and usually the smaller of the two. Apple Silicon has one pool. The M1 Max's 400 GB/s bus is shared by CPU, GPU, and Neural Engine, so a 64 GB machine can hand almost all of that to a model with no copy across a bus.

Apple stopped selling this one, which is exactly why it is interesting. The Studio exists for this workload. It carries the widest memory buses and the highest capacities Apple sells, and it runs at full clocks indefinitely. A used M1 Max at 64 GB still gives you 400 GB/s and a hard 103B ceiling, and neither number degrades with age the way a battery does.

Common questions

Can the M1 Max Mac Studio run a 70B model?

Yes, at 64 GB. A 70B model at Q4_K_M needs about 38 GB including an 8K context, and 64 GB of unified memory leaves about 56 GB for weights once macOS takes its share. At 32 GB it does not fit at any quantization worth running.

How much unified memory should I get with the M1 Max Mac Studio?

Memory is the only spec that changes what you can run at all. 32 GB holds about a 51B model at Q4; 64 GB holds about 103B. It is soldered, so this is a one-time decision, and it is the upgrade worth paying for before core count.

How fast are local LLMs on the M1 Max?

Token generation is bandwidth-bound, so M1 Max throughput scales with its 400 GB/s memory bus. Divide bandwidth by the size of the weights actually read per token to get the ceiling, then expect roughly half of that in practice. A 7B model at Q4 reads about 4 GB per token pass, so the M1 Max lands in the tens of tokens per second and a 70B model lands in the single digits.

Is the 32-core GPU worth it over the 24-core on the M1 Max?

Not for LLMs. Both bins run the same 400 GB/s memory bus and take the same memory options, and token generation is bound by bandwidth rather than GPU cores. The extra cores show up in image generation and video work, not in tokens per second.

Is a used M1 Max Mac Studio still worth buying for local AI?

For inference, the specs that matter do not age: 400 GB/s and up to 64 GB of unified memory are the same numbers today as they were in 2021. A used M1 Max at the top memory option usually beats a new base-tier machine at the same price on both. Check the battery and the display, not the silicon.

Run these models on your Mac Studio

ToolPiper downloads, manages, and runs local models on Apple Silicon. Free, and nothing leaves the machine.