What LLMs Can Run on Your Mac?

Chip
Unified Memory
Bandwidth 307 GB/s 16-core GPU, the lowest of 2 bins at this memory
Available for Models ~21 GB

Everything this chip runs, machine by machine: MacBook Pro 14" M5 Pro, MacBook Pro 16" M5 Pro, Mac mini M5 Pro

Model Compatibility

Showing 6563 of 6563 models

General · goekdeniz-guelmez · 2026-07-31

Q8_0 Fit unconfirmed
at least 2.2 GiB9% of RAM~71 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-07-28

Q8_0 Fit unconfirmed
at least 2.6 GiB11% of RAM~60 tok/sEstimated2.7B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-06-24

Q8_0 Fit unconfirmed
at least 0.2 GiB1% of RAM~700 tok/sEstimated0.23B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · radixark · 2026-07-27

Q8_0 Fit unconfirmed
at least 2.2 GiB9% of RAM~71 tok/sEstimated2.25B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ktruestory · 2026-05-28

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ma7ee7 · 2026-07-30

Q8_0 Fit unconfirmed
at least 3.1 GiB13% of RAM~51 tok/sEstimated3.13B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · weiboai · 2026-06-12

Q8_0 Fit unconfirmed
at least 3.0 GiB13% of RAM~52 tok/sEstimated3.09B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openbmb · 2026-05-21

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · internscience · 2026-07-13

Q8_0 Fit unconfirmed
at least 4.4 GiB18% of RAM~35 tok/sEstimated4.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-03-31

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nanbeige · 2026-07-21

Q8_0 Fit unconfirmed
at least 4.1 GiB17% of RAM~39 tok/sEstimated4.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · openonerec · 2026-06-09

Q8_0 Fit unconfirmed
at least 0.8 GiB3% of RAM~201 tok/sEstimated0.8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 0.9 GiB4% of RAM~184 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 0.9 GiB4% of RAM~184 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 2.2 GiB9% of RAM~71 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-28

Q8_0 Fit unconfirmed
at least 2.2 GiB9% of RAM~71 tok/sEstimated2.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · cagrigungor · 2026-08-06

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~600 tok/sEstimated0.27B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2026-04-06

Q8_0 Fit unconfirmed
at least 3.3 GiB14% of RAM~47 tok/sEstimated3.4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-20

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · ibm-granite · 2026-04-16

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · frontiersmind · 2026-08-03

Q8_0 Fit unconfirmed
at least 0.6 GiB3% of RAM~248 tok/sEstimated0.65B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-12-25

Q8_0 Fit unconfirmed
at least 2.5 GiB10% of RAM~63 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 2.5 GiB10% of RAM~63 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Liquid AI · 2026-01-06

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2026-01-04

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2026-01-05

Q8_0 Fit unconfirmed
at least 1.6 GiB7% of RAM~101 tok/sEstimated1.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · nvidia · 2026-03-02

Q8_0 Fit unconfirmed
at least 3.7 GiB16% of RAM~42 tok/sEstimated3.83B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · davidau · 2026-02-02

Q8_0 Fit unconfirmed
at least 4.2 GiB18% of RAM~37 tok/sEstimated4.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-10-28

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~455 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 4.6 GiB19% of RAM~35 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2026-02-27

Q8_0 Fit unconfirmed
at least 4.6 GiB19% of RAM~35 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Google · 2026-03-02

Q8_0 Fit unconfirmed
at least 5.0 GiB21% of RAM~31 tok/sEstimated5.12B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-22

Q8_0 Fit unconfirmed
at least 2.5 GiB10% of RAM~63 tok/sEstimated2.57B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-30

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-10-22

Q8_0 Fit unconfirmed
at least 2.9 GiB12% of RAM~54 tok/sEstimated3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openonerec · 2025-12-30

Q8_0 Fit unconfirmed
at least 2.1 GiB9% of RAM~75 tok/sEstimated2.13B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · bytedance · 2025-10-28

Q8_0 Fit unconfirmed
at least 1.4 GiB6% of RAM~112 tok/sEstimated1.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lgai-exaone · 2025-07-11

Q8_0 Fit unconfirmed
at least 1.3 GiB5% of RAM~126 tok/sEstimated1.28B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-08-22

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-09-03

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-08-25

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI · 2025-07-10

Q8_0 Fit unconfirmed
at least 0.7 GiB3% of RAM~217 tok/sEstimated0.74B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~101 tok/sEstimated1.58B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI · 2025-08-12

Q8_0 Fit unconfirmed
at least 0.4 GiB2% of RAM~357 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · baseten · 2025-09-12

Q8_0 Fit unconfirmed
at least 3.1 GiB13% of RAM~50 tok/sEstimated3.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · hmellor · 2025-07-22

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0 Fit unconfirmed
at least 3.1 GiB13% of RAM~50 tok/sEstimated3.19B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · uzlm · 2025-09-03

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · amd · 2025-05-17

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~107 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-09-16

Q8_0 Fit unconfirmed
at least 3.3 GiB14% of RAM~47 tok/sEstimated3.4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · stefanruseti · 2025-06-04

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · typhoon-ai · 2025-09-23

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · z-lab · 2026-01-04

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · bananamind · 2026-07-17

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~13,328 tok/sEstimated0.01B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lgai-exaone · 2025-03-12

Q8_0 Fit unconfirmed
at least 2.4 GiB10% of RAM~67 tok/sEstimated2.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · maliosdark · 2026-07-09

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~2,970 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Alibaba · 2025-08-05

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-08-05

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0 Fit unconfirmed
at least 0.7 GiB3% of RAM~214 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · viorikaai-org · 2026-07-05

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~3,540 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · fableforge-ai · 2026-07-05

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~104 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · ibm-granite · 2025-04-09

Q8_0 Fit unconfirmed
at least 2.5 GiB10% of RAM~63 tok/sEstimated2.53B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ibm-granite · 2025-10-07

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~456 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · raidium · 2026-06-15

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~7,184 tok/sEstimated0.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Embedding · taide · 2026-06-12

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~531 tok/sEstimated0.3B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huggingfacetb · 2025-07-08

Q8_0 Fit unconfirmed
at least 3.0 GiB13% of RAM~52 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lukebailey181pub · 2026-04-21

Q8_0 Fit unconfirmed
at least 6.8 GiB28% of RAM~23 tok/sEstimated6.91B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · menlo · 2025-06-25

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ravichandranj · 2026-02-13

Q8_0 Fit unconfirmed
at least 6.3 GiB26% of RAM~25 tok/sEstimated6.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · darthcrawl · 2026-05-07

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~4,259 tok/sEstimated0.04B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · DeepSeek · 2025-01-20

Q8_0 Fit unconfirmed
at least 1.7 GiB7% of RAM~90 tok/sEstimated1.78B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · Microsoft · 2025-04-29

Q8_0 Fit unconfirmed
at least 3.8 GiB16% of RAM~42 tok/sEstimated3.84B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · naver-hyperclovax · 2025-04-22

Q8_0 Fit unconfirmed
at least 3.6 GiB15% of RAM~43 tok/sEstimated3.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · pfnet · 2025-02-05

Q8_0 Fit unconfirmed
at least 1.3 GiB5% of RAM~125 tok/sEstimated1.29B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · preparebuddy · 2026-06-02

Q8_0 Fit unconfirmed
at least 3.0 GiB13% of RAM~52 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · sapientinc · 2026-05-17

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~136 tok/sEstimated1.18B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · huggingfacetb · 2025-06-19

Q8_0 Fit unconfirmed
at least 3.0 GiB13% of RAM~52 tok/sEstimated3.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · k-intelligence · 2025-07-03

Q8_0 Fit unconfirmed
at least 2.3 GiB9% of RAM~70 tok/sEstimated2.31B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · lgai-exaone · 2024-12-01

Q8_0 Fit unconfirmed
at least 2.4 GiB10% of RAM~67 tok/sEstimated2.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 0.7 GiB3% of RAM~214 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 2.0 GiB8% of RAM~79 tok/sEstimated2.03B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-09-23

Q8_0 Fit unconfirmed
at least 4.3 GiB18% of RAM~36 tok/sEstimated4.41B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · shaungves · 2026-06-19

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · deepreinforce-ai · 2026-06-21

Q8_0 Fit unconfirmed
at least 8.0 GiB33% of RAM~20 tok/sEstimated8.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · onnx-community · 2025-04-28

Q8_0 Fit unconfirmed
at least 0.4 GiB2% of RAM~358 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · openbmb · 2025-06-05

Q8_0 Fit unconfirmed
at least 0.4 GiB2% of RAM~371 tok/sEstimated0.43B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-05-31

Q8_0 Fit unconfirmed
at least 0.5 GiB2% of RAM~326 tok/sEstimated0.49B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-05-31

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~104 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2024-09-15

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~104 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Alibaba · 2025-01-26

Q8_0 Fit unconfirmed
at least 3.7 GiB15% of RAM~43 tok/sEstimated3.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 0.6 GiB2% of RAM~270 tok/sEstimated0.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-28

Q8_0 Fit unconfirmed
at least 1.7 GiB7% of RAM~93 tok/sEstimated1.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · sunbird · 2026-07-19

Q8_0 Fit unconfirmed
at least 5.0 GiB21% of RAM~32 tok/sEstimated5.1B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Google · 2022-03-02

Q8_0 Fit unconfirmed
at least 0.0 GiB0% of RAM~14,498 tok/sEstimated0.01B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · huihui-ai · 2024-10-01

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~107 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Coding · mahiatlinux · 2026-06-30

Q8_0 Fit unconfirmed
at least 4.6 GiB19% of RAM~35 tok/sEstimated4.66B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · Microsoft · 2025-02-19

Q8_0 Fit unconfirmed
at least 3.8 GiB16% of RAM~42 tok/sEstimated3.84B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · obliteratus · 2026-04-15

Q8_0 Fit unconfirmed
at least 7.8 GiB33% of RAM~20 tok/sEstimated8B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Alibaba · 2025-04-27

Q8_0 Fit unconfirmed
at least 3.9 GiB16% of RAM~40 tok/sEstimated4.02B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · supralabs · 2026-06-03

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~3,105 tok/sEstimated0.05B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · swiss-ai · 2026-04-12

Q8_0 Fit unconfirmed
at least 3.7 GiB16% of RAM~42 tok/sEstimated3.83B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · 8f-ai

Q8_0 Fit unconfirmed
at least 0.7 GiB3% of RAM~214 tok/sEstimated0.75B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · ath-maas

Q8_0 Fit unconfirmed
at least 0.8 GiB3% of RAM~189 tok/sEstimated0.85B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · adamlucek

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Chat · alignmentresearch

Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~2,141 tok/sEstimated0.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · andrew0425

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 0.1 GiB0% of RAM~2,513 tok/sEstimated0.06B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · daremodels

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~104 tok/sEstimated1.54B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
at least 2.2 GiB9% of RAM~73 tok/sEstimated2.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 2.2 GiB9% of RAM~73 tok/sEstimated2.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · dingdust

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · dingdust

Q8_0 Fit unconfirmed
at least 0.8 GiB3% of RAM~189 tok/sEstimated0.85B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ephemeralyou

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · etherll

Q8_0 Fit unconfirmed
at least 0.7 GiB3% of RAM~217 tok/sEstimated0.74B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Reasoning · healshsj

Q8_0 Fit unconfirmed
at least 1.1 GiB4% of RAM~149 tok/sEstimated1.08B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 0.9 GiB4% of RAM~184 tok/sEstimated0.87B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · kamilamila

Q8_0 Fit unconfirmed
at least 0.6 GiB3% of RAM~259 tok/sEstimated0.62B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper
Q8_0 Fit unconfirmed
at least 2.6 GiB11% of RAM~60 tok/sEstimated2.7B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · lemonelabs

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~133 tok/sEstimated1.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI

Q8_0 Fit unconfirmed
at least 1.1 GiB5% of RAM~137 tok/sEstimated1.17B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~133 tok/sEstimated1.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI

Q8_0 Fit unconfirmed
at least 2.6 GiB11% of RAM~60 tok/sEstimated2.7B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · Liquid AI

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Coding · Liquid AI

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI

Q8_0 Fit unconfirmed
at least 1.6 GiB7% of RAM~101 tok/sEstimated1.6B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI

Q8_0 Fit unconfirmed
at least 0.4 GiB2% of RAM~358 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · Liquid AI

Q8_0 Fit unconfirmed
at least 0.4 GiB2% of RAM~358 tok/sEstimated0.45B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · meddies

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · menlo

Q8_0 Fit unconfirmed
at least 1.7 GiB7% of RAM~93 tok/sEstimated1.72B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · mihaipopa-1

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~454 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Multimodal · mirilai

Q8_0 Fit unconfirmed
at least 2.0 GiB8% of RAM~80 tok/sEstimated2B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · muxodious

Q8_0 Fit unconfirmed
at least 2.6 GiB11% of RAM~60 tok/sEstimated2.7B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · novacorp

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~130 tok/sEstimated1.24B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · novacorp

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~107 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · novachronoai

Q8_0 Fit unconfirmed
at least 1.2 GiB5% of RAM~133 tok/sEstimated1.21B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · novaciano

Q8_0 Fit unconfirmed
at least 1.5 GiB6% of RAM~107 tok/sEstimated1.5B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

General · ordenwills

Q8_0 Fit unconfirmed
at least 0.3 GiB1% of RAM~455 tok/sEstimated0.35B params
Estimate limits

Unvalidated forecasts for an 8K text context under a declared full-GPU scenario, at the lowest GPU bin sold with the chosen memory. Memory is reported as its components: exact artifact bytes and a KV allocation derived from a declared attention layout where the catalog carries them, a parameter proxy with no error bound where it does not, and the runtime's working memory, which a browser cannot measure. A fit is confirmed only when every component is known, so a model whose known memory fits reads "Fit unconfirmed" until ToolPiper measures it on your Mac, and one whose known memory alone exceeds the budget reads "Won't fit". Estimated speed ignores prompt length and prefill, and can overstate long-context performance. No measured speed error range is available. The recommendation score weighs declared preferences, not measured answer quality, and becomes a range where evidence is missing.

Known memory fits; the runtime's working memory is not measured.

Run with ToolPiper

Already know the repository? Paste a HuggingFace URL or ID to jump to its fit check and refresh its file sizes.

Frequently Asked Questions

How much RAM do I need to run LLMs on a Mac?

It depends on the model size and quantization. A 7B parameter model at Q4 quantization needs about 5 GB of RAM, while a 70B model needs 40+ GB. Apple Silicon Macs use unified memory, so your entire RAM pool is available for model weights — no separate VRAM required.

Can I run a 70B model on a MacBook Air?

Not comfortably. A 70B model at Q4 quantization needs about 40 GB of RAM. The MacBook Air maxes out at 24-32 GB depending on the generation. You'd need a Mac Studio or MacBook Pro with 48+ GB for a 70B model to run well.

What's the fastest LLM I can run on my Mac?

Speed depends on your chip's memory bandwidth and the model size. Smaller models (3-7B) run fastest — expect 40-70+ tokens per second on M2 Pro or better. Use the calculator above to see estimated speeds for your specific Mac.

What does quantization mean for model quality?

Quantization reduces model precision to use less memory. Q8 (8-bit) is nearly lossless. Q4 (4-bit) reduces memory by ~75% with minor quality loss — it's the sweet spot for most users. Q2 (2-bit) saves the most memory but noticeably degrades output quality.

How is Apple Silicon different from NVIDIA for LLMs?

Apple Silicon uses unified memory — CPU and GPU share the same RAM pool. A Mac with 32 GB can load a 28 GB model directly. On NVIDIA systems, you're limited by GPU VRAM (typically 8-24 GB on consumer cards), even if the PC has 64 GB of system RAM.

Does ToolPiper use GPU or CPU for inference?

ToolPiper uses Metal GPU acceleration via llama.cpp for LLM inference on Apple Silicon. The GPU and CPU share unified memory, so there's no data transfer overhead. The Neural Engine (ANE) is used for specific tasks like super-resolution and pose detection.

Can I run multiple models at the same time?

Yes, if you have enough RAM. ToolPiper manages model loading and can keep multiple models in memory simultaneously. When memory gets tight, it automatically evicts the least recently used model to make room for a new one.

What's the difference between GGUF and other formats?

GGUF is the standard format for running quantized models with llama.cpp (and ToolPiper). It supports all quantization levels and runs on CPU+GPU. MLX is Apple's format optimized for Apple Silicon. AWQ and GPTQ are NVIDIA-focused formats that don't run natively on Mac.

Browse the whole catalog

A cross-section of all 6563 models. Open any one to see which Macs run it, then follow its related models to work through the rest of the catalog.

Model database updated: 2026-09-09 · 6563 models