Apple · apple

Apple M5

Apple M5 has 12 GB of unified memory at 154 GB/s — about 8.37 GiB usable after driver and compositor overhead. 526 of 2118 indexed models fit at 128K context with f16 KV. Note only 9 GB of its 12 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
12 GB
LPDDR5X-9600
Bandwidth
154 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 385vision language 65video 12audio asr 31audio tts 18embedding 14image 1

What fits at 128K context

largest quantization that fits, per model · 526 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
GLM-OCRI1-IQ4_XS1.3B0.47 GiB8.00 GiB9.00 GiB0.00 GiB14±8.3%
G9v3-3BQ4_K_L3.0B1.96 GiB6.50 GiB9.00 GiB0.00 GiB14±8.3%
Qwen3.5-9B-CoderI1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Qwopus3.5-9B-v3.5Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Qwythos-9B-Claude-Mythos-5-1M-MTPI1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedI1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Qwen3.5-9B-Fable-5-v1I1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Qwythos-9B-v2I1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
PINQWEN-3.5-9B-1M-BF16I1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Openprose-2-FlashI1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Qwen3.5-9B-Nikusui-v1I1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Ornstein-3.5-9B-V1.5I1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Ornith-1.0-9B-heretic-MTPI1-Q3_K_M9.4B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Tess-4-9BI1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
dotwebs-1I1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
liftQ3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Hemlock-Qwopus3.5-9B-CoderI1-Q3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Qwen3.5-9B-DeepSeek-V4-FlashQ3_K_M9.7B4.41 GiB4.00 GiB9.00 GiB0.00 GiB14±8.3%
Wan2.1-T2V-14BQ4_014.3B8.41 GiB0.00 GiB9.00 GiB0.00 GiB14±8.3%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseI1-Q4_18.1B4.91 GiB3.50 GiB8.99 GiB0.01 GiB14±8.3%
Fara1.5-9BIQ3_M9.4B4.40 GiB4.00 GiB8.98 GiB0.02 GiB14±8.3%
QwenPaw-Flash-9BIQ3_M9.4B4.40 GiB4.00 GiB8.98 GiB0.02 GiB14±8.3%
grug-9bIQ3_M9.4B4.40 GiB4.00 GiB8.98 GiB0.02 GiB14±8.3%
OmniCoder-9BIQ3_M9.4B4.40 GiB4.00 GiB8.98 GiB0.02 GiB14±8.3%
Ornith-1.0-9BIQ3_M9.2B4.40 GiB4.00 GiB8.98 GiB0.02 GiB14±8.3%
Qwen3.5-9B-NeoIQ3_M9.7B4.40 GiB4.00 GiB8.98 GiB0.02 GiB14±8.3%
InternVL3_5-14BQ4_K_M15.1B8.38 GiB0.00 GiB8.98 GiB0.02 GiB14±8.3%
gemma-4-E4B-itQ6_K8.0B6.59 GiB1.82 GiB8.97 GiB0.03 GiB14±8.3%
HunyuanVideo-1.5Q8_08.3B8.38 GiB0.00 GiB8.97 GiB0.03 GiB14±8.3%
Teuken-7B-instruct-research-v0.4I1-Q4_K_S7.5B4.38 GiB4.00 GiB8.97 GiB0.03 GiB14±8.3%
glm4.1v-9b-base-sftI1-IQ2_XS10.3B3.38 GiB5.00 GiB8.97 GiB0.03 GiB14±8.3%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB8.95 GiB0.05 GiB15±8.3%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB8.95 GiB0.05 GiB15±8.3%
EXAONE-4.0-1.2B-abliteratedI1-Q4_11.5B0.91 GiB7.50 GiB8.94 GiB0.06 GiB14±8.3%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopIQ3_M12.1B5.33 GiB3.00 GiB8.94 GiB0.06 GiB15±8.3%
gemma-4-E4B-it-hereticQ6_K8.0B6.55 GiB1.82 GiB8.93 GiB0.07 GiB14±8.3%
gemma-4-E2B-it-uncensoredBF165.1B7.49 GiB0.90 GiB8.93 GiB0.07 GiB14±8.3%
Vikhr-Gemma-2B-instructQ4_12.6B1.64 GiB6.73 GiB8.92 GiB0.08 GiB14±8.3%
Vero-Qwen35-9B-BaseI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Vero-Qwen35-9BI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Morphos-9BI1-Q3_K_M9.0B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwable-9B-Claude-Fable-5-hereticI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Holo-3.1-9BI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwable-9B-Claude-Fable-5I1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-imabari-v2I1-Q3_K_M9.7B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-abliterated-v2-MAXI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwable-9B-Claude-Fable-5-StraTAI1-Q3_K_M9.0B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q3_K_M9.0B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-RpRMax-v1I1-Q3_K_M9.7B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
AdQWENistrator-9BI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
cajal-9b-v2-fullI1-Q3_K_M9.0B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-ultra-uncensored-hereticQ3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Holo-3.1-9B-CoderI1-Q3_K_M9.0B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
PlutoI1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Holo-3.1-9B-abliterated-rdoI1-Q3_K_M9.0B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-Uncensored-cyber-v3Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
Qwen3.5-9B-BaseI1-Q3_K_M9.7B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
qwen3.5-9b-nsfw-captioning-v5I1-Q3_K_M9.4B4.31 GiB4.00 GiB8.89 GiB0.11 GiB15±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing489.78 tok/s264.15636.369
Text generation16.62 tok/s9.6727.929
Benchmarked· n=9

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M5 run?
526 of 2118 indexed open-weight models fit a Apple M5 at 131,072 context with f16 KV cache, the largest being GLM-OCR at I1-IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M5 actually have?
Its nameplate is 12 GB, but about 8.37 GiB is available to a model once driver and compositor overhead is accounted for, and only 9 GB of the pool can be allocated to the GPU at all.
Is a Apple M5 fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.