Apple · apple

Apple M5 Pro

Apple M5 Pro has 24 GB of unified memory at 307 GB/s — about 16.74 GiB usable after driver and compositor overhead. 1775 of 2118 indexed models fit at 128K context with q4_0 KV. Note only 18 GB of its 24 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
24 GB
LPDDR5X-9600
Bandwidth
307 GB/s
256-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1502vision language 170video 16audio asr 39image 1embedding 26audio tts 21

What fits at 128K context

largest quantization that fits, per model · 1775 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
internlm2-math-plus-20bI1-Q4_K_S19.9B10.62 GiB6.75 GiB17.98 GiB0.02 GiB14±8.3%
Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-PreservedQ4_K_S27.4B15.11 GiB2.25 GiB17.97 GiB0.03 GiB14±8.3%
Qwen3.5-27B-uncensored-heretic-v2-Native-MTP-PreservedQ4_K_S27.4B15.11 GiB2.25 GiB17.97 GiB0.03 GiB14±8.3%
Wan2.1-VACE-14BQ8_017.3B17.38 GiB0.00 GiB17.97 GiB0.03 GiB14±8.3%
Phi-3-mini-4k-instructKV unresolvedIQ2_XS3.8B3.91 GiB13.50 GiB17.97 GiB0.03 GiB14±8.3%
SOLAR-10.7B-Instruct-v1.0-uncensoredQ8_010.7B10.62 GiB6.75 GiB17.96 GiB0.04 GiB14±8.3%
Nous-Hermes-2-SOLAR-10.7BQ8_010.7B10.62 GiB6.75 GiB17.96 GiB0.04 GiB14±8.3%
SOLAR-10.7B-Instruct-v1.0Q8_010.7B10.62 GiB6.75 GiB17.96 GiB0.04 GiB14±8.3%
InternVL3_5-30B-A3BQ4_K_M30.8B17.35 GiB0.00 GiB17.95 GiB0.05 GiB14±8.3%
reka-flash-3.1I1-Q4_K_M20.9B12.68 GiB4.64 GiB17.94 GiB0.06 GiB14±8.3%
reka-flash-3Q4_K_M20.9B12.68 GiB4.64 GiB17.94 GiB0.06 GiB14±8.3%
Snowpiercer-15B-v4Q5_K_L15.0B10.31 GiB7.03 GiB17.94 GiB0.06 GiB14±8.3%
spoomplesmaxx-v2.1-30BI1-IQ2_S28.9B8.28 GiB9.00 GiB17.94 GiB0.06 GiB14±8.3%
Huihui-granite-4.1-30b-abliteratedI1-IQ2_S28.9B8.28 GiB9.00 GiB17.94 GiB0.06 GiB14±8.3%
granite-4.1-30b-hereticI1-IQ2_S28.9B8.28 GiB9.00 GiB17.94 GiB0.06 GiB14±8.3%
UncensoredLM-DeepSeek-R1-Distill-Qwen-14BQ6_K14.2B10.87 GiB6.47 GiB17.94 GiB0.06 GiB14±8.3%
Mistral-MOE-4X7B-Dark-MultiVerse-Uncensored-Enhanced32-24BMoEQ4_K_S24.2B12.84 GiB4.50 GiB17.93 GiB0.07 GiB8±37%
EuroLLM-22B-Instruct-2512IQ3_M22.6B9.72 GiB7.59 GiB17.93 GiB0.07 GiB14±8.3%
OLMoE-1B-7B-0924-InstructMoEF166.9B12.89 GiB4.50 GiB17.91 GiB0.09 GiB20±37%
Magistry-24B-v1.1Q3_K_L23.6B11.60 GiB5.63 GiB17.90 GiB0.10 GiB14±8.3%
gemma-4-26B-A4B-itMoEQ4_K_M26.5B15.87 GiB1.49 GiB17.89 GiB0.11 GiB14±8.3%
NVIDIA-Nemotron-Nano-12B-v2Q5_K_L12.3B8.55 GiB8.72 GiB17.89 GiB0.11 GiB14±8.3%
OmniAtlas-Qwen3-30B-A3BI1-Q4_K_M31.7B17.28 GiB0.00 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3-Omni-30B-A3B-InstructQ4_K_M35.3B17.28 GiB0.00 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3-Omni-30B-A3B-CaptionerI1-Q4_K_M31.7B17.28 GiB0.00 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3-Omni-30B-A3B-ThinkingQ4_K_M31.7B17.28 GiB0.00 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3.8-27BQ4_K_S27.8B15.01 GiB2.25 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3.6-27BQ4_K_S27.8B15.01 GiB2.25 GiB17.88 GiB0.12 GiB14±8.3%
NousCoder-14BQ6_K_L14.8B11.64 GiB5.63 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3-14B-abliteratedQ6_K_L14.8B11.64 GiB5.63 GiB17.88 GiB0.12 GiB14±8.3%
Josiefied-Qwen3-14B-abliterated-v3Q6_K_L14.8B11.64 GiB5.63 GiB17.88 GiB0.12 GiB14±8.3%
Hermes-4-14BQ6_K_L14.8B11.64 GiB5.63 GiB17.88 GiB0.12 GiB14±8.3%
Qwen3.6-34B-80L-Fable-5-HereticI1-IQ3_M33.4B14.45 GiB2.81 GiB17.87 GiB0.13 GiB14±8.3%
Phi-3.5-MoE-instructMoEKV unresolvedIQ2_M41.9B12.82 GiB4.50 GiB17.87 GiB0.13 GiB20±37%
Qwen3.6-27B-Heretic2-Uncensored-Finetune-ThinkingIQ4_NL27.4B15.01 GiB2.25 GiB17.87 GiB0.13 GiB14±8.3%
Le-Chaton-Slim-23BMoEI1-Q4_123.3B13.65 GiB3.66 GiB17.87 GiB0.13 GiB19±37%
grug-27bQ4_027.4B15.00 GiB2.25 GiB17.87 GiB0.13 GiB14±8.3%
Carnice-V2-27bQ4_027.4B15.00 GiB2.25 GiB17.87 GiB0.13 GiB14±8.3%
Fara1.5-27BQ4_027.4B15.00 GiB2.25 GiB17.87 GiB0.13 GiB14±8.3%
Trinity-2-Codestral-22B-v0.2IQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
Cydonia-v1.3-Magnum-v4-22BI1-IQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
Mistral-Small-22B-ArliAI-RPMax-v1.1I1-IQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
Mistral-Small-Drummer-22BIQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
magnum-v4-22bI1-IQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
Mistral-Small-Instruct-2409IQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
Codestral-22B-v0.1IQ3_M22.2B9.37 GiB7.88 GiB17.86 GiB0.14 GiB14±8.3%
Qwen3.5-35B-A3BMoEUD-IQ4_NL36.0B16.60 GiB0.70 GiB17.86 GiB0.14 GiB49±37%
Qwen3.6-28BMoEI1-Q4_128.2B16.60 GiB0.70 GiB17.85 GiB0.15 GiB47±37%
Qwen3.5-28BMoEI1-Q4_128.7B16.60 GiB0.70 GiB17.85 GiB0.15 GiB47±37%
dolphin-2.9.1-mixtral-1x22bMoEI1-IQ3_M22.2B9.37 GiB7.88 GiB17.85 GiB0.15 GiB8±37%
North-Mini-Code-1.0MoEQ4_030.5B16.32 GiB1.00 GiB17.85 GiB0.15 GiB39±37%
Voxtral-Small-24B-2507Q3_K_L24.3B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Devstral-Small-2-24B-Instruct-2512Q3_K_L24.0B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Transformed-Journey-24BI1-Q3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Mergedonia-AETHER-24B-v1aI1-Q3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Mergedonia-AETHER-24B-v1bI1-Q3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Slimaki-Tavern-24B-v1.3I1-Q3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Maginum-Cydoms-24BI1-Q3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Maginum-Cydoms-24B-absolute-heresyI1-Q3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
Dolphin3.0-R1-Mistral-24BQ3_K_L23.6B11.55 GiB5.63 GiB17.84 GiB0.16 GiB14±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing964.18 tok/s431.141304.669
Text generation37.70 tok/s21.3760.049
Benchmarked· n=9

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M5 Pro run?
1775 of 2118 indexed open-weight models fit a Apple M5 Pro at 131,072 context with q4_0 KV cache, the largest being internlm2-math-plus-20b at I1-Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M5 Pro actually have?
Its nameplate is 24 GB, but about 16.74 GiB is available to a model once driver and compositor overhead is accounted for, and only 18 GB of the pool can be allocated to the GPU at all.
Is a Apple M5 Pro fast for local AI?
Its memory bandwidth is 307 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.