Apple · apple

Apple M5 Pro

Apple M5 Pro has 48 GB of unified memory at 307 GB/s — about 33.48 GiB usable after driver and compositor overhead. 2023 of 2118 indexed models fit at 32K context with q8_0 KV. Note only 36 GB of its 48 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
48 GB
LPDDR5X-9600
Bandwidth
307 GB/s
256-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1737vision language 182video 16audio tts 21image 2embedding 26audio asr 39

What fits at 32K context

largest quantization that fits, per model · 2023 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
command-r-35b-writer-v2I1-IQ3_XS35.0B14.05 GiB21.25 GiB35.96 GiB0.04 GiB7±8.3%
Mistral-Small-4-119B-2603MoEUD-IQ2_M119B34.99 GiB0.37 GiB35.95 GiB0.05 GiB34±37%
Apertus-70B-Instruct-2509IQ3_M70.6B29.84 GiB5.31 GiB35.88 GiB0.12 GiB7±8.3%
Hunyuan-A13B-InstructMoEQ3_K_S80.4B33.10 GiB2.13 GiB35.78 GiB0.22 GiB7±8.3%
Assistant_Pepe_70BQ3_K_S70.6B29.77 GiB5.31 GiB35.75 GiB0.25 GiB7±8.3%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedIQ2_S109B31.98 GiB3.19 GiB35.74 GiB0.26 GiB19±37%
Maenad-70BI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
DeepSeek-R1-Distill-Llama-70B-Uncensored-v2-Unbiased-ReasonerI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Rombos-LLM-70b-Llama-3.3I1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
L3.3-Electra-R1-70bI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
L3.3-70B-Magnum-v4-SEIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Latxa-Llama-3.1-70B-Instruct-v2I1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Llama-3.3_70_b_uncensored_continuedI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Llama-3.3-70B-Instruct-abliteratedI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
grok-oss-Revenant-70BI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Llama-3.1-Nemotron-70B-Instruct-HFI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
L3.3-70B-Euryale-v2.3I1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Hermes-3-Llama-3.1-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Hermes-4-70B-hereticI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Llama-3.3-70B-InstructIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Llama-3.1-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Anubis-70B-v1.2IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Hermes-4-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Golem-70B-v1bI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
DeepSeek-R1-Distill-Llama-70B-abliteratedI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
DeepSeek-R1-Distill-Llama-70B-hereticI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
DeepSeek-R1-Distill-Llama-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Legion-V2.1-LLaMa-70BI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Tess-R1-Limerick-Llama-3.1-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
SEMIKONG-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
functionary-medium-v3.2KV unresolvedIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Llama-3.1-WhiteRabbitNeo-2-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Infinity-Instruct-7M-Gen-Llama3_1-70BI1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
New-Dawn-Llama-3-70B-32K-v1.0I1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Meta-Llama-3-70B-Instruct-abliterated-v3.5I1-IQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Athene-70BIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
L3.3-70B-Magnum-DiamondIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Meta-Llama-3-70B-InstructIQ3_M70.6B29.74 GiB5.31 GiB35.73 GiB0.27 GiB7±8.3%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPQ4_K_M27.8B34.04 GiB1.06 GiB35.71 GiB0.29 GiB7±8.3%
llm-surgery-dark-arts-gpt-oss-60b-96a12MoEI1-Q4_160.9B34.76 GiB0.41 GiB35.71 GiB0.29 GiB20±37%
Gemma-4-31B-Isometry-RPQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Prosopon-31BQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Gemma-4-Novelist-Eclipse-31BQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Giftige-Blume-31B-v1-StyleSwapQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
G4-MeroMero-31B-StyleSwapQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Gemma-4-31B-StyleTune-heretic-araQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Pantheon-Reasoning-31B-1.1Q8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Gemma-4-31B-StyleTuneQ8_032.7B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Barcenas-StyleTune-31B-FableQ8_032.1B31.79 GiB3.28 GiB35.70 GiB0.30 GiB7±8.3%
Qwen2.5-VL-72B-InstructUD-IQ3_XXS73.4B29.67 GiB5.31 GiB35.66 GiB0.34 GiB7±8.3%
CalmeRys-78B-Orpo-v0.1I1-IQ2_M78.0B29.27 GiB5.71 GiB35.66 GiB0.34 GiB7±8.3%
calme-2.3-rys-78bIQ2_M78.0B29.27 GiB5.71 GiB35.66 GiB0.34 GiB7±8.3%
Kimi-Dev-72BUD-IQ3_XXS72.7B29.67 GiB5.31 GiB35.66 GiB0.34 GiB7±8.3%
Rombo-LLM-V3.0-Qwen-72bI1-IQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
Qwen2.5-72B-Instruct-abliteratedI1-IQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
Qwen2.5-72B-Instruct-abliterated-v2I1-IQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
HuatuoGPT-o1-72BIQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
MiroThinker-v1.0-72BI1-IQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
EVA-Qwen2.5-72B-v0.2IQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
Qwen2.5-Math-72B-InstructIQ3_XXS72.7B29.66 GiB5.31 GiB35.65 GiB0.35 GiB7±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M5 Pro run?
2023 of 2118 indexed open-weight models fit a Apple M5 Pro at 32,768 context with q8_0 KV cache, the largest being command-r-35b-writer-v2 at I1-IQ3_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M5 Pro actually have?
Its nameplate is 48 GB, but about 33.48 GiB is available to a model once driver and compositor overhead is accounted for, and only 36 GB of the pool can be allocated to the GPU at all.
Is a Apple M5 Pro fast for local AI?
Its memory bandwidth is 307 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.