Apple · apple

Apple M3 Pro

Apple M3 Pro has 18 GB of unified memory at 154 GB/s — about 12.56 GiB usable after driver and compositor overhead. 1854 of 2118 indexed models fit at 4K context with q8_0 KV. Note only 14 GB of its 18 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
18 GB
LPDDR5-6400
Bandwidth
154 GB/s
192-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1591video 15vision language 160audio asr 39audio tts 21image 2embedding 26

What fits at 4K context

largest quantization that fits, per model · 1854 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
dolphin-2.9.1-mixtral-1x22bMoEI1-Q4_K_M22.2B12.42 GiB0.46 GiB13.50 GiB0.00 GiB6±37%
Wan2.2-S2V-14BQ4_K_M16.3B12.91 GiB0.00 GiB13.49 GiB0.01 GiB10±8.3%
GLM-4-32B-0414-Korean-CultureI1-IQ3_XS32.6B12.72 GiB0.13 GiB13.49 GiB0.01 GiB10±8.3%
GLM-Z1-32B-0414IQ3_XS32.6B12.72 GiB0.13 GiB13.49 GiB0.01 GiB10±8.3%
GLM-4-32B-0414IQ3_XS32.6B12.72 GiB0.13 GiB13.49 GiB0.01 GiB10±8.3%
grug-27bIQ3_M27.4B12.73 GiB0.13 GiB13.47 GiB0.03 GiB10±8.3%
Carnice-V2-27bIQ3_M27.4B12.73 GiB0.13 GiB13.47 GiB0.03 GiB10±8.3%
Fara1.5-27BIQ3_M27.4B12.73 GiB0.13 GiB13.47 GiB0.03 GiB10±8.3%
granite-4.0-h-tinyMoEBF166.9B12.94 GiB0.02 GiB13.47 GiB0.03 GiB32±37%
granite-4.0-h-tiny-baseMoEBF166.9B12.94 GiB0.02 GiB13.47 GiB0.03 GiB32±37%
Fallen-Gemma3-27B-v1Q3_K_M27.4B12.51 GiB0.36 GiB13.46 GiB0.04 GiB10±8.3%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-Q4_121.8B12.77 GiB0.12 GiB13.45 GiB0.05 GiB10±8.3%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-Q4_121.8B12.77 GiB0.12 GiB13.45 GiB0.05 GiB10±8.3%
ERNIE-4.5-21B-A3B-ThinkingI1-Q4_121.8B12.77 GiB0.12 GiB13.45 GiB0.05 GiB10±8.3%
gemma-4-26B-A4B-itMoEUD-IQ4_NL26.5B12.68 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
Goetia-26B-A4B-v1.4MoEI1-Q3_K_M26.0B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
G4-Moonlight-Dusk-26B-A4BMoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
Chimera-X-26B-A4BMoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
Gemma-4-26B-A4B-StyleTuneMoEI1-Q3_K_M26.5B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-Q3_K_M25.8B12.67 GiB0.24 GiB13.45 GiB0.05 GiB10±8.3%
gpt-oss-20bMoEF1621.5B12.85 GiB0.06 GiB13.44 GiB0.06 GiB27±37%
gpt-oss-safeguard-20bMoEF1621.5B12.85 GiB0.06 GiB13.44 GiB0.06 GiB27±37%
Huihui-gemma-3n-E4B-it-abliteratedF167.8B12.80 GiB0.06 GiB13.44 GiB0.06 GiB10±8.3%
gemma-3n-E4B-itF167.8B12.80 GiB0.06 GiB13.44 GiB0.06 GiB10±8.3%
Qwen3-VL-8B-Instruct-HereticI1-Q6_K8.8B12.53 GiB0.30 GiB13.41 GiB0.09 GiB10±8.3%
MiniCPM-V-4_5Q6_K8.7B12.53 GiB0.30 GiB13.41 GiB0.09 GiB10±8.3%
granite-4.0-h-smallMoEIQ3_XS32.2B12.83 GiB0.03 GiB13.40 GiB0.10 GiB26±37%
c4ai-command-r-08-2024Q2_K_L32.3B12.40 GiB0.33 GiB13.40 GiB0.10 GiB10±8.3%
Gemma-4-31B-Isometry-RPI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Gemma-4-Dark-Gemistry-31BI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Prosopon-31BI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Gemma-4-Novelist-Eclipse-31BI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Giftige-Blume-31B-v1-StyleSwapI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
G4-MeroMero-31B-StyleSwapI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Gemma-4-31B-StyleTune-heretic-araI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Pantheon-Reasoning-31B-1.1I1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Gemma-4-31B-StyleTuneI1-IQ3_XXS32.7B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Barcenas-StyleTune-31B-FableI1-IQ3_XXS32.1B11.81 GiB0.95 GiB13.40 GiB0.10 GiB10±8.3%
Qwen-AgentWorld-35B-A3BMoEUD-IQ3_XXS34.7B12.80 GiB0.04 GiB13.39 GiB0.11 GiB46±37%
Ornith-1.0-35BMoEUD-IQ3_XXS34.7B12.80 GiB0.04 GiB13.39 GiB0.11 GiB46±37%
Qwen3.6-27B-A3B-CoderMoEI1-Q3_K_L26.7B12.80 GiB0.04 GiB13.39 GiB0.11 GiB39±37%
Skywork-R1V3-38BIQ3_XS38.4B12.76 GiB0.00 GiB13.38 GiB0.12 GiB10±8.3%
OpenBuddy-R1-0528-Distill-Qwen3-32B-Preview0-QATQ2_K_L32.8B12.21 GiB0.53 GiB13.38 GiB0.12 GiB10±8.3%
KAT-DevQ2_K_L32.8B12.20 GiB0.53 GiB13.38 GiB0.12 GiB10±8.3%
Qwen3-VL-32B-InstructQ2_K_L33.4B12.20 GiB0.53 GiB13.38 GiB0.12 GiB10±8.3%
DeepSWE-PreviewQ2_K_L32.8B12.20 GiB0.53 GiB13.38 GiB0.12 GiB10±8.3%
Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-IQ3_S30.0B12.43 GiB0.39 GiB13.38 GiB0.12 GiB25±37%
LFM2-24B-A2BMoEQ4_023.8B12.77 GiB0.04 GiB13.37 GiB0.13 GiB37±37%
Luna-7B-A4BMoEF166.7B12.51 GiB0.30 GiB13.37 GiB0.13 GiB10±37%
lingbot-world-v2-14b-causal-fastQ5_K_M18.5B12.78 GiB0.00 GiB13.37 GiB0.13 GiB10±8.3%
Gemma4-Gutenberg-31BIQ2_M31.3B11.78 GiB0.95 GiB13.37 GiB0.13 GiB10±8.3%
gemma-4-31B-itIQ2_M31.3B11.78 GiB0.95 GiB13.37 GiB0.13 GiB10±8.3%
Gemma4-Gutenberg-31B-HereticIQ2_M31.3B11.78 GiB0.95 GiB13.37 GiB0.13 GiB10±8.3%
Equinox-31BIQ2_M31.3B11.78 GiB0.95 GiB13.37 GiB0.13 GiB10±8.3%
gemma-4-31B-it-SDFT-Heretic-RPIQ2_M30.7B11.78 GiB0.95 GiB13.37 GiB0.13 GiB10±8.3%
North-Mini-Code-1.0MoEQ3_K_S30.5B12.63 GiB0.20 GiB13.36 GiB0.14 GiB35±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing339.31 tok/s305.24343.177
Text generation17.53 tok/s16.9530.517
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M3 Pro run?
1854 of 2118 indexed open-weight models fit a Apple M3 Pro at 4,096 context with q8_0 KV cache, the largest being dolphin-2.9.1-mixtral-1x22b at I1-Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Pro actually have?
Its nameplate is 18 GB, but about 12.56 GiB is available to a model once driver and compositor overhead is accounted for, and only 14 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Pro fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.