Apple · apple

Apple M5

Apple M5 has 24 GB of unified memory at 154 GB/s — about 16.74 GiB usable after driver and compositor overhead. 1864 of 2118 indexed models fit at 32K context with q8_0 KV. Note only 18 GB of its 24 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
24 GB
LPDDR5X-9600
Bandwidth
154 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1589video 16vision language 171embedding 26audio tts 21audio asr 39image 2

What fits at 32K context

largest quantization that fits, per model · 1864 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
llm-jp-4-32b-a3b-thinkingMoEIQ4_XS32.1B16.37 GiB1.06 GiB17.99 GiB0.01 GiB22±37%
OLMo-2-0325-32BQ3_K_S32.2B13.09 GiB4.25 GiB17.99 GiB0.01 GiB7±8.3%
granite-8b-code-instruct-4kF168.1B15.01 GiB2.39 GiB17.99 GiB0.01 GiB7±8.3%
granite-8b-code-base-4kF168.1B15.01 GiB2.39 GiB17.99 GiB0.01 GiB7±8.3%
MythoMax-L2-13bI1-IQ2_S13.0B4.11 GiB13.28 GiB17.98 GiB0.02 GiB7±8.3%
Llama-2-7b-chat-hfQ5_K_M6.7B8.91 GiB8.50 GiB17.98 GiB0.02 GiB7±8.3%
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-ThinkingIQ4_XS23.4B11.99 GiB5.38 GiB17.97 GiB0.03 GiB7±8.3%
Wan2.1-VACE-14BQ8_017.3B17.38 GiB0.00 GiB17.97 GiB0.03 GiB7±8.3%
gemma-2-27b-itIQ4_XS27.2B13.80 GiB3.48 GiB17.96 GiB0.04 GiB7±8.3%
magnum-v4-27bIQ4_XS27.2B13.80 GiB3.48 GiB17.96 GiB0.04 GiB7±8.3%
InternVL3_5-30B-A3BQ4_K_M30.8B17.35 GiB0.00 GiB17.95 GiB0.05 GiB7±8.3%
Muse-Glimmer-30BQ4_K_L29.8B17.05 GiB0.27 GiB17.95 GiB0.05 GiB7±8.3%
Huihui-Qwen3-4B-Instruct-2507-abliteratedF324.0B14.99 GiB2.39 GiB17.94 GiB0.06 GiB7±8.3%
OpenCaption-4B-VL-SFT-v1.0F324.4B14.99 GiB2.39 GiB17.94 GiB0.06 GiB7±8.3%
gemma-4-A4B-98e-v7-coder-itMoEQ6_K20.5B16.58 GiB0.82 GiB17.94 GiB0.06 GiB7±8.3%
gemma-4-A4B-98e-v7-coderx-itMoEQ6_K20.5B16.58 GiB0.82 GiB17.94 GiB0.06 GiB7±8.3%
glm-4-9b-chat-1mQ5_K_M9.5B6.72 GiB10.63 GiB17.94 GiB0.06 GiB7±8.3%
North-Mini-Code-1.0MoEQ4_K_S30.5B16.81 GiB0.60 GiB17.94 GiB0.06 GiB25±37%
Ministral-8B-Instruct-2410F168.0B14.95 GiB2.39 GiB17.92 GiB0.08 GiB7±8.3%
Olmo-3.1-32B-InstructQ3_K_L32.2B15.75 GiB1.51 GiB17.91 GiB0.09 GiB7±8.3%
Olmo-3.1-32B-ThinkQ3_K_L32.2B15.75 GiB1.51 GiB17.91 GiB0.09 GiB7±8.3%
Olmo-3-32B-ThinkQ3_K_L32.2B15.75 GiB1.51 GiB17.91 GiB0.09 GiB7±8.3%
Ornith-1.0-35B-uncensored-hereticMoEQ3_K_L35.1B17.02 GiB0.33 GiB17.91 GiB0.09 GiB33±37%
Gemma-4-31B-Isometry-RPI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Gemma-4-Dark-Gemistry-31BI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Prosopon-31BI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Gemma-4-Novelist-Eclipse-31BI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Giftige-Blume-31B-v1-StyleSwapI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
G4-MeroMero-31B-StyleSwapI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Gemma-4-31B-StyleTune-heretic-araI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Pantheon-Reasoning-31B-1.1I1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Gemma-4-31B-StyleTuneI1-IQ3_M32.7B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
Barcenas-StyleTune-31B-FableI1-IQ3_M32.1B14.00 GiB3.28 GiB17.91 GiB0.09 GiB7±8.3%
spoomplesmaxx-v2.1-30BI1-Q3_K_M28.9B13.00 GiB4.25 GiB17.91 GiB0.09 GiB7±8.3%
Huihui-granite-4.1-30b-abliteratedI1-Q3_K_M28.9B13.00 GiB4.25 GiB17.91 GiB0.09 GiB7±8.3%
granite-4.1-30b-hereticI1-Q3_K_M28.9B13.00 GiB4.25 GiB17.91 GiB0.09 GiB7±8.3%
granite-4.1-30bQ3_K_M28.9B13.00 GiB4.25 GiB17.91 GiB0.09 GiB7±8.3%
Qwen3-53B-A3B-2507-THINKING-TOTAL-RECALL-v2-MASTER-CODERMoEI1-IQ2_XS53.0B14.56 GiB2.79 GiB17.90 GiB0.10 GiB15±37%
Qwen3-Coder-Next-Opus-4.6-Reasoning-DistilledMoEIQ1_M16.95 GiB0.40 GiB17.89 GiB0.11 GiB35±37%
NousCoder-14BQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
spoomplesmaxx-mini-14BQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
qwen3-14b-code-reasoning-conversationalQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Claria-14bQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
NTX-2.1-ProQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Qwen3-14B-UncensoredQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Qwen3-14BQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Qwen3-14B-abliteratedQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
FrogMini-14B-2510Q8_014.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Josiefied-Qwen3-14B-abliterated-v3Q8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Hermes-4-14BQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Slava-Qwen3-14B-SerbianQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Qwen3-14B-BaseQ8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Huihui-Qwen3-14B-abliterated-v2Q8_014.8B14.62 GiB2.66 GiB17.89 GiB0.11 GiB7±8.3%
Qwen3.6-27B-Fable-5-ExperimentalQ4_K_M27.8B16.20 GiB1.06 GiB17.88 GiB0.12 GiB7±8.3%
OmniAtlas-Qwen3-30B-A3BI1-Q4_K_M31.7B17.28 GiB0.00 GiB17.88 GiB0.12 GiB7±8.3%
Qwen3-Omni-30B-A3B-InstructQ4_K_M35.3B17.28 GiB0.00 GiB17.88 GiB0.12 GiB7±8.3%
Qwen3-Omni-30B-A3B-CaptionerI1-Q4_K_M31.7B17.28 GiB0.00 GiB17.88 GiB0.12 GiB7±8.3%
Qwen3-Omni-30B-A3B-ThinkingQ4_K_M31.7B17.28 GiB0.00 GiB17.88 GiB0.12 GiB7±8.3%
codegeex4-all-9bQ5_K_M9.4B6.65 GiB10.63 GiB17.88 GiB0.12 GiB7±8.3%
Yi-34B-200K-DARE-megamerge-v8I1-IQ3_XS34.4B13.26 GiB3.98 GiB17.88 GiB0.12 GiB7±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M5 run?
1864 of 2118 indexed open-weight models fit a Apple M5 at 32,768 context with q8_0 KV cache, the largest being llm-jp-4-32b-a3b-thinking at IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M5 actually have?
Its nameplate is 24 GB, but about 16.74 GiB is available to a model once driver and compositor overhead is accounted for, and only 18 GB of the pool can be allocated to the GPU at all.
Is a Apple M5 fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.