Apple · apple

Apple M2 Max

Apple M2 Max has 96 GB of unified memory at 410 GB/s — about 66.96 GiB usable after driver and compositor overhead. 2053 of 2118 indexed models fit at 128K context with q8_0 KV. Note only 72 GB of its 96 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
96 GB
LPDDR5-6400
Bandwidth
410 GB/s
512-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 186text 1763image 2audio asr 39audio tts 21video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 2053 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedIQ4_NL109B58.67 GiB12.75 GiB71.99 GiB0.01 GiB10±37%
Mistral-Medium-3.5-128BQ2_K_L128B47.90 GiB23.38 GiB71.98 GiB0.02 GiB5±8.3%
Laguna-S-2.1MoEUD-Q4_K_M118B68.10 GiB3.26 GiB71.93 GiB0.07 GiB18±37%
v6-Finch-14B-HFIQ3_M14.1B6.49 GiB64.81 GiB71.90 GiB0.10 GiB5±8.3%
Qwen3.5-122B-A10BMoEQ4_K_S125B69.66 GiB1.59 GiB71.84 GiB0.16 GiB22±37%
Qwen3-VL-235B-A22B-ThinkingMoEUD-IQ1_S236B58.65 GiB12.48 GiB71.72 GiB0.28 GiB10±37%
Gemma-4-Novelist-Eclipse-31BBF1632.7B59.82 GiB11.25 GiB71.70 GiB0.30 GiB5±8.3%
Gemma-4-31B-StyleTuneBF1632.7B59.82 GiB11.25 GiB71.70 GiB0.30 GiB5±8.3%
calme-2.3-rys-78bQ4_K_L78.0B48.08 GiB22.84 GiB71.60 GiB0.40 GiB5±8.3%
GLM-4.7-REAP-218B-A32BMoEIQ1_M218B46.56 GiB24.44 GiB71.59 GiB0.41 GiB6±37%
Qwen3-VL-235B-A22B-InstructMoEUD-IQ1_S236B58.49 GiB12.48 GiB71.56 GiB0.44 GiB10±37%
GLM-4.5-Air-DerestrictedMoEIQ4_NL110B58.73 GiB12.22 GiB71.53 GiB0.47 GiB10±37%
GLM-4.5-AirMoEIQ4_NL110B58.73 GiB12.22 GiB71.53 GiB0.47 GiB10±37%
Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16MoEQ8_035.1B69.57 GiB1.33 GiB71.46 GiB0.54 GiB23±37%
c4ai-command-r-08-2024F1632.3B60.17 GiB10.63 GiB71.46 GiB0.54 GiB5±8.3%
granite-4.1-30bBF1628.9B53.77 GiB17.00 GiB71.43 GiB0.57 GiB5±8.3%
gpt-oss-120b-Uncensored-xCloudMoEI1-Q4_1117B68.42 GiB2.40 GiB71.36 GiB0.64 GiB21±37%
gpt-oss-120b-abliteratedMoEI1-Q4_1117B68.42 GiB2.40 GiB71.36 GiB0.64 GiB21±37%
step-3.5-flashIQ2_XXS199B44.74 GiB26.05 GiB71.36 GiB0.64 GiB5±8.3%
Meta-Llama-3-70B-InstructQ5_170.6B49.37 GiB21.25 GiB71.29 GiB0.71 GiB5±8.3%
Llama-3.1-70BQ5_170.6B49.36 GiB21.25 GiB71.29 GiB0.71 GiB5±8.3%
Qwen3.5-122B-A10B-hereticMoEI1-Q4_K_M123B69.11 GiB1.59 GiB71.28 GiB0.72 GiB22±37%
Step-3.7-FlashIQ1_M201B44.41 GiB26.05 GiB71.03 GiB0.97 GiB5±8.3%
Llama-2-13b-chat-hfQ5_K_M13.0B17.19 GiB53.13 GiB70.91 GiB1.09 GiB5±8.3%
Hunyuan-A13B-InstructMoEQ6_K80.4B61.75 GiB8.50 GiB70.80 GiB1.20 GiB5±8.3%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEIQ4_NL124B64.39 GiB5.84 GiB70.78 GiB1.22 GiB15±37%
Behemoth-X-123B-v2IQ3_XS123B46.70 GiB23.38 GiB70.78 GiB1.22 GiB5±8.3%
Mistral-Small-4-119B-2603MoEUD-Q4_K_M119B68.70 GiB1.49 GiB70.77 GiB1.23 GiB23±37%
CalmeRys-78B-Orpo-v0.1Q4_K_M78.0B47.22 GiB22.84 GiB70.74 GiB1.26 GiB5±8.3%
Qwen3-235B-A22B-abliteratedMoEI1-IQ2_XXS235B57.55 GiB12.48 GiB70.62 GiB1.38 GiB10±37%
L3-DARKEST-PLANET-16.5BQ8_016.5B50.95 GiB18.86 GiB70.40 GiB1.60 GiB5±8.3%
GLM-4.6VMoEQ4_0108B57.48 GiB12.22 GiB70.28 GiB1.72 GiB10±37%
deepseek-llm-67b-chatI1-Q5_K_M67.4B44.38 GiB25.23 GiB70.27 GiB1.73 GiB5±8.3%
deepseek-llm-67b-baseI1-Q5_K_M67.4B44.38 GiB25.23 GiB70.27 GiB1.73 GiB5±8.3%
openbuddy-deepseek-67b-v15.3-4kI1-Q5_K_M67.4B44.38 GiB25.23 GiB70.26 GiB1.74 GiB5±8.3%
DeepSeek-Coder-V2-Instruct-0724MoEIQ2_S236B65.07 GiB4.48 GiB70.14 GiB1.86 GiB17±37%
DeepSeek-V2.5MoEIQ2_S236B65.07 GiB4.48 GiB70.14 GiB1.86 GiB17±37%
DeepSeek-Coder-V2-InstructMoEIQ2_S236B65.07 GiB4.48 GiB70.14 GiB1.86 GiB17±37%
Assistant_Pepe_70BQ5_K_L70.6B48.20 GiB21.25 GiB70.13 GiB1.87 GiB5±8.3%
c4ai-command-r-plus-08-2024IQ4_XS104B52.34 GiB17.00 GiB70.07 GiB1.93 GiB5±8.3%
Step-3.5-Flash-REAP-121B-A11BI1-IQ3_XXS121B43.40 GiB26.05 GiB70.02 GiB1.98 GiB5±8.3%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ3_XS139B53.00 GiB16.47 GiB70.00 GiB2.00 GiB9±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ3_XS139B53.00 GiB16.47 GiB70.00 GiB2.00 GiB9±37%
MiMo-V2-FlashMoEKV unresolvedIQ1_M310B61.31 GiB7.97 GiB69.87 GiB2.13 GiB14±37%
HuatuoGPT-o1-72BQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Rombo-LLM-V3.0-Qwen-72bQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
EVA-Qwen2.5-72B-v0.2Q5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
MiroThinker-v1.0-72BQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Qwen2.5-72BQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
magnum-v4-72bQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
KAT-Dev-72B-ExpQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Homer-v1.0-Qwen2.5-72BQ5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Chuluun-Qwen2.5-72B-v0.01Q5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Qwen2.5-VL-72B-InstructQ5_K_S73.4B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Tower-Plus-72B-ultra-uncensored-hereticI1-Q5_K_S72.7B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
UI-TARS-72B-DPOQ5_K_S73.4B47.85 GiB21.25 GiB69.78 GiB2.22 GiB5±8.3%
Mixtral-8x22B-Instruct-v0.1MoEIQ3_XS141B54.23 GiB14.88 GiB69.72 GiB2.28 GiB6±37%
gpt-oss-20b-hereticMoEIQ4_NL20.9B67.58 GiB1.60 GiB69.72 GiB2.28 GiB14±37%
Mixtral-8x22B-v0.1MoEIQ3_XS141B54.23 GiB14.88 GiB69.71 GiB2.29 GiB6±37%
Mixtral-8x22B-v0.1MoEIQ3_XS141B54.23 GiB14.88 GiB69.71 GiB2.29 GiB6±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M2 Max run?
2053 of 2118 indexed open-weight models fit a Apple M2 Max at 131,072 context with q8_0 KV cache, the largest being Llama-4-Scout-17B-16E-Instruct at IQ4_NL. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M2 Max actually have?
Its nameplate is 96 GB, but about 66.96 GiB is available to a model once driver and compositor overhead is accounted for, and only 72 GB of the pool can be allocated to the GPU at all.
Is a Apple M2 Max fast for local AI?
Its memory bandwidth is 410 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.