Apple · apple

Apple M2

Apple M2 has 8 GB of unified memory at 102 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1334 of 2118 indexed models fit at 8K context with q4_0 KV. Note only 6 GB of its 8 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
LPDDR5-6400
Bandwidth
102 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 97text 1148embedding 26audio asr 38audio tts 19image 1video 5

What fits at 8K context

largest quantization that fits, per model · 1334 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwythos-9B-v2Q4_K_S9.7B5.34 GiB0.07 GiB6.00 GiB0.00 GiB15±8.3%
Tess-4-9BQ4_K_S9.7B5.34 GiB0.07 GiB6.00 GiB0.00 GiB15±8.3%
NVIDIA-Nemotron-Nano-9B-v2IQ4_XS8.9B4.91 GiB0.49 GiB6.00 GiB0.00 GiB15±8.3%
gemma-7bI1-Q3_K_L8.5B4.39 GiB0.98 GiB5.99 GiB0.01 GiB15±8.3%
Smilodon-9B-v1I1-IQ4_XS10.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
bella-bartender-v2I1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Gemma-2-9B-It-SPPO-Iter3I1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Gemma-SEA-LION-v3-9B-ITI1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
G2-Darkest-Writer-9B-v1I1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Tiger-Gemma-9B-v3I1-IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
gemma-2-9b-it-abliteratedIQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
gemma-2-9b-itIQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Tiger-Gemma-9B-v1IQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
magnum-v4-9bIQ4_XS9.2B4.83 GiB0.58 GiB5.99 GiB0.01 GiB15±8.3%
Qwen3-16B-A3BMoEIQ2_M16.0B5.24 GiB0.21 GiB5.99 GiB0.01 GiB35±37%
Fimbulvetr-11B-v2I1-Q3_K_M10.7B4.98 GiB0.42 GiB5.99 GiB0.01 GiB15±8.3%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopIQ3_M12.1B5.33 GiB0.05 GiB5.99 GiB0.01 GiB15±8.3%
Ministral-3-14B-Instruct-2512Q2_K_L13.9B5.03 GiB0.35 GiB5.99 GiB0.01 GiB15±8.3%
Ministral-3-14B-Reasoning-2512Q2_K_L13.9B5.03 GiB0.35 GiB5.99 GiB0.01 GiB15±8.3%
Hunyuan-7B-InstructQ5_K_L7.5B5.12 GiB0.28 GiB5.99 GiB0.01 GiB15±8.3%
Phi-3-medium-128k-instructQ2_K_L14.0B4.94 GiB0.44 GiB5.99 GiB0.01 GiB15±8.3%
Phi-3-medium-4k-instructQ2_K_L14.0B4.94 GiB0.44 GiB5.99 GiB0.01 GiB15±8.3%
Gemma-4-E4B-it-Minecraft-MT-en-zh-v0.1I1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-Queen-it-qat-q4_0-unquantizedI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
Gemma-4-E4B-Luchador-RudoI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
supergemma4-e4b-abliteratedI1-Q5_K_M7.5B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
Gemma-4-E4B-AbliteratedI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-itQ5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-ultra-uncensored-hereticQ5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-The-DECKARD-Claude-Opus-Expresso-Universe-HERETIC-UNCENSORED-ThinkingI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-The-DECKARD-Expresso-Universe-HERETIC-UNCENSORED-ThinkingI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-hereticI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-Claude-Opus-4.5-HERETIC-UNCENSORED-ThinkingI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
Huihui-gemma-4-E4B-it-abliteratedI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-Uncensored-MAXI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
Darkidol-Gemma-4-E4B-itI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-abliteratedI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-itQ5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-hereticQ5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-Agentic-Opus-Reasoning-GeminiCLI-mlx-4bitQ5_K_M7.5B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
OpenMedResearch-Gemma-4E4NI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
Reasoning-Medical0.1-E4B-sftI1-Q5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4BQ5_K_M8.0B5.37 GiB0.05 GiB5.98 GiB0.02 GiB15±8.3%
dolphincoder-starcoder2-15bKV unresolvedI1-IQ2_M16.0B5.16 GiB0.18 GiB5.98 GiB0.02 GiB15±8.3%
starcoder2-15bKV unresolvedIQ2_M16.0B5.16 GiB0.18 GiB5.98 GiB0.02 GiB15±8.3%
spoomplesmaxx-mini-14BI1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
vanilla-cn-roleplay-0.2I1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Claria-14bI1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
NTX-2.1-ProI1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Qwen3-14B-UncensoredI1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
FrogMini-14B-2510I1-Q2_K_S5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Qwen3-14B-abliteratedI1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Hermes-4-14BQ2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Slava-Qwen3-14B-SerbianI1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Huihui-Qwen3-14B-abliterated-v2I1-Q2_K_S14.8B5.02 GiB0.35 GiB5.98 GiB0.02 GiB15±8.3%
Apertus-8B-Instruct-2509Q4_K_L8.1B5.08 GiB0.28 GiB5.98 GiB0.02 GiB15±8.3%
Kuwutu-7B-CYOA-v2Q5_K_M7.6B5.08 GiB0.32 GiB5.98 GiB0.02 GiB15±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing147.27 tok/s115.58180.497
Text generation12.18 tok/s7.6716.967
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M2 run?
1334 of 2118 indexed open-weight models fit a Apple M2 at 8,192 context with q4_0 KV cache, the largest being Qwythos-9B-v2 at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M2 actually have?
Its nameplate is 8 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for, and only 6 GB of the pool can be allocated to the GPU at all.
Is a Apple M2 fast for local AI?
Its memory bandwidth is 102 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.
Apple M2 — what AI models can it run locally? — ossmodeldb