Apple · apple

Apple M2

Apple M2 has 8 GB of unified memory at 102 GB/s — about 5.58 GiB usable after driver and compositor overhead. 716 of 2118 indexed models fit at 128K context with q4_0 KV. Note only 6 GB of its 8 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
LPDDR5-6400
Bandwidth
102 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 560vision language 75audio tts 19embedding 21audio asr 36video 5

What fits at 128K context

largest quantization that fits, per model · 716 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nanbeige4.2-3BQ4_K_S4.2B2.33 GiB3.09 GiB5.99 GiB0.01 GiB15±8.3%
Qwen3.5-4BQ8_04.7B4.30 GiB1.13 GiB5.99 GiB0.01 GiB15±8.3%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseIQ4_XS8.1B4.42 GiB0.98 GiB5.99 GiB0.01 GiB15±8.3%
Dolphin3.0-Llama3.2-3BIQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Llama-Doctor-3.2-3B-InstructI1-IQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Llama-Song-Stream-3B-InstructIQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3.2-3B-Instruct-roleplay-tunedI1-IQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3.2-3B-Instruct-heretic-ablitered-uncensoredI1-IQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
llama-3.2-Korean-Bllossom-3BIQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3.2-3B-InstructIQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
llama-3.2-3b-instructIQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Llama3.2-3B-creative-writer-v0.1I1-IQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Firefly-V3.2I1-IQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Firefly-V3I1-IQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
Hermes-3-Llama-3.2-3BIQ3_M3.2B1.49 GiB3.94 GiB5.99 GiB0.01 GiB15±8.3%
orpheus-3b-0.1-pretrainedQ2_K3.8B1.49 GiB3.94 GiB5.98 GiB0.02 GiB15±8.3%
EVA-Yi-1.5-9B-32K-V1I1-IQ1_M8.8B2.03 GiB3.38 GiB5.98 GiB0.02 GiB15±8.3%
Yi-Coder-9B-ChatIQ1_M8.8B2.03 GiB3.38 GiB5.98 GiB0.02 GiB15±8.3%
FrickFritz-4BQ8_04.7B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
Newton-bot-3-VLM-mini-4BQ8_04.7B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
qwen3.5-4b-agentic-coder-v4Q8_04.7B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
Myth-4BQ8_04.3B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
Qwen3.5-4B-UncensoredQ8_04.7B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
JOSIE-2-4B-PreviewQ8_04.7B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
Surogate-3.5-4BQ8_05.3B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
Qwopus3.5-4B-v3Q8_04.7B4.29 GiB1.13 GiB5.98 GiB0.02 GiB15±8.3%
internlm3-8b-instructQ3_K_S8.8B3.72 GiB1.69 GiB5.98 GiB0.02 GiB15±8.3%
AMD-OLMo-1B-SFT-DPOQ6_K_L1.2B0.92 GiB4.50 GiB5.97 GiB0.03 GiB15±8.3%
granite-vision-4.1-4bQ6_K4.0B2.60 GiB2.81 GiB5.97 GiB0.03 GiB15±8.3%
Parable-Granite-4.1-3B-Claude-Fable-5I1-Q6_K3.4B2.60 GiB2.81 GiB5.97 GiB0.03 GiB15±8.3%
granite-4.0-microQ6_K3.4B2.60 GiB2.81 GiB5.97 GiB0.03 GiB15±8.3%
granite-4.1-3bQ6_K3.4B2.60 GiB2.81 GiB5.97 GiB0.03 GiB15±8.3%
granite-4.0-micro-baseQ6_K3.4B2.60 GiB2.81 GiB5.97 GiB0.03 GiB15±8.3%
Tini-Cybersec-8B-A1BMoEQ4_18.5B5.00 GiB0.42 GiB5.96 GiB0.04 GiB31±37%
SuperGemma-4-12b-abliteratedI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-hereticI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-hereticI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-it-uncensored-hereticI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Grug-12BI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Aura-Medium-v1-BF16I1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-it-Esper4I1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-it-GuardpointI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Gemma-4-12B-it-AEON-Abliterated-K4-BF16I1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-it-Tachibana-AgentI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12b-marvin-gutenberg-rp-v2I1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12b-crownelius-writerI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliteratedI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12b-asterion-agenticI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Huihui-gemma-4-12B-agentic-fable5-abliteratedI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
g4-12b-it-trismegistusI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma4-12b-it-asimovI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
FabGemmaI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliteratedI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-it-abliterated-uncensoredI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Gemma-4-12b-it-AbliteratedI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-Queen-it-qat-q4_0-unquantizedI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-it-heretic_decensoredI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
Iris-12B-gemma-4-it-qatI1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
gemma-4-12B-coder-fable5-composer2.5-v1I1-IQ1_M12.0B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
G4-Starry-Ocean-12BI1-IQ1_M11.9B2.98 GiB2.38 GiB5.96 GiB0.04 GiB15±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing147.27 tok/s115.58180.497
Text generation12.18 tok/s7.6716.967
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M2 run?
716 of 2118 indexed open-weight models fit a Apple M2 at 131,072 context with q4_0 KV cache, the largest being Nanbeige4.2-3B at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M2 actually have?
Its nameplate is 8 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for, and only 6 GB of the pool can be allocated to the GPU at all.
Is a Apple M2 fast for local AI?
Its memory bandwidth is 102 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.