Apple · apple

Apple M4

Apple M4 has 12 GB of unified memory at 120 GB/s — about 8.37 GiB usable after driver and compositor overhead. 1209 of 2118 indexed models fit at 32K context with f16 KV. Note only 9 GB of its 12 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
12 GB
LPDDR5X-7500
Bandwidth
120 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1013video 12vision language 99audio asr 38embedding 26audio tts 20image 1

What fits at 32K context

largest quantization that fits, per model · 1209 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3-16B-A3BMoEUD-IQ2_M16.0B5.45 GiB3.00 GiB9.00 GiB0.00 GiB14±37%
Wan2.1-T2V-14BQ4_014.3B8.41 GiB0.00 GiB9.00 GiB0.00 GiB11±8.3%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ2_M21.8B6.68 GiB1.75 GiB9.00 GiB0.00 GiB11±8.3%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ2_M21.8B6.68 GiB1.75 GiB9.00 GiB0.00 GiB11±8.3%
ERNIE-4.5-21B-A3B-ThinkingI1-IQ2_M21.8B6.68 GiB1.75 GiB9.00 GiB0.00 GiB11±8.3%
OpenCoder-1.5B-InstructQ8_01.9B1.89 GiB6.56 GiB8.99 GiB0.01 GiB11±8.3%
legitus-instruct-v1I1-Q4_08.1B4.38 GiB4.00 GiB8.99 GiB0.01 GiB12±8.3%
Apertus-8B-Instruct-2509I1-Q4_08.1B4.38 GiB4.00 GiB8.99 GiB0.01 GiB12±8.3%
ERNIE-4.5-21B-A3B-PTIQ2_M21.9B6.67 GiB1.75 GiB8.99 GiB0.01 GiB11±8.3%
YuYu1015-Ornith-1.0-9B-abliteratedUD-Q6_K9.4B7.40 GiB1.00 GiB8.99 GiB0.01 GiB11±8.3%
Mistral-7B-v0.3Q4_K_L7.2B4.40 GiB4.00 GiB8.99 GiB0.01 GiB11±8.3%
Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-IQ2_S21.3B6.87 GiB1.50 GiB8.99 GiB0.01 GiB12±8.3%
Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-ThinkingI1-IQ2_S21.3B6.87 GiB1.50 GiB8.99 GiB0.01 GiB12±8.3%
InternVL3_5-14BQ4_K_M15.1B8.38 GiB0.00 GiB8.98 GiB0.02 GiB11±8.3%
SOLAR-10.7B-Instruct-v1.0I1-IQ1_M10.7B2.39 GiB6.00 GiB8.98 GiB0.02 GiB11±8.3%
Fimbulvetr-11B-v2I1-IQ1_M10.7B2.39 GiB6.00 GiB8.98 GiB0.02 GiB11±8.3%
MiniCPM-Llama3-V-2_5IQ4_NL8.5B4.38 GiB4.00 GiB8.97 GiB0.03 GiB11±8.3%
Llama3-ChatQA-1.5-8BIQ4_NL8.0B4.38 GiB4.00 GiB8.97 GiB0.03 GiB11±8.3%
grok-oss-Apollyon-8BIQ4_NL8.0B4.38 GiB4.00 GiB8.97 GiB0.03 GiB11±8.3%
Turkish-Llama-8b-Instruct-v0.1IQ4_NL8.0B4.38 GiB4.00 GiB8.97 GiB0.03 GiB11±8.3%
OLMoE-1B-7B-0924-InstructMoEI1-Q5_K_S6.9B4.45 GiB4.00 GiB8.97 GiB0.03 GiB11±37%
HunyuanVideo-1.5Q8_08.3B8.38 GiB0.00 GiB8.97 GiB0.03 GiB12±8.3%
Teuken-7B-instruct-research-v0.4Q8_07.5B7.38 GiB1.00 GiB8.97 GiB0.03 GiB11±8.3%
LFM2-24B-A2BMoEQ2_K_L23.8B7.78 GiB0.63 GiB8.97 GiB0.03 GiB31±37%
Olmo-3-7B-InstructIQ3_XXS7.3B2.70 GiB5.69 GiB8.97 GiB0.03 GiB11±8.3%
Olmo-3-7B-ThinkI1-IQ3_XXS7.3B2.70 GiB5.69 GiB8.97 GiB0.03 GiB11±8.3%
Voxtral-Mini-3B-2507Q8_04.7B4.66 GiB3.75 GiB8.97 GiB0.03 GiB11±8.3%
Assistant_Pepe_8BQ4_K_S4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Falcon3-10B-InstructIQ2_M10.3B3.35 GiB5.00 GiB8.96 GiB0.04 GiB12±8.3%
Foundation-Sec-8B-InstructI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Foundation-Sec-8B-Instruct-hereticI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Qwen3-VL-Embedding-8BQ3_K_L8.1B3.88 GiB4.50 GiB8.96 GiB0.04 GiB12±8.3%
dolphin-2.9-llama3-8bQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
saiga_llama3_8bQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Meta-Llama-3-8B-InstructQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Meta-Llama-3-8BQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama-3.1-Tulu-3-8BQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama-3-Groq-8B-Tool-UseQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
llama3.1-heretic-unsensoredI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Dolphin3.0-Llama3.1-8BQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
dolphin-2.9.4-llama3.1-8bQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Dolphin3.0-Llama3.1-8B-abliteratedQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
LLAMA-3_8B_Unaligned_BETAQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama-3.1-8BQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Anubis-Mini-8B-v1I1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Meta-Llama-3-8BQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Deepseek-R1-Distill-NSFW-RPv1Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
L3.1-Dark-Reasoning-LewdPlay-evo-Hermes-R1-Uncensored-8B-hereticI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
L3.1-Dark-Reasoning-LewdPlay-evo-Hermes-R1-Uncensored-8BI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Gluon-8BI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama3.3-8B-Instruct-Thinking-Heretic-Uncensored-Claude-4.5-Opus-High-ReasoningI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama3.3-8B-Instruct-Thinking-Claude-4.5-Opus-High-ReasoningI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Hypnos-i1-8BQ4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
grok-oss-Apollyon-8B-hereticI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama-3.3-8B-Instruct-128K-JbliteratedI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
PsyCoPref-Llama3-8BI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama3.1-GptDeluxe-8BI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
SelfCite-8B-CC-SFTI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
ai-girlfriend-v2I1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
Llama-3.3-8B-Instruct-128K_AbliteratedI1-Q4_K_S8.0B4.37 GiB4.00 GiB8.96 GiB0.04 GiB12±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M4 run?
1209 of 2118 indexed open-weight models fit a Apple M4 at 32,768 context with f16 KV cache, the largest being Qwen3-16B-A3B at UD-IQ2_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M4 actually have?
Its nameplate is 12 GB, but about 8.37 GiB is available to a model once driver and compositor overhead is accounted for, and only 9 GB of the pool can be allocated to the GPU at all.
Is a Apple M4 fast for local AI?
Its memory bandwidth is 120 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.