Apple · apple
Apple M4
Apple M4 has 12 GB of unified memory at 120 GB/s — about 8.37 GiB usable after driver and compositor overhead. 1535 of 2118 indexed models fit at 32K context with q4_0 KV. Note only 9 GB of its 12 GB is allocatable to the GPU.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
12 GB
LPDDR5X-7500
Bandwidth
120 GB/s
128-bit bus
Tensor FP16
—
dense
TDP
—
text 1321video 12vision language 114image 2audio asr 39audio tts 21embedding 26
What fits at 32K context
largest quantization that fits, per model · 1535 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Le-Chaton-Slim-23BMoE | I1-Q2_K_S | 23.3B | 7.52 GiB | 0.91 GiB | 9.00 GiB | 0.00 GiB | 19±37% |
| Wan2.1-T2V-14B | Q4_0 | 14.3B | 8.41 GiB | 0.00 GiB | 9.00 GiB | 0.00 GiB | 11±8.3% |
| GLM-4.7-Flash-REAP-23B-A3BMoE | UD-IQ2_M | 23.0B | 7.97 GiB | 0.46 GiB | 8.99 GiB | 0.01 GiB | 31±37% |
| Qwen3-VL-30B-A3B-InstructMoE | UD-TQ1_0 | 31.1B | 7.60 GiB | 0.84 GiB | 8.99 GiB | 0.01 GiB | 28±37% |
| GigaChat3-10B-A1.8B-baseMoE | Q6_K | 11.5B | 8.18 GiB | 0.26 GiB | 8.99 GiB | 0.01 GiB | 37±37% |
| Falcon3-7B-Instruct | Q8_0 | 7.5B | 7.38 GiB | 0.98 GiB | 8.98 GiB | 0.02 GiB | 12±8.3% |
| Gemma-The-Writer-N-Restless-Quill-10B-Uncensored | I1-Q5_K_S | 10.0B | 6.55 GiB | 1.84 GiB | 8.98 GiB | 0.02 GiB | 11±8.3% |
| gemma-3n-E2B-it | F16 | 5.4B | 8.31 GiB | 0.12 GiB | 8.98 GiB | 0.02 GiB | 11±8.3% |
| InternVL3_5-14B | Q4_K_M | 15.1B | 8.38 GiB | 0.00 GiB | 8.98 GiB | 0.02 GiB | 11±8.3% |
| GLM-4.7-Flash-hereticMoE | IQ2_XXS | 29.9B | 7.95 GiB | 0.46 GiB | 8.98 GiB | 0.02 GiB | 33±37% |
| Phi-3-mini-4k-instructKV unresolved | IQ3_XXS | 3.8B | 5.05 GiB | 3.38 GiB | 8.98 GiB | 0.02 GiB | 11±8.3% |
| Qwen3-VL-30B-A3B-ThinkingMoE | UD-TQ1_0 | 31.1B | 7.59 GiB | 0.84 GiB | 8.98 GiB | 0.02 GiB | 28±37% |
| Qwen3-30B-A3B-Thinking-2507MoE | UD-TQ1_0 | 30.5B | 7.59 GiB | 0.84 GiB | 8.98 GiB | 0.02 GiB | 28±37% |
| Kimi-VL-A3B-InstructMoE | I1-IQ4_XS | 16.4B | 8.15 GiB | 0.27 GiB | 8.98 GiB | 0.02 GiB | 32±37% |
| Moonlight-16B-A3B-InstructMoE | IQ4_XS | 16.0B | 8.15 GiB | 0.27 GiB | 8.98 GiB | 0.02 GiB | 32±37% |
| Ministral-3-14B-Instruct-2512-BF16-abliterated | IQ4_XS | 13.9B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Ministral-3-14B-abliterated | IQ4_XS | 13.9B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Ministral-3-14B-Reasoning-2512-Uncensored | IQ4_XS | 13.9B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Qwen3-30B-A3BMoE | IQ2_XXS | 30.5B | 7.59 GiB | 0.84 GiB | 8.97 GiB | 0.03 GiB | 28±37% |
| Pantheon-Proto-RP-1.8-30B-A3BMoE | IQ2_XXS | 30.5B | 7.59 GiB | 0.84 GiB | 8.97 GiB | 0.03 GiB | 28±37% |
| HunyuanVideo-1.5 | Q8_0 | 8.3B | 8.38 GiB | 0.00 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Forsaken-Void-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Silver-Siren-ST-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Tess-3-Mistral-Nemo-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| KrakenSakura-Maelstrom-12B-v1 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| MN-12B-Runeweaver-RP-RU | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Impish_Bloodmoon_12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Wayfarer-2-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Wayfarer-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Muse-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Vikhr-Nemo-12B-Instruct-R-21-09-24 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-Base-2407 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| writing-roleplay-20k-context-nemo-12b-v1.0 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Dans-PersonalityEngine-V1.3.0-12b | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| mini-magnum-12b-v1.1 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Lumimaid-v0.2-12B | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| MN-Violet-Lotus-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Rocinante-X-12B-v1-Heretic-Uncensored | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Heretica-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Violet_Twilight-v0.2 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Lumimaid-Magnum-v4-12B | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| arcee-fusion-lumaid-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-NeMo-12B-Abliterated | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Captain-Eris_Violet-V0.420-12B | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Rocinante-X-12B-v1-absolute-heresy | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Rocinante-X-12B-v1 | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-Gutenberg-Doppel-12B | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-Instruct-2407 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| magnum-v4-12b | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| MN-12b-RP-Ink | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-12B-ArliAI-RPMax-v1.1 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-2407-12B-Thinking-Claude-Gemini-GPT5.2-Uncensored-HERETIC | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Dans-SakuraKaze-V1.0.0-12b | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-Inst-2407-12B-Thinking-Uncensored-HERETIC-HI-Claude-Opus | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mistral-Nemo-Instruct-2407-12B-Thinking-M-Claude-Opus-High-Reasoning | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| MN-12B-Mag-Mell-R1 | Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Mordant-12B-Think | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Riverfish-Rocinante-12B-SFT-DPO | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| MN-Violet-Lotus-12B-Heretic | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
| Himeyuri-Magnum-12B-HereticMerge | I1-Q4_K_M | 12.2B | 6.96 GiB | 1.41 GiB | 8.97 GiB | 0.03 GiB | 12±8.3% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a Apple M4 run?
- 1535 of 2118 indexed open-weight models fit a Apple M4 at 32,768 context with q4_0 KV cache, the largest being Le-Chaton-Slim-23B at I1-Q2_K_S. That covers text, vision-language, image, video and speech models.
- How much usable memory does a Apple M4 actually have?
- Its nameplate is 12 GB, but about 8.37 GiB is available to a model once driver and compositor overhead is accounted for, and only 9 GB of the pool can be allocated to the GPU at all.
- Is a Apple M4 fast for local AI?
- Its memory bandwidth is 120 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.