Apple · apple

Apple M1

Apple M1 has 16 GB of unified memory at 68 GB/s — about 11.16 GiB usable after driver and compositor overhead. 1712 of 2118 indexed models fit at 64K context with q4_0 KV. Note only 12 GB of its 16 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
LPDDR4X-4266
Bandwidth
68 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 149video 15text 1460image 2audio asr 39embedding 26audio tts 21

What fits at 64K context

largest quantization that fits, per model · 1712 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3.5-35B-A3BMoEIQ2_S36.0B11.09 GiB0.35 GiB12.00 GiB0.00 GiB22±37%
Qwen3.6-35B-A3BMoEIQ2_S36.0B11.09 GiB0.35 GiB12.00 GiB0.00 GiB22±37%
Wan2.1-VACE-14BQ5_K_S17.3B11.41 GiB0.00 GiB12.00 GiB0.00 GiB5±8.3%
medgemma-27b-itI1-Q2_K28.8B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
gemma-3-27b-it-abliterated-refined-visionI1-Q2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
gemma-3-27b-it-abliteratedQ2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
Nidum-Gemma-3-27B-it-UncensoredI1-Q2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
gemma-3-27b-itQ2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
AtomicGPT-gemma3-27bI1-Q2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
Unbound-v1.12.0-27BI1-Q2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
Mira-v1.12-Ties-27BI1-Q2_K27.4B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
Medgamma27BI1-Q2_K27.0B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
medgemma-27b-text-itQ2_K27.0B9.78 GiB1.58 GiB11.99 GiB0.01 GiB5±8.3%
Qwythos-9B-v2Q4_K_M9.7B10.84 GiB0.56 GiB11.99 GiB0.01 GiB5±8.3%
GLM-4.7-Flash-REAP-23B-A3BMoEQ3_K_M23.0B10.50 GiB0.93 GiB11.99 GiB0.01 GiB13±37%
Phi-4-reasoning-plusQ4_K_S14.7B7.86 GiB3.52 GiB11.99 GiB0.01 GiB5±8.3%
Phi-4-reasoningQ4_K_S14.7B7.86 GiB3.52 GiB11.99 GiB0.01 GiB5±8.3%
phi-4Q4_K_S14.7B7.86 GiB3.52 GiB11.99 GiB0.01 GiB5±8.3%
dolphin-2.6-mixtral-8x7bMoEI1-IQ1_S46.7B9.15 GiB2.25 GiB11.98 GiB0.02 GiB7±37%
xLAM-8x7b-rMoEIQ1_S46.7B9.15 GiB2.25 GiB11.98 GiB0.02 GiB7±37%
deepseek-coder-6.7B-kexerI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
Magicoder-S-DS-6.7BI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
deepseek-coder-6.7b-baseI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
WizardLM-7B-UncensoredI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
Llama-2-7B-32K-InstructI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
Luna-AI-Llama2-UncensoredI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
Swallow-7b-NVE-instruct-hfI1-IQ3_XXS6.7B2.41 GiB9.00 GiB11.98 GiB0.02 GiB5±8.3%
Rocinante-XL-16B-v1Q3_K_M16.1B7.59 GiB3.80 GiB11.98 GiB0.02 GiB5±8.3%
North-Mini-Code-1.0MoEUD-IQ3_XXS30.5B10.90 GiB0.55 GiB11.98 GiB0.02 GiB16±37%
Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoEI1-IQ1_M42.4B9.08 GiB2.36 GiB11.97 GiB0.03 GiB9±37%
InternVL3_5-30B-A3BIQ3_XXS30.8B11.38 GiB0.00 GiB11.97 GiB0.03 GiB5±8.3%
gemma-2-27b-itIQ2_XS27.2B7.82 GiB3.46 GiB11.97 GiB0.03 GiB5±8.3%
magnum-v4-27bIQ2_XS27.2B7.82 GiB3.46 GiB11.97 GiB0.03 GiB5±8.3%
Llama-3.2-8X3B-MOE-Dark-Champion-Instruct-uncensored-abliterated-18.4BMoEIQ4_XS18.4B9.44 GiB1.97 GiB11.96 GiB0.04 GiB7±37%
Qwen3-VL-32B-Instruct-ultra-uncensored-hereticI1-IQ1_S33.4B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
Huihui-Qwen3-VL-32B-Instruct-abliteratedI1-IQ1_S33.4B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
ColorGUI-32BI1-IQ1_S33.4B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
Qwen3-32B-UncensoredI1-IQ1_S32.8B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
Qwen3-32B-abliteratedI1-IQ1_S32.8B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
AReaL-boba-2-32BI1-IQ1_S32.8B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
Assistant_Pepe_32BI1-IQ1_S32.8B6.82 GiB4.50 GiB11.96 GiB0.04 GiB5±8.3%
NVIDIA-Nemotron-Nano-12B-v2Q4_K_M12.3B6.98 GiB4.36 GiB11.96 GiB0.04 GiB5±8.3%
gemma-4-A4B-98e-v6-coder-itMoEIQ4_NL20.5B10.63 GiB0.79 GiB11.96 GiB0.04 GiB5±8.3%
gemma-4-A4B-98e-v7-coder-itMoEIQ4_NL20.5B10.63 GiB0.79 GiB11.96 GiB0.04 GiB5±8.3%
gemma-4-A4B-98e-v7-coderx-itMoEIQ4_NL20.5B10.63 GiB0.79 GiB11.96 GiB0.04 GiB5±8.3%
EVA-abliterated-TIES-Qwen2.5-14BI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Neuron-V1-14B-InstructI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Ektome-Qwen2.5-Coder-14B-Instruct-PristinelyUncensoredI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Qwen2.5-14B-Instruct-1M-abliteratedI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
DeepCoder-14B-PreviewQ4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Deepseeker-Kunou-Qwen2.5-14bI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
SuperNova-MediusQ4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
14B-Qwen2.5-Kunou-v1I1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Sugoi-14B-Ultra-HFI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Qwen2.5-14B-Instruct-abliterated-v2Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Qwen2.5-14B-Instruct-UncensoredQ4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
Qwen2.5-Coder-14B-Instruct-abliteratedQ4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
OpenCodeReasoning-Nemotron-14BQ4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
DeepSeek-R1-Distill-Qwen-14B-abliterated-v2I1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
C1-TachuI1-Q4_K_S14.8B7.98 GiB3.38 GiB11.96 GiB0.04 GiB5±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M1 run?
1712 of 2118 indexed open-weight models fit a Apple M1 at 65,536 context with q4_0 KV cache, the largest being Qwen3.5-35B-A3B at IQ2_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M1 actually have?
Its nameplate is 16 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for, and only 12 GB of the pool can be allocated to the GPU at all.
Is a Apple M1 fast for local AI?
Its memory bandwidth is 68 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.