Apple · apple

Apple M1

Apple M1 has 8 GB of unified memory at 68 GB/s — about 5.58 GiB usable after driver and compositor overhead. 725 of 2118 indexed models fit at 64K context with q8_0 KV. Note only 6 GB of its 8 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
LPDDR4X-4266
Bandwidth
68 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 566vision language 77audio asr 37audio tts 19embedding 21video 5

What fits at 64K context

largest quantization that fits, per model · 725 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen2.5-3B-Instruct-abliteratedI1-Q5_K_S3.1B4.24 GiB1.20 GiB6.00 GiB0.00 GiB10±8.3%
Fara1.5-9BQ3_K_S9.4B4.35 GiB1.06 GiB6.00 GiB0.00 GiB10±8.3%
QwenPaw-Flash-9BQ3_K_S9.4B4.35 GiB1.06 GiB6.00 GiB0.00 GiB10±8.3%
grug-9bQ3_K_S9.4B4.35 GiB1.06 GiB6.00 GiB0.00 GiB10±8.3%
OmniCoder-9BQ3_K_S9.4B4.35 GiB1.06 GiB6.00 GiB0.00 GiB10±8.3%
Ornith-1.0-9BQ3_K_S9.2B4.35 GiB1.06 GiB6.00 GiB0.00 GiB10±8.3%
Qwen3.5-9B-NeoQ3_K_S9.7B4.35 GiB1.06 GiB6.00 GiB0.00 GiB10±8.3%
gemma-3-12b-itUD-IQ1_M12.2B3.03 GiB2.37 GiB6.00 GiB0.00 GiB10±8.3%
Voxtral-Mini-3B-2507IQ2_M4.7B1.45 GiB3.98 GiB6.00 GiB0.00 GiB10±8.3%
Llama-3.2-3B-Instruct-roleplay-tunedIQ4_XS3.2B1.71 GiB3.72 GiB5.99 GiB0.01 GiB10±8.3%
Llama-3.2-3B-Instruct-heretic-ablitered-uncensoredIQ4_XS3.2B1.71 GiB3.72 GiB5.99 GiB0.01 GiB10±8.3%
Llama3.2-3B-creative-writer-v0.1IQ4_XS3.2B1.71 GiB3.72 GiB5.99 GiB0.01 GiB10±8.3%
Firefly-V3.2IQ4_XS3.2B1.71 GiB3.72 GiB5.99 GiB0.01 GiB10±8.3%
Firefly-V3IQ4_XS3.2B1.71 GiB3.72 GiB5.99 GiB0.01 GiB10±8.3%
orpheus-3b-0.1-pretrainedIQ3_XS3.8B1.71 GiB3.72 GiB5.99 GiB0.01 GiB10±8.3%
Dolphin3.0-Llama3.2-3BIQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
Llama-Doctor-3.2-3B-InstructI1-IQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
Llama-Song-Stream-3B-InstructIQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
llama-3.2-Korean-Bllossom-3BIQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
Llama-3.2-3B-InstructIQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
llama-3.2-3b-instructIQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
Hermes-3-Llama-3.2-3BIQ4_XS3.2B1.70 GiB3.72 GiB5.98 GiB0.02 GiB10±8.3%
Darwin-36B-OpusMoEIQ2_M34.7B4.75 GiB0.66 GiB5.97 GiB0.03 GiB26±37%
Qwen3-VL-2B-ThinkingQ8_02.1B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
Qwen3-VL-Reranker-2BQ8_02.1B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
Qwen3-VL-2B-InstructQ8_02.1B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
Qwen3-VL-Embedding-2BQ8_02.1B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
Atomight-V2.5-1.7BQ8_01.7B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
OpenCaption-2B-VL-SFT-v1.0Q8_02.1B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
OpenClaude-1.7B-MergedQ8_01.7B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
gaon-1.7b-v2-instructQ8_01.7B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
gaon-1.7b-v2-translateQ8_01.7B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
Lightning-1.7BQ8_01.7B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
DorsetHeatwaveLLM2Q8_01.7B1.71 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
llama-3.2-3b-instruct-bnb-4bitQ3_K_L3.3B1.69 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
Llama-3.2-3BQ3_K_L3.2B1.69 GiB3.72 GiB5.97 GiB0.03 GiB10±8.3%
AMD-OLMo-1B-SFT-DPOQ8_01.2B1.17 GiB4.25 GiB5.96 GiB0.04 GiB10±8.3%
Yi-6B-ChatI1-Q4_K_S6.1B3.26 GiB2.13 GiB5.96 GiB0.04 GiB10±8.3%
Yi-1.5-6B-ChatQ4_K_S6.1B3.26 GiB2.13 GiB5.96 GiB0.04 GiB10±8.3%
Qwen3-1.7BQ6_K_L2.0B1.70 GiB3.72 GiB5.96 GiB0.04 GiB10±8.3%
glm-4v-9bQ4_K_S13.9B5.36 GiB0.00 GiB5.96 GiB0.04 GiB10±8.3%
Vero-Qwen35-9B-BaseI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Vero-Qwen35-9BI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Morphos-9BI1-Q3_K_M9.0B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwable-9B-Claude-Fable-5-hereticI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Holo-3.1-9BI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwable-9B-Claude-Fable-5I1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwen3.5-9B-imabari-v2I1-Q3_K_M9.7B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwen3.5-9B-abliterated-v2-MAXI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwable-9B-Claude-Fable-5-StraTAI1-Q3_K_M9.0B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q3_K_M9.0B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwen3.5-9B-RpRMax-v1I1-Q3_K_M9.7B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
AdQWENistrator-9BI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
cajal-9b-v2-fullI1-Q3_K_M9.0B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Qwen3.5-9B-ultra-uncensored-hereticQ3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Holo-3.1-9B-CoderI1-Q3_K_M9.0B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
PlutoI1-Q3_K_M9.4B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
Holo-3.1-9B-abliterated-rdoI1-Q3_K_M9.0B4.31 GiB1.06 GiB5.95 GiB0.05 GiB10±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing117.61 tok/s113.81124.168
Text generation10.96 tok/s7.8614.148
Benchmarked· n=8

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M1 run?
725 of 2118 indexed open-weight models fit a Apple M1 at 65,536 context with q8_0 KV cache, the largest being Qwen2.5-3B-Instruct-abliterated at I1-Q5_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M1 actually have?
Its nameplate is 8 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for, and only 6 GB of the pool can be allocated to the GPU at all.
Is a Apple M1 fast for local AI?
Its memory bandwidth is 68 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.