AMD · datacenter

Instinct MI210

Instinct MI210 has 64 GB of VRAM at 1638 GB/s — about 59.52 GiB usable after driver and compositor overhead. 2027 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
64 GB
HBM2e
Bandwidth
1638 GB/s
4096-bit bus
Tensor FP16
181 TF
dense
TDP
300 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1739vision language 184image 2audio asr 39audio tts 21video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 2027 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
MythoMax-L2-Kimiko-v2-13bIQ3_S13.0B5.45 GiB53.13 GiB59.52 GiB0.00 GiB18±26.5%
MythoMax-L2-13bI1-IQ3_S13.0B5.45 GiB53.13 GiB59.52 GiB0.00 GiB18±26.5%
Meta-Llama-3-70B-InstructQ4_070.6B37.23 GiB21.25 GiB59.51 GiB0.01 GiB18±26.5%
Huihui-GLM-4.5-Air-abliterated-lossytensorsMoEI1-IQ3_XS110B46.34 GiB12.22 GiB59.48 GiB0.04 GiB33±37%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEUD-IQ3_S124B52.74 GiB5.84 GiB59.48 GiB0.04 GiB51±37%
OLMo-2-1124-13B-InstructIQ3_XS13.7B5.40 GiB53.13 GiB59.47 GiB0.05 GiB18±26.5%
GLM-4.6VMoEUD-IQ3_XXS108B46.26 GiB12.22 GiB59.41 GiB0.11 GiB33±37%
deepseek-llm-67b-chatI1-Q3_K_L67.4B33.13 GiB25.23 GiB59.37 GiB0.15 GiB18±26.5%
deepseek-llm-67b-baseI1-Q3_K_L67.4B33.13 GiB25.23 GiB59.37 GiB0.15 GiB18±26.5%
openbuddy-deepseek-67b-v15.3-4kI1-Q3_K_L67.4B33.13 GiB25.23 GiB59.37 GiB0.15 GiB18±26.5%
Mixtral-8x22B-v0.1MoEIQ2_M141B43.50 GiB14.88 GiB59.34 GiB0.18 GiB21±37%
codellama-13b-oasst-sft-v10Q3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
chronos-hermes-13b-v2Q3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
WhiteRabbitNeo-13B-v1Q3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
CodeLlama-13b-Instruct-hfQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
Orca-2-13b-Alpaca-UncensoredI1-IQ3_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
WizardLM-13B-UncensoredI1-IQ3_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
WizardCoder-Python-13B-V1.0I1-IQ3_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
Guanaco-13B-UncensoredI1-IQ3_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
Llama-2-13b-chat-hfQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
Wizard-Vicuna-13B-UncensoredQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
WizardLM-13b-V1.0-UncensoredQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
WizardLM-1.0-Uncensored-Llama2-13bQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
speechless-llama2-hermes-orca-platypus-wizardlm-13bQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
mythalion-13bQ3_K_S13.0B5.27 GiB53.13 GiB59.34 GiB0.18 GiB18±26.5%
Kimi-Dev-72BIQ4_XS72.7B37.02 GiB21.25 GiB59.30 GiB0.22 GiB18±26.5%
Qwen2.5-VL-72B-InstructIQ4_XS73.4B37.02 GiB21.25 GiB59.30 GiB0.22 GiB18±26.5%
Rombo-LLM-V3.0-Qwen-72bI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Qwen2.5-72B-Instruct-abliteratedI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Qwen2.5-72B-Instruct-abliterated-v2I1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
HuatuoGPT-o1-72BIQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
MiroThinker-v1.0-72BI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
EVA-Qwen2.5-72B-v0.2IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Qwen2.5-Math-72B-InstructIQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Qwen2.5-72B-InstructIQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Malaysian-Qwen2.5-72B-InstructI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Qwen2.5-72BI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
magnum-v4-72bI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
KAT-Dev-72B-ExpIQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Homer-v1.0-Qwen2.5-72BIQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Tower-Plus-72B-ultra-uncensored-hereticI1-IQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Chronos-Platinum-72BIQ4_XS72.7B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
UI-TARS-72B-DPOIQ4_XS73.4B36.98 GiB21.25 GiB59.26 GiB0.26 GiB18±26.5%
Qwen3-Coder-NextMoEUD-Q5_K_S79.7B51.99 GiB6.38 GiB59.25 GiB0.27 GiB54±37%
NSFW_13B_sftQ2_K13.3B5.18 GiB53.13 GiB59.25 GiB0.27 GiB18±26.5%
Apertus-70B-Instruct-2509Q3_K_L70.6B36.87 GiB21.25 GiB59.20 GiB0.32 GiB18±26.5%
CalmeRys-78B-Orpo-v0.1I1-IQ3_M78.0B35.33 GiB22.84 GiB59.20 GiB0.32 GiB18±26.5%
calme-2.3-rys-78bIQ3_M78.0B35.33 GiB22.84 GiB59.20 GiB0.32 GiB18±26.5%
Devstral-2-123B-Instruct-2512IQ2_XS125B34.75 GiB23.38 GiB59.18 GiB0.34 GiB18±26.5%
Mistral-Medium-3.5-128BI1-IQ2_XS128B34.75 GiB23.38 GiB59.18 GiB0.34 GiB18±26.5%
XORTRON-NXTXPRTXXLI1-IQ2_XS128B34.75 GiB23.38 GiB59.18 GiB0.34 GiB18±26.5%
gemma-4-31B-it-Mystery-Fine-Tune-HERETIC-UNCENSORED-ThinkingQ6_K31.3B46.94 GiB11.25 GiB59.17 GiB0.35 GiB18±26.5%
Chuluun-Qwen2.5-72B-v0.01Q3_K_L72.7B36.79 GiB21.25 GiB59.07 GiB0.45 GiB18±26.5%
Qwen3-72B-SynthesisQ3_K_L72.7B36.79 GiB21.25 GiB59.07 GiB0.45 GiB18±26.5%
GLM-4.5-Air-REAP-82B-A12BMoEQ4_081.9B45.78 GiB12.22 GiB58.93 GiB0.59 GiB31±37%
Laguna-S-2.1MoEUD-IQ4_NL118B54.71 GiB3.26 GiB58.90 GiB0.62 GiB64±37%
Qwen3.5-88BMoEI1-Q5_K_S87.7B56.33 GiB1.59 GiB58.85 GiB0.67 GiB74±37%
CodeLlama-70b-Instruct-hfI1-Q4_K_S69.0B36.55 GiB21.25 GiB58.83 GiB0.69 GiB18±26.5%
CodeLlama-70b-Python-hfI1-Q4_K_S69.0B36.55 GiB21.25 GiB58.83 GiB0.69 GiB18±26.5%
Nous-Hermes-Llama2-70bI1-Q4_K_S69.0B36.55 GiB21.25 GiB58.83 GiB0.69 GiB18±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Instinct MI210 run?
2027 of 2118 indexed open-weight models fit a Instinct MI210 at 131,072 context with q8_0 KV cache, the largest being MythoMax-L2-Kimiko-v2-13b at IQ3_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Instinct MI210 actually have?
Its nameplate is 64 GB, but about 59.52 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Instinct MI210 fast for local AI?
Its memory bandwidth is 1638 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.