AMD · datacenter

Instinct MI210

Instinct MI210 has 64 GB of VRAM at 1638 GB/s — about 59.52 GiB usable after driver and compositor overhead. 2061 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
64 GB
HBM2e
Bandwidth
1638 GB/s
4096-bit bus
Tensor FP16
181 TF
dense
TDP
300 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1772vision language 185image 2audio asr 39audio tts 21video 16embedding 26

What fits at 16K context

largest quantization that fits, per model · 2061 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
gpt-oss-120b-uncensored-bf16MoEQ3_K_L117B58.30 GiB0.31 GiB59.50 GiB0.02 GiB95±37%
Qwen3.5-REAP-212B-A17BMoEIQ2_XS212B58.30 GiB0.25 GiB59.50 GiB0.02 GiB91±37%
MiniMax-M2.7MoEUD-IQ1_M229B56.53 GiB2.06 GiB59.48 GiB0.04 GiB76±37%
dots.llm1.instMoEUD-IQ2_XXS143B50.31 GiB8.23 GiB59.48 GiB0.04 GiB42±37%
gpt-oss-120bMoEQ2_K120B58.27 GiB0.31 GiB59.47 GiB0.05 GiB95±37%
gpt-oss-safeguard-120bMoEQ2_K120B58.27 GiB0.31 GiB59.47 GiB0.05 GiB95±37%
Mistral-Medium-3.5-128BIQ3_M128B55.44 GiB2.92 GiB59.42 GiB0.10 GiB18±26.5%
command-a-plus-05-2026-bf16MoEIQ2_XXS219B57.98 GiB0.49 GiB59.37 GiB0.15 GiB73±37%
Hypernova-60B-2605MoEQ8_058.7B58.13 GiB0.28 GiB59.30 GiB0.22 GiB82±37%
GLM-4.6VMoEIQ4_NL108B56.83 GiB1.53 GiB59.28 GiB0.24 GiB66±37%
Qwen3.5-88BMoEI1-Q5_K_M87.7B58.10 GiB0.20 GiB59.23 GiB0.29 GiB87±37%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEUD-Q3_K_M124B57.47 GiB0.73 GiB59.10 GiB0.42 GiB82±37%
GLM-4.5-Air-DerestrictedMoEIQ4_XS110B56.63 GiB1.53 GiB59.09 GiB0.43 GiB66±37%
GLM-4.5-AirMoEIQ4_XS110B56.63 GiB1.53 GiB59.09 GiB0.43 GiB66±37%
Behemoth-X-123B-v2Q3_K_M123B55.04 GiB2.92 GiB59.02 GiB0.50 GiB18±26.5%
Mistral-Large-Instruct-2411Q3_K_M123B55.04 GiB2.92 GiB59.02 GiB0.50 GiB18±26.5%
Qwen3.5-99BMoEI1-Q4_199.0B57.88 GiB0.20 GiB59.00 GiB0.52 GiB91±37%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ3_S139B56.04 GiB2.06 GiB58.98 GiB0.54 GiB67±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ3_S139B56.04 GiB2.06 GiB58.98 GiB0.54 GiB67±37%
MiniMax-M2.7-BF16-ultra-uncensored-hereticMoEI1-IQ2_XXS229B55.99 GiB2.06 GiB58.93 GiB0.59 GiB76±37%
MiniMax-M2.1MoEI1-IQ2_XXS229B55.99 GiB2.06 GiB58.93 GiB0.59 GiB76±37%
MiniMax-M2.5MoEI1-IQ2_XXS229B55.99 GiB2.06 GiB58.93 GiB0.59 GiB76±37%
GLM-4.5VMoEI1-Q4_0108B56.39 GiB1.53 GiB58.84 GiB0.68 GiB67±37%
Solar-Open2-250BMoEIQ1_M250B56.30 GiB1.59 GiB58.82 GiB0.70 GiB84±37%
Qwen3.5-122B-A10BMoEUD-IQ4_XS125B57.67 GiB0.20 GiB58.80 GiB0.72 GiB97±37%
GLM-4.5-Air-REAP-82B-A12BMoEQ5_K_S81.9B56.33 GiB1.53 GiB58.79 GiB0.73 GiB59±37%
DeepSeek-Coder-V2-Instruct-0724MoEIQ2_XXS236B57.28 GiB0.56 GiB58.78 GiB0.74 GiB88±37%
DeepSeek-V2.5MoEIQ2_XXS236B57.28 GiB0.56 GiB58.78 GiB0.74 GiB88±37%
DeepSeek-Coder-V2-InstructMoEIQ2_XXS236B57.28 GiB0.56 GiB58.78 GiB0.74 GiB88±37%
c4ai-command-r-plus-08-2024Q4_K_S104B55.55 GiB2.13 GiB58.75 GiB0.77 GiB18±26.5%
Qwen3-Coder-30B-A3B-InstructMoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Salience-1.5-FlashMoEBF1631.1B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Qwen3-VL-30B-A3B-InstructMoEBF1631.1B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Qwen3-VL-30B-A3B-ThinkingMoEBF1631.1B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
MiroThinker-v1.0-30BMoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Qwen3-30B-A3BMoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Qwen3-30B-A3B-Thinking-2507-Claude-4.5-Sonnet-High-Reasoning-DistillMoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Qwen3-30B-A3B-Instruct-2507MoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Qwen3-30B-A3B-Thinking-2507MoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Pantheon-Proto-RP-1.8-30B-A3BMoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Tongyi-DeepResearch-30B-A3BMoEBF1630.5B56.90 GiB0.80 GiB58.59 GiB0.93 GiB72±37%
Mistral-MOE-4X7B-Dark-MultiVerse-Uncensored-Enhanced32-24BMoEQ6_K24.2B56.49 GiB1.06 GiB58.49 GiB1.03 GiB10±37%
Step-3.5-Flash-REAP-121B-A11BI1-Q3_K_M121B53.75 GiB3.74 GiB58.42 GiB1.10 GiB18±26.5%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedIQ4_XS109B55.78 GiB1.59 GiB58.30 GiB1.22 GiB67±37%
CalmeRys-78B-Orpo-v0.1Q5_K_M78.0B54.30 GiB2.86 GiB58.19 GiB1.33 GiB18±26.5%
calme-2.3-rys-78bQ5_K_M78.0B54.30 GiB2.86 GiB58.19 GiB1.33 GiB18±26.5%
North-Mini-Code-1.0MoEBF1630.5B56.81 GiB0.38 GiB58.08 GiB1.44 GiB76±37%
Llama-3.3-70B-InstructQ6_K_L70.6B54.39 GiB2.66 GiB58.07 GiB1.45 GiB18±26.5%
Anubis-70B-v1.2Q6_K_L70.6B54.39 GiB2.66 GiB58.07 GiB1.45 GiB18±26.5%
Tess-R1-Limerick-Llama-3.1-70BQ6_K_L70.6B54.39 GiB2.66 GiB58.07 GiB1.45 GiB18±26.5%
Infinity-Instruct-7M-Gen-Llama3_1-70BQ6_K_L70.6B54.39 GiB2.66 GiB58.07 GiB1.45 GiB18±26.5%
Athene-70BQ6_K_L70.6B54.39 GiB2.66 GiB58.07 GiB1.45 GiB18±26.5%
openPangu-2.0-FlashMoEKV unresolvedQ4_K_M100B56.71 GiB0.43 GiB58.05 GiB1.47 GiB95±37%
Qwen3-Omni-30B-A3B-InstructBF1635.3B56.90 GiB0.00 GiB57.85 GiB1.67 GiB18±26.5%
Qwen3-Omni-30B-A3B-ThinkingBF1631.7B56.90 GiB0.00 GiB57.85 GiB1.67 GiB18±26.5%
InternVL3_5-30B-A3BBF1630.8B56.90 GiB0.00 GiB57.84 GiB1.68 GiB18±26.5%
Hy-MT2-30B-A3BMoEBF1630.1B56.03 GiB0.80 GiB57.72 GiB1.80 GiB73±37%
Apertus-70B-Instruct-2509Q6_K70.6B53.95 GiB2.66 GiB57.68 GiB1.84 GiB18±26.5%
Meta-Llama-3-70B-InstructQ6_K70.6B53.92 GiB2.66 GiB57.60 GiB1.92 GiB18±26.5%
calme-2.4-llama3-70bQ6_K70.6B53.91 GiB2.66 GiB57.59 GiB1.93 GiB18±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Instinct MI210 run?
2061 of 2118 indexed open-weight models fit a Instinct MI210 at 16,384 context with q8_0 KV cache, the largest being gpt-oss-120b-uncensored-bf16 at Q3_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a Instinct MI210 actually have?
Its nameplate is 64 GB, but about 59.52 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Instinct MI210 fast for local AI?
Its memory bandwidth is 1638 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.