AMD · workstation

Radeon Pro W6400

Radeon Pro W6400 has 4 GB of VRAM at 128 GB/s — about 3.72 GiB usable after driver and compositor overhead. 269 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
128 GB/s
64-bit bus
Tensor FP16
dense
TDP
50 W
$229 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 189vision language 21audio asr 29embedding 13audio tts 15video 2

What fits at 128K context

largest quantization that fits, per model · 269 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Dolphin3.0-Llama3.2-1BIQ4_XS1.2B0.69 GiB2.13 GiB3.72 GiB0.00 GiB28±26.5%
OneLLM-Doey-ChatQA-V1-Llama-3.2-1BI1-IQ4_XS1.2B0.69 GiB2.13 GiB3.72 GiB0.00 GiB28±26.5%
Llama-3.2-1B-Instruct-hereticI1-IQ4_XS1.2B0.69 GiB2.13 GiB3.72 GiB0.00 GiB28±26.5%
Llama-3.2-1B-InstructIQ4_XS1.2B0.69 GiB2.13 GiB3.72 GiB0.00 GiB28±26.5%
Llama-3.2-1B-Instruct-UncensoredI1-IQ4_XS1.2B0.69 GiB2.13 GiB3.72 GiB0.00 GiB28±26.5%
smollm-360M-instruct-add-basicsIQ2_XXS362M0.19 GiB2.66 GiB3.72 GiB0.00 GiB28±26.5%
Qwen2.5-Coder-1.5B-Instruct-abliteratedIQ4_XS1.8B0.96 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
Qwen2.5-1.5B-Instruct-uncensoredIQ4_XS1.8B0.96 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
VibeThinker-1.5BIQ4_XS1.8B0.96 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
DeepSeek-R1-Distill-Qwen-1.5B-uncensoredIQ4_XS1.8B0.96 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
DeepSeek-R1-Distill-Qwen-1.5B-Fully-UncensoredIQ4_XS1.8B0.96 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
LFM2-1.2BQ4_K_M1.2B0.69 GiB2.13 GiB3.71 GiB0.01 GiB28±26.5%
DeepScaleR-1.5B-PreviewIQ4_XS1.8B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Nemotron-Research-Reasoning-Qwen-1.5BIQ4_XS1.8B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
DeepSeek-R1-Distill-Qwen-1.5BIQ4_XS1.8B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
dots.ocrI1-IQ4_XS3.0B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
SEX_ROLEPLAY-3.2-1BI1-IQ3_XS1.5B0.68 GiB2.13 GiB3.71 GiB0.01 GiB28±26.5%
Llama-3.2-1B-Instruct-abliteratedI1-IQ3_XS1.5B0.68 GiB2.13 GiB3.71 GiB0.01 GiB28±26.5%
Novaciano-3.2-1BI1-IQ3_XS1.5B0.68 GiB2.13 GiB3.71 GiB0.01 GiB28±26.5%
Imp-RPG.System-1BI1-IQ3_XS1.5B0.68 GiB2.13 GiB3.71 GiB0.01 GiB28±26.5%
LFM2-2.6BQ5_K_M2.6B1.74 GiB1.06 GiB3.71 GiB0.01 GiB29±26.5%
Dolphin3.0-Qwen2.5-1.5BQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Qwen2.5-1.5B-VibeThinker-heretic-uncensored-abliteratedI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-Coder-OBLITERATEDI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-Coder-AbliteratedI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Qwen2.5-1.5B-hereticI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Qwen2.5-Coder-1.5B-Unsensored-DPOI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
ShellWhisperer-1.5BQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Qwen2.5-1.5BQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
PiCo-1BI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Fourier-Qwen2-VL-2B-0.67I1-Q4_12.2B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
FableForge-1.5BI1-Q4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-MedicalQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-ScienceQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-LegalQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-CoderQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-FinanceQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
NEXUS-SecurityQ4_11.5B0.95 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
gemma-4-E2B-it-abliteratedI1-IQ2_S5.1B2.33 GiB0.48 GiB3.70 GiB0.02 GiB28±26.5%
gemma-4-E2B-it-qat-q4_0-unquantized-hereticI1-IQ2_S5.1B2.33 GiB0.48 GiB3.70 GiB0.02 GiB28±26.5%
Huihui-gemma-4-E2B-it-qat-q4_0-unquantized-abliteratedI1-IQ2_S5.1B2.33 GiB0.48 GiB3.70 GiB0.02 GiB28±26.5%
Gemma4_E2B_Abliterated_Baked_HF_ReadyI1-IQ2_S5.1B2.33 GiB0.48 GiB3.70 GiB0.02 GiB28±26.5%
dolphin-2_6-phi-2Q8_02.8B2.75 GiB0.00 GiB3.70 GiB0.02 GiB29±26.5%
alduin-4b-it-baseI1-IQ2_S4.3B1.36 GiB1.42 GiB3.69 GiB0.03 GiB29±26.5%
gemma-2bQ4_12.5B1.56 GiB1.20 GiB3.69 GiB0.03 GiB29±26.5%
moonshine-streaming-mediumF16266M0.50 GiB2.32 GiB3.69 GiB0.03 GiB28±26.5%
gemma-3-4b-it-roleplay-tuned-v1I1-IQ2_S4.3B1.35 GiB1.42 GiB3.68 GiB0.04 GiB29±26.5%
gemma-3-4b-it-roleplay-tuned-v2I1-IQ2_S4.3B1.35 GiB1.42 GiB3.68 GiB0.04 GiB29±26.5%
gemma-3-4b-it-heretic-uncensored-abliterated-ExtremeI1-IQ2_S4.3B1.35 GiB1.42 GiB3.68 GiB0.04 GiB29±26.5%
Gemma3-4B-CodeCenturionI1-IQ2_S4.3B1.35 GiB1.42 GiB3.68 GiB0.04 GiB29±26.5%
ArrowMint-Gemma3-4B-YUKI-v0.1I1-IQ2_S4.3B1.35 GiB1.42 GiB3.68 GiB0.04 GiB29±26.5%
granite-4.0-h-microQ5_13.2B2.25 GiB0.53 GiB3.68 GiB0.04 GiB29±26.5%
granite-4.0-h-3b-arI1-Q5_K_M3.4B2.25 GiB0.53 GiB3.68 GiB0.04 GiB29±26.5%
granite-4.0-7B-A1B-Creative-v0.1MoEI1-Q2_K6.7B2.28 GiB0.53 GiB3.68 GiB0.04 GiB50±37%
Qwen2.5-1.5B-Instruct-abliteratedI1-Q4_K_M1.5B0.92 GiB1.86 GiB3.68 GiB0.04 GiB29±26.5%
Qwen2.5-Math-1.5B-InstructQ4_K_M1.5B0.92 GiB1.86 GiB3.68 GiB0.04 GiB29±26.5%
Qwen2-VL-2B-InstructQ4_K_M2.2B0.92 GiB1.86 GiB3.68 GiB0.04 GiB29±26.5%
Qwen2-1.5B-InstructQ4_K_M1.5B0.92 GiB1.86 GiB3.68 GiB0.04 GiB29±26.5%
Qwen2-1.5BQ4_K1.5B0.92 GiB1.86 GiB3.68 GiB0.04 GiB29±26.5%
Qwen2.5-Coder-1.5BQ4_K_M1.5B0.92 GiB1.86 GiB3.68 GiB0.04 GiB29±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Radeon Pro W6400 run?
269 of 2118 indexed open-weight models fit a Radeon Pro W6400 at 131,072 context with q8_0 KV cache, the largest being Dolphin3.0-Llama3.2-1B at IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon Pro W6400 actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon Pro W6400 fast for local AI?
Its memory bandwidth is 128 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.