AMD · workstation

Radeon Pro W6400

Radeon Pro W6400 has 4 GB of VRAM at 128 GB/s — about 3.72 GiB usable after driver and compositor overhead. 484 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
128 GB/s
64-bit bus
Tensor FP16
dense
TDP
50 W
$229 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 367vision language 42audio asr 35audio tts 18embedding 20video 2

What fits at 32K context

largest quantization that fits, per model · 484 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
LFM2-2.6BQ8_02.6B2.55 GiB0.27 GiB3.72 GiB0.00 GiB28±26.5%
LFM2-2.6B-TranscriptQ8_02.6B2.55 GiB0.27 GiB3.72 GiB0.00 GiB28±26.5%
LFM2-VL-3BQ8_03.0B2.55 GiB0.27 GiB3.72 GiB0.00 GiB28±26.5%
Llama-Doctor-3.2-3B-InstructI1-IQ2_XXS3.2B0.95 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
Llama-3.2-3B-Instruct-roleplay-tunedI1-IQ2_XXS3.2B0.95 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
Llama-3.2-3B-Instruct-heretic-ablitered-uncensoredI1-IQ2_XXS3.2B0.95 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
Llama3.2-3B-creative-writer-v0.1I1-IQ2_XXS3.2B0.95 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
Firefly-V3.2I1-IQ2_XXS3.2B0.95 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
Firefly-V3I1-IQ2_XXS3.2B0.95 GiB1.86 GiB3.72 GiB0.00 GiB28±26.5%
G9v3-3BQ4_K_L3.0B1.96 GiB0.86 GiB3.71 GiB0.01 GiB28±26.5%
Holo-3.1-4BI1-IQ3_M5.2B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
AfriqueQwen3.5-4BI1-IQ3_M5.2B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
TimeOmni-1-4BI1-IQ3_M5.2B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
granite-4.0-7B-A1B-Creative-v0.1MoEI1-IQ3_S6.7B2.71 GiB0.13 GiB3.71 GiB0.01 GiB77±37%
SmolLM3-3BIQ4_XS3.1B1.61 GiB1.20 GiB3.71 GiB0.01 GiB29±26.5%
Qwen3.5-4B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGI1-IQ4_XS4.5B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
Qwen3.5-4B-SOMPOA-heresy-v2I1-IQ4_XS4.5B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
Qwen3.5-4B-SOMPOA-heresyI1-IQ4_XS4.5B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
Qwen3.5-4B-Safety-ThinkingI1-IQ4_XS4.2B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
Huihui-Qwen3.5-4B-abliteratedI1-IQ4_XS4.5B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
Darkidol-Ballad-4BI1-IQ4_XS4.5B2.27 GiB0.53 GiB3.71 GiB0.01 GiB29±26.5%
Qwen3-1.7BIQ3_M2.0B0.96 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Nanbeige4.1-3BQ3_K_S3.9B1.73 GiB1.06 GiB3.71 GiB0.01 GiB29±26.5%
OpenClaude-1.7B-MergedIQ4_XS1.7B0.96 GiB1.86 GiB3.71 GiB0.01 GiB28±26.5%
Gemma-3-4b-it-Uncensored-DBL-XI1-IQ4_XS4.7B2.29 GiB0.50 GiB3.71 GiB0.01 GiB29±26.5%
AMD-OLMo-1B-SFT-DPOQ4_K_M1.2B0.68 GiB2.13 GiB3.71 GiB0.01 GiB28±26.5%
granite-speech-4.1-2bQ4_K_M2.3B1.49 GiB1.33 GiB3.71 GiB0.01 GiB28±26.5%
granite-4.0-1b-speechQ4_K_M2.3B1.49 GiB1.33 GiB3.71 GiB0.01 GiB28±26.5%
Darwin-4B-ChimeraI1-IQ4_XS4.0B2.11 GiB0.68 GiB3.70 GiB0.02 GiB29±26.5%
granite-3.1-2b-instructQ4_12.5B1.48 GiB1.33 GiB3.70 GiB0.02 GiB28±26.5%
Unlimited-OCRMoEKV unresolvedQ4_K_M3.3B1.82 GiB1.00 GiB3.70 GiB0.02 GiB36±37%
dolphin-2_6-phi-2Q8_02.8B2.75 GiB0.00 GiB3.70 GiB0.02 GiB29±26.5%
Qwen3-VL-Reranker-2BIQ4_XS2.1B0.95 GiB1.86 GiB3.70 GiB0.02 GiB28±26.5%
Atomight-V2.5-1.7BIQ4_XS1.7B0.95 GiB1.86 GiB3.70 GiB0.02 GiB28±26.5%
OpenCaption-2B-VL-SFT-v1.0IQ4_XS2.1B0.95 GiB1.86 GiB3.70 GiB0.02 GiB28±26.5%
gaon-1.7b-v2-instructIQ4_XS1.7B0.95 GiB1.86 GiB3.70 GiB0.02 GiB28±26.5%
gaon-1.7b-v2-translateIQ4_XS1.7B0.95 GiB1.86 GiB3.70 GiB0.02 GiB28±26.5%
Qwen3.5-4B-NSFW-ARA-Heretic-LiteroticaI1-Q3_K_L4.2B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Qwen3.5-4B-RpRMax-v1I1-Q3_K_L4.7B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Holo-3.1-4B-uncensored-hereticI1-Q3_K_L4.5B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
GRaPE-2-MiniI1-Q3_K_L4.7B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Qwen3.5-DPO-4B-2I1-Q3_K_L4.2B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Qwen3.5-4B-BaseQ3_K_L4.7B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Huihui-Qwen3.5-4B-Claude-4.6-Opus-abliteratedI1-Q3_K_L4.7B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Qwopus3.5-4B-v3-hereticI1-Q3_K_L4.5B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Aureth-4B-Qwen3.5I1-Q3_K_L4.5B2.26 GiB0.53 GiB3.70 GiB0.02 GiB29±26.5%
Parable-Granite-4.1-3B-Claude-Fable-5I1-IQ3_S3.4B1.47 GiB1.33 GiB3.70 GiB0.02 GiB29±26.5%
granite-3.1-3b-a800m-instructMoEQ4_K_S3.3B1.77 GiB1.06 GiB3.70 GiB0.02 GiB30±37%
Llama-3.2-3B-Instruct-abliteratedI1-IQ1_S3.6B0.93 GiB1.86 GiB3.70 GiB0.02 GiB29±26.5%
Supertron2-Reranker-2BI1-IQ4_XS2.1B0.94 GiB1.86 GiB3.69 GiB0.03 GiB29±26.5%
Uni-MuMER-Qwen3-VL-2BI1-IQ4_XS2.1B0.94 GiB1.86 GiB3.69 GiB0.03 GiB29±26.5%
Qwen3-VL-2B-ThinkingIQ4_XS2.1B0.94 GiB1.86 GiB3.69 GiB0.03 GiB29±26.5%
Qwen3-VL-2B-InstructIQ4_XS2.1B0.94 GiB1.86 GiB3.69 GiB0.03 GiB29±26.5%
Lightning-1.7BIQ4_XS1.7B0.94 GiB1.86 GiB3.69 GiB0.03 GiB29±26.5%
DorsetHeatwaveLLM2I1-IQ4_XS1.7B0.94 GiB1.86 GiB3.69 GiB0.03 GiB29±26.5%
gemma-3-4b-it-qatQ4_03.9B2.35 GiB0.42 GiB3.69 GiB0.03 GiB29±26.5%
granite-4.0-microQ3_K_S3.4B1.46 GiB1.33 GiB3.69 GiB0.03 GiB29±26.5%
granite-4.1-3bQ3_K_S3.4B1.46 GiB1.33 GiB3.69 GiB0.03 GiB29±26.5%
granite-4.0-micro-baseQ3_K_S3.4B1.46 GiB1.33 GiB3.69 GiB0.03 GiB29±26.5%
granite-3.3-2b-instructQ4_K_L2.5B1.46 GiB1.33 GiB3.69 GiB0.03 GiB29±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Radeon Pro W6400 run?
484 of 2118 indexed open-weight models fit a Radeon Pro W6400 at 32,768 context with q8_0 KV cache, the largest being LFM2-2.6B at Q8_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon Pro W6400 actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon Pro W6400 fast for local AI?
Its memory bandwidth is 128 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.