AMD · workstation
Radeon Pro W7700
Radeon Pro W7700 has 16 GB of VRAM at 576 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1670 of 2118 indexed models fit at 128K context with q4_0 KV.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
576 GB/s
256-bit bus
Tensor FP16
—
dense
TDP
190 W
$999 MSRP
text 1417vision language 151video 15audio asr 39embedding 26audio tts 21image 1
What fits at 128K context
largest quantization that fits, per model · 1670 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Ling-liteMoE | Q5_K_L | 16.8B | 12.02 GiB | 1.97 GiB | 14.88 GiB | 0.00 GiB | 53±37% |
| granite-4.0-h-smallMoE | Q3_K_S | 32.2B | 13.43 GiB | 0.56 GiB | 14.88 GiB | 0.00 GiB | 62±37% |
| NVIDIA-Nemotron-Nano-12B-v2 | Q3_K_S | 12.3B | 5.19 GiB | 8.72 GiB | 14.88 GiB | 0.00 GiB | 26±26.5% |
| Huihui-gemma-4-26B-A4B-it-abliteratedMoE | UD-IQ4_NL | 26.5B | 12.50 GiB | 1.49 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| Apriel-1.6-15b-Thinker | I1-Q3_K_L | 14.9B | 7.18 GiB | 6.75 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| gemma-4-A4B-98e-v7-coder-itMoE | Q4_K_L | 20.5B | 12.50 GiB | 1.49 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| gemma-4-A4B-98e-v6-coder-itMoE | Q4_K_L | 20.5B | 12.50 GiB | 1.49 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| gemma-4-A4B-98e-v7-coderx-itMoE | Q4_K_L | 20.5B | 12.50 GiB | 1.49 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| reka-flash-3.1 | I1-Q3_K_S | 20.9B | 9.25 GiB | 4.64 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| reka-flash-3 | Q3_K_S | 20.9B | 9.25 GiB | 4.64 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| Snowpiercer-15B-v4-heretic | I1-Q3_K_M | 15.0B | 6.89 GiB | 7.03 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| Snowpiercer-15B-v4 | Q3_K_M | 15.0B | 6.89 GiB | 7.03 GiB | 14.87 GiB | 0.01 GiB | 26±26.5% |
| Fallen-Gemma3-27B-v1 | IQ3_M | 27.4B | 11.69 GiB | 2.23 GiB | 14.86 GiB | 0.02 GiB | 26±26.5% |
| Parable-Granite-4.1-8B-Claude-Fable-5 | Q8_0 | 8.4B | 8.30 GiB | 5.63 GiB | 14.86 GiB | 0.02 GiB | 26±26.5% |
| Qwen3.6-27B-A3B-CoderMoE | I1-IQ4_XS | 26.7B | 13.25 GiB | 0.70 GiB | 14.85 GiB | 0.03 GiB | 86±37% |
| Nemotron-Mini-4B-Instruct | Q5_K_M | 4.2B | 9.43 GiB | 4.50 GiB | 14.85 GiB | 0.03 GiB | 26±26.5% |
| Phi-4-reasoning | Q3_K_M | 14.7B | 6.86 GiB | 7.03 GiB | 14.85 GiB | 0.03 GiB | 26±26.5% |
| Phi-4-reasoning-plus | Q3_K_M | 14.7B | 6.86 GiB | 7.03 GiB | 14.85 GiB | 0.03 GiB | 26±26.5% |
| phi-4 | Q3_K_M | 14.7B | 6.86 GiB | 7.03 GiB | 14.85 GiB | 0.03 GiB | 26±26.5% |
| Homunculus | Q5_K_M | 12.5B | 8.27 GiB | 5.63 GiB | 14.85 GiB | 0.03 GiB | 26±26.5% |
| Magistry-24B-v1.1 | IQ2_M | 23.6B | 8.19 GiB | 5.63 GiB | 14.84 GiB | 0.04 GiB | 26±26.5% |
| granite-20b-code-instruct-8k | Q5_K_L | 20.1B | 13.86 GiB | 0.00 GiB | 14.84 GiB | 0.04 GiB | 26±26.5% |
| dolphin-2.9.2-Phi-3-MediumKV unresolved | Q3_K_L | 14.0B | 6.84 GiB | 7.03 GiB | 14.83 GiB | 0.05 GiB | 26±26.5% |
| gemma-2-27b-it | IQ2_XXS | 27.2B | 7.10 GiB | 6.70 GiB | 14.83 GiB | 0.05 GiB | 26±26.5% |
| GLM-4.7-Flash-hereticMoE | IQ3_XXS | 29.9B | 12.06 GiB | 1.86 GiB | 14.83 GiB | 0.05 GiB | 59±37% |
| Qwen3-Coder-30B-A3B-InstructMoE | Q2_K_L | 30.5B | 10.55 GiB | 3.38 GiB | 14.82 GiB | 0.06 GiB | 43±37% |
| Qwen3-VL-30B-A3B-InstructMoE | Q2_K_L | 31.1B | 10.55 GiB | 3.38 GiB | 14.82 GiB | 0.06 GiB | 43±37% |
| Qwen3-VL-30B-A3B-ThinkingMoE | Q2_K_L | 31.1B | 10.55 GiB | 3.38 GiB | 14.82 GiB | 0.06 GiB | 43±37% |
| Qwen3-30B-A3BMoE | Q2_K_L | 30.5B | 10.55 GiB | 3.38 GiB | 14.82 GiB | 0.06 GiB | 43±37% |
| Qwen3-30B-A3B-Instruct-2507MoE | Q2_K_L | 30.5B | 10.55 GiB | 3.38 GiB | 14.82 GiB | 0.06 GiB | 43±37% |
| Qwen3-30B-A3B-Thinking-2507MoE | Q2_K_L | 30.5B | 10.55 GiB | 3.38 GiB | 14.82 GiB | 0.06 GiB | 43±37% |
| Rocinante-XL-16B-v1 | IQ3_XXS | 16.1B | 6.28 GiB | 7.59 GiB | 14.82 GiB | 0.06 GiB | 26±26.5% |
| Pantheon-Reasoning-26B-A4B-1.1MoE | Q3_K_M | 26.5B | 12.42 GiB | 1.49 GiB | 14.80 GiB | 0.08 GiB | 26±26.5% |
| ERNIE-4.5-21B-A3B-Thinking | Q4_0 | 21.8B | 11.90 GiB | 1.97 GiB | 14.79 GiB | 0.09 GiB | 26±26.5% |
| ERNIE-4.5-21B-A3B-PT | Q4_0 | 21.9B | 11.90 GiB | 1.97 GiB | 14.79 GiB | 0.09 GiB | 26±26.5% |
| GLM-4.7-FlashMoE | UD-IQ3_XXS | 31.2B | 12.02 GiB | 1.86 GiB | 14.79 GiB | 0.09 GiB | 59±37% |
| Salience-1.5-FlashMoE | Q2_K_L | 31.1B | 10.51 GiB | 3.38 GiB | 14.78 GiB | 0.10 GiB | 43±37% |
| Qwen3.6-27B-Heretic2-Thinking | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.6-27B-Uncensored-Aggressive | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen-3.5-Opus-GLM-27B | I1-IQ3_S | 26.9B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.6-27B-abliterated | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| KoQweopus-3.5-27B-experimental | I1-IQ3_S | 27.8B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Webcoda-AI-27B | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.5-27B-imabari-v2 | I1-IQ3_S | 27.8B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.5-27B-uncensored-heretic-v1 | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Carnice-V2-27b | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.5-Queen-27B | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| GRaPE-2-Pro | I1-IQ3_S | 27.8B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Darwin-28B-REASON | I1-IQ3_S | 26.9B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated | I1-IQ3_S | 27.8B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.5-27B-WebNovel-Writer-zh | I1-IQ3_S | 26.9B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Qwen3.5-27B_Homebrew-v2 | I1-IQ3_S | 27.4B | 11.57 GiB | 2.25 GiB | 14.78 GiB | 0.10 GiB | 26±26.5% |
| Llama-3.2-8X3B-MOE-Dark-Champion-Instruct-uncensored-abliterated-18.4BMoE | Q4_K_S | 18.4B | 9.93 GiB | 3.94 GiB | 14.78 GiB | 0.10 GiB | 30±37% |
| granite-20b-code-base-8k | I1-Q5_K_M | 20.1B | 13.79 GiB | 0.00 GiB | 14.77 GiB | 0.11 GiB | 26±26.5% |
| granite-34b-code-base-8k | I1-IQ3_S | 33.7B | 13.79 GiB | 0.00 GiB | 14.77 GiB | 0.11 GiB | 26±26.5% |
| Ling-mini-2.0MoE | Q6_K | 16.3B | 12.47 GiB | 1.41 GiB | 14.76 GiB | 0.12 GiB | 75±37% |
| SOLAR-10.7B-Instruct-v1.0-uncensored | Q5_K_M | 10.7B | 7.08 GiB | 6.75 GiB | 14.76 GiB | 0.12 GiB | 26±26.5% |
| Nous-Hermes-2-SOLAR-10.7B | Q5_K_M | 10.7B | 7.08 GiB | 6.75 GiB | 14.76 GiB | 0.12 GiB | 26±26.5% |
| SOLAR-10.7B-Instruct-v1.0 | I1-Q5_K_M | 10.7B | 7.08 GiB | 6.75 GiB | 14.76 GiB | 0.12 GiB | 26±26.5% |
| Skywork-R1V3-38B | IQ3_M | 38.4B | 13.79 GiB | 0.00 GiB | 14.76 GiB | 0.12 GiB | 26±26.5% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a Radeon Pro W7700 run?
- 1670 of 2118 indexed open-weight models fit a Radeon Pro W7700 at 131,072 context with q4_0 KV cache, the largest being Ling-lite at Q5_K_L. That covers text, vision-language, image, video and speech models.
- How much usable memory does a Radeon Pro W7700 actually have?
- Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a Radeon Pro W7700 fast for local AI?
- Its memory bandwidth is 576 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.