AMD · consumer

Radeon RX 7700

Radeon RX 7700 has 16 GB of VRAM at 624 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1670 of 2118 indexed models fit at 128K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
624 GB/s
256-bit bus
Tensor FP16
dense
TDP
263 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1417vision language 151video 15audio asr 39embedding 26audio tts 21image 1

What fits at 128K context

largest quantization that fits, per model · 1670 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Ling-liteMoEQ5_K_L16.8B12.02 GiB1.97 GiB14.88 GiB0.00 GiB57±37%
granite-4.0-h-smallMoEQ3_K_S32.2B13.43 GiB0.56 GiB14.88 GiB0.00 GiB67±37%
NVIDIA-Nemotron-Nano-12B-v2Q3_K_S12.3B5.19 GiB8.72 GiB14.88 GiB0.00 GiB28±26.5%
Huihui-gemma-4-26B-A4B-it-abliteratedMoEUD-IQ4_NL26.5B12.50 GiB1.49 GiB14.87 GiB0.01 GiB28±26.5%
Apriel-1.6-15b-ThinkerI1-Q3_K_L14.9B7.18 GiB6.75 GiB14.87 GiB0.01 GiB28±26.5%
gemma-4-A4B-98e-v7-coder-itMoEQ4_K_L20.5B12.50 GiB1.49 GiB14.87 GiB0.01 GiB28±26.5%
gemma-4-A4B-98e-v6-coder-itMoEQ4_K_L20.5B12.50 GiB1.49 GiB14.87 GiB0.01 GiB28±26.5%
gemma-4-A4B-98e-v7-coderx-itMoEQ4_K_L20.5B12.50 GiB1.49 GiB14.87 GiB0.01 GiB28±26.5%
reka-flash-3.1I1-Q3_K_S20.9B9.25 GiB4.64 GiB14.87 GiB0.01 GiB28±26.5%
reka-flash-3Q3_K_S20.9B9.25 GiB4.64 GiB14.87 GiB0.01 GiB28±26.5%
Snowpiercer-15B-v4-hereticI1-Q3_K_M15.0B6.89 GiB7.03 GiB14.87 GiB0.01 GiB28±26.5%
Snowpiercer-15B-v4Q3_K_M15.0B6.89 GiB7.03 GiB14.87 GiB0.01 GiB28±26.5%
Fallen-Gemma3-27B-v1IQ3_M27.4B11.69 GiB2.23 GiB14.86 GiB0.02 GiB28±26.5%
Parable-Granite-4.1-8B-Claude-Fable-5Q8_08.4B8.30 GiB5.63 GiB14.86 GiB0.02 GiB28±26.5%
Qwen3.6-27B-A3B-CoderMoEI1-IQ4_XS26.7B13.25 GiB0.70 GiB14.85 GiB0.03 GiB92±37%
Nemotron-Mini-4B-InstructQ5_K_M4.2B9.43 GiB4.50 GiB14.85 GiB0.03 GiB28±26.5%
Phi-4-reasoningQ3_K_M14.7B6.86 GiB7.03 GiB14.85 GiB0.03 GiB28±26.5%
Phi-4-reasoning-plusQ3_K_M14.7B6.86 GiB7.03 GiB14.85 GiB0.03 GiB28±26.5%
phi-4Q3_K_M14.7B6.86 GiB7.03 GiB14.85 GiB0.03 GiB28±26.5%
HomunculusQ5_K_M12.5B8.27 GiB5.63 GiB14.85 GiB0.03 GiB28±26.5%
Magistry-24B-v1.1IQ2_M23.6B8.19 GiB5.63 GiB14.84 GiB0.04 GiB28±26.5%
granite-20b-code-instruct-8kQ5_K_L20.1B13.86 GiB0.00 GiB14.84 GiB0.04 GiB28±26.5%
dolphin-2.9.2-Phi-3-MediumKV unresolvedQ3_K_L14.0B6.84 GiB7.03 GiB14.83 GiB0.05 GiB28±26.5%
gemma-2-27b-itIQ2_XXS27.2B7.10 GiB6.70 GiB14.83 GiB0.05 GiB28±26.5%
GLM-4.7-Flash-hereticMoEIQ3_XXS29.9B12.06 GiB1.86 GiB14.83 GiB0.05 GiB64±37%
Qwen3-Coder-30B-A3B-InstructMoEQ2_K_L30.5B10.55 GiB3.38 GiB14.82 GiB0.06 GiB47±37%
Qwen3-VL-30B-A3B-InstructMoEQ2_K_L31.1B10.55 GiB3.38 GiB14.82 GiB0.06 GiB47±37%
Qwen3-VL-30B-A3B-ThinkingMoEQ2_K_L31.1B10.55 GiB3.38 GiB14.82 GiB0.06 GiB47±37%
Qwen3-30B-A3BMoEQ2_K_L30.5B10.55 GiB3.38 GiB14.82 GiB0.06 GiB47±37%
Qwen3-30B-A3B-Instruct-2507MoEQ2_K_L30.5B10.55 GiB3.38 GiB14.82 GiB0.06 GiB47±37%
Qwen3-30B-A3B-Thinking-2507MoEQ2_K_L30.5B10.55 GiB3.38 GiB14.82 GiB0.06 GiB47±37%
Rocinante-XL-16B-v1IQ3_XXS16.1B6.28 GiB7.59 GiB14.82 GiB0.06 GiB28±26.5%
Pantheon-Reasoning-26B-A4B-1.1MoEQ3_K_M26.5B12.42 GiB1.49 GiB14.80 GiB0.08 GiB28±26.5%
ERNIE-4.5-21B-A3B-ThinkingQ4_021.8B11.90 GiB1.97 GiB14.79 GiB0.09 GiB28±26.5%
ERNIE-4.5-21B-A3B-PTQ4_021.9B11.90 GiB1.97 GiB14.79 GiB0.09 GiB28±26.5%
GLM-4.7-FlashMoEUD-IQ3_XXS31.2B12.02 GiB1.86 GiB14.79 GiB0.09 GiB64±37%
Salience-1.5-FlashMoEQ2_K_L31.1B10.51 GiB3.38 GiB14.78 GiB0.10 GiB47±37%
Qwen3.6-27B-Heretic2-ThinkingI1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.6-27B-Uncensored-AggressiveI1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen-3.5-Opus-GLM-27BI1-IQ3_S26.9B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.6-27B-abliteratedI1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
KoQweopus-3.5-27B-experimentalI1-IQ3_S27.8B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Webcoda-AI-27BI1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.5-27B-imabari-v2I1-IQ3_S27.8B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.5-27B-uncensored-heretic-v1I1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Carnice-V2-27bI1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.5-Queen-27BI1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
GRaPE-2-ProI1-IQ3_S27.8B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Darwin-28B-REASONI1-IQ3_S26.9B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliteratedI1-IQ3_S27.8B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.5-27B-WebNovel-Writer-zhI1-IQ3_S26.9B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Qwen3.5-27B_Homebrew-v2I1-IQ3_S27.4B11.57 GiB2.25 GiB14.78 GiB0.10 GiB28±26.5%
Llama-3.2-8X3B-MOE-Dark-Champion-Instruct-uncensored-abliterated-18.4BMoEQ4_K_S18.4B9.93 GiB3.94 GiB14.78 GiB0.10 GiB32±37%
granite-20b-code-base-8kI1-Q5_K_M20.1B13.79 GiB0.00 GiB14.77 GiB0.11 GiB28±26.5%
granite-34b-code-base-8kI1-IQ3_S33.7B13.79 GiB0.00 GiB14.77 GiB0.11 GiB28±26.5%
Ling-mini-2.0MoEQ6_K16.3B12.47 GiB1.41 GiB14.76 GiB0.12 GiB81±37%
SOLAR-10.7B-Instruct-v1.0-uncensoredQ5_K_M10.7B7.08 GiB6.75 GiB14.76 GiB0.12 GiB28±26.5%
Nous-Hermes-2-SOLAR-10.7BQ5_K_M10.7B7.08 GiB6.75 GiB14.76 GiB0.12 GiB28±26.5%
SOLAR-10.7B-Instruct-v1.0I1-Q5_K_M10.7B7.08 GiB6.75 GiB14.76 GiB0.12 GiB28±26.5%
Skywork-R1V3-38BIQ3_M38.4B13.79 GiB0.00 GiB14.76 GiB0.12 GiB28±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Radeon RX 7700 run?
1670 of 2118 indexed open-weight models fit a Radeon RX 7700 at 131,072 context with q4_0 KV cache, the largest being Ling-lite at Q5_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 7700 actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 7700 fast for local AI?
Its memory bandwidth is 624 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.