AMD · consumer

Radeon RX 6700

Radeon RX 6700 has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1527 of 2118 indexed models fit at 32K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
175 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1314vision language 113image 2audio asr 39video 12audio tts 21embedding 26

What fits at 32K context

largest quantization that fits, per model · 1527 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
gemma-4-12B-it-uncensored-hereticQ4_K_S12.0B7.66 GiB0.69 GiB9.30 GiB0.00 GiB24±26.5%
gemma-7bI1-Q3_K_L8.5B4.39 GiB3.94 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Nexa-AI-4x4B-InstructMoEI1-Q4_112.1B7.12 GiB1.27 GiB9.29 GiB0.01 GiB23±37%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ3_XXS21.8B7.88 GiB0.49 GiB9.29 GiB0.01 GiB24±26.5%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ3_XXS21.8B7.88 GiB0.49 GiB9.29 GiB0.01 GiB24±26.5%
ERNIE-4.5-21B-A3B-ThinkingI1-IQ3_XXS21.8B7.88 GiB0.49 GiB9.29 GiB0.01 GiB24±26.5%
Ministral-3-14B-Instruct-2512IQ4_XS13.9B6.92 GiB1.41 GiB9.28 GiB0.02 GiB24±26.5%
Ministral-3-14B-Reasoning-2512IQ4_XS13.9B6.92 GiB1.41 GiB9.28 GiB0.02 GiB24±26.5%
gemma-4-12B-it-hereticQ5_K_S12.0B7.64 GiB0.69 GiB9.28 GiB0.02 GiB24±26.5%
Apriel-1.6-15b-ThinkerI1-Q3_K_M14.9B6.65 GiB1.69 GiB9.28 GiB0.02 GiB24±26.5%
Qwen3-30B-A3B-Instruct-2507MoEUD-TQ1_030.5B7.54 GiB0.84 GiB9.28 GiB0.02 GiB62±37%
GLM-4.6V-FlashQ6_K_L10.3B7.98 GiB0.35 GiB9.27 GiB0.03 GiB24±26.5%
GLM-Z1-9B-0414Q6_K_L9.4B7.98 GiB0.35 GiB9.27 GiB0.03 GiB24±26.5%
GLM-4-9B-0414Q6_K_L9.4B7.98 GiB0.35 GiB9.27 GiB0.03 GiB24±26.5%
Goetia-26B-A4B-v1.4MoEI1-IQ1_S26.0B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
G4-Moonlight-Dusk-26B-A4BMoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
Chimera-X-26B-A4BMoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
Gemma-4-26B-A4B-StyleTuneMoEI1-IQ1_S26.5B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-IQ1_S25.8B7.95 GiB0.43 GiB9.27 GiB0.03 GiB24±26.5%
Tiger-Gemma-12B-v3Q4_112.8B7.63 GiB0.69 GiB9.27 GiB0.03 GiB24±26.5%
AfriqueGemma-12BI1-Q4_112.2B7.63 GiB0.69 GiB9.27 GiB0.03 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-IQ4_XS13.9B6.90 GiB1.41 GiB9.26 GiB0.04 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16IQ4_XS13.9B6.90 GiB1.41 GiB9.26 GiB0.04 GiB24±26.5%
Ministral-3-14B-Reasoning-2512-UncensoredI1-IQ4_XS13.9B6.90 GiB1.41 GiB9.26 GiB0.04 GiB24±26.5%
FrickFritz-4BF164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
Newton-bot-3-VLM-mini-4BF164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
qwen3.5-4b-agentic-coder-v4F164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
Myth-4BF164.3B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
Qwen3.5-4B-UncensoredF164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
JOSIE-2-4B-PreviewF164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
Qwen3.5-4BBF164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
Surogate-3.5-4BF165.3B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
Qwopus3.5-4B-v3BF164.7B8.07 GiB0.28 GiB9.26 GiB0.04 GiB24±26.5%
LFM2-8B-A1BMoEQ8_08.3B8.26 GiB0.11 GiB9.26 GiB0.04 GiB68±37%
Qwen3-VL-8B-Instruct-HereticI1-IQ3_S8.8B7.06 GiB1.27 GiB9.26 GiB0.04 GiB24±26.5%
Qwen3-Coder-30B-A3B-InstructMoEUD-TQ1_030.5B7.52 GiB0.84 GiB9.26 GiB0.04 GiB62±37%
Janus-Pro-7BI1-Q4_17.4B4.10 GiB4.22 GiB9.25 GiB0.05 GiB24±26.5%
deepseek-math-7b-instructQ4_16.9B4.10 GiB4.22 GiB9.25 GiB0.05 GiB24±26.5%
gemma-4-12B-it-qat-q4_0-unquantized-uncensored-hereticNVFP412.0B7.60 GiB0.69 GiB9.24 GiB0.06 GiB24±26.5%
Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-IQ1_M30.0B6.68 GiB1.65 GiB9.24 GiB0.06 GiB39±37%
starcoder2-15bKV unresolvedQ3_K_M16.0B7.54 GiB0.70 GiB9.24 GiB0.06 GiB24±26.5%
GLM-Z1-32B-0414UD-IQ1_M32.6B7.71 GiB0.54 GiB9.24 GiB0.06 GiB24±26.5%
GLM-4-32B-0414UD-IQ1_M32.6B7.71 GiB0.54 GiB9.24 GiB0.06 GiB24±26.5%
Rocinante-XL-16B-v1I1-IQ3_XS16.1B6.39 GiB1.90 GiB9.24 GiB0.06 GiB24±26.5%
Qwen3.6-28BMoEI1-IQ2_S28.2B8.16 GiB0.18 GiB9.24 GiB0.06 GiB104±37%
Qwen3.5-28BMoEI1-IQ2_S28.7B8.16 GiB0.18 GiB9.24 GiB0.06 GiB104±37%
dolphin-2.9.3-mistral-7B-32kQ8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
Mistral-7B-v0.3Q8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
Mistral-7B-Instruct-v0.3-ParasiteQ8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
Mistral-7B-Instruct-v0.3-JbliteratedQ8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
Mistral-7B-Instruct-v0.3Q8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
Mistral-7B-v0.3-Chinese-ChatQ8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
mistral-7b-v0.3-bnb-4bitQ8_07.5B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
Mathstral-7B-v0.1Q8_07.2B7.17 GiB1.13 GiB9.24 GiB0.06 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.09 it/s1.923.5627
Benchmarked· n=27

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6700 run?
1527 of 2118 indexed open-weight models fit a Radeon RX 6700 at 32,768 context with q4_0 KV cache, the largest being gemma-4-12B-it-uncensored-heretic at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6700 actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6700 fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.