AMD · consumer

Radeon RX 6700

Radeon RX 6700 has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1424 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
175 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1216vision language 110embedding 26video 12audio asr 38audio tts 21image 1

What fits at 32K context

largest quantization that fits, per model · 1424 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
granite-3.1-8b-instructQ5_18.2B5.71 GiB2.66 GiB9.30 GiB0.00 GiB24±26.5%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-Q4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-Q4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-Q4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
Floppa-12B-Gemma3-UncensoredI1-Q4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
gemma-3-12b-it-hereticI1-Q4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
gemma-3-12b-it-abliteratedQ4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
gemma-3-12b-itQ4_112.2B7.04 GiB1.31 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-IQ3_S13.9B5.68 GiB2.66 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Reasoning-2512-UncensoredI1-IQ3_S13.9B5.68 GiB2.66 GiB9.30 GiB0.00 GiB24±26.5%
Ling-liteMoEQ3_K_S16.8B7.47 GiB0.93 GiB9.30 GiB0.00 GiB54±37%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q3_K_L15.7B7.88 GiB0.50 GiB9.29 GiB0.01 GiB64±37%
DeepSeek-Coder-V2-Lite-InstructMoEQ3_K_L15.7B7.88 GiB0.50 GiB9.29 GiB0.01 GiB64±37%
DeepSeek-V2-Lite-ChatMoEQ3_K_L15.7B7.88 GiB0.50 GiB9.29 GiB0.01 GiB64±37%
DeepSeek-V2-Lite-Chat-Uncensored-Unbiased-ReasonerMoEQ3_K_L15.7B7.88 GiB0.50 GiB9.29 GiB0.01 GiB64±37%
Qwen3.5-4B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGF164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Agents-A1-4B-Heretic-ARA-Refusals8BF164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Agents-A1-4BF164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3.5-4B-NSFW-ARA-Heretic-LiteroticaF164.2B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3.5-4B-SOMPOA-heresy-v2F164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Huihui-Qwen3.5-4B-abliteratedF164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3.5-4B-Safety-ThinkingF164.2B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
GRaPE-2-MiniF164.7B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3.5-4B-BaseF164.7B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Fara1.5-4BBF164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3.5-4BBF164.7B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
AREX-TurboBF164.5B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
Huihui-Qwen3.5-4B-Claude-4.6-Opus-abliteratedF164.7B7.85 GiB0.53 GiB9.29 GiB0.01 GiB24±26.5%
DeepSeek-V2-Lite-Chat-UncensoredMoEQ3_K_L15.7B7.87 GiB0.50 GiB9.29 GiB0.01 GiB64±37%
Qwen3-15B-A2B-BaseMoEQ3_K_L15.6B7.59 GiB0.80 GiB9.29 GiB0.01 GiB63±37%
NVIDIA-Nemotron-Nano-9B-v2IQ2_S8.9B4.62 GiB3.72 GiB9.28 GiB0.02 GiB24±26.5%
Nemotron-3-Embed-8B-BF16Q6_K8.0B6.08 GiB2.26 GiB9.28 GiB0.02 GiB24±26.5%
Gemma-4-12B-StyleTuneI1-IQ4_NL13.0B7.02 GiB1.31 GiB9.28 GiB0.02 GiB24±26.5%
gemma-4-12b-heretic-styletune-headI1-IQ4_NL12.0B7.02 GiB1.31 GiB9.28 GiB0.02 GiB24±26.5%
syrian-gemma-12bI1-IQ4_NL13.0B7.02 GiB1.31 GiB9.28 GiB0.02 GiB24±26.5%
Devstral-Small-2-24B-Instruct-2512UD-IQ1_M24.0B5.60 GiB2.66 GiB9.28 GiB0.02 GiB24±26.5%
Mistral-Small-3.2-24B-Instruct-2506UD-IQ1_M24.0B5.60 GiB2.66 GiB9.28 GiB0.02 GiB24±26.5%
Devstral-Small-2507UD-IQ1_M23.6B5.60 GiB2.66 GiB9.28 GiB0.02 GiB24±26.5%
Devstral-Small-2505UD-IQ1_M23.6B5.60 GiB2.66 GiB9.28 GiB0.02 GiB24±26.5%
Magistral-Small-2507UD-IQ1_M23.6B5.60 GiB2.66 GiB9.28 GiB0.02 GiB24±26.5%
Mistral-Small-3.1-24B-Instruct-2503UD-IQ1_M24.0B5.60 GiB2.66 GiB9.28 GiB0.02 GiB24±26.5%
granite-4.1-8bQ5_K_S8.8B5.68 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3.6-28BMoEI1-IQ2_XS28.2B8.04 GiB0.33 GiB9.27 GiB0.03 GiB92±37%
Qwen3.5-28BMoEI1-IQ2_XS28.7B8.04 GiB0.33 GiB9.27 GiB0.03 GiB92±37%
Ministral-3-14B-abliteratedQ3_K_S13.9B5.66 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16Q3_K_S13.9B5.66 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Ministral-3-14B-Instruct-2512Q3_K_S13.9B5.66 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Ministral-3-14B-Reasoning-2512Q3_K_S13.9B5.66 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Forsaken-Void-12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Silver-Siren-ST-12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Tess-3-Mistral-Nemo-12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
KrakenSakura-Maelstrom-12B-v1Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
MN-12B-Runeweaver-RP-RUI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Dans-PersonalityEngine-V1.3.0-12bI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Impish_Bloodmoon_12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Wayfarer-2-12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Wayfarer-12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
Muse-12BI1-Q3_K_M12.2B5.67 GiB2.66 GiB9.27 GiB0.03 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.09 it/s1.923.5627
Benchmarked· n=27

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6700 run?
1424 of 2118 indexed open-weight models fit a Radeon RX 6700 at 32,768 context with q8_0 KV cache, the largest being granite-3.1-8b-instruct at Q5_1. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6700 actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6700 fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.