AMD · consumer

Radeon RX 6700

Radeon RX 6700 has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1678 of 2118 indexed models fit at 8K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
175 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1455vision language 123audio tts 21embedding 26video 12audio asr 39image 2

What fits at 8K context

largest quantization that fits, per model · 1678 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-ThinkingI1-IQ2_M23.4B7.64 GiB0.71 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-Q4_113.9B7.99 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16Q4_113.9B7.99 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Instruct-2512Q4_113.9B7.99 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Reasoning-2512-UncensoredI1-Q4_113.9B7.99 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Reasoning-2512Q4_113.9B7.99 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Qwen2.5-14BQ4_014.8B7.93 GiB0.42 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
NousCoder-14BQ4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
spoomplesmaxx-mini-14BI1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
vanilla-cn-roleplay-0.2I1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Claria-14bI1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
qwen3-14b-code-reasoning-conversationalQ4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
NTX-2.1-ProI1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Qwen3-14B-UncensoredI1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Qwen3-14BQ4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
FrogMini-14B-2510I1-Q4_K_S7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Qwen3-14B-abliteratedQ4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Josiefied-Qwen3-14B-abliterated-v3Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Hermes-4-14BQ4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Slava-Qwen3-14B-SerbianI1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Qwen3-14B-BaseQ4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
Huihui-Qwen3-14B-abliterated-v2I1-Q4_K_S14.8B7.98 GiB0.35 GiB9.30 GiB0.00 GiB24±26.5%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q4_015.7B8.32 GiB0.07 GiB9.29 GiB0.01 GiB81±37%
L3-DARKEST-PLANET-16.5BQ3_K_M16.5B7.73 GiB0.62 GiB9.29 GiB0.01 GiB24±26.5%
QwQ-32BUD-IQ1_M32.8B7.73 GiB0.56 GiB9.29 GiB0.01 GiB24±26.5%
Gemma-4-31B-Isometry-RPI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Gemma-4-Dark-Gemistry-31BI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Prosopon-31BI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Gemma-4-Novelist-Eclipse-31BI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Giftige-Blume-31B-v1-StyleSwapI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
G4-MeroMero-31B-StyleSwapI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Gemma-4-31B-StyleTune-heretic-araI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Pantheon-Reasoning-31B-1.1I1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Gemma-4-31B-StyleTuneI1-IQ1_M32.7B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Barcenas-StyleTune-31B-FableI1-IQ1_M32.1B7.63 GiB0.68 GiB9.29 GiB0.01 GiB24±26.5%
Octen-Embedding-8BQ8_07.6B8.04 GiB0.32 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3-VL-32B-InstructUD-IQ1_M33.4B7.73 GiB0.56 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3-VL-32B-ThinkingUD-IQ1_M33.4B7.73 GiB0.56 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3-32BUD-IQ1_M32.8B7.73 GiB0.56 GiB9.29 GiB0.01 GiB24±26.5%
GLM-4.7-Flash-DerestrictedMoEI1-IQ2_XS31.2B8.26 GiB0.12 GiB9.29 GiB0.01 GiB93±37%
Huihui-GLM-4.7-Flash-abliteratedMoEI1-IQ2_XS31.2B8.26 GiB0.12 GiB9.29 GiB0.01 GiB93±37%
dolphincoder-starcoder2-15bKV unresolvedIQ4_XS16.0B8.12 GiB0.18 GiB9.28 GiB0.02 GiB24±26.5%
OLMo-2-1124-7B-InstructQ8_07.3B7.23 GiB1.13 GiB9.28 GiB0.02 GiB24±26.5%
gemma-4-A4B-98e-v7-coder-itMoEQ2_K20.5B8.22 GiB0.17 GiB9.27 GiB0.03 GiB24±26.5%
gemma-4-A4B-98e-v7-coderx-itMoEQ2_K20.5B8.22 GiB0.17 GiB9.27 GiB0.03 GiB24±26.5%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-Q3_K_L8.0B8.05 GiB0.28 GiB9.27 GiB0.03 GiB24±26.5%
INTELLECT-2UD-IQ1_M32.8B7.71 GiB0.56 GiB9.27 GiB0.03 GiB24±26.5%
DeepSeek-Coder-V2-Lite-InstructMoEIQ4_NL15.7B8.29 GiB0.07 GiB9.27 GiB0.03 GiB81±37%
DeepSeek-V2-Lite-ChatMoEIQ4_NL15.7B8.29 GiB0.07 GiB9.27 GiB0.03 GiB81±37%
Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-IQ1_S39.5B8.09 GiB0.21 GiB9.26 GiB0.04 GiB24±26.5%
Phi-4-reasoning-plusQ4_K_S14.7B7.86 GiB0.44 GiB9.26 GiB0.04 GiB24±26.5%
Phi-4-reasoningQ4_K_S14.7B7.86 GiB0.44 GiB9.26 GiB0.04 GiB24±26.5%
phi-4Q4_K_S14.7B7.86 GiB0.44 GiB9.26 GiB0.04 GiB24±26.5%
GLM-4.7-Flash-REAP-23B-A3BMoEQ2_K_L23.0B8.24 GiB0.12 GiB9.26 GiB0.04 GiB83±37%
Nexa-AI-4x4B-InstructMoEI1-Q5_K_M12.1B8.03 GiB0.32 GiB9.25 GiB0.05 GiB25±37%
gemma-3n-E2B-itF165.4B8.31 GiB0.04 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3-VL-30B-A3B-ThinkingMoEIQ2_S31.1B8.14 GiB0.21 GiB9.25 GiB0.05 GiB88±37%
MiroThinker-v1.0-30BMoEIQ2_S30.5B8.14 GiB0.21 GiB9.25 GiB0.05 GiB88±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.09 it/s1.923.5627
Benchmarked· n=27

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6700 run?
1678 of 2118 indexed open-weight models fit a Radeon RX 6700 at 8,192 context with q4_0 KV cache, the largest being MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking at I1-IQ2_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6700 actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6700 fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.
Radeon RX 6700 — what AI models can it run locally? — ossmodeldb