AMD · consumer

Radeon RX 6700

Radeon RX 6700 has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1684 of 2118 indexed models fit at 4K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
175 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1460audio asr 39vision language 124audio tts 21video 12embedding 26image 2

What fits at 4K context

largest quantization that fits, per model · 1684 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Voxtral-Mini-4B-Realtime-2602KV unresolvedF164.4B8.27 GiB0.11 GiB9.30 GiB0.00 GiB24±26.5%
CycleGRPO-4BF164.8B8.23 GiB0.16 GiB9.29 GiB0.01 GiB24±26.5%
Jan-v3-4B-base-instructBF164.4B8.22 GiB0.16 GiB9.29 GiB0.01 GiB24±26.5%
Jan-code-4bBF164.4B8.22 GiB0.16 GiB9.29 GiB0.01 GiB24±26.5%
Magistral-Small-2509-VisionQ2_K_M24.0B8.09 GiB0.18 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3-15B-A2B-BaseMoEQ4_K_S15.6B8.33 GiB0.05 GiB9.29 GiB0.01 GiB97±37%
Llama3.2-24B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-Q3_K_M18.0B8.25 GiB0.12 GiB9.28 GiB0.02 GiB67±37%
Mistral-7B-v0.1KV unresolvedQ6_K7.2B8.20 GiB0.14 GiB9.28 GiB0.02 GiB24±26.5%
Phi-3-medium-4k-instructQ4_K_L14.0B8.09 GiB0.22 GiB9.27 GiB0.03 GiB24±26.5%
Qwen2.5-Coder-7B-InstructQ4_07.6B8.25 GiB0.06 GiB9.27 GiB0.03 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16Q4_K_L13.9B8.14 GiB0.18 GiB9.27 GiB0.03 GiB24±26.5%
Kepler-8B-Instruct-v2Q4_07.6B8.25 GiB0.06 GiB9.27 GiB0.03 GiB24±26.5%
EuroLLM-22B-Instruct-2512Q2_K22.6B8.06 GiB0.24 GiB9.27 GiB0.03 GiB24±26.5%
IQuest-Coder-V1-40B-InstructI1-IQ1_S39.8B7.91 GiB0.35 GiB9.26 GiB0.04 GiB24±26.5%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q4_015.7B8.32 GiB0.03 GiB9.26 GiB0.04 GiB82±37%
gemma-2-27b-itIQ2_XS27.2B7.82 GiB0.40 GiB9.26 GiB0.04 GiB24±26.5%
magnum-v4-27bIQ2_XS27.2B7.82 GiB0.40 GiB9.26 GiB0.04 GiB24±26.5%
Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoEI1-IQ1_S42.4B8.21 GiB0.15 GiB9.25 GiB0.05 GiB91±37%
zeta-2.1Q8_08.3B8.17 GiB0.14 GiB9.25 GiB0.05 GiB24±26.5%
zeta-2Q8_08.3B8.17 GiB0.14 GiB9.25 GiB0.05 GiB24±26.5%
Forsaken-Void-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Silver-Siren-ST-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Tess-3-Mistral-Nemo-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
KrakenSakura-Maelstrom-12B-v1Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
MN-12B-Runeweaver-RP-RUI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Impish_Bloodmoon_12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Wayfarer-2-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Wayfarer-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Muse-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Vikhr-Nemo-12B-Instruct-R-21-09-24Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-Base-2407Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
writing-roleplay-20k-context-nemo-12b-v1.0Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Dans-PersonalityEngine-V1.3.0-12bI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
mini-magnum-12b-v1.1Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Lumimaid-v0.2-12BQ5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
MN-Violet-Lotus-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Rocinante-X-12B-v1-Heretic-UncensoredI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Heretica-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Violet_Twilight-v0.2Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Lumimaid-Magnum-v4-12BQ5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
arcee-fusion-lumaid-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-NeMo-12B-AbliteratedI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Captain-Eris_Violet-V0.420-12BI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Rocinante-X-12B-v1-absolute-heresyI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Rocinante-X-12B-v1I1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-Gutenberg-Doppel-12BQ5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-Instruct-2407Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
magnum-v4-12bQ5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
MN-12b-RP-InkQ5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-12B-ArliAI-RPMax-v1.1Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-2407-12B-Thinking-Claude-Gemini-GPT5.2-Uncensored-HERETICI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Dans-SakuraKaze-V1.0.0-12bI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-Inst-2407-12B-Thinking-Uncensored-HERETIC-HI-Claude-OpusI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mistral-Nemo-Instruct-2407-12B-Thinking-M-Claude-Opus-High-ReasoningI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
MN-12B-Mag-Mell-R1Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Mordant-12B-ThinkI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
Riverfish-Rocinante-12B-SFT-DPOI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
MN-Violet-Lotus-12B-HereticI1-Q5_K_M12.2B8.13 GiB0.18 GiB9.25 GiB0.05 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.09 it/s1.923.5627
Benchmarked· n=27

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6700 run?
1684 of 2118 indexed open-weight models fit a Radeon RX 6700 at 4,096 context with q4_0 KV cache, the largest being granite-20b-code-instruct-8k at IQ3_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6700 actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6700 fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.