AMD · consumer

Radeon RX 6700

Radeon RX 6700 has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1146 of 2118 indexed models fit at 128K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
175 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 950vision language 99embedding 26audio asr 38video 12audio tts 20image 1

What fits at 128K context

largest quantization that fits, per model · 1146 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Bonsai-8B-unpackedIQ3_XXS8.2B3.30 GiB5.06 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
dolphin-2.9.3-mistral-7B-32kI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Mistral-7B-v0.3Q4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Mistral-7B-Instruct-v0.3-ParasiteI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Mistral-7B-Instruct-v0.3-JbliteratedI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Mistral-7B-Instruct-v0.3Q4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Mistral-7B-v0.3-Chinese-ChatQ4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
mistral-7b-v0.3-bnb-4bitQ4_K_S7.5B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Mathstral-7B-v0.1Q4_K_S7.2B3.86 GiB4.50 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-8B-Instruct-2410Q6_K8.0B6.14 GiB2.23 GiB9.30 GiB0.00 GiB24±26.5%
OLMoE-1B-7B-0924-InstructMoEI1-Q4_K_M6.9B3.92 GiB4.50 GiB9.30 GiB0.00 GiB22±37%
openchat-3.5-0106KV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
dolphin-2.6-mistral-7bQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Silicon-Maid-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
SciPhi-Self-RAG-Mistral-7B-32kKV unresolvedI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
dolphin-2.2.1-mistral-7bKV unresolvedI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
OpenChat-3.5-7B-Qwen-v2.0KV unresolvedI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
CapybaraHermes-2.5-Mistral-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
dolphin-2.8-mistral-7b-v02Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
openchat-3.5-1210KV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Mistral-7B-v0.2Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
OpenHermes-2.5-Mistral-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Hermes-Trismegistus-Mistral-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Mistral-7B-OpenOrcaKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
dolphin-2.1-mistral-7bKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
OpenHermes-2-Mistral-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
dolphin-2.6-mistral-7b-dpo-laserQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Mistral-7B-Instruct-v0.1KV unresolvedI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Mistral-7B-Instruct-v0.2I1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
ContextualKunoichi_KTO-7BI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
xLAM-7b-rI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
mistral-7b-uncensoredKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
MegaBeam-Mistral-7B-512kQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Yarn-Mistral-7b-128kKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Ninja-v1-RP-WIPKV unresolvedI1-Q4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
BioMistral-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Mistral-7B-Instruct-v0.2-code-ftKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
SpydazWeb_AI_CyberTron_Ultra_7bKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
MetaMath-Cybertron-StarlingKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
zephyr-7b-betaKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
japanese-stablelm-instruct-gamma-7bKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Kunoichi-DPO-v2-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
dolphin-2.0-mistral-7bKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
SciPhi-Mistral-7B-32kKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Kimiko-Mistral-7BKV unresolvedQ4_K_S7.2B3.86 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
MiMo-VL-7B-RLI1-IQ3_S8.3B3.30 GiB5.06 GiB9.29 GiB0.01 GiB24±26.5%
Kuwutu-7B-CYOA-v2I1-IQ3_S7.6B3.30 GiB5.06 GiB9.29 GiB0.01 GiB24±26.5%
Ling-liteMoEQ2_K16.8B6.43 GiB1.97 GiB9.29 GiB0.01 GiB38±37%
NVIDIA-Nemotron-3-Nano-4B-BF16Q3_K_L4.0B2.46 GiB5.91 GiB9.29 GiB0.01 GiB24±26.5%
granite-3.1-8b-instructIQ2_M8.2B2.73 GiB5.63 GiB9.29 GiB0.01 GiB24±26.5%
llm-jp-4-8b-instructIQ3_M8.6B3.85 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Falcon3-10B-InstructI1-IQ2_XXS10.3B2.70 GiB5.63 GiB9.29 GiB0.01 GiB24±26.5%
Anubis-Mini-8B-v1Q3_K_M8.0B3.85 GiB4.50 GiB9.29 GiB0.01 GiB24±26.5%
Carnice-Qwen3.6-MoE-35B-A3BMoEI1-IQ1_M36.0B7.67 GiB0.70 GiB9.28 GiB0.02 GiB76±37%
Qwen35B-Agent-R2-AbliteratedMoEI1-IQ1_M34.7B7.67 GiB0.70 GiB9.28 GiB0.02 GiB76±37%
Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEI1-IQ1_M36.0B7.67 GiB0.70 GiB9.28 GiB0.02 GiB76±37%
Darwin-35B-A3B-OpusMoEI1-IQ1_M36.0B7.67 GiB0.70 GiB9.28 GiB0.02 GiB76±37%
Qwen35B-Agent-R2MoEI1-IQ1_M34.7B7.67 GiB0.70 GiB9.28 GiB0.02 GiB76±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.09 it/s1.923.5627
Benchmarked· n=27

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6700 run?
1146 of 2118 indexed open-weight models fit a Radeon RX 6700 at 131,072 context with q4_0 KV cache, the largest being Bonsai-8B-unpacked at IQ3_XXS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6700 actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6700 fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.