AMD · consumer

Radeon RX 6650 XT

Radeon RX 6650 XT has 8 GB of VRAM at 280 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1174 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
280 GB/s
128-bit bus
Tensor FP16
dense
TDP
180 W
$399 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
embedding 26text 988vision language 94video 7audio asr 38image 1audio tts 20

What fits at 64K context

largest quantization that fits, per model · 1174 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nemotron-3-Embed-8B-BF16IQ4_XS8.0B4.11 GiB2.39 GiB7.44 GiB0.00 GiB27±26.5%
DeepSeek-V2-Lite-Chat-Uncensored-Unbiased-ReasonerMoEQ2_K15.7B5.99 GiB0.53 GiB7.43 GiB0.01 GiB66±37%
Qwen3-VL-Embedding-8BIQ4_XS8.1B3.97 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q2_K15.7B5.99 GiB0.53 GiB7.43 GiB0.01 GiB66±37%
DeepSeek-Coder-V2-Lite-InstructMoEQ2_K15.7B5.99 GiB0.53 GiB7.43 GiB0.01 GiB66±37%
DeepSeek-V2-Lite-ChatMoEQ2_K15.7B5.99 GiB0.53 GiB7.43 GiB0.01 GiB66±37%
qwen-indic-v1IQ4_XS7.6B3.97 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Falcon3-10B-InstructQ2_K10.3B3.65 GiB2.81 GiB7.43 GiB0.01 GiB27±26.5%
Mistral-7B-Instruct-v0.3-ParasiteI1-Q4_17.2B4.24 GiB2.25 GiB7.43 GiB0.01 GiB27±26.5%
Mistral-7B-Instruct-v0.3-JbliteratedI1-Q4_17.2B4.24 GiB2.25 GiB7.43 GiB0.01 GiB27±26.5%
Mistral-7B-Instruct-v0.3Q4_17.2B4.24 GiB2.25 GiB7.43 GiB0.01 GiB27±26.5%
Mistral-7B-v0.3-Chinese-ChatQ4_17.2B4.24 GiB2.25 GiB7.43 GiB0.01 GiB27±26.5%
LFM2-24B-A2BMoEIQ2_XS23.8B6.17 GiB0.35 GiB7.43 GiB0.01 GiB84±37%
Nexa-AI-4x4B-InstructMoEI1-Q2_K_S12.1B3.99 GiB2.53 GiB7.43 GiB0.01 GiB22±37%
Qwen3-VL-4B-ThinkingQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-VL-4B-Instruct-Uncensored-abliteratedQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
PopiT-Qwen3-4B-Medical-SFT-1128Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-VL-4B-Instruct-Unredacted-MAXQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Zubr1.0-VL-4BQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-VL-4B-Thinking-Unredacted-MAXQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Huihui-Qwen3-VL-4B-Instruct-abliteratedQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-VL-4B-Instruct-UncensoredQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-VL-4B-InstructQ8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Logics-Parsing-v2Q8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Jan-v1-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
OpenCaption-4B-VL-SFT-v1.0Q8_04.4B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Parable-Qwen3-4B-Claude-Fable-5Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Z-Image-Engineer-V6Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4b-Z-Image-Turbo-AbliteratedV1Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Huihui-Qwen3-4B-Instruct-2507-abliteratedQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Josiefied-Qwen3-4B-abliterated-v2Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-Thinking-2507Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Jan-nano-128kQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-Instruct-2507Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Jan-nanoQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-abliteratedQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-Qwen3.6-plus-Reasoning-DistilledQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-Kimi2.5-Reasoning-DistilledQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Neuron-4B-InstructQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
ChineseErrorCorrector4-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-Instruct_NSFW-V2.1Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
FastContext-1.0-4B-SFT-abliteratedQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
FastContext-1.0-4B-SFTQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-abliterated-v2Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
LocoOperator-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
fable-traces-abliteratedQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Nexa-AI-4B-InstructQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
fable-tracesQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-Instruct-2507-hereticQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Lumen-4B-InstructQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
CyberSecQwen-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-HereticLM-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Huihui-Qwen3-4B-abliterated-v2Q8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-4B-BaseQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-Reranker-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Ternary-Bonsai-4B-unpackedQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Octen-Embedding-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
Qwen3-Embedding-4BQ8_04.0B3.99 GiB2.53 GiB7.43 GiB0.01 GiB27±26.5%
DeepSeek-V2-Lite-Chat-UncensoredMoEQ2_K15.7B5.98 GiB0.53 GiB7.43 GiB0.01 GiB66±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.97 it/s2.634.85138
Benchmarked· n=138

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6650 XT run?
1174 of 2118 indexed open-weight models fit a Radeon RX 6650 XT at 65,536 context with q4_0 KV cache, the largest being Nemotron-3-Embed-8B-BF16 at IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6650 XT actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6650 XT fast for local AI?
Its memory bandwidth is 280 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.