AMD · consumer

Radeon RX 6650 XT

Radeon RX 6650 XT has 8 GB of VRAM at 280 GB/s — about 7.44 GiB usable after driver and compositor overhead. 598 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
280 GB/s
128-bit bus
Tensor FP16
dense
TDP
180 W
$399 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 454embedding 16vision language 71audio asr 31video 7image 1audio tts 18

What fits at 128K context

largest quantization that fits, per model · 598 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nanbeige4.1-3BQ4_K_M3.9B2.28 GiB4.25 GiB7.44 GiB0.00 GiB27±26.5%
gte-largeQ4_K_S335M0.19 GiB6.38 GiB7.44 GiB0.00 GiB27±26.5%
nomic-embed-codeQ2_K_L7.1B2.77 GiB3.72 GiB7.44 GiB0.00 GiB27±26.5%
Qwythos-9B-v2IQ3_XS9.7B4.37 GiB2.13 GiB7.44 GiB0.00 GiB27±26.5%
Tess-4-9BIQ3_XS9.7B4.37 GiB2.13 GiB7.44 GiB0.00 GiB27±26.5%
Yi-6B-ChatI1-IQ3_XXS6.1B2.25 GiB4.25 GiB7.42 GiB0.02 GiB27±26.5%
Wan2.1-T2V-1.3BQ4_01.4B6.50 GiB0.00 GiB7.42 GiB0.02 GiB27±26.5%
granite-4.1-3bUD-IQ2_M3.4B1.20 GiB5.31 GiB7.41 GiB0.03 GiB27±26.5%
Parable-Granite-4.1-3B-Claude-Fable-5I1-Q2_K_S3.4B1.20 GiB5.31 GiB7.41 GiB0.03 GiB27±26.5%
Fara1.5-9BQ3_K_S9.4B4.35 GiB2.13 GiB7.41 GiB0.03 GiB27±26.5%
QwenPaw-Flash-9BQ3_K_S9.4B4.35 GiB2.13 GiB7.41 GiB0.03 GiB27±26.5%
grug-9bQ3_K_S9.4B4.35 GiB2.13 GiB7.41 GiB0.03 GiB27±26.5%
OmniCoder-9BQ3_K_S9.4B4.35 GiB2.13 GiB7.41 GiB0.03 GiB27±26.5%
Ornith-1.0-9BQ3_K_S9.2B4.35 GiB2.13 GiB7.41 GiB0.03 GiB27±26.5%
Qwen3.5-9B-NeoQ3_K_S9.7B4.35 GiB2.13 GiB7.41 GiB0.03 GiB27±26.5%
SmolLM3-3BQ4_K_S3.1B1.69 GiB4.78 GiB7.39 GiB0.05 GiB27±26.5%
granite-4.0-h-microBF163.2B5.95 GiB0.53 GiB7.38 GiB0.06 GiB27±26.5%
Tini-Cybersec-8B-A1BMoEQ5_K_L8.5B5.68 GiB0.80 GiB7.38 GiB0.06 GiB53±37%
granite-3.3-2b-instructQ3_K_M2.5B1.17 GiB5.31 GiB7.38 GiB0.06 GiB27±26.5%
granite-3.2-2b-instructQ3_K_M2.5B1.17 GiB5.31 GiB7.38 GiB0.06 GiB27±26.5%
granite-3.1-2b-instructQ3_K_M2.5B1.17 GiB5.31 GiB7.38 GiB0.06 GiB27±26.5%
granite-vision-3.2-2bQ3_K_M3.0B1.17 GiB5.31 GiB7.38 GiB0.06 GiB27±26.5%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseI1-IQ4_NL8.1B4.58 GiB1.86 GiB7.38 GiB0.06 GiB27±26.5%
granite-4.0-microIQ2_M3.4B1.16 GiB5.31 GiB7.37 GiB0.07 GiB27±26.5%
Vero-Qwen35-9B-BaseI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Vero-Qwen35-9BI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Morphos-9BI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwable-9B-Claude-Fable-5-hereticI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Holo-3.1-9BI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwable-9B-Claude-Fable-5I1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-imabari-v2I1-Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-abliterated-v2-MAXI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwable-9B-Claude-Fable-5-StraTAI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-RpRMax-v1I1-Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
AdQWENistrator-9BI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
cajal-9b-v2-fullI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-ultra-uncensored-hereticQ3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Holo-3.1-9B-CoderI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
PlutoI1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Holo-3.1-9B-abliterated-rdoI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-Uncensored-cyber-v3Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-BaseI1-Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
qwen3.5-9b-nsfw-captioning-v5I1-Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Miss_MARTHA-9B-Qwen3.5-OmniI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Huihui-Qwen3.5-9B-abliteratedQ3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-DS-v4-Flash-v3.0Q3_K_M9.4B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-Claude-Opus-4.6-DistillQ3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-DeepSeek-V4-FlashI1-Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Huihui-Qwen3.5-9B-Claude-4.6-Opus-abliteratedI1-Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9BQ3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Qwen3.5-9B-GLM5.1-Distill-v1Q3_K_M9.7B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Katarau-9B-ru-RP-nsfwI1-Q3_K_M9.0B4.31 GiB2.13 GiB7.37 GiB0.07 GiB27±26.5%
Crow-9B-HERETIC-4.6I1-Q3_K_M9.4B4.30 GiB2.13 GiB7.36 GiB0.08 GiB27±26.5%
Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGI1-Q3_K_M9.4B4.30 GiB2.13 GiB7.36 GiB0.08 GiB27±26.5%
Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSOREDI1-Q3_K_M9.4B4.30 GiB2.13 GiB7.36 GiB0.08 GiB27±26.5%
Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCTI1-Q3_K_M9.4B4.30 GiB2.13 GiB7.36 GiB0.08 GiB27±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.97 it/s2.634.85138
Benchmarked· n=138

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6650 XT run?
598 of 2118 indexed open-weight models fit a Radeon RX 6650 XT at 131,072 context with q8_0 KV cache, the largest being Nanbeige4.1-3B at Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6650 XT actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6650 XT fast for local AI?
Its memory bandwidth is 280 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.
Radeon RX 6650 XT — what AI models can it run locally? — ossmodeldb