AMD · consumer

Radeon RX 7900 XT

Radeon RX 7900 XT has 20 GB of VRAM at 800 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1869 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6
Bandwidth
800 GB/s
320-bit bus
Tensor FP16
dense
TDP
315 W
$899 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1594vision language 171video 16embedding 26audio tts 21audio asr 39image 2

What fits at 32K context

largest quantization that fits, per model · 1869 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16MoEQ2_K_L31.6B16.85 GiB0.86 GiB18.60 GiB0.00 GiB97±37%
Qwen3-Coder-NextMoEIQ1_M79.7B16.11 GiB1.59 GiB18.60 GiB0.00 GiB90±37%
Qwen3-Next-80B-A3B-ThinkingMoEIQ1_M81.3B16.11 GiB1.59 GiB18.60 GiB0.00 GiB90±37%
Qwen3-Next-80B-A3B-InstructMoEIQ1_M81.3B16.11 GiB1.59 GiB18.60 GiB0.00 GiB90±37%
Nemotron-Cascade-2-30B-A3BMoEIQ2_XS31.6B16.85 GiB0.86 GiB18.59 GiB0.01 GiB97±37%
Mistral-MOE-4X7B-Dark-MultiVerse-Uncensored-Enhanced32-24BMoEQ5_K_S24.2B15.53 GiB2.13 GiB18.59 GiB0.01 GiB16±37%
gemma-4-26B-A4B-itMoEQ5_K_S26.5B16.88 GiB0.82 GiB18.59 GiB0.01 GiB28±26.5%
Laguna-XS-2.1MoEIQ4_XS33.4B16.96 GiB0.73 GiB18.59 GiB0.01 GiB111±37%
next-8bF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
next-ocrF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-GLM-4.7-Flash-Heretic-Uncensored-ThinkingF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Midas-FableAgent-8BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-Heretic-1.3.0F168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-Instruct-Unredacted-MAXF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Poe-8B-GLM5-Opus4.6-Sonnet4.5-Kimi-Grok-Gemini-3-pro-preview-HERETICF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen-3-VL-8B-Instruct-hereticF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
nsfwcaption-qwen3-vl-8b-v3-safetensorsF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-ThinkingBF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Huihui-Qwen3-VL-8B-Instruct-abliteratedF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-Reranker-8BF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Salience-1-9BF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-Instruct-Uncensored-V2F168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
GRaPE-2-FlashF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Jan-v2-VL-highF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Jan-v2-VL-medF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-InstructBF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Parable-Qwen3-8B-Claude-Fable-5F168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
ReasonCritic-7BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Finch-8B-KTOF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
DeepSeek-R1-0528-Qwen3-8BBF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Finch-8BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
mythos-9b-unhinged-hereticF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
MathSmith-hc-Qwen3-8BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-VL-8B-Thinking-Unredacted-MAXBF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
MiroThinker-v1.0-8BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
mythos-9b-unhingedF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Maestro1-9BBF168.8B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
qwen3-8b-claude-agentic-fable5F168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-8B-DeepSeek-v3.2-Speciale-DistillBF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Ektome-Qwen3-8B-PristinelyUncensoredF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
mythos-9b-mergedF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
qwen3-8b-apostateF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Josiefied-Qwen3-8B-abliterated-v1F168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
tmax-8bF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-8BBF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Marco-DeepResearch-8BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-8B-abliteratedBF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Qwen3-8B-UncensoredF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
LMT-60-8BF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Nemotron-Orchestrator-8BBF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
DS-R1-Qwen3-8B-ArliAI-RpR-v4-SmallBF168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
Huihui-Qwen3-8B-abliterated-v2F168.2B15.26 GiB2.39 GiB18.59 GiB0.01 GiB28±26.5%
deepseek-coder-33b-instructIQ3_S33.3B13.49 GiB4.12 GiB18.58 GiB0.02 GiB28±26.5%
MiniCPM-o-4_5F169.4B15.26 GiB2.39 GiB18.58 GiB0.02 GiB28±26.5%
Qwen3-Reranker-8BF168.2B15.26 GiB2.39 GiB18.58 GiB0.02 GiB28±26.5%
Ternary-Bonsai-8B-unpackedF168.2B15.26 GiB2.39 GiB18.58 GiB0.02 GiB28±26.5%
Bonsai-8B-unpackedBF168.2B15.26 GiB2.39 GiB18.58 GiB0.02 GiB28±26.5%
Trinity-MiniMoEQ5_K_M26.1B17.36 GiB0.33 GiB18.58 GiB0.02 GiB105±37%
glm-4-9b-chat-abliteratedQ5_K_L9.4B7.01 GiB10.63 GiB18.58 GiB0.02 GiB28±26.5%
glm-4-9b-chatQ5_K_L9.4B7.01 GiB10.63 GiB18.58 GiB0.02 GiB28±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation11.45 it/s7.7516.22328
Prompt processing3219.16 tok/s2738.953754.6863
Text generation101.20 tok/s99.80107.4539
Benchmarked· n=328

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 7900 XT run?
1869 of 2118 indexed open-weight models fit a Radeon RX 7900 XT at 32,768 context with q8_0 KV cache, the largest being NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 at Q2_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 7900 XT actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 7900 XT fast for local AI?
Its memory bandwidth is 800 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.