AMD · consumer

Radeon RX 7900 XT

Radeon RX 7900 XT has 20 GB of VRAM at 800 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1928 of 2118 indexed models fit at 32K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6
Bandwidth
800 GB/s
320-bit bus
Tensor FP16
dense
TDP
315 W
$899 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1653vision language 171audio asr 39video 16image 2audio tts 21embedding 26

What fits at 32K context

largest quantization that fits, per model · 1928 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Aurora-Code-1MoEIQ4_XS34.7B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
grug-35bMoEIQ4_XS35.1B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
WorldSim-Opus-3.6-35B-A3BMoEIQ4_XS35.1B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
Qwen3.6-35B-A3B-AnkoMoEIQ4_XS35.1B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
KAT-Coder-V2.5-DevMoEIQ4_XS34.7B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
Ornith-1.0-35BMoEIQ4_XS34.7B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
Nex-N2-miniMoEIQ4_XS35.1B17.51 GiB0.18 GiB18.59 GiB0.01 GiB136±37%
EXAONE-4.5-33BIQ4_XS34.4B16.79 GiB0.80 GiB18.59 GiB0.01 GiB28±26.5%
Nous-Hermes-2-Yi-34BQ3_K_M34.4B15.49 GiB2.11 GiB18.59 GiB0.01 GiB28±26.5%
OrionStar-Yi-34B-Chat-LlamaQ3_K_M34.4B15.49 GiB2.11 GiB18.59 GiB0.01 GiB28±26.5%
Nous-Capybara-limarpv3-34BQ3_K_M34.4B15.49 GiB2.11 GiB18.59 GiB0.01 GiB28±26.5%
Noromaid-v0.4-Mixtral-Instruct-8x7b-ZlossMoEQ2_K46.7B16.51 GiB1.13 GiB18.57 GiB0.03 GiB46±37%
Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-Q4_030.0B16.00 GiB1.65 GiB18.56 GiB0.04 GiB60±37%
TildeOpen-30B-Instruct-LVI1-IQ4_XS30.7B15.46 GiB2.11 GiB18.55 GiB0.05 GiB28±26.5%
spoomplesmaxx-v2.1-30BI1-Q4_028.9B15.29 GiB2.25 GiB18.55 GiB0.05 GiB28±26.5%
Huihui-granite-4.1-30b-abliteratedI1-Q4_028.9B15.29 GiB2.25 GiB18.55 GiB0.05 GiB28±26.5%
granite-4.1-30b-hereticI1-Q4_028.9B15.29 GiB2.25 GiB18.55 GiB0.05 GiB28±26.5%
Goetia-26B-A4B-v1.4MoEI1-Q5_K_S26.0B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
G4-Moonlight-Dusk-26B-A4BMoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
Chimera-X-26B-A4BMoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
Gemma-4-26B-A4B-StyleTuneMoEI1-Q5_K_S26.5B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-Q5_K_S25.8B17.22 GiB0.43 GiB18.55 GiB0.05 GiB28±26.5%
granite-4.0-h-smallMoEQ4_032.2B17.51 GiB0.14 GiB18.54 GiB0.06 GiB76±37%
Delphi-25B-SimpleRL-MathI1-IQ2_M25.0B8.13 GiB9.41 GiB18.52 GiB0.08 GiB28±26.5%
Carnice-Qwen3.6-MoE-35B-A3BMoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen35B-Agent-R2-AbliteratedMoEI1-IQ4_XS34.7B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Darwin-35B-A3B-OpusMoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen35B-Agent-R2MoEI1-IQ4_XS34.7B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Carnice-MoE-35B-A3BMoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
spoomplesmaxx-flash-35B-A3MoEI1-IQ4_XS35.1B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliteratedMoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-35B-A3B-Uncensored-AggressiveMoEI1-IQ4_XS35.1B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-35B-A3B-abliterated-MAXMoEI1-IQ4_XS35.1B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Huihui-Qwen3.6-35B-A3B-abliteratedMoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwopus3.6-35B-A3B-v1MoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-35B-A3B-StyleTuneMoEI1-IQ4_XS35.1B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-35B-A3B-abliteratedMoEI1-IQ4_XS35.1B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-35B-A3B-abliterated-v4MoEIQ4_XS34.7B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
0GM-1.0-35B-A3B-0427MoEI1-IQ4_XS36.0B17.44 GiB0.18 GiB18.52 GiB0.08 GiB136±37%
Qwen3.6-27B-A3B-CoderMoEI1-Q5_K_M26.7B17.44 GiB0.18 GiB18.52 GiB0.08 GiB116±37%
Phi-3.5-mini-instructF323.8B14.24 GiB3.38 GiB18.52 GiB0.08 GiB28±26.5%
Phi-3-mini-128k-instructF323.8B14.24 GiB3.38 GiB18.52 GiB0.08 GiB28±26.5%
Phi-3-mini-4k-instructF323.8B14.24 GiB3.38 GiB18.52 GiB0.08 GiB28±26.5%
InternVL3_5-30B-A3BQ4_K_L30.8B17.57 GiB0.00 GiB18.51 GiB0.09 GiB28±26.5%
Salience-1.5-FlashMoEQ4_K_S31.1B16.77 GiB0.84 GiB18.51 GiB0.09 GiB91±37%
Qwen3-VL-30B-A3B-ThinkingMoEQ4_K_S31.1B16.75 GiB0.84 GiB18.49 GiB0.11 GiB91±37%
MiroThinker-v1.0-30BMoEQ4_K_S30.5B16.75 GiB0.84 GiB18.49 GiB0.11 GiB91±37%
Qwen3-30B-A3BMoEQ4_K_S30.5B16.75 GiB0.84 GiB18.49 GiB0.11 GiB91±37%
Qwen3-30B-A3B-Instruct-2507MoEQ4_K_S30.5B16.75 GiB0.84 GiB18.49 GiB0.11 GiB91±37%
Qwen3-30B-A3B-Thinking-2507MoEQ4_K_S30.5B16.75 GiB0.84 GiB18.49 GiB0.11 GiB91±37%
Pantheon-Proto-RP-1.8-30B-A3BMoEQ4_K_S30.5B16.75 GiB0.84 GiB18.49 GiB0.11 GiB91±37%
granite-4.1-30bQ4_028.9B15.23 GiB2.25 GiB18.48 GiB0.12 GiB29±26.5%
Tongyi-DeepResearch-30B-A3BMoEQ4_K_S30.5B16.75 GiB0.84 GiB18.48 GiB0.12 GiB91±37%
Qwen3-48B-A4B-Savant-Commander-Distill-12X-Closed-Open-Heretic-UncensoredMoEI1-Q3_K_L33.6B16.30 GiB1.27 GiB18.47 GiB0.13 GiB57±37%
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16UD-IQ3_S33.0B17.53 GiB0.00 GiB18.47 GiB0.13 GiB28±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation11.45 it/s7.7516.22328
Prompt processing3219.16 tok/s2738.953754.6863
Text generation101.20 tok/s99.80107.4539
Benchmarked· n=328

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 7900 XT run?
1928 of 2118 indexed open-weight models fit a Radeon RX 7900 XT at 32,768 context with q4_0 KV cache, the largest being Aurora-Code-1 at IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 7900 XT actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 7900 XT fast for local AI?
Its memory bandwidth is 800 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.