AMD · consumer

Radeon RX 9070 XT

Radeon RX 9070 XT has 16 GB of VRAM at 640 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1384 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
640 GB/s
256-bit bus
Tensor FP16
dense
TDP
304 W
$599 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1148vision language 135embedding 26video 15audio tts 21audio asr 38image 1

What fits at 128K context

largest quantization that fits, per model · 1384 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEQ4_K_S23.6B12.65 GiB1.33 GiB14.88 GiB0.00 GiB78±37%
MiniCPM-Llama3-V-2_5BF168.5B5.43 GiB8.50 GiB14.87 GiB0.01 GiB29±26.5%
North-Mini-Code-1.0MoEIQ3_XXS30.5B12.10 GiB1.89 GiB14.87 GiB0.01 GiB64±37%
INTELLECT-1-InstructI1-IQ2_XXS10.2B2.77 GiB11.16 GiB14.87 GiB0.01 GiB29±26.5%
Qwen3-VL-30B-A3B-InstructMoEUD-TQ1_031.1B7.60 GiB6.38 GiB14.87 GiB0.01 GiB31±37%
Mistral-7B-v0.1KV unresolvedIQ4_XS7.2B5.43 GiB8.50 GiB14.86 GiB0.02 GiB29±26.5%
Kuwutu-7B-CYOA-v2Q4_K_M7.6B4.37 GiB9.56 GiB14.86 GiB0.02 GiB29±26.5%
Smilodon-9B-v1I1-IQ1_M10.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
bella-bartender-v2I1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
Gemma-2-9B-It-SPPO-Iter3I1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
Gemma-SEA-LION-v3-9B-ITI1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
G2-Darkest-Writer-9B-v1I1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
Tiger-Gemma-9B-v3I1-IQ1_M9.2B2.37 GiB11.55 GiB14.86 GiB0.02 GiB29±26.5%
Qwen3-VL-30B-A3B-ThinkingMoEUD-TQ1_031.1B7.59 GiB6.38 GiB14.86 GiB0.02 GiB31±37%
Qwen3-30B-A3B-Thinking-2507MoEUD-TQ1_030.5B7.59 GiB6.38 GiB14.86 GiB0.02 GiB31±37%
NVIDIA-Nemotron-3-Nano-4B-BF16Q4_K_M4.0B2.77 GiB11.16 GiB14.86 GiB0.02 GiB29±26.5%
Qwen3-30B-A3BMoEIQ2_XXS30.5B7.59 GiB6.38 GiB14.86 GiB0.02 GiB31±37%
Pantheon-Proto-RP-1.8-30B-A3BMoEIQ2_XXS30.5B7.59 GiB6.38 GiB14.86 GiB0.02 GiB31±37%
SmolLM2-1.7B-Instruct-UncensoredQ5_K_M1.8B1.21 GiB12.75 GiB14.85 GiB0.03 GiB29±26.5%
Qianfan-OCRQ8_04.7B4.38 GiB9.56 GiB14.85 GiB0.03 GiB29±26.5%
MiMo-VL-7B-RLI1-Q4_K_M8.3B4.36 GiB9.56 GiB14.85 GiB0.03 GiB29±26.5%
Qwen3-VL-Embedding-8BQ4_K_M8.1B4.36 GiB9.56 GiB14.85 GiB0.03 GiB29±26.5%
qwen-indic-v1I1-Q4_K_M7.6B4.36 GiB9.56 GiB14.85 GiB0.03 GiB29±26.5%
Qwen3-Embedding-8BQ4_K_M7.6B4.36 GiB9.56 GiB14.85 GiB0.03 GiB29±26.5%
Aya-Medikal-V2I1-Q5_K_M8.0B5.40 GiB8.50 GiB14.85 GiB0.03 GiB29±26.5%
CycleGRPO-4BQ8_04.8B4.37 GiB9.56 GiB14.85 GiB0.03 GiB29±26.5%
Jan-v3-4B-base-instructQ8_04.4B4.37 GiB9.56 GiB14.84 GiB0.04 GiB29±26.5%
Jan-code-4bQ8_04.4B4.37 GiB9.56 GiB14.84 GiB0.04 GiB29±26.5%
Qwen3-4B-BaseQ8_04.0B4.37 GiB9.56 GiB14.84 GiB0.04 GiB29±26.5%
rnj-1-instructQ5_K_S8.3B5.40 GiB8.50 GiB14.84 GiB0.04 GiB29±26.5%
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-IQ1_M13.9B3.26 GiB10.63 GiB14.84 GiB0.04 GiB29±26.5%
Ministral-3-14B-Reasoning-2512-UncensoredI1-IQ1_M13.9B3.26 GiB10.63 GiB14.84 GiB0.04 GiB29±26.5%
Bonsai-8B-unpackedIQ4_XS8.2B4.35 GiB9.56 GiB14.84 GiB0.04 GiB29±26.5%
zeta-2Q5_K_S8.3B5.40 GiB8.50 GiB14.84 GiB0.04 GiB29±26.5%
L3-Dark-Planet-8BQ5_K_S8.0B5.40 GiB8.50 GiB14.84 GiB0.04 GiB29±26.5%
granite-20b-code-instruct-8kQ5_K_L20.1B13.86 GiB0.00 GiB14.84 GiB0.04 GiB29±26.5%
dolphincoder-starcoder2-15bKV unresolvedI1-Q4_K_S16.0B8.53 GiB5.31 GiB14.84 GiB0.04 GiB29±26.5%
Huihui-Qwen3.5-35B-A3B-abliteratedMoEI1-IQ3_XXS36.0B12.60 GiB1.33 GiB14.84 GiB0.04 GiB85±37%
Qwen3.5-35B-A3B-BaseMoEI1-IQ3_XXS36.0B12.60 GiB1.33 GiB14.84 GiB0.04 GiB85±37%
Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEI1-IQ3_XXS36.0B12.60 GiB1.33 GiB14.84 GiB0.04 GiB85±37%
Goetia-26B-A4B-v1.4MoEI1-IQ3_XS26.0B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
G4-Moonlight-Dusk-26B-A4BMoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
Chimera-X-26B-A4BMoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
Gemma-4-26B-A4B-StyleTuneMoEI1-IQ3_XS26.5B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-IQ3_XS25.8B11.13 GiB2.81 GiB14.83 GiB0.05 GiB29±26.5%
Yi-Coder-1.5B-ChatQ6_K1.5B1.19 GiB12.75 GiB14.83 GiB0.05 GiB29±26.5%
Yi-Coder-1.5BQ6_K1.5B1.19 GiB12.75 GiB14.83 GiB0.05 GiB29±26.5%
Luna-7B-A4BMoEI1-Q5_K_S6.7B4.35 GiB9.56 GiB14.83 GiB0.05 GiB19±37%
Qwen3.8-27BUD-IQ2_M27.8B9.61 GiB4.25 GiB14.82 GiB0.06 GiB29±26.5%
Aurora-Code-1MoEI1-IQ3_M34.7B12.59 GiB1.33 GiB14.82 GiB0.06 GiB85±37%
Qwen2.5-Coder-7B-InstructQ5_K_M7.6B10.14 GiB3.72 GiB14.81 GiB0.07 GiB29±26.5%
Ministral-3-8B-Instruct-2512Q4_K_M8.9B4.84 GiB9.03 GiB14.81 GiB0.07 GiB29±26.5%
Ministral-3-8B-Reasoning-2512Q4_K_M8.9B4.84 GiB9.03 GiB14.81 GiB0.07 GiB29±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation5.76 it/s2.5411.88118
Prompt processing4275.83 tok/s3591.414809.8730
Text generation95.24 tok/s86.48108.0130
Benchmarked· n=118

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 9070 XT run?
1384 of 2118 indexed open-weight models fit a Radeon RX 9070 XT at 131,072 context with q8_0 KV cache, the largest being Qwen3.6-35B-A3B-REAM-160-ru-agent at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 9070 XT actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 9070 XT fast for local AI?
Its memory bandwidth is 640 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.