AMD · consumer

Radeon RX 6800 XT

Radeon RX 6800 XT has 16 GB of VRAM at 512 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1813 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
512 GB/s
256-bit bus
Tensor FP16
dense
TDP
300 W
$649 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1551vision language 159video 15audio asr 39image 2embedding 26audio tts 21

What fits at 32K context

largest quantization that fits, per model · 1813 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Goetia-26B-A4B-v1.4MoEI1-Q3_K_L26.0B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
G4-Moonlight-Dusk-26B-A4BMoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
Chimera-X-26B-A4BMoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
Gemma-4-26B-A4B-StyleTuneMoEI1-Q3_K_L26.5B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-Q3_K_L25.8B13.17 GiB0.82 GiB14.88 GiB0.00 GiB23±26.5%
EXAONE-4.0-32BIQ3_XS32.0B12.37 GiB1.51 GiB14.88 GiB0.00 GiB23±26.5%
Gemma-3-27B-MeditronFOI1-IQ3_M28.8B12.25 GiB1.65 GiB14.88 GiB0.00 GiB23±26.5%
Salience-1.5-FlashMoEI1-IQ3_S31.1B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Huihui-Qwen3-VL-30B-A3B-Instruct-abliteratedMoEI1-IQ3_S31.1B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-30B-A3B-Gemini-Pro-High-Reasoning-2507-ABLITERATED-UNCENSOREDMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
MiroThinker-v1.0-30BMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-30B-A3B-YOYO-V5MoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-30B-A3B-Thinking-2507-Claude-4.5-Sonnet-High-Reasoning-DistillMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Huihui-Qwen3-30B-A3B-Thinking-2507-abliteratedMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Huihui-Qwen3-30B-A3B-Instruct-2507-abliteratedMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-30B-A3B-abliterated-eroticMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
UncensoredLM-DeepSeek-R1-Distill-Qwen-14BQ6_K14.2B10.87 GiB3.05 GiB14.87 GiB0.01 GiB23±26.5%
Huihui-Qwen3-Coder-30B-A3B-Instruct-abliteratedMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-Coder-30B-A3B-Instruct-RTPurboMoEI1-IQ3_S30.5B12.39 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-Coder-30B-A3B-InstructMoEQ3_K_S30.5B12.38 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-VL-30B-A3B-InstructMoEQ3_K_S31.1B12.38 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Qwen3-30B-A3B-abliteratedMoEQ3_K_S30.5B12.38 GiB1.59 GiB14.87 GiB0.01 GiB57±37%
Fimbulvetr-11B-v2Q8_010.7B10.74 GiB3.19 GiB14.86 GiB0.02 GiB23±26.5%
MiniCPM-V-4_5Q5_18.7B11.53 GiB2.39 GiB14.86 GiB0.02 GiB23±26.5%
granite-20b-code-instruct-8kQ5_K_L20.1B13.86 GiB0.00 GiB14.84 GiB0.04 GiB23±26.5%
Qwen3-48B-A4B-Savant-Commander-Distill-12X-Closed-Open-Heretic-UncensoredMoEI1-Q2_K33.6B11.53 GiB2.39 GiB14.83 GiB0.05 GiB37±37%
Qwen3.6-VL-REAP-26B-A3BMoEIQ4_XS26.6B13.59 GiB0.33 GiB14.82 GiB0.06 GiB97±37%
diffusiongemma-26B-A4B-it-HERETIC-UncensoredMoEIQ4_XS25.8B13.11 GiB0.82 GiB14.81 GiB0.07 GiB23±26.5%
NVIDIA-Nemotron-Nano-12B-v2Q6_K_L12.3B9.72 GiB4.12 GiB14.81 GiB0.07 GiB23±26.5%
Goetia-26B-A4B-v1.3-Absolute-Heretic-ARAMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.81 GiB0.07 GiB23±26.5%
G4-MeroMero-26B-A4B-it-uncensored-hereticMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.81 GiB0.07 GiB23±26.5%
EVE-26b-XENO-HATMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
Gemma-4-26B-A4B-Animus-V14.1-FFT-hereticMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
G4-MeroMero-26B-A4BMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliteratedMoEIQ4_XS26.5B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
G4-Dark-Soul-26B-A4BMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-SOMPOA-heresyMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-heretic-ara-v2MoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-Heretic-StableMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-Uncensored-MAXMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-ultra-uncensored-hereticMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-uncensored-hereticMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-ara-abliteratedMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
Huihui-gemma-4-26B-A4B-it-abliteratedMoEIQ4_XS26.5B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-abliteratedMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4B-it-heretic-araMoEIQ4_XS25.8B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
gemma-4-26B-A4BMoEIQ4_XS26.5B13.10 GiB0.82 GiB14.80 GiB0.08 GiB23±26.5%
GRM-2.6-Plus-0628Q3_K_S27.8B12.78 GiB1.06 GiB14.80 GiB0.08 GiB23±26.5%
ThinkingCap-Qwen3.6-27BQ3_K_S27.4B12.78 GiB1.06 GiB14.80 GiB0.08 GiB23±26.5%
Tess-4-27BQ3_K_S27.8B12.78 GiB1.06 GiB14.80 GiB0.08 GiB23±26.5%
Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoEI1-IQ2_XS42.4B11.68 GiB2.22 GiB14.80 GiB0.08 GiB49±37%
Mistral-MOE-4X7B-Dark-MultiVerse-Uncensored-Enhanced32-24BMoEQ3_K_L24.2B11.73 GiB2.13 GiB14.79 GiB0.09 GiB13±37%
Skyfall-31B-v4.2-hereticI1-Q2_K_S31.4B10.18 GiB3.59 GiB14.78 GiB0.10 GiB23±26.5%
Skyfall-31B-v4.2I1-Q2_K_S31.4B10.18 GiB3.59 GiB14.78 GiB0.10 GiB23±26.5%
ERNIE-4.5-21B-A3B-ThinkingQ4_121.8B12.93 GiB0.93 GiB14.78 GiB0.10 GiB23±26.5%
ERNIE-4.5-21B-A3B-PTQ4_121.9B12.93 GiB0.93 GiB14.78 GiB0.10 GiB23±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation8.45 it/s5.9310.04215
Benchmarked· n=215

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6800 XT run?
1813 of 2118 indexed open-weight models fit a Radeon RX 6800 XT at 32,768 context with q8_0 KV cache, the largest being Goetia-26B-A4B-v1.4 at I1-Q3_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6800 XT actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6800 XT fast for local AI?
Its memory bandwidth is 512 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.