NVIDIA · consumer

GeForce RTX 5080

GeForce RTX 5080 has 16 GB of VRAM at 960 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1870 of 2118 indexed models fit at 4K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR7
Bandwidth
960 GB/s
256-bit bus
Tensor FP16
225 TF
dense
TDP
360 W
$999 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1604vision language 163video 15audio asr 39image 2embedding 26audio tts 21

What fits at 4K context

largest quantization that fits, per model · 1870 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
magnum-v2-32bIQ3_M32.5B13.69 GiB0.28 GiB14.87 GiB0.01 GiB49±12.9%
Gemma-4-E4B-it-Minecraft-MT-en-zh-v0.1F168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
supergemma4-e4b-abliteratedF167.5B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
Gemma-4-E4B-LuchadorBF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
Gemma-4-E4B-AbliteratedF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-it-ultra-uncensored-hereticBF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-it-The-DECKARD-Claude-Opus-Expresso-Universe-HERETIC-UNCENSORED-ThinkingF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-it-The-DECKARD-Expresso-Universe-HERETIC-UNCENSORED-ThinkingF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
Huihui-gemma-4-E4B-it-abliteratedF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
Darkidol-Gemma-4-E4B-itF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-it-Claude-Opus-4.5-HERETIC-UNCENSORED-ThinkingF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-it-hereticBF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-it-Uncensored-MAXBF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4B-Agentic-Opus-Reasoning-GeminiCLI-mlx-4bitF167.5B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
OpenMedResearch-Gemma-4E4NF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
Reasoning-Medical0.1-E4B-sftF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
gemma-4-E4BF168.0B14.02 GiB0.03 GiB14.87 GiB0.01 GiB48±12.9%
Ling-liteMoEQ6_K16.8B14.02 GiB0.06 GiB14.87 GiB0.01 GiB165±37%
HarmonicHarlequin_v5-20BI1-IQ3_XXS33.3B11.74 GiB2.29 GiB14.87 GiB0.01 GiB48±12.9%
Qwen3.6-27B-A3B-CoderMoEI1-Q4_K_S26.7B14.04 GiB0.02 GiB14.86 GiB0.02 GiB214±37%
Gemma-4-Novelist-Eclipse-31BQ2_K_L32.7B13.48 GiB0.51 GiB14.86 GiB0.02 GiB49±12.9%
Gemma-4-31B-StyleTuneQ2_K_L32.7B13.48 GiB0.51 GiB14.86 GiB0.02 GiB49±12.9%
gemma-2-27b-itQ3_K_L27.2B13.52 GiB0.40 GiB14.86 GiB0.02 GiB49±12.9%
magnum-v4-27bQ3_K_L27.2B13.52 GiB0.40 GiB14.86 GiB0.02 GiB49±12.9%
Yi-34B-200K-DARE-megamerge-v8IQ3_XS34.4B13.71 GiB0.26 GiB14.85 GiB0.03 GiB49±12.9%
Nous-Hermes-2-Yi-34BI1-IQ3_XS34.4B13.71 GiB0.26 GiB14.85 GiB0.03 GiB49±12.9%
NSFW_13B_sftQ8_013.3B13.13 GiB0.88 GiB14.85 GiB0.03 GiB49±12.9%
G4-MeroMero-26B-A4B-it-uncensored-hereticMoEQ3_K_M25.8B13.93 GiB0.13 GiB14.84 GiB0.04 GiB48±12.9%
Gemma-4-Gembrain-X-Core-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-Gembrain-X-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-31B-Isometry-Fabled-PersonaI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Versipellis-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma4-Gutenberg-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
G4-MeroMero-31B-uncensored-hereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-Novelist-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Wanabi-Gemma4-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
G4-Alice-v1.2-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Agares-31B-v1I1-IQ3_M30.7B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma4-Gutenberg-31B-HereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-Ortenzya-The-Creative-Wordsmith-31B-it-uncensored-hereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-Gemsicle-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-Gembrain-31B-it-uncensored-hereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Melinoe-Gemma4-31B-VL-hereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
G4-MeroMero-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Glistening-Gem-31B-v1.0I1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Melinoe-Gemma4-31B-VLI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-31B-Storymaxxed3I1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliteratedI1-IQ3_M32.7B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-31B-Queen-it-qat-q4_0-unquantizedI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-31B-it-qat-q4_0-unquantized-hereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-AssGuard-31BI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
copywriter-gemma4-31bI1-IQ3_M32.7B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-31B-heretic-finetuneI1-IQ3_M30.7B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Gemma-4-Garnet-V2-31B-it-ultra-uncensored-hereticI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-31B-it-abliterated-v3I1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-31B-it-noloopI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Webs-Sejong-31B-v7I1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
Lilith-31B-v1.0I1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
JGOS-31B-ThinkI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
gemma-4-31B-MergemaxxedI1-IQ3_M31.3B13.43 GiB0.51 GiB14.82 GiB0.06 GiB49±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation17.40 it/s8.7221.8270
Prompt processing7405.11 tok/s6126.628738.2335
Text generation180.83 tok/s175.11184.3015
Benchmarked· n=70

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 5080 run?
1870 of 2118 indexed open-weight models fit a GeForce RTX 5080 at 4,096 context with q4_0 KV cache, the largest being magnum-v2-32b at IQ3_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 5080 actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 5080 fast for local AI?
Its memory bandwidth is 960 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.