NVIDIA · consumer

GeForce RTX 5090

GeForce RTX 5090 has 32 GB of VRAM at 1792 GB/s — about 29.76 GiB usable after driver and compositor overhead. 1859 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
GDDR7
Bandwidth
1792 GB/s
512-bit bus
Tensor FP16
419 TF
dense
TDP
575 W
$1999 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1580audio tts 21vision language 176video 16image 1audio asr 39embedding 26

What fits at 128K context

largest quantization that fits, per model · 1859 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
solar-pro-preview-instructKV unresolvedQ2_K22.1B7.65 GiB21.25 GiB29.76 GiB0.00 GiB44±12.9%
Seed-OSS-36B-InstructUD-IQ2_M36.2B11.86 GiB17.00 GiB29.75 GiB0.01 GiB44±12.9%
Qwen3-TTS-12Hz-0.6B-BaseF32915M28.88 GiB0.00 GiB29.72 GiB0.04 GiB44±12.9%
Gemma-4-Novelist-Eclipse-31BQ4_032.7B17.57 GiB11.25 GiB29.70 GiB0.06 GiB44±12.9%
Gemma-4-31B-StyleTuneQ4_032.7B17.57 GiB11.25 GiB29.70 GiB0.06 GiB44±12.9%
Noromaid-v0.4-Mixtral-Instruct-8x7b-ZlossMoEIQ3_M46.7B20.35 GiB8.50 GiB29.68 GiB0.08 GiB50±37%
spoomplesmaxx-v2.1-30BI1-IQ3_S28.9B11.74 GiB17.00 GiB29.65 GiB0.11 GiB44±12.9%
Huihui-granite-4.1-30b-abliteratedI1-IQ3_S28.9B11.74 GiB17.00 GiB29.65 GiB0.11 GiB44±12.9%
granite-4.1-30b-hereticI1-IQ3_S28.9B11.74 GiB17.00 GiB29.65 GiB0.11 GiB44±12.9%
Qwen3-Coder-Next-Opus-4.6-Reasoning-DistilledMoEQ2_K27.26 GiB1.59 GiB29.64 GiB0.12 GiB175±37%
Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliteratedI1-Q2_K_S36.2B11.74 GiB17.00 GiB29.64 GiB0.12 GiB44±12.9%
Hermes-4.3-36B-hereticI1-Q2_K_S36.2B11.74 GiB17.00 GiB29.64 GiB0.12 GiB44±12.9%
Skyfall-31B-v4.2Q3_K_M31.4B14.37 GiB14.34 GiB29.63 GiB0.13 GiB44±12.9%
GLM-Z1-Rumination-32B-0414Q2_K_L33.1B12.53 GiB16.20 GiB29.63 GiB0.13 GiB44±12.9%
granite-4.1-30bQ3_K_S28.9B11.71 GiB17.00 GiB29.62 GiB0.14 GiB44±12.9%
Smilodon-9B-v1F1610.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
bella-bartender-v2F169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedF169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
Gemma-2-9B-It-SPPO-Iter3BF169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
G2-Darkest-Writer-Dirty-Shirley-9B-v2F169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
G2-Darkest-Writer-9B-v1F169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
Gemma-SEA-LION-v3-9B-ITF169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
Tiger-Gemma-9B-v3F169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
gemma-2-9b-itF169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
magnum-v4-9bF169.2B17.22 GiB11.55 GiB29.61 GiB0.15 GiB44±12.9%
Hermes-4.3-36BIQ2_M36.2B11.68 GiB17.00 GiB29.58 GiB0.18 GiB44±12.9%
Devstral-Small-2-24B-Instruct-2512Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Transformed-Journey-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magistry-24B-v1.1I1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Mergedonia-AETHER-24B-v1aI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Mergedonia-AETHER-24B-v1bI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Slimaki-Tavern-24B-v1.3I1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Maginum-Cydoms-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Maginum-Cydoms-24B-absolute-heresyI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Dolphin3.0-Mistral-24BQ6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Dolphin3.0-R1-Mistral-24BQ6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Cydonia_VistralQ6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Mistral-Small-3.2-24B-Instruct-2506-ultra-uncensored-hereticI1-Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Huihui-Mistral-Small-3.2-24B-Instruct-2506-abliterated-llamacppfixedI1-Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Dans-PersonalityEngine-V1.2.0-24bI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Mistral-Small-3_2-24B-Instruct-2506-antislop.v2I1-Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Mistral-Small-3.2-24B-Instruct-2506Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Dans-PersonalityEngine-V1.3.0-24bI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Devstral-Small-2507Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Goetia-24B-v1.1I1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Devstral-Small-2505Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
MS3.2-PaintedFantasy-v3-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
RP-Spectrum-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
MS3.2-PaintedFantasy-v4.1-24B-ultra-uncensored-heretic-v2I1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magidonia-24B-v4.3-heretic-v1.2I1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magidonia-24B-v4.3-absolute-heresyI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
MagiSeek-Pro-V1I1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magistral-Small-2509Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magistral-Small-2507Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Cogidonia-v2-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magidonia-24B-v4.3I1-Q6_K18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Precog-24B-v1I1-Q6_K18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
experiment024bI1-Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Magidonia-24B-v4.2.0Q6_K23.6B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
Berthier-Mistral-Military-24BI1-Q6_K24.0B18.02 GiB10.63 GiB29.56 GiB0.20 GiB44±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation21.32 it/s11.9034.75172
Prompt processing13493.29 tok/s10927.3414983.7050
Text generation288.98 tok/s280.78298.5234
Benchmarked· n=172

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 5090 run?
1859 of 2118 indexed open-weight models fit a GeForce RTX 5090 at 131,072 context with q8_0 KV cache, the largest being solar-pro-preview-instruct at Q2_K. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 5090 actually have?
Its nameplate is 32 GB, but about 29.76 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 5090 fast for local AI?
Its memory bandwidth is 1792 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.