NVIDIA · consumer

GeForce RTX 3090 Ti

GeForce RTX 3090 Ti has 24 GB of VRAM at 1008 GB/s — about 22.32 GiB usable after driver and compositor overhead. 1947 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
24 GB
GDDR6X
Bandwidth
1008 GB/s
384-bit bus
Tensor FP16
160 TF
dense
TDP
450 W
$1999 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1670audio asr 39vision language 173image 2video 16audio tts 21embedding 26

What fits at 64K context

largest quantization that fits, per model · 1947 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
EuroLLM-22B-Instruct-2512Q6_K_L22.6B17.65 GiB3.80 GiB22.31 GiB0.01 GiB34±12.9%
Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingQ8_021.3B20.61 GiB0.84 GiB22.31 GiB0.01 GiB34±12.9%
Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-ThinkingQ8_021.3B20.61 GiB0.84 GiB22.31 GiB0.01 GiB34±12.9%
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16UD-Q4_K_S33.0B21.47 GiB0.00 GiB22.31 GiB0.01 GiB34±12.9%
Yi-34B-200K-DARE-megamerge-v8I1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
dolphin-2.9.1-yi-1.5-34bI1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
OrionStar-Yi-34B-Chat-LlamaI1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
Yi-34B-200K-LlamafiedI1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
Nous-Hermes-2-Yi-34BI1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
Merged-RP-Stew-V2-34BI1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
Capybara-Tess-Yi-34B-200KI1-IQ4_XS34.4B17.21 GiB4.22 GiB22.31 GiB0.01 GiB34±12.9%
Huihui-GLM-4.7-Flash-abliterated-57BMoEI1-Q2_K57.3B19.11 GiB2.35 GiB22.30 GiB0.02 GiB86±37%
Voxtral-Small-24B-2507Q6_K24.3B18.57 GiB2.81 GiB22.30 GiB0.02 GiB34±12.9%
NSFW_13B_sftQ4_K_S13.3B7.39 GiB14.06 GiB22.30 GiB0.02 GiB34±12.9%
Gemma-The-Writer-N-Restless-Quill-10B-UncensoredIQ4_XS10.0B17.99 GiB3.46 GiB22.30 GiB0.02 GiB34±12.9%
spoomplesmaxx-v2.1-30BI1-Q4_128.9B16.88 GiB4.50 GiB22.29 GiB0.03 GiB34±12.9%
Huihui-granite-4.1-30b-abliteratedI1-Q4_128.9B16.88 GiB4.50 GiB22.29 GiB0.03 GiB34±12.9%
granite-4.1-30b-hereticI1-Q4_128.9B16.88 GiB4.50 GiB22.29 GiB0.03 GiB34±12.9%
granite-4.1-30bQ4_128.9B16.88 GiB4.50 GiB22.29 GiB0.03 GiB34±12.9%
OLMo-2-1124-13B-InstructQ4_K_S13.7B7.37 GiB14.06 GiB22.28 GiB0.04 GiB34±12.9%
Gemma4-Gutenberg-31BQ4_K_M31.3B18.25 GiB3.14 GiB22.28 GiB0.04 GiB34±12.9%
gemma-4-31B-itQ4_K_M31.3B18.25 GiB3.14 GiB22.28 GiB0.04 GiB34±12.9%
Gemma4-Gutenberg-31B-HereticQ4_K_M31.3B18.25 GiB3.14 GiB22.28 GiB0.04 GiB34±12.9%
Equinox-31BQ4_K_M31.3B18.25 GiB3.14 GiB22.28 GiB0.04 GiB34±12.9%
gemma-4-31B-it-SDFT-Heretic-RPQ4_K_M30.7B18.25 GiB3.14 GiB22.28 GiB0.04 GiB34±12.9%
Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-IQ4_XS39.5B19.72 GiB1.69 GiB22.27 GiB0.05 GiB34±12.9%
Qwen3.6-35B-A3BMoEUD-Q4_K_M36.0B21.11 GiB0.35 GiB22.26 GiB0.06 GiB166±37%
Qwen3.5-35B-A3BMoEQ4_K_L36.0B21.11 GiB0.35 GiB22.26 GiB0.06 GiB166±37%
Pantheon-Reasoning-27BQ5_K_L27.8B20.26 GiB1.13 GiB22.24 GiB0.08 GiB34±12.9%
Qwen3.5-27BQ5_K_L27.8B20.26 GiB1.13 GiB22.24 GiB0.08 GiB34±12.9%
codellama-13b-oasst-sft-v10Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
chronos-hermes-13b-v2Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
WhiteRabbitNeo-13B-v1Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
CodeLlama-13b-Instruct-hfQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
Orca-2-13b-Alpaca-UncensoredI1-Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
WizardLM-13B-UncensoredI1-Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
WizardCoder-Python-13B-V1.0I1-Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
Guanaco-13B-UncensoredI1-Q4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
Llama-2-13b-chat-hfQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
Wizard-Vicuna-13B-UncensoredQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
WizardLM-13b-V1.0-UncensoredQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
WizardLM-1.0-Uncensored-Llama2-13bQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
speechless-llama2-hermes-orca-platypus-wizardlm-13bQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
mythalion-13bQ4_K_M13.0B7.33 GiB14.06 GiB22.23 GiB0.09 GiB34±12.9%
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16MoEQ4_K_S31.6B20.51 GiB0.91 GiB22.20 GiB0.12 GiB126±37%
Gemma-4-31B-StyleTuneQ5_K32.7B18.17 GiB3.14 GiB22.20 GiB0.12 GiB34±12.9%
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-ThinkingQ5_K_M23.4B15.64 GiB5.70 GiB22.19 GiB0.13 GiB34±12.9%
Salience-1.5-FlashMoEQ5_K_S31.1B19.69 GiB1.69 GiB22.17 GiB0.15 GiB97±37%
c4ai-command-r-08-2024Q4_K_M32.3B18.44 GiB2.81 GiB22.17 GiB0.15 GiB34±12.9%
Gemma-4-Gembrain-X-Core-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Gemma-4-Gembrain-X-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Gemma-4-31B-Isometry-Fabled-PersonaI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Versipellis-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
G4-MeroMero-31B-uncensored-hereticI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Gemma-4-Novelist-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Wanabi-Gemma4-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
G4-Alice-v1.2-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Agares-31B-v1I1-Q4_130.7B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
gemma-4-Ortenzya-The-Creative-Wordsmith-31B-it-uncensored-hereticI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
Gemma-4-Gemsicle-31BI1-Q4_131.3B18.14 GiB3.14 GiB22.17 GiB0.15 GiB34±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation18.14 it/s13.3722.67393
Benchmarked· n=393

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 3090 Ti run?
1947 of 2118 indexed open-weight models fit a GeForce RTX 3090 Ti at 65,536 context with q4_0 KV cache, the largest being EuroLLM-22B-Instruct-2512 at Q6_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3090 Ti actually have?
Its nameplate is 24 GB, but about 22.32 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3090 Ti fast for local AI?
Its memory bandwidth is 1008 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.