NVIDIA · consumer
GeForce RTX 3080 Ti
GeForce RTX 3080 Ti has 20 GB of VRAM at 760 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1798 of 2118 indexed models fit at 64K context with q8_0 KV.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6X
Bandwidth
760 GB/s
320-bit bus
Tensor FP16
136 TF
dense
TDP
350 W
$1199 MSRP
text 1524vision language 170audio tts 21image 2video 16audio asr 39embedding 26
What fits at 64K context
largest quantization that fits, per model · 1798 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| GLM-4.7-FlashMoE | Q4_0 | 31.2B | 16.03 GiB | 1.76 GiB | 18.60 GiB | 0.00 GiB | 82±37% |
| GLM-4.7-Flash-hereticMoE | IQ4_NL | 29.9B | 16.03 GiB | 1.76 GiB | 18.60 GiB | 0.00 GiB | 82±37% |
| Gemma4-Gutenberg-31B | IQ2_M | 31.3B | 11.78 GiB | 5.94 GiB | 18.60 GiB | 0.00 GiB | 31±12.9% |
| gemma-4-31B-it | IQ2_M | 31.3B | 11.78 GiB | 5.94 GiB | 18.60 GiB | 0.00 GiB | 31±12.9% |
| Gemma4-Gutenberg-31B-Heretic | IQ2_M | 31.3B | 11.78 GiB | 5.94 GiB | 18.60 GiB | 0.00 GiB | 31±12.9% |
| Equinox-31B | IQ2_M | 31.3B | 11.78 GiB | 5.94 GiB | 18.60 GiB | 0.00 GiB | 31±12.9% |
| gemma-4-31B-it-SDFT-Heretic-RP | IQ2_M | 30.7B | 11.78 GiB | 5.94 GiB | 18.60 GiB | 0.00 GiB | 31±12.9% |
| dolphin-2.9.3-mistral-7B-32k | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Mistral-7B-Instruct-v0.3-Parasite | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Mistral-7B-Instruct-v0.3-Jbliterated | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Mistral-7B-Instruct-v0.3 | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Mistral-7B-v0.3 | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Mistral-7B-v0.3-Chinese-Chat | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| mistral-7b-v0.3-bnb-4bit | BF16 | 7.5B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Mathstral-7B-v0.1 | F16 | 7.2B | 13.50 GiB | 4.25 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Tess-4-9B | F16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Crow-9B-HERETIC-4.6 | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSORED | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCT | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwable-9B-Claude-Fable-5-heretic | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Holo-3.1-9B | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwable-9B-Claude-Fable-5 | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwythos-9B-Claude-Mythos-5-1M-uncensored-heretic | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-ultra-uncensored-heretic | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-abliterated-v2-MAX | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwable-9B-Claude-Fable-5-StraTA | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Uncensored-cyber-v3 | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| NaNovel-9B | F16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| OmniCoder-9B-Claude-Opus-High-Reasoning-Distill | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Fable-5-v1 | BF16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwable-9B-Claude-Fable-5-OBLITERATED | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Unredacted-MAX | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Huihui-Qwen3.5-9B-abliterated | F16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-abliterated | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Pluto | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| QwenPaw-Flash-9B | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Holo-3.1-9B-Coder | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Holo-3.1-9B-abliterated-rdo | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| MaralGPT-Mythos-9B-2606 | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| qwen3.5-9b-nsfw-captioning-v5 | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Fara1.5-9B | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| grug-9b | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| OmniCoder-9B | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Base | F16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-DS-v4-Flash-v3.0 | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Star-Trek-TNG-DS9-Heretic-Uncensored-Thinking | BF16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-heretic-v2 | F16 | 9.4B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Claude-Opus-4.6-Distill | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Miss_MARTHA-9B-Qwen3.5-Omni | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-DeepSeek-V4-Flash | F16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Huihui-Qwen3.5-9B-Claude-4.6-Opus-abliterated | F16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.5-9B-Neo | BF16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwopus3.5-9B-v3.5 | BF16 | 9.7B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Katarau-9B-ru-RP-nsfw | F16 | 9.0B | 16.69 GiB | 1.06 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| t5-v1_1-xxl | F32 | 4.8B | 17.74 GiB | 0.00 GiB | 18.59 GiB | 0.01 GiB | 31±12.9% |
| Qwen3.6-27B-Fable-5-Experimental | IQ4_NL | 27.8B | 15.59 GiB | 2.13 GiB | 18.58 GiB | 0.02 GiB | 31±12.9% |
| OpenChat-3.5-7B-Qwen-v2.0KV unresolved | F16 | 7.2B | 13.49 GiB | 4.25 GiB | 18.58 GiB | 0.02 GiB | 31±12.9% |
| ContextualKunoichi_KTO-7B | F16 | 7.2B | 13.49 GiB | 4.25 GiB | 18.58 GiB | 0.02 GiB | 31±12.9% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a GeForce RTX 3080 Ti run?
- 1798 of 2118 indexed open-weight models fit a GeForce RTX 3080 Ti at 65,536 context with q8_0 KV cache, the largest being GLM-4.7-Flash at Q4_0. That covers text, vision-language, image, video and speech models.
- How much usable memory does a GeForce RTX 3080 Ti actually have?
- Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a GeForce RTX 3080 Ti fast for local AI?
- Its memory bandwidth is 760 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.