GeForce GTX 1080 Ti
GeForce GTX 1080 Ti has 11 GB of VRAM at 484 GB/s — about 10.23 GiB usable after driver and compositor overhead. 1519 of 2118 indexed models fit at 16K context with f16 KV.
What fits at 16K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Janus-Pro-7B | I1-IQ2_XXS | 7.4B | 1.90 GiB | 7.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| deepseek-coder-7b-instruct-v1.5 | I1-IQ2_XXS | 6.9B | 1.90 GiB | 7.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| reka-flash-3.1 | I1-IQ2_XS | 20.9B | 7.29 GiB | 2.06 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| reka-flash-3 | IQ2_XS | 20.9B | 7.29 GiB | 2.06 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| Fara1.5-9B | Q8_0 | 9.4B | 8.89 GiB | 0.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| QwenPaw-Flash-9B | Q8_0 | 9.4B | 8.89 GiB | 0.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| grug-9b | Q8_0 | 9.4B | 8.89 GiB | 0.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| OmniCoder-9B | Q8_0 | 9.4B | 8.89 GiB | 0.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| Qwen3.5-9B-Neo | Q8_0 | 9.7B | 8.89 GiB | 0.50 GiB | 10.23 GiB | 0.00 GiB | 37±12.9% |
| Bielik-11B-v2.3-Instruct | Q4_K_M | 11.2B | 6.26 GiB | 3.13 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| Darwin-35B-A3B-OpusMoE | IQ2_XXS | 36.0B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| Aurora-Code-1MoE | IQ2_XXS | 34.7B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| grug-35b-v2MoE | IQ2_XXS | 35.1B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| grug-35bMoE | IQ2_XXS | 35.1B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| WorldSim-Opus-3.6-35B-A3BMoE | IQ2_XXS | 35.1B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| Qwen3.6-35B-A3B-AnkoMoE | IQ2_XXS | 35.1B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| KAT-Coder-V2.5-DevMoE | IQ2_XXS | 34.7B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| Ornith-1.0-35BMoE | IQ2_XXS | 34.7B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| Nex-N2-miniMoE | IQ2_XXS | 35.1B | 9.11 GiB | 0.31 GiB | 10.22 GiB | 0.01 GiB | 159±37% |
| Ling-liteMoE | IQ4_XS | 16.8B | 8.55 GiB | 0.88 GiB | 10.22 GiB | 0.01 GiB | 89±37% |
| Snowpiercer-15B-v4-heretic | I1-Q3_K_S | 15.0B | 6.25 GiB | 3.13 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| Snowpiercer-15B-v4 | Q3_K_S | 15.0B | 6.25 GiB | 3.13 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| gemma-4-19B-A4B-it-INSTRUCT-Heretic-UncensoredMoE | I1-IQ3_M | 19.0B | 8.51 GiB | 0.92 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| gemma-4-19B-A4B-it-The-DECKARD-Heretic-Uncensored-ThinkingMoE | I1-IQ3_M | 19.0B | 8.51 GiB | 0.92 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| gemma-4-19b-a4b-it-REAP-hereticMoE | I1-IQ3_M | 19.0B | 8.51 GiB | 0.92 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| Gemma-4-19BMoE | I1-IQ3_M | 19.0B | 8.51 GiB | 0.92 GiB | 10.22 GiB | 0.01 GiB | 37±12.9% |
| Nemotron-Mini-4B-Instruct | Q3_K_L | 4.2B | 7.40 GiB | 2.00 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSORED | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCT | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwable-9B-Claude-Fable-5-heretic | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Holo-3.1-9B | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwable-9B-Claude-Fable-5 | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Fable-5-Quad-Stock | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-ultra-uncensored-heretic | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-abliterated-v2-MAX | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwable-9B-Claude-Fable-5-StraTA | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Uncensored-cyber-v3 | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| NaNovel-9B | Q8_0 | 9.7B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| OmniCoder-9B-Claude-Opus-High-Reasoning-Distill | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwable-9B-Claude-Fable-5-OBLITERATED | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Unredacted-MAX | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Huihui-Qwen3.5-9B-abliterated | Q8_0 | 9.7B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-abliterated | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Pluto | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Holo-3.1-9B-Coder | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Holo-3.1-9B-abliterated-rdo | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| MaralGPT-Mythos-9B-2606 | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-heretic-v2 | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| qwen3.5-9b-nsfw-captioning-v5 | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-abliterated | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Base | Q8_0 | 9.7B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9b-Sushi-Coder-RL | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-DS-v4-Flash-v3.0 | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Ornith-1.0-9B-Uncensored | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| YuYu1015-Ornith-1.0-9B-abliterated | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwythos-9B-v2-Heretic | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING-X8b | Q8_0 | 9.4B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Qwen3.5-9B-Claude-Opus-4.6-Distill | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
| Miss_MARTHA-9B-Qwen3.5-Omni | Q8_0 | 9.0B | 8.87 GiB | 0.50 GiB | 10.21 GiB | 0.02 GiB | 37±12.9% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Image generation | 3.19 it/s | 2.12–3.64 | 422 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.
Questions people ask
- What AI models can a GeForce GTX 1080 Ti run?
- 1519 of 2118 indexed open-weight models fit a GeForce GTX 1080 Ti at 16,384 context with f16 KV cache, the largest being Janus-Pro-7B at I1-IQ2_XXS. That covers text, vision-language, image, video and speech models.
- How much usable memory does a GeForce GTX 1080 Ti actually have?
- Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a GeForce GTX 1080 Ti fast for local AI?
- Its memory bandwidth is 484 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.