NVIDIA · consumer

GeForce RTX 3080 Laptop

GeForce RTX 3080 Laptop has 16 GB of VRAM at 448 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1673 of 2118 indexed models fit at 128K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
448 GB/s
256-bit bus
Tensor FP16
dense
TDP
150 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 152text 1419audio asr 39video 15embedding 26audio tts 21image 1

What fits at 128K context

largest quantization that fits, per model · 1673 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
gemma-3-12b-it-ultra-uncensored-hereticQ8_012.2B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingQ8_012.2B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
Floppa-12B-Gemma3-UncensoredQ8_012.2B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
gemma-3-12b-it-hereticQ8_012.2B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
gemma-3-12b-it-abliteratedQ8_012.2B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
gemma-3-12b-it-abliterated-v2Q8_011.8B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
gemma-3-12b-itQ8_012.2B11.65 GiB2.38 GiB14.88 GiB0.00 GiB23±12.9%
GRM-2.6-Plus-0628IQ3_XXS27.8B11.76 GiB2.25 GiB14.87 GiB0.01 GiB23±12.9%
ThinkingCap-Qwen3.6-27BIQ3_XXS27.4B11.76 GiB2.25 GiB14.87 GiB0.01 GiB23±12.9%
Tess-4-27BIQ3_XXS27.8B11.76 GiB2.25 GiB14.87 GiB0.01 GiB23±12.9%
NousCoder-14BQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
spoomplesmaxx-mini-14BI1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
vanilla-cn-roleplay-0.2I1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Claria-14bI1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
qwen3-14b-code-reasoning-conversationalQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
NTX-2.1-ProI1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Qwen3-14B-UncensoredI1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Qwen3-14B-Claude-4.5-Opus-High-Reasoning-DistillQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Qwen3-14BQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
FrogMini-14B-2510I1-Q4_K_M8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Qwen3-14B-abliteratedQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Josiefied-Qwen3-14B-abliterated-v3Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Hermes-4-14BQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Slava-Qwen3-14B-SerbianI1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Qwen3-14B-BaseQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Huihui-Qwen3-14B-abliterated-v2I1-Q4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Qwen3-14B-InstructQ4_K_M14.8B8.38 GiB5.63 GiB14.87 GiB0.01 GiB23±12.9%
Phi-3-medium-128k-instructQ3_K_L14.0B6.98 GiB7.03 GiB14.87 GiB0.01 GiB23±12.9%
Phi-3-medium-4k-instructI1-Q3_K_L14.0B6.98 GiB7.03 GiB14.87 GiB0.01 GiB23±12.9%
Pantheon-Reasoning-26B-A4B-1.1MoEQ3_K_L26.5B12.59 GiB1.49 GiB14.86 GiB0.02 GiB23±12.9%
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresyMoEI1-Q4_023.0B12.18 GiB1.86 GiB14.85 GiB0.03 GiB51±37%
Pantheon-Reasoning-27BI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-PreservedI1-IQ3_S27.4B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwen3.6-27B-Fable-5-ExperimentalI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwable-5-27B-CoderI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
EVE-27b-XENO-HAT-DeepSeek-V4-FlashI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
EVE-27B-XENO-HATI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Godoter-27BI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Reasoning-Medical-27BI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwopus3.6-27B-v2-abliteratedI1-IQ3_S27.4B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-BF16I1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Reasoning-Medical0.1-27BI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Huihui-ThinkingCap-Qwen3.6-27B-abliteratedI1-IQ3_S27.4B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Semancer-27BI1-IQ3_S27.8B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Darwin-28B-CoderI1-IQ3_S26.9B11.74 GiB2.25 GiB14.85 GiB0.03 GiB23±12.9%
Qwen3-VL-8B-Instruct-HereticI1-Q4_K_S8.8B8.94 GiB5.06 GiB14.84 GiB0.04 GiB23±12.9%
Huihui-gemma-4-31B-it-abliterated-v2UD-IQ2_XXS32.7B8.00 GiB5.95 GiB14.84 GiB0.04 GiB23±12.9%
Rocinante-XL-16B-v1I1-IQ3_XS16.1B6.39 GiB7.59 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.5-27B-Engineer-Deckard-GeminiI1-IQ3_M27.7B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.5-27B-HERETIC-Polaris-Advanced-Thinking-Alpha-uncensoredI1-IQ3_M27.4B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-ThinkingI1-IQ3_M27.4B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.6-27B-Heretic2-ThinkingI1-IQ3_M27.4B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.6-27B-Uncensored-AggressiveI1-IQ3_M27.4B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen-3.5-Opus-GLM-27BI1-IQ3_M26.9B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.6-27B-abliteratedI1-IQ3_M27.4B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
KoQweopus-3.5-27B-experimentalI1-IQ3_M27.8B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Webcoda-AI-27BI1-IQ3_M27.4B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Qwen3.5-27B-imabari-v2I1-IQ3_M27.8B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
Huihui-Qwen3.5-27B-abliteratedI1-IQ3_M27.8B11.72 GiB2.25 GiB14.83 GiB0.05 GiB23±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation8.89 it/s6.8810.56245
Benchmarked· n=245

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 3080 Laptop run?
1673 of 2118 indexed open-weight models fit a GeForce RTX 3080 Laptop at 131,072 context with q4_0 KV cache, the largest being gemma-3-12b-it-ultra-uncensored-heretic at Q8_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3080 Laptop actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3080 Laptop fast for local AI?
Its memory bandwidth is 448 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.