NVIDIA · consumer

GeForce RTX 3050

GeForce RTX 3050 has 6 GB of VRAM at 168 GB/s — about 5.58 GiB usable after driver and compositor overhead. 949 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
168 GB/s
96-bit bus
Tensor FP16
27 TF
dense
TDP
70 W
$179 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 782vision language 82audio asr 38audio tts 19embedding 25video 3

What fits at 32K context

largest quantization that fits, per model · 949 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3.5-9B-CoderI1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Qwythos-9B-Claude-Mythos-5-1M-MTPI1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedI1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Qwen3.5-9B-Fable-5-v1I1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Qwythos-9B-v2I1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
PINQWEN-3.5-9B-1M-BF16I1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Openprose-2-FlashI1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Qwen3.5-9B-Nikusui-v1I1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Ornstein-3.5-9B-V1.5I1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Ornith-1.0-9B-heretic-MTPI1-IQ3_M9.4B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Tess-4-9BI1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
dotwebs-1I1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
liftIQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
Hemlock-Qwopus3.5-9B-CoderI1-IQ3_M9.7B4.21 GiB0.53 GiB5.58 GiB0.00 GiB26±12.9%
gemma-4-E4B-uncensoredI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma-4-E4B-it-qat-heretic_decensoredI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma4-e4b-mahou-nsfwI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma-4-E4B-it-mentalchat16kI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma4-E4B-it-abliteratedI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
gemma-4-E4B-it-OBLITERATEDI1-Q3_K_M8.0B4.49 GiB0.27 GiB5.58 GiB0.00 GiB25±12.9%
MiMo-VL-7B-RLI1-IQ2_XS8.3B2.36 GiB2.39 GiB5.58 GiB0.00 GiB26±12.9%
Kuwutu-7B-CYOA-v2I1-IQ2_XS7.6B2.36 GiB2.39 GiB5.58 GiB0.00 GiB26±12.9%
Surogate-3.5-2BF162.8B4.58 GiB0.20 GiB5.57 GiB0.01 GiB25±12.9%
Aya-Medikal-V2I1-IQ2_XS8.0B2.60 GiB2.13 GiB5.57 GiB0.01 GiB26±12.9%
canary-qwen-2.5bBF162.6B4.73 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
EXAONE-Deep-7.8BQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
EXAONE-3.5-7.8B-InstructQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
Qwen3-TTS-12Hz-0.6B-BaseQ4_K_M915M4.72 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
VoxCPM2F162.3B4.72 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
gemma-4-E4B-it-hereticQ4_K_S8.0B4.48 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
orpheus-3b-0.1-pretrainedQ6_K3.8B2.90 GiB1.86 GiB5.57 GiB0.01 GiB25±12.9%
OLMoE-1B-7B-0924-InstructMoEI1-IQ3_XS6.9B2.67 GiB2.13 GiB5.56 GiB0.02 GiB27±37%
t5-v1_1-xxlQ2_K4.8B4.72 GiB0.00 GiB5.56 GiB0.02 GiB26±12.9%
granite-3.1-8b-instructTQ2_08.2B2.07 GiB2.66 GiB5.56 GiB0.02 GiB26±12.9%
Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING-X8bF169.4B4.19 GiB0.53 GiB5.56 GiB0.02 GiB26±12.9%
Ministral-3-8B-Instruct-2512UD-IQ2_XXS8.9B2.46 GiB2.26 GiB5.55 GiB0.03 GiB26±12.9%
Firefly-v4Q8_05.1B4.63 GiB0.13 GiB5.55 GiB0.03 GiB25±12.9%
gemma-4-E2B-it-Uncensored-MAXQ8_05.1B4.63 GiB0.13 GiB5.55 GiB0.03 GiB25±12.9%
gemma-4-E2B-it-uncensoredQ8_05.1B4.63 GiB0.13 GiB5.55 GiB0.03 GiB25±12.9%
gemma-4-E2B-it-abliteratedQ8_05.1B4.63 GiB0.13 GiB5.55 GiB0.03 GiB25±12.9%
gemma-4-E2B-it-heretic-araQ8_05.1B4.63 GiB0.13 GiB5.55 GiB0.03 GiB25±12.9%
gemma-4-E2BQ8_05.1B4.63 GiB0.13 GiB5.55 GiB0.03 GiB25±12.9%
Ministral-3-8B-Reasoning-2512UD-IQ2_XXS8.9B2.45 GiB2.26 GiB5.55 GiB0.03 GiB26±12.9%
LFM2-8B-A1BMoEQ4_K_S8.3B4.56 GiB0.20 GiB5.55 GiB0.03 GiB66±37%
qwen-indic-v1I1-IQ2_XS7.6B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
EVA-Yi-1.5-9B-32K-V1I1-Q2_K8.8B3.12 GiB1.59 GiB5.54 GiB0.04 GiB26±12.9%
Yi-Coder-9B-ChatQ2_K8.8B3.12 GiB1.59 GiB5.54 GiB0.04 GiB26±12.9%
Yi-1.5-9B-ChatQ2_K8.8B3.12 GiB1.59 GiB5.54 GiB0.04 GiB26±12.9%
next-8bI1-IQ2_XXS8.2B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Supertron2-Reranker-8BI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
next-ocrI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Qwen3-VL-8B-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Midas-FableAgent-8BI1-IQ2_XXS8.2B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Qwen3-VL-8B-Heretic-1.3.0I1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Qwen3-VL-8B-Thinking-Unredacted-MAXI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Qwen3-VL-8B-Instruct-Minecraft-MT-en-zhI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Qwen-3-VL-8B-Instruct-hereticI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
Poe-8B-GLM5-Opus4.6-Sonnet4.5-Kimi-Grok-Gemini-3-pro-preview-HERETICI1-IQ2_XXS8.8B2.32 GiB2.39 GiB5.54 GiB0.04 GiB26±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation0.31 it/s0.212.479
Benchmarked· n=9

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 3050 run?
949 of 2118 indexed open-weight models fit a GeForce RTX 3050 at 32,768 context with q8_0 KV cache, the largest being Qwen3.5-9B-Coder at I1-IQ3_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3050 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3050 fast for local AI?
Its memory bandwidth is 168 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.