NVIDIA · consumer

GeForce RTX 3050

GeForce RTX 3050 has 6 GB of VRAM at 168 GB/s — about 5.58 GiB usable after driver and compositor overhead. 937 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
168 GB/s
96-bit bus
Tensor FP16
27 TF
dense
TDP
70 W
$179 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 771audio asr 38audio tts 19vision language 81embedding 25video 3

What fits at 64K context

largest quantization that fits, per model · 937 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
granite-3.3-8b-instructUD-IQ1_M8.2B1.93 GiB2.81 GiB5.57 GiB0.01 GiB26±12.9%
gemma-4-E4B-uncensoredI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma-4-E4B-it-qat-heretic_decensoredI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma4-e4b-mahou-nsfwI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma-4-E4B-it-mentalchat16kI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma4-E4B-it-abliteratedI1-Q3_K_M7.9B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
gemma-4-E4B-it-OBLITERATEDI1-Q3_K_M8.0B4.49 GiB0.27 GiB5.57 GiB0.01 GiB25±12.9%
Teuken-7B-instruct-research-v0.4I1-Q4_07.5B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
canary-qwen-2.5bBF162.6B4.73 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
EXAONE-Deep-7.8BQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
EXAONE-3.5-7.8B-InstructQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
Qwen3-TTS-12Hz-0.6B-BaseQ4_K_M915M4.72 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
VoxCPM2F162.3B4.72 GiB0.00 GiB5.57 GiB0.01 GiB26±12.9%
Qwen3.5-9B-CoderI1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Qwythos-9B-Claude-Mythos-5-1M-MTPI1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedI1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Qwen3.5-9B-Fable-5-v1I1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Qwythos-9B-v2I1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
PINQWEN-3.5-9B-1M-BF16I1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Openprose-2-FlashI1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Qwen3.5-9B-Nikusui-v1I1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Ornstein-3.5-9B-V1.5I1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Ornith-1.0-9B-heretic-MTPI1-IQ3_S9.4B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Tess-4-9BI1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
dotwebs-1I1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
Hemlock-Qwopus3.5-9B-CoderI1-IQ3_S9.7B4.17 GiB0.56 GiB5.57 GiB0.01 GiB26±12.9%
gemma-4-E4B-it-hereticQ4_K_S8.0B4.48 GiB0.27 GiB5.56 GiB0.02 GiB26±12.9%
Luna-7B-A4BMoEI1-IQ2_M6.7B2.22 GiB2.53 GiB5.56 GiB0.02 GiB19±37%
t5-v1_1-xxlQ2_K4.8B4.72 GiB0.00 GiB5.56 GiB0.02 GiB26±12.9%
Qwen3-VL-4B-Instruct-Unredacted-MAXI1-Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-VL-4B-ThinkingQ4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-VL-4B-Thinking-Unredacted-MAXI1-Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Zubr1.0-VL-4BI1-Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-VL-4B-Instruct-Uncensored-abliteratedQ4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
PopiT-Qwen3-4B-Medical-SFT-1128Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Huihui-Qwen3-VL-4B-Instruct-abliteratedI1-Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-VL-4B-Instruct-UncensoredI1-Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-VL-4B-InstructQ4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
OpenCaption-4B-VL-SFT-v1.0I1-Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Parable-Qwen3-4B-Claude-Fable-5I1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Logics-Parsing-v2Q4_K_S4.4B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Jan-v1-4BQ4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-4b-Z-Image-Turbo-AbliteratedV1I1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Z-Image-Engineer-V6Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-4BQ4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Huihui-Qwen3-4B-Instruct-2507-abliteratedQ4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Josiefied-Qwen3-4B-abliterated-v2Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-4B-Thinking-2507Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Jan-nano-128kQ4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-4B-Instruct-2507Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Jan-nanoQ4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-4B-abliteratedQ4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Neuron-4B-InstructI1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
ChineseErrorCorrector4-4BI1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
FastContext-1.0-4B-SFT-abliteratedI1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
Qwen3-4B-Instruct_NSFW-V2.1I1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
FastContext-1.0-4B-SFTI1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
fable-traces-abliteratedI1-Q4_K_S4.0B2.22 GiB2.53 GiB5.56 GiB0.02 GiB25±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation0.31 it/s0.212.479
Benchmarked· n=9

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 3050 run?
937 of 2118 indexed open-weight models fit a GeForce RTX 3050 at 65,536 context with q4_0 KV cache, the largest being granite-3.3-8b-instruct at UD-IQ1_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3050 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3050 fast for local AI?
Its memory bandwidth is 168 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.