NVIDIA · consumer

GeForce RTX 3060 Ti

GeForce RTX 3060 Ti has 8 GB of VRAM at 448 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1180 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
448 GB/s
256-bit bus
Tensor FP16
65 TF
dense
TDP
200 W
$399 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 993vision language 94embedding 26audio asr 38video 8image 1audio tts 20

What fits at 64K context

largest quantization that fits, per model · 1180 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
granite-8b-code-instruct-4kI1-IQ4_XS8.1B4.07 GiB2.53 GiB7.44 GiB0.00 GiB48±12.9%
granite-8b-code-base-4kI1-IQ4_XS8.1B4.07 GiB2.53 GiB7.44 GiB0.00 GiB48±12.9%
OmniAtlas-Qwen3-30B-A3BI1-IQ1_M31.7B6.59 GiB0.00 GiB7.44 GiB0.00 GiB48±12.9%
Qwen3-Omni-30B-A3B-CaptionerI1-IQ1_M31.7B6.59 GiB0.00 GiB7.44 GiB0.00 GiB48±12.9%
Grug-12BQ3_K_S12.0B5.33 GiB1.26 GiB7.43 GiB0.01 GiB48±12.9%
gemma-4-12B-it-Esper4Q3_K_S12.0B5.33 GiB1.26 GiB7.43 GiB0.01 GiB48±12.9%
gemma-4-12B-itQ3_K_S12.0B5.33 GiB1.26 GiB7.43 GiB0.01 GiB48±12.9%
salamandra-7b-instruct-2606I1-Q4_K_S7.8B4.35 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
stable-code-3bI1-Q2_K2.8B1.01 GiB5.63 GiB7.43 GiB0.01 GiB48±12.9%
saiga_llama3_8bQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
MiniCPM-Llama3-V-2_5Q4_08.5B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Llama3-ChatQA-1.5-8BQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Mistral-7B-v0.1KV unresolvedQ3_K_S7.2B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Llama-3-Groq-8B-Tool-UseQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Dolphin3.0-Llama3.1-8B-abliteratedQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Llama-3.1-8BQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Llama-3.1-8B-Lexi-Uncensored-V2Q4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Llama-3.1-Swallow-8B-Instruct-v0.5Q4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Turkish-Llama-8b-Instruct-v0.1Q4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Infinity-Instruct-7M-Gen-Llama3_1-8BQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Meta-Llama-3-8B-InstructQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
openchat-3.6-8b-20240522Q4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
L3-8B-Stheno-v3.3-32KQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
NeuralDaredevil-8B-abliteratedQ4_08.0B4.34 GiB2.25 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-9B-CoderI1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Qwythos-9B-Claude-Mythos-5-1M-MTPI1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedI1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-9B-Fable-5-v1I1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Qwythos-9B-v2I1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
PINQWEN-3.5-9B-1M-BF16I1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Openprose-2-FlashI1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-9B-Nikusui-v1I1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Ornstein-3.5-9B-V1.5I1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Ornith-1.0-9B-heretic-MTPI1-Q5_K_S9.4B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Tess-4-9BI1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
dotwebs-1I1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
liftQ5_09.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Hemlock-Qwopus3.5-9B-CoderI1-Q5_K_S9.7B6.03 GiB0.56 GiB7.43 GiB0.01 GiB48±12.9%
Vero-Qwen35-9B-BaseI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Vero-Qwen35-9BI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3-Embedding-8BIQ4_XS7.6B4.06 GiB2.53 GiB7.42 GiB0.02 GiB48±12.9%
Octen-Embedding-8BIQ4_XS7.6B4.06 GiB2.53 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Morphos-9BI1-Q5_K_M9.0B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwable-9B-Claude-Fable-5-hereticI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Holo-3.1-9BI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwable-9B-Claude-Fable-5I1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3.5-9B-imabari-v2I1-Q5_K_M9.7B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3.5-9B-abliterated-v2-MAXI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwable-9B-Claude-Fable-5-StraTAI1-Q5_K_M9.0B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q5_K_M9.0B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3.5-9B-RpRMax-v1I1-Q5_K_M9.7B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
AdQWENistrator-9BI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
cajal-9b-v2-fullI1-Q5_K_M9.0B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3.5-9B-ultra-uncensored-hereticQ5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Holo-3.1-9B-CoderI1-Q5_K_M9.0B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
PlutoI1-Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Holo-3.1-9B-abliterated-rdoI1-Q5_K_M9.0B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
Qwen3.5-9B-Uncensored-cyber-v3Q5_K_M9.4B6.02 GiB0.56 GiB7.42 GiB0.02 GiB48±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation8.53 it/s7.029.83916
Benchmarked· n=916

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 3060 Ti run?
1180 of 2118 indexed open-weight models fit a GeForce RTX 3060 Ti at 65,536 context with q4_0 KV cache, the largest being granite-8b-code-instruct-4k at I1-IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3060 Ti actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3060 Ti fast for local AI?
Its memory bandwidth is 448 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.