HuggingFaceTB · text

SmolVLM2-2.2B-Instruct

HuggingFaceTB/SmolVLM2-2.2B-Instruct

SmolVLM2-2.2B-Instruct at Q4_K_M is exactly 1,112,602,656 bytes (1.04 GiB / 1.11 GB) — an effective 3.962 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
2.2B
Architecture
llama
24 layers
Context
8,192
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S0.42 GiB445,612,7681.587mradermacher
I1-IQ1_M0.44 GiB477,463,2641.700mradermacher
I1-IQ2_XXS0.49 GiB530,547,4241.889mradermacher
I1-IQ2_XS0.54 GiB576,160,4802.051mradermacher
I1-IQ2_S0.57 GiB615,901,9202.193mradermacher
I1-IQ2_M0.61 GiB658,369,2482.344mradermacher
I1-Q2_K_S0.61 GiB658,377,4402.344mradermacher
Q2_K0.66 GiB707,921,9842.521second-state
I1-Q2_K0.66 GiB707,922,6562.521mradermacher
I1-IQ3_XXS0.67 GiB723,643,1042.577mradermacher
I1-IQ3_XS0.73 GiB782,660,3202.787mradermacher
Q3_K_S0.76 GiB820,408,3842.921second-state
I1-Q3_K_S0.76 GiB820,409,0562.921mradermacher
I1-IQ3_S0.76 GiB820,409,0562.921mradermacher
I1-IQ3_M0.80 GiB853,832,4163.040mradermacher
Q3_K_M0.84 GiB903,770,1763.218second-state
I1-Q3_K_M0.84 GiB903,770,8483.218mradermacher
Q3_K_L0.91 GiB976,121,9203.476second-state
I1-Q3_K_L0.91 GiB976,122,5923.476mradermacher
I1-IQ4_XS0.93 GiB994,237,1523.540mradermacher
Q4_00.98 GiB1,047,722,0483.731second-state
I1-IQ4_NL0.98 GiB1,047,722,7203.731mradermacher
I1-Q4_00.98 GiB1,050,868,4483.742mradermacher
Q4_K_S0.98 GiB1,056,110,6563.760second-state
I1-Q4_K_S0.98 GiB1,056,111,3283.760mradermacher
Q4_K_M1.04 GiB1,112,602,6563.962ggml-org
Q4_K_M1.04 GiB1,112,602,6883.962second-state
I1-Q4_K_M1.04 GiB1,112,603,3603.962mradermacher
I1-Q4_11.08 GiB1,154,693,8564.112mradermacher
Q5_01.18 GiB1,261,664,3204.492second-state
Q5_K_S1.18 GiB1,261,664,3204.492second-state
I1-Q5_K_S1.18 GiB1,261,664,9924.492mradermacher
Q5_K_M1.21 GiB1,295,087,6804.611second-state
I1-Q5_K_M1.21 GiB1,295,088,3524.611mradermacher
Q6_K1.39 GiB1,488,977,9845.302second-state
I1-Q6_K1.39 GiB1,488,978,6565.302mradermacher
Q8_01.80 GiB1,927,933,9846.865ggml-org
Q8_01.80 GiB1,927,934,0166.865second-state
F163.38 GiB3,627,118,62412.915ggml-org
F163.38 GiB3,627,118,65612.915second-state

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.18 GiB. The real file is 1.04 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
24
Attention heads
KV heads
Head dim
64
Hidden size
2048
Vocab
49,280
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does SmolVLM2-2.2B-Instruct need?
Q4_K_M is exactly 1,112,602,656 bytes (1.04 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of SmolVLM2-2.2B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.