HuggingFaceTB · text

SmolVLM2-256M-Video-Instruct

HuggingFaceTB/SmolVLM2-256M-Video-Instruct

SmolVLM2-256M-Video-Instruct at I1-IQ1_S is exactly 98,386,880 bytes (0.09 GiB / 0.10 GB) — an effective 3.069 bits per weight, not the nominal 1. Its KV cache at 32K is 0.70 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
256M
Architecture
llama
30 layers
Context
8,192
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S0.09 GiB98,386,8803.069mradermacher
I1-IQ1_M0.09 GiB98,946,7523.086mradermacher
I1-IQ2_XXS0.09 GiB99,879,8723.115mradermacher
I1-IQ2_XS0.09 GiB100,626,3683.139mradermacher
I1-IQ2_S0.09 GiB100,895,9363.147mradermacher
I1-IQ2_M0.09 GiB101,642,4323.170mradermacher
I1-Q2_K_S0.10 GiB102,181,5683.187mradermacher
I1-IQ3_XXS0.10 GiB103,011,0083.213mradermacher
I1-Q2_K0.10 GiB104,255,1683.252mradermacher
I1-Q3_K_S0.10 GiB104,255,1683.252mradermacher
I1-IQ3_XS0.10 GiB104,255,1683.252mradermacher
I1-IQ3_S0.10 GiB104,255,1683.252mradermacher
I1-IQ3_M0.10 GiB106,266,5603.315mradermacher
I1-IQ4_XS0.10 GiB106,950,8483.336mradermacher
I1-IQ4_NL0.10 GiB107,780,2883.362mradermacher
I1-Q4_00.10 GiB107,946,1763.367mradermacher
I1-Q3_K_M0.10 GiB109,563,5843.417mradermacher
I1-Q3_K_L0.11 GiB113,586,3683.543mradermacher
I1-Q4_10.11 GiB116,189,8883.624mradermacher
I1-Q4_K_S0.11 GiB121,641,1523.794mradermacher
I1-Q4_K_M0.12 GiB125,055,6803.901mradermacher
I1-Q5_K_S0.12 GiB131,350,2084.097mradermacher
I1-Q5_K_M0.12 GiB133,479,1044.163mradermacher
I1-Q6_K0.16 GiB168,628,9285.260mradermacher
Q8_00.16 GiB175,056,3525.460ggml-org
F160.31 GiB327,811,55210.225ggml-org

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.09 GiB0.09 GiB30 / 0 / 0
8,1920.18 GiB0.18 GiB30 / 0 / 0
16,3840.35 GiB0.35 GiB30 / 0 / 0
32,7680.70 GiB0.70 GiB30 / 0 / 0
65,5361.41 GiB1.41 GiB30 / 0 / 0
131,0722.81 GiB2.81 GiB30 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts I1-IQ1_S at roughly 0.13 GiB. The real file is 0.09 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
30
Attention heads
9
KV heads
3
Head dim
64
Hidden size
576
Vocab
49,280
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does SmolVLM2-256M-Video-Instruct need?
I1-IQ1_S is exactly 98,386,880 bytes (0.09 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is SmolVLM2-256M-Video-Instruct's KV cache?
0.70 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of SmolVLM2-256M-Video-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.
SmolVLM2-256M-Video-Instruct — VRAM requirements, exact quant sizes — ossmodeldb