microsoft · text

GELab-Zero-4B-preview-Sico-Evolution

microsoft/GELab-Zero-4B-preview-Sico-Evolution

GELab-Zero-4B-preview-Sico-Evolution at Q4_K_M is exactly 2,497,282,464 bytes (2.33 GiB / 2.50 GB) — an effective 4.502 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
4.4B
Architecture
qwen3vl
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S0.98 GiB1,055,257,7921.902mradermacher
I1-IQ1_M1.05 GiB1,127,019,7122.032mradermacher
I1-IQ2_XXS1.16 GiB1,246,622,9122.247mradermacher
I1-IQ2_XS1.26 GiB1,354,101,9522.441mradermacher
I1-IQ2_S1.32 GiB1,417,303,2322.555mradermacher
I1-IQ2_M1.41 GiB1,512,985,7922.727mradermacher
I1-Q2_K_S1.46 GiB1,563,456,1922.818mradermacher
Q2_K1.55 GiB1,669,501,3443.010mradermacher
I1-Q2_K1.55 GiB1,669,501,6323.010mradermacher
I1-IQ3_XXS1.56 GiB1,670,190,2723.011mradermacher
I1-IQ3_XS1.69 GiB1,814,377,1523.271mradermacher
Q3_K_S1.76 GiB1,886,998,9443.402mradermacher
I1-Q3_K_S1.76 GiB1,886,999,2323.402mradermacher
I1-IQ3_S1.77 GiB1,899,532,9923.424mradermacher
I1-IQ3_M1.83 GiB1,962,898,1123.538mradermacher
Q3_K_M1.93 GiB2,075,619,7443.742mradermacher
I1-Q3_K_M1.93 GiB2,075,620,0323.742mradermacher
Q3_K_L2.09 GiB2,239,787,4244.038mradermacher
I1-Q3_K_L2.09 GiB2,239,787,7124.038mradermacher
I1-IQ4_XS2.11 GiB2,270,753,4724.093mradermacher
IQ4_XS2.13 GiB2,286,317,9844.122mradermacher
I1-Q4_02.21 GiB2,375,774,9124.283mradermacher
I1-IQ4_NL2.22 GiB2,381,345,4724.293mradermacher
Q4_K_S2.22 GiB2,383,311,2644.296mradermacher
I1-Q4_K_S2.22 GiB2,383,311,5524.296mradermacher
Q4_K_M2.33 GiB2,497,282,4644.502mradermacher
I1-Q4_K_M2.33 GiB2,497,282,7524.502mradermacher
I1-Q4_12.42 GiB2,596,631,2324.681mradermacher
Q5_K_S2.63 GiB2,823,713,1845.090mradermacher
I1-Q5_K_S2.63 GiB2,823,713,4725.090mradermacher
Q5_K_M2.69 GiB2,889,515,4245.209mradermacher
I1-Q5_K_M2.69 GiB2,889,515,7125.209mradermacher
Q6_K3.08 GiB3,306,262,9445.960mradermacher
I1-Q6_K3.08 GiB3,306,263,2325.960mradermacher
Q8_03.99 GiB4,280,406,9447.716mradermacher
F167.50 GiB8,051,286,94414.514mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.32 GiB. The real file is 2.33 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does GELab-Zero-4B-preview-Sico-Evolution need?
Q4_K_M is exactly 2,497,282,464 bytes (2.33 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of GELab-Zero-4B-preview-Sico-Evolution should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.