bigcode · text

starcoder2-3b

bigcode/starcoder2-3b

starcoder2-3b at Q4_K_M is exactly 1,848,976,448 bytes (1.72 GiB / 1.85 GB) — an effective 4.881 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
3.0B
Architecture
starcoder2
30 layers
Context
16,384
native (config.json)
License
bigcode-openrail-m

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K1.07 GiB1,149,187,1363.034second-state
Q2_K1.14 GiB1,223,505,5363.230DevQuasar-9
Q2_K1.14 GiB1,223,505,6003.230QuantFactory
Q3_K_S1.22 GiB1,307,554,8803.452second-state
Q3_K_S1.27 GiB1,366,537,8563.608DevQuasar-9
Q3_K_S1.27 GiB1,366,537,9203.608QuantFactory
Q3_K_M1.41 GiB1,513,047,1043.994second-state
Q3_K_M1.46 GiB1,562,592,8964.125DevQuasar-9
Q3_K_M1.46 GiB1,562,592,9604.125QuantFactory
Q3_K_L1.56 GiB1,678,591,0404.431second-state
Q4_01.59 GiB1,709,888,5764.514second-state
Q3_K_L1.62 GiB1,737,574,0164.587DevQuasar-9
Q3_K_L1.62 GiB1,737,574,0804.587QuantFactory
Q4_K_S1.62 GiB1,743,311,9364.602second-state
Q4_01.63 GiB1,748,817,6004.617QuantFactory
Q4_K_S1.64 GiB1,763,366,5284.655DevQuasar-9
Q4_K_S1.64 GiB1,763,366,5924.655QuantFactory
Q4_K_M1.72 GiB1,848,976,4484.881second-state
Q4_K_M1.76 GiB1,887,905,4084.984DevQuasar-9
Q4_K_M1.76 GiB1,887,905,4724.984QuantFactory
Q4_11.80 GiB1,928,713,9205.092QuantFactory
Q5_01.95 GiB2,088,555,5845.514second-state
Q5_K_S1.95 GiB2,088,555,5845.514second-state
Q5_K_S1.96 GiB2,108,610,1765.567DevQuasar-9
Q5_K_S1.96 GiB2,108,610,2405.567QuantFactory
Q5_01.96 GiB2,108,610,2405.567QuantFactory
Q5_K_M2.01 GiB2,160,206,9125.703second-state
Q5_K_M2.03 GiB2,180,261,5045.756DevQuasar-9
Q5_K_M2.03 GiB2,180,261,5685.756QuantFactory
Q5_12.13 GiB2,288,506,5606.042QuantFactory
Q6_K2.32 GiB2,490,889,2806.576second-state
Q6_K2.32 GiB2,490,889,8566.576DevQuasar-9
Q6_K2.32 GiB2,490,889,9206.576QuantFactory
Q8_03.00 GiB3,224,556,6088.513second-state
Q8_03.00 GiB3,224,557,1848.513DevQuasar-9
Q8_03.00 GiB3,224,557,2488.513QuantFactory
F165.65 GiB6,064,559,74416.010DevQuasar-9

KV cache by context

unresolved

This model declares a 4,096-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.

A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.59 GiB. The real file is 1.72 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
30
Attention heads
24
KV heads
2
Head dim
128
Hidden size
3072
Vocab
49,152
Sliding window
4096
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does starcoder2-3b need?
Q4_K_M is exactly 1,848,976,448 bytes (1.72 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of starcoder2-3b should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.