Model comparison

stable-code-3b vs gemma-3-1b-it

At Q4_K_M, gemma-3-1b-it is the smaller download — 806,058,240 bytes against 1,708,595,200. At long context the gap widens: gemma-3-1b-it's KV cache at 32K is 68.3× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

stable-code-3bgemma-3-1b-it
Parameters2.8B1000M
Architecturestablelmgemma3
Layers3226
Native context16,38432,768
Mixture of expertsnono
Quantizations published4828
Smallest quantization0.63 GiB0.52 GiB
Q4_K_M1.59 GiB0.75 GiB
Licenceothergemma

KV cache by context

the term that decides long-context viability
Contextstable-code-3bgemma-3-1b-itRatio
4,0961.25 GiB0.04 GiB33.68×
8,1922.50 GiB0.05 GiB47.41×
16,3845.00 GiB0.08 GiB59.53×
32,76810.00 GiB0.15 GiB68.27×
65,53620.00 GiB0.27 GiB73.67×
131,07240.00 GiB0.52 GiB76.70×