Model comparison

gemma-2-2b-it-abliterated vs Qwen3-8B

At Q4_K_M, gemma-2-2b-it-abliterated is the smaller download — 1,708,582,784 bytes against 5,027,783,488. At long context the gap widens: gemma-2-2b-it-abliterated's KV cache at 32K is 2.4× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

gemma-2-2b-it-abliteratedQwen3-8B
Parameters2.6B8.2B
Architecturegemma2qwen3
Layers2636
Native context8,19240,960
Mixture of expertsnono
Quantizations published2951
Smallest quantization1.15 GiB2.12 GiB
Q4_K_M1.59 GiB4.68 GiB
Licencegemmaapache-2.0

KV cache by context

the term that decides long-context viability
Contextgemma-2-2b-it-abliteratedQwen3-8BRatio
4,0960.41 GiB0.56 GiB1.38×
8,1920.63 GiB1.13 GiB1.77×
16,3841.04 GiB2.25 GiB2.16×
32,7681.85 GiB4.50 GiB2.43×
65,5363.48 GiB9.00 GiB2.59×
131,0726.73 GiB18.00 GiB2.68×