Model comparison

gemma-3-4b-it-abliterated vs Qwen2.5-VL-7B-Instruct

At Q4_K_M, gemma-3-4b-it-abliterated is the smaller download — 2,489,894,304 bytes against 4,683,072,032. At long context the gap widens: gemma-3-4b-it-abliterated's KV cache at 32K is 2.2× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

gemma-3-4b-it-abliteratedQwen2.5-VL-7B-Instruct
Parameters4.3B8.3B
Architecturegemma3qwen2vl
Layers3428
Native context131,072128,000
Mixture of expertsnono
Quantizations published3550
Smallest quantization1.43 GiB1.93 GiB
Q4_K_M2.32 GiB4.36 GiB
Licencegemmaapache-2.0

KV cache by context

the term that decides long-context viability
Contextgemma-3-4b-it-abliteratedQwen2.5-VL-7B-InstructRatio
4,0960.25 GiB0.22 GiB1.13×
8,1920.33 GiB0.44 GiB1.34×
16,3840.48 GiB0.88 GiB1.81×
32,7680.79 GiB1.75 GiB2.20×
65,5361.42 GiB3.50 GiB2.46×
131,0722.67 GiB7.00 GiB2.62×