Model comparison

codegeex4-all-9b vs Qwen3-8B

At Q4_K_M, Qwen3-8B is the smaller download — 5,027,783,488 bytes against 6,250,923,136. At long context the gap widens: Qwen3-8B's KV cache at 32K is 4.4× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

codegeex4-all-9bQwen3-8B
Parameters9.4B8.2B
Architecturechatglmqwen3
Layers4036
Native context40,960
Mixture of expertsnono
Quantizations published5351
Smallest quantization2.89 GiB2.12 GiB
Q4_K_M5.82 GiB4.68 GiB
Licenceotherapache-2.0

KV cache by context

the term that decides long-context viability
Contextcodegeex4-all-9bQwen3-8BRatio
4,0962.50 GiB0.56 GiB4.44×
8,1925.00 GiB1.13 GiB4.44×
16,38410.00 GiB2.25 GiB4.44×
32,76820.00 GiB4.50 GiB4.44×
65,53640.00 GiB9.00 GiB4.44×
131,07280.00 GiB18.00 GiB4.44×