Model comparison

GLM-5.2 vs Qwen3-235B-A22B

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: GLM-5.2's KV cache at 32K is 2.1× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

GLM-5.2Qwen3-235B-A22B
Parameters753B235B
Architectureglm-dsaqwen3moe
Layers7894
Native context1,048,57640,960
Mixture of expertsyes, 256 expertsyes, 128 experts
Quantizations published2035
Smallest quantization169.33 GiB79.81 GiB
Q4_K_M132.39 GiB
Licencemitapache-2.0

KV cache by context

the term that decides long-context viability
ContextGLM-5.2Qwen3-235B-A22BRatio
4,0960.34 GiB0.73 GiB2.14×
8,1920.69 GiB1.47 GiB2.14×
16,3841.37 GiB2.94 GiB2.14×
32,7682.74 GiB5.88 GiB2.14×
65,5365.48 GiB11.75 GiB2.14×
131,07210.97 GiB23.50 GiB2.14×