Model comparison

GLM-4.7-REAP-218B-A32B vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 182.6× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

GLM-4.7-REAP-218B-A32BDeepSeek-V4-Flash
Parameters218B291B
Architectureglm4moedeepseek4
Layers9243
Native context202,7521,048,576
Mixture of expertsyes, 96 expertsyes, 256 experts
Quantizations published4512
Smallest quantization44.60 GiB76.87 GiB
Q4_K_M122.98 GiB
Licencemitmit

KV cache by context

the term that decides long-context viability
ContextGLM-4.7-REAP-218B-A32BDeepSeek-V4-FlashRatio
4,0961.44 GiB0.06 GiB22.82×
8,1922.88 GiB0.06 GiB45.64×
16,3845.75 GiB0.06 GiB91.29×
32,76811.50 GiB0.06 GiB182.57×
65,53623.00 GiB0.06 GiB365.15×
131,07246.00 GiB0.06 GiB730.29×