Model comparison

DeepSeek-V4-Flash vs Qwen3.5-REAP-212B-A17B

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 14.9× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

DeepSeek-V4-FlashQwen3.5-REAP-212B-A17B
Parameters291B212B
Architecturedeepseek4qwen35moe
Layers4360
Native context1,048,576262,144
Mixture of expertsyes, 256 expertsyes, 267 experts
Quantizations published128
Smallest quantization76.87 GiB52.41 GiB
Q4_K_M119.56 GiB
Licencemitapache-2.0

KV cache by context

the term that decides long-context viability
ContextDeepSeek-V4-FlashQwen3.5-REAP-212B-A17BRatio
4,0960.06 GiB0.12 GiB1.86×
8,1920.06 GiB0.23 GiB3.72×
16,3840.06 GiB0.47 GiB7.44×
32,7680.06 GiB0.94 GiB14.89×
65,5360.06 GiB1.88 GiB29.77×
131,0720.06 GiB3.75 GiB59.54×