Model comparison

dots.llm1.inst vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 492.2× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

dots.llm1.instDeepSeek-V4-Flash
Parameters143B291B
Architecturedots1deepseek4
Layers6243
Native context32,7681,048,576
Mixture of expertsyes, 128 expertsyes, 256 experts
Quantizations published4612
Smallest quantization44.04 GiB76.87 GiB
Q4_K_M88.00 GiB
Licencemitmit

KV cache by context

the term that decides long-context viability
Contextdots.llm1.instDeepSeek-V4-FlashRatio
4,0963.88 GiB0.06 GiB61.52×
8,1927.75 GiB0.06 GiB123.04×
16,38415.50 GiB0.06 GiB246.08×
32,76831.00 GiB0.06 GiB492.16×
65,53662.00 GiB0.06 GiB984.31×
131,072124.00 GiB0.06 GiB1968.62×