Model comparison

Inkling vs Kimi-K2.7-Code

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: Kimi-K2.7-Code's KV cache at 32K is 3.8× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

InklingKimi-K2.7-Code
Parameters952B1059B
Architectureinklingdeepseek2
Layers6661
Native context262,144
Mixture of expertsyes, 256 expertsyes, 384 experts
Quantizations published719
Smallest quantization210.69 GiB283.04 GiB
Q4_K_M
Licenceapache-2.0other

KV cache by context

the term that decides long-context viability
ContextInklingKimi-K2.7-CodeRatio
4,0961.03 GiB0.27 GiB3.85×
8,1922.06 GiB0.54 GiB3.85×
16,3844.13 GiB1.07 GiB3.85×
32,7688.25 GiB2.14 GiB3.85×
65,53616.50 GiB4.29 GiB3.85×
131,07233.00 GiB8.58 GiB3.85×