Model comparison

Hermes-4-405B vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 250.0× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

Hermes-4-405BDeepSeek-V4-Flash
Parameters406B291B
Architecturellamadeepseek4
Layers12643
Native context131,0721,048,576
Mixture of expertsnoyes, 256 experts
Quantizations published1612
Smallest quantization81.23 GiB76.87 GiB
Q4_K_M226.38 GiB
Licencellama3mit

KV cache by context

the term that decides long-context viability
ContextHermes-4-405BDeepSeek-V4-FlashRatio
4,0961.97 GiB0.06 GiB31.26×
8,1923.94 GiB0.06 GiB62.51×
16,3847.88 GiB0.06 GiB125.02×
32,76815.75 GiB0.06 GiB250.05×
65,53631.50 GiB0.06 GiB500.09×
131,07263.00 GiB0.06 GiB1000.19×