Model comparison

Solar-Open2-250B vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 95.3× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

Solar-Open2-250BDeepSeek-V4-Flash
Parameters250B291B
Architecturesolar-open2deepseek4
Layers4843
Native context1,048,5761,048,576
Mixture of expertsyes, 320 expertsyes, 256 experts
Quantizations published1312
Smallest quantization52.67 GiB76.87 GiB
Q4_K_M141.39 GiB
Licenceothermit

KV cache by context

the term that decides long-context viability
ContextSolar-Open2-250BDeepSeek-V4-FlashRatio
4,0960.75 GiB0.06 GiB11.91×
8,1921.50 GiB0.06 GiB23.81×
16,3843.00 GiB0.06 GiB47.63×
32,7686.00 GiB0.06 GiB95.26×
65,53612.00 GiB0.06 GiB190.51×
131,07224.00 GiB0.06 GiB381.02×
Solar-Open2-250B vs DeepSeek-V4-Flash — size, memory and hardware fit — ossmodeldb