Model comparison

Qwen3.5-99B vs gpt-oss-20b

At Q4_K_M, gpt-oss-20b is the smaller download — 11,624,759,488 bytes against 60,194,980,672.

From the file· summed bytes, KV per layer

Side by side

Qwen3.5-99Bgpt-oss-20b
Parameters99.0B21.5B
Architectureqwen35moegpt-oss
Layers4824
Native context262,144131,072
Mixture of expertsyes, 205 expertsyes, 32 experts
Quantizations published3615
Smallest quantization19.31 GiB10.68 GiB
Q4_K_M56.06 GiB10.83 GiB
Licenceotherapache-2.0

KV cache by context

the term that decides long-context viability
ContextQwen3.5-99Bgpt-oss-20bRatio
4,0960.09 GiB0.11 GiB1.19×
8,1920.19 GiB0.21 GiB1.09×
16,3840.38 GiB0.39 GiB1.05×
32,7680.75 GiB0.77 GiB1.02×
65,5361.50 GiB1.52 GiB1.01×
131,0723.00 GiB3.02 GiB1.01×