Best models for long context

At long context the KV cache often exceeds the weights. These models keep it small through sliding-window attention, latent attention, or hybrid linear-attention layers.

From the file· live filter over real dataFrom the file· 40 models

How this is ranked

Ranked by KV cache bytes per token, computed per layer from the real architecture. Lower is better, and the spread between models is far larger than most people expect.

Best models for long context

#ModelParamsSmallest quantKV per token
1gemma-4-E4B-it-assistant79M0.07 GiB2.0 KiB
2gemma-3-270m-it268M0.17 GiB3.0 KiB
3functiongemma-270m-it268M0.17 GiB3.0 KiB
4gemma-3-270m268M0.22 GiB3.0 KiB
5gemma-3-270m-it-qat268M0.22 GiB3.0 KiB
6Qwen3.6-27B-DFlash1.7B0.60 GiB4.0 KiB
7Qwen3.6-35B-A3B-DFlash386M0.14 GiB4.0 KiB
8privacy-filter-nemotronMoE1.4B2.62 GiB4.0 KiB
9privacy-filter-multilingualMoE1.4B2.62 GiB4.0 KiB
10gemma-4-31B-it-DFlash1.5B0.54 GiB4.0 KiB
11Qwen3.5-9B-DFlash1.3B0.45 GiB4.0 KiB
12gemma-4-26B-A4B-it-DFlash430M0.16 GiB4.0 KiB
13Qwen3.5-4B-DFlash634M0.23 GiB4.0 KiB
14gemma4-12B-it-DFlash727M0.38 GiB4.0 KiB
15next-1b1000M0.75 GiB4.0 KiB
16gemma-4-E2B-it5.1B2.13 GiB7.0 KiB
17gemma-4-E2B-it-qat-q4_0-unquantized5.1B3.12 GiB7.0 KiB
18gemma-4-E2B-it-ultra-uncensored-heretic5.1B2.12 GiB7.0 KiB
19gemma-4-E2B-it-Uncensored-MAX5.1B2.78 GiB7.0 KiB
20gemma-4-E2B5.1B2.78 GiB7.0 KiB
21gemma-4-E2B-it-uncensored5.1B2.30 GiB7.0 KiB
22gemma-4-E2B-it-qat-q4_0-unquantized-heretic5.1B2.16 GiB7.0 KiB
23gemma-4-E2B-it-heretic-ara5.1B2.78 GiB7.0 KiB
24Gemma4_E2B_Abliterated_Baked_HF_Ready5.1B2.16 GiB7.0 KiB
25Miril-Drone-2B-15.1B2.43 GiB7.0 KiB
26Firefly-v45.1B2.78 GiB7.0 KiB
27BartaLens-E2B5.1B2.78 GiB7.0 KiB
28Huihui-gemma-4-E2B-it-qat-q4_0-unquantized-abliterated5.1B2.16 GiB7.0 KiB
29gemma-4-E2B-it-abliterated5.1B2.16 GiB7.0 KiB
30gemma-4-E2B-it5.1B2.13 GiB7.0 KiB
31Barcenas-E2B5.1B2.78 GiB7.0 KiB
32rp-model_E2B_v5.15.1B2.78 GiB7.0 KiB
33gemma-4-12B-it-qat-q4_0-unquantized-assistant423M0.30 GiB8.0 KiB
34gemma-4-26B-A4B-it-assistant420M0.30 GiB8.0 KiB
35gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant420M0.27 GiB8.0 KiB
36gemma-4-12B-it-assistant423M0.30 GiB8.0 KiB
37Qwen3.5-0.8B873M0.31 GiB12.0 KiB
38LFM2.5-1.2B-Instruct1.2B0.45 GiB12.0 KiB
39Qwen3.5-2B2.3B0.72 GiB12.0 KiB
40Qwen2.5-0.5B-Instruct494M0.31 GiB12.0 KiB
Spec sheetFrom the fileFrom the filewhat these mean