Best models for long context
At long context the KV cache often exceeds the weights. These models keep it small through sliding-window attention, latent attention, or hybrid linear-attention layers.
From the file· live filter over real dataFrom the file· 40 models
How this is ranked
Ranked by KV cache bytes per token, computed per layer from the real architecture. Lower is better, and the spread between models is far larger than most people expect.
Best models for long context
Best local models for codingBest local reasoning modelsBest models for an 8GB cardBest permissively licensed modelsBest models under 4GBLargest models with a runnable quantizationBest local speech recognition modelsBest local text-to-speech modelsBest local vision-language modelsBest local image generation models