Can I run Kimi-K2.6 on a GeForce RTX 2080 Ti?
Not at these settings. No indexed quantization of Kimi-K2.6 fits GeForce RTX 2080 Ti at any context we compute, with f16 KV. The smallest shipped quantization is 193.16 GiB in weights alone, against 10.23 GiB usable. CPU offload can still run it, slowly.
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| BF16 | 1912.15 GiB | 1913.3 | 1913.6 | 1914.1 | 1915.2 | 1917.3 | 1921.6 |
| Q4_0 | 543.62 GiB | 544.8 | 545.0 | 545.6 | 546.6 | 548.8 | 553.1 |
| IQ3_M | 454.53 GiB | 455.7 | 455.9 | 456.5 | 457.6 | 459.7 | 464.0 |
| Q3_K_L | 454.05 GiB | 455.2 | 455.5 | 456.0 | 457.1 | 459.2 | 463.5 |
| Q3_K_M | 434.81 GiB | 436.0 | 436.2 | 436.8 | 437.8 | 440.0 | 444.3 |
| IQ3_XS | 434.41 GiB | 435.6 | 435.8 | 436.4 | 437.4 | 439.6 | 443.9 |
| Q3_K_S | 413.86 GiB | 415.0 | 415.3 | 415.8 | 416.9 | 419.0 | 423.3 |
| IQ3_XXS | 397.34 GiB | 398.5 | 398.8 | 399.3 | 400.4 | 402.5 | 406.8 |
| Q2_K_L | 334.66 GiB | 335.8 | 336.1 | 336.6 | 337.7 | 339.8 | 344.1 |
| Q2_K | 333.59 GiB | 334.7 | 335.0 | 335.5 | 336.6 | 338.8 | 343.0 |
| IQ2_M | 318.27 GiB | 319.4 | 319.7 | 320.2 | 321.3 | 323.4 | 327.7 |
| IQ2_S | 287.58 GiB | 288.7 | 289.0 | 289.5 | 290.6 | 292.8 | 297.0 |
| IQ2_XS | 282.35 GiB | 283.5 | 283.8 | 284.3 | 285.4 | 287.5 | 291.8 |
| IQ2_XXS | 252.81 GiB | 254.0 | 254.2 | 254.8 | 255.8 | 258.0 | 262.3 |
| IQ1_M | 216.46 GiB | 217.6 | 217.9 | 218.4 | 219.5 | 221.6 | 225.9 |
| IQ1_S | 193.16 GiB | 194.3 | 194.6 | 195.1 | 196.2 | 198.3 | 202.6 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, this model uses latent attention and allocates no V cache at all, so any formula reading num_key_value_heads overstates its cache by more than an order of magnitude.