Can I run Kimi-K2.5 on a Apple M5 Max?
Not at these settings. No indexed quantization of Kimi-K2.5 fits Apple M5 Max at any context we compute, with q8_0 KV. The smallest shipped quantization is 195.86 GiB in weights alone, against 25.11 GiB usable. CPU offload can still run it, slowly.
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| BF16 | 1912.15 GiB | 1912.9 | 1913.1 | 1913.3 | 1913.9 | 1915.1 | 1917.3 |
| Q6_K | 785.02 GiB | 785.8 | 785.9 | 786.2 | 786.8 | 787.9 | 790.2 |
| Q5_K_M | 678.69 GiB | 679.5 | 679.6 | 679.9 | 680.5 | 681.6 | 683.9 |
| Q5_K_S | 658.45 GiB | 659.2 | 659.4 | 659.7 | 660.2 | 661.4 | 663.6 |
| Q4_1 | 598.99 GiB | 599.8 | 599.9 | 600.2 | 600.8 | 601.9 | 604.2 |
| Q4_K_M | 578.58 GiB | 579.4 | 579.5 | 579.8 | 580.3 | 581.5 | 583.8 |
| Q4_0 | 549.26 GiB | 550.0 | 550.2 | 550.5 | 551.0 | 552.2 | 554.4 |
| Q8_0 | 543.62 GiB | 544.4 | 544.5 | 544.8 | 545.4 | 546.5 | 548.8 |
| Q4_K_S | 543.24 GiB | 544.0 | 544.2 | 544.4 | 545.0 | 546.1 | 548.4 |
| Q4_K_L | 541.31 GiB | 542.1 | 542.2 | 542.5 | 543.1 | 544.2 | 546.5 |
| IQ4_NL | 539.67 GiB | 540.4 | 540.6 | 540.9 | 541.4 | 542.6 | 544.9 |
| IQ4_XS | 510.00 GiB | 510.8 | 510.9 | 511.2 | 511.8 | 512.9 | 515.2 |
| Q3_K_M | 456.14 GiB | 456.9 | 457.1 | 457.3 | 457.9 | 459.0 | 461.3 |
| Q3_K_L | 454.05 GiB | 454.8 | 455.0 | 455.3 | 455.8 | 457.0 | 459.2 |
| IQ3_M | 434.81 GiB | 435.6 | 435.7 | 436.0 | 436.6 | 437.7 | 440.0 |
| Q3_K_S | 414.28 GiB | 415.1 | 415.2 | 415.5 | 416.0 | 417.2 | 419.5 |
| IQ3_XS | 391.23 GiB | 392.0 | 392.1 | 392.4 | 393.0 | 394.1 | 396.4 |
| UD-IQ3_XXS | 386.32 GiB | 387.1 | 387.2 | 387.5 | 388.1 | 389.2 | 391.5 |
| IQ3_S | 377.51 GiB | 378.3 | 378.4 | 378.7 | 379.3 | 380.4 | 382.7 |
| IQ3_XXS | 376.80 GiB | 377.6 | 377.7 | 378.0 | 378.6 | 379.7 | 382.0 |
| Q2_K_L | 348.36 GiB | 349.1 | 349.3 | 349.6 | 350.1 | 351.3 | 353.6 |
| Q2_K | 348.11 GiB | 348.9 | 349.0 | 349.3 | 349.9 | 351.0 | 353.3 |
| UD-IQ2_M | 321.54 GiB | 322.3 | 322.5 | 322.7 | 323.3 | 324.4 | 326.7 |
| IQ2_S | 311.72 GiB | 312.5 | 312.6 | 312.9 | 313.5 | 314.6 | 316.9 |
| UD-IQ2_XXS | 304.30 GiB | 305.1 | 305.2 | 305.5 | 306.1 | 307.2 | 309.5 |
| IQ2_M | 300.77 GiB | 301.5 | 301.7 | 302.0 | 302.5 | 303.7 | 306.0 |
| UD-IQ1_M | 279.94 GiB | 280.7 | 280.9 | 281.1 | 281.7 | 282.8 | 285.1 |
| IQ2_XS | 263.67 GiB | 264.4 | 264.6 | 264.9 | 265.4 | 266.6 | 268.9 |
| IQ2_XXS | 262.75 GiB | 263.5 | 263.7 | 263.9 | 264.5 | 265.7 | 267.9 |
| UD-IQ1_S | 256.97 GiB | 257.7 | 257.9 | 258.2 | 258.7 | 259.9 | 262.2 |
| UD-TQ1_0 | 223.09 GiB | 223.9 | 224.0 | 224.3 | 224.9 | 226.0 | 228.3 |
| IQ1_M | 204.50 GiB | 205.3 | 205.4 | 205.7 | 206.3 | 207.4 | 209.7 |
| IQ1_S | 195.86 GiB | 196.6 | 196.8 | 197.1 | 197.6 | 198.8 | 201.0 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, this model uses latent attention and allocates no V cache at all, so any formula reading num_key_value_heads overstates its cache by more than an order of magnitude.