Can I run MiMo-V2.5-Pro on a Apple M3 Pro?
Not at these settings. No indexed quantization of MiMo-V2.5-Pro fits Apple M3 Pro at any context we compute, with q4_0 KV. The smallest shipped quantization is 197.49 GiB in weights alone, against 12.56 GiB usable. CPU offload can still run it, slowly.
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| BF16 | 1906.14 GiB | 1907.1 | 1907.5 | 1908.3 | 1909.8 | 1912.9 | 1919.1 |
| Q8_0 | 1012.92 GiB | 1013.9 | 1014.3 | 1015.1 | 1016.6 | 1019.7 | 1025.8 |
| UD-Q6_K | 788.42 GiB | 789.4 | 789.8 | 790.6 | 792.1 | 795.2 | 801.3 |
| UD-Q5_K_M | 705.94 GiB | 706.9 | 707.3 | 708.1 | 709.6 | 712.7 | 718.9 |
| Q5_K_M | 704.85 GiB | 705.8 | 706.2 | 707.0 | 708.5 | 711.6 | 717.8 |
| UD-Q5_K_S | 664.21 GiB | 665.2 | 665.6 | 666.4 | 667.9 | 671.0 | 677.1 |
| Q4_1 | 596.26 GiB | 597.3 | 597.6 | 598.4 | 600.0 | 603.0 | 609.2 |
| UD-Q4_K_M | 586.37 GiB | 587.4 | 587.8 | 588.5 | 590.1 | 593.1 | 599.3 |
| Q4_K_M | 585.99 GiB | 587.0 | 587.4 | 588.1 | 589.7 | 592.8 | 598.9 |
| Q4_K_S | 557.32 GiB | 558.3 | 558.7 | 559.5 | 561.0 | 564.1 | 570.2 |
| UD-Q4_K_S | 548.12 GiB | 549.1 | 549.5 | 550.3 | 551.8 | 554.9 | 561.0 |
| Q4_0 | 539.02 GiB | 540.0 | 540.4 | 541.2 | 542.7 | 545.8 | 551.9 |
| IQ4_NL | 537.62 GiB | 538.6 | 539.0 | 539.8 | 541.3 | 544.4 | 550.5 |
| IQ4_XS | 508.09 GiB | 509.1 | 509.5 | 510.2 | 511.8 | 514.9 | 521.0 |
| UD-IQ4_NL | 466.84 GiB | 467.8 | 468.2 | 469.0 | 470.5 | 473.6 | 479.8 |
| UD-IQ4_XS | 457.00 GiB | 458.0 | 458.4 | 459.1 | 460.7 | 463.8 | 469.9 |
| IQ3_M | 454.54 GiB | 455.5 | 455.9 | 456.7 | 458.2 | 461.3 | 467.5 |
| Q3_K_L | 452.86 GiB | 453.9 | 454.2 | 455.0 | 456.5 | 459.6 | 465.8 |
| Q3_K_M | 434.60 GiB | 435.6 | 436.0 | 436.8 | 438.3 | 441.4 | 447.5 |
| IQ3_XS | 434.58 GiB | 435.6 | 436.0 | 436.7 | 438.3 | 441.3 | 447.5 |
| UD-Q3_K_M | 428.10 GiB | 429.1 | 429.5 | 430.3 | 431.8 | 434.9 | 441.0 |
| Q3_K_S | 414.26 GiB | 415.3 | 415.6 | 416.4 | 417.9 | 421.0 | 427.2 |
| IQ3_XXS | 397.96 GiB | 399.0 | 399.3 | 400.1 | 401.6 | 404.7 | 410.9 |
| UD-IQ3_XXS | 384.32 GiB | 385.3 | 385.7 | 386.5 | 388.0 | 391.1 | 397.2 |
| UD-IQ3_S | 351.94 GiB | 352.9 | 353.3 | 354.1 | 355.6 | 358.7 | 364.9 |
| IQ3_S | 350.82 GiB | 351.8 | 352.2 | 353.0 | 354.5 | 357.6 | 363.7 |
| Q2_K_L | 334.58 GiB | 335.6 | 336.0 | 336.7 | 338.3 | 341.3 | 347.5 |
| Q2_K | 333.73 GiB | 334.7 | 335.1 | 335.9 | 337.4 | 340.5 | 346.6 |
| IQ2_M | 321.02 GiB | 322.0 | 322.4 | 323.2 | 324.7 | 327.8 | 333.9 |
| IQ2_S | 297.46 GiB | 298.5 | 298.8 | 299.6 | 301.1 | 304.2 | 310.4 |
| UD-IQ2_M | 295.48 GiB | 296.5 | 296.9 | 297.6 | 299.2 | 302.2 | 308.4 |
| UD-IQ2_XXS | 295.37 GiB | 296.4 | 296.8 | 297.5 | 299.1 | 302.1 | 308.3 |
| IQ2_XS | 285.38 GiB | 286.4 | 286.8 | 287.5 | 289.1 | 292.1 | 298.3 |
| UD-IQ1_M | 283.21 GiB | 284.2 | 284.6 | 285.4 | 286.9 | 290.0 | 296.1 |
| IQ2_XXS | 256.27 GiB | 257.3 | 257.6 | 258.4 | 260.0 | 263.0 | 269.2 |
| IQ1_M | 220.44 GiB | 221.4 | 221.8 | 222.6 | 224.1 | 227.2 | 233.4 |
| IQ1_S | 197.49 GiB | 198.5 | 198.9 | 199.6 | 201.2 | 204.3 | 210.4 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, most of this model's layers cache only a 128-token window rather than the full context.