Can I run GLM-5 on a Apple M5 Max?
Not at these settings. No indexed quantization of GLM-5 fits Apple M5 Max at any context we compute, with q4_0 KV. The smallest shipped quantization is 164.05 GiB in weights alone, against 25.11 GiB usable. CPU offload can still run it, slowly.
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| BF16 | 1404.42 GiB | 1405.1 | 1405.2 | 1405.4 | 1405.8 | 1406.6 | 1408.1 |
| Q8_0 | 746.31 GiB | 747.0 | 747.1 | 747.3 | 747.7 | 748.5 | 750.0 |
| Q6_K | 576.84 GiB | 577.5 | 577.6 | 577.8 | 578.2 | 579.0 | 580.5 |
| Q5_K_M | 498.42 GiB | 499.1 | 499.2 | 499.4 | 499.8 | 500.6 | 502.1 |
| Q5_K_S | 484.04 GiB | 484.7 | 484.8 | 485.0 | 485.4 | 486.2 | 487.7 |
| Q4_1 | 440.24 GiB | 440.9 | 441.0 | 441.2 | 441.6 | 442.4 | 443.9 |
| Q4_K_M | 424.57 GiB | 425.3 | 425.4 | 425.6 | 425.9 | 426.7 | 428.3 |
| Q4_K_S | 398.95 GiB | 399.6 | 399.7 | 399.9 | 400.3 | 401.1 | 402.6 |
| Q4_0 | 397.81 GiB | 398.5 | 398.6 | 398.8 | 399.2 | 400.0 | 401.5 |
| IQ4_NL | 396.67 GiB | 397.4 | 397.5 | 397.7 | 398.0 | 398.8 | 400.4 |
| IQ4_XS | 375.23 GiB | 375.9 | 376.0 | 376.2 | 376.6 | 377.4 | 378.9 |
| Q3_K_M | 335.56 GiB | 336.3 | 336.4 | 336.5 | 336.9 | 337.7 | 339.2 |
| Q3_K_S | 303.87 GiB | 304.6 | 304.7 | 304.9 | 305.2 | 306.0 | 307.6 |
| UD-IQ3_XXS | 283.62 GiB | 284.3 | 284.4 | 284.6 | 285.0 | 285.8 | 287.3 |
| Q2_K_L | 257.21 GiB | 257.9 | 258.0 | 258.2 | 258.6 | 259.4 | 260.9 |
| Q2_K | 257.00 GiB | 257.7 | 257.8 | 258.0 | 258.4 | 259.1 | 260.7 |
| UD-IQ2_M | 237.27 GiB | 238.0 | 238.1 | 238.3 | 238.6 | 239.4 | 241.0 |
| UD-IQ2_XXS | 224.52 GiB | 225.2 | 225.3 | 225.5 | 225.9 | 226.7 | 228.2 |
| UD-IQ1_M | 208.48 GiB | 209.2 | 209.3 | 209.5 | 209.9 | 210.6 | 212.2 |
| UD-IQ1_S | 189.71 GiB | 190.4 | 190.5 | 190.7 | 191.1 | 191.9 | 193.4 |
| UD-TQ1_0 | 164.05 GiB | 164.7 | 164.8 | 165.0 | 165.4 | 166.2 | 167.7 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, this model uses latent attention and allocates no V cache at all, so any formula reading num_key_value_heads overstates its cache by more than an order of magnitude.