Can I run spoomplesmaxx-mini-14B on a Radeon RX 6500 XT?
Not at these settings. No indexed quantization of spoomplesmaxx-mini-14B fits Radeon RX 6500 XT at any context we compute, with q8_0 KV. The smallest shipped quantization is 3.33 GiB in weights alone, against 3.72 GiB usable. CPU offload can still run it, slowly.
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| Q8_0 | 14.62 GiB | 15.9 | 16.2 | 16.9 | 18.2 | 20.9 | 26.2 |
| I1-Q6_K | 11.29 GiB | 12.6 | 12.9 | 13.6 | 14.9 | 17.6 | 22.9 |
| Q6_K | 11.29 GiB | 12.6 | 12.9 | 13.6 | 14.9 | 17.6 | 22.9 |
| I1-Q5_K_M | 9.79 GiB | 11.1 | 11.4 | 12.1 | 13.4 | 16.1 | 21.4 |
| Q5_K_M | 9.79 GiB | 11.1 | 11.4 | 12.1 | 13.4 | 16.1 | 21.4 |
| I1-Q5_K_S | 9.56 GiB | 10.9 | 11.2 | 11.8 | 13.2 | 15.8 | 21.1 |
| Q5_K_S | 9.56 GiB | 10.9 | 11.2 | 11.8 | 13.2 | 15.8 | 21.1 |
| I1-Q4_1 | 8.74 GiB | 10.0 | 10.4 | 11.0 | 12.4 | 15.0 | 20.3 |
| I1-Q4_K_M | 8.38 GiB | 9.7 | 10.0 | 10.7 | 12.0 | 14.7 | 20.0 |
| Q4_K_M | 8.38 GiB | 9.7 | 10.0 | 10.7 | 12.0 | 14.7 | 20.0 |
| I1-Q4_K_S | 7.98 GiB | 9.3 | 9.6 | 10.3 | 11.6 | 14.3 | 19.6 |
| Q4_K_S | 7.98 GiB | 9.3 | 9.6 | 10.3 | 11.6 | 14.3 | 19.6 |
| I1-Q4_0 | 7.96 GiB | 9.2 | 9.6 | 10.2 | 11.6 | 14.2 | 19.5 |
| I1-IQ4_NL | 7.95 GiB | 9.2 | 9.6 | 10.2 | 11.6 | 14.2 | 19.5 |
| IQ4_XS | 7.62 GiB | 8.9 | 9.2 | 9.9 | 11.2 | 13.9 | 19.2 |
| I1-IQ4_XS | 7.55 GiB | 8.8 | 9.2 | 9.8 | 11.2 | 13.8 | 19.1 |
| I1-Q3_K_L | 7.36 GiB | 8.7 | 9.0 | 9.6 | 11.0 | 13.6 | 18.9 |
| Q3_K_L | 7.36 GiB | 8.7 | 9.0 | 9.6 | 11.0 | 13.6 | 18.9 |
| I1-Q3_K_M | 6.82 GiB | 8.1 | 8.4 | 9.1 | 10.4 | 13.1 | 18.4 |
| Q3_K_M | 6.82 GiB | 8.1 | 8.4 | 9.1 | 10.4 | 13.1 | 18.4 |
| I1-IQ3_M | 6.41 GiB | 7.7 | 8.0 | 8.7 | 10.0 | 12.7 | 18.0 |
| I1-IQ3_S | 6.23 GiB | 7.5 | 7.8 | 8.5 | 9.8 | 12.5 | 17.8 |
| I1-Q3_K_S | 6.20 GiB | 7.5 | 7.8 | 8.5 | 9.8 | 12.5 | 17.8 |
| Q3_K_S | 6.20 GiB | 7.5 | 7.8 | 8.5 | 9.8 | 12.5 | 17.8 |
| I1-IQ3_XS | 5.94 GiB | 7.2 | 7.6 | 8.2 | 9.6 | 12.2 | 17.5 |
| I1-IQ3_XXS | 5.53 GiB | 6.8 | 7.2 | 7.8 | 9.2 | 11.8 | 17.1 |
| I1-Q2_K | 5.36 GiB | 6.7 | 7.0 | 7.6 | 9.0 | 11.6 | 16.9 |
| Q2_K | 5.36 GiB | 6.7 | 7.0 | 7.6 | 9.0 | 11.6 | 16.9 |
| I1-Q2_K_S | 5.02 GiB | 6.3 | 6.6 | 7.3 | 8.6 | 11.3 | 16.6 |
| I1-IQ2_M | 4.96 GiB | 6.2 | 6.6 | 7.2 | 8.6 | 11.2 | 16.5 |
| I1-IQ2_S | 4.62 GiB | 5.9 | 6.2 | 6.9 | 8.2 | 10.9 | 16.2 |
| I1-IQ2_XS | 4.37 GiB | 5.7 | 6.0 | 6.7 | 8.0 | 10.6 | 16.0 |
| I1-IQ2_XXS | 4.00 GiB | 5.3 | 5.6 | 6.3 | 7.6 | 10.3 | 15.6 |
| I1-IQ1_M | 3.59 GiB | 4.9 | 5.2 | 5.9 | 7.2 | 9.9 | 15.2 |
| I1-IQ1_S | 3.33 GiB | 4.6 | 5.0 | 5.6 | 7.0 | 9.6 | 14.9 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, the KV cache grows linearly with context and, past about 32K, becomes larger than the weights for many models.