Can I run MATE-3B on a GeForce RTX 5050?
Yes. The best fit is I1-Q5_K_M at 131,072 context with f16 KV — 7.38 GiB of 7.44 GiB usable, leaving 0.06 GiB headroom. Expect roughly 35 tokens/sec (modeled, ±12.9%).
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| F16 | 5.75 GiB | 6.7 | 6.8 | 7.1 | 7.7 | 8.8 | 11.1 |
| Q8_0 | 3.06 GiB | 4.0 | 4.2 | 4.4 | 5.0 | 6.1 | 8.4 |
| I1-Q6_K | 2.36 GiB | 3.3 | 3.5 | 3.7 | 4.3 | 5.4 | 7.7 |
| Q6_K | 2.36 GiB | 3.3 | 3.5 | 3.7 | 4.3 | 5.4 | 7.7 |
| I1-Q5_K_M | 2.07 GiB | 3.0 | 3.2 | 3.4 | 4.0 | 5.1 | 7.4 |
| Q5_K_M | 2.07 GiB | 3.0 | 3.2 | 3.4 | 4.0 | 5.1 | 7.4 |
| I1-Q5_K_S | 2.02 GiB | 3.0 | 3.1 | 3.4 | 4.0 | 5.1 | 7.3 |
| Q5_K_S | 2.02 GiB | 3.0 | 3.1 | 3.4 | 4.0 | 5.1 | 7.3 |
| I1-Q4_1 | 1.86 GiB | 2.8 | 3.0 | 3.2 | 3.8 | 4.9 | 7.2 |
| I1-Q4_K_M | 1.80 GiB | 2.8 | 2.9 | 3.2 | 3.7 | 4.9 | 7.1 |
| Q4_K_M | 1.80 GiB | 2.8 | 2.9 | 3.2 | 3.7 | 4.9 | 7.1 |
| I1-Q4_K_S | 1.71 GiB | 2.7 | 2.8 | 3.1 | 3.6 | 4.8 | 7.0 |
| Q4_K_S | 1.71 GiB | 2.7 | 2.8 | 3.1 | 3.6 | 4.8 | 7.0 |
| I1-Q4_0 | 1.70 GiB | 2.7 | 2.8 | 3.1 | 3.6 | 4.8 | 7.0 |
| I1-IQ4_NL | 1.70 GiB | 2.7 | 2.8 | 3.1 | 3.6 | 4.8 | 7.0 |
| IQ4_XS | 1.63 GiB | 2.6 | 2.7 | 3.0 | 3.6 | 4.7 | 6.9 |
| I1-IQ4_XS | 1.62 GiB | 2.6 | 2.7 | 3.0 | 3.6 | 4.7 | 6.9 |
| I1-Q3_K_L | 1.59 GiB | 2.5 | 2.7 | 3.0 | 3.5 | 4.7 | 6.9 |
| Q3_K_L | 1.59 GiB | 2.5 | 2.7 | 3.0 | 3.5 | 4.7 | 6.9 |
| I1-Q3_K_M | 1.48 GiB | 2.4 | 2.6 | 2.9 | 3.4 | 4.5 | 6.8 |
| Q3_K_M | 1.48 GiB | 2.4 | 2.6 | 2.9 | 3.4 | 4.5 | 6.8 |
| I1-IQ3_M | 1.39 GiB | 2.3 | 2.5 | 2.8 | 3.3 | 4.4 | 6.7 |
| I1-IQ3_S | 1.36 GiB | 2.3 | 2.5 | 2.7 | 3.3 | 4.4 | 6.7 |
| I1-Q3_K_S | 1.35 GiB | 2.3 | 2.4 | 2.7 | 3.3 | 4.4 | 6.7 |
| Q3_K_S | 1.35 GiB | 2.3 | 2.4 | 2.7 | 3.3 | 4.4 | 6.7 |
| I1-IQ3_XS | 1.30 GiB | 2.2 | 2.4 | 2.7 | 3.2 | 4.4 | 6.6 |
| I1-IQ3_XXS | 1.19 GiB | 2.1 | 2.3 | 2.6 | 3.1 | 4.3 | 6.5 |
| I1-Q2_K | 1.19 GiB | 2.1 | 2.3 | 2.6 | 3.1 | 4.2 | 6.5 |
| Q2_K | 1.19 GiB | 2.1 | 2.3 | 2.6 | 3.1 | 4.2 | 6.5 |
| I1-Q2_K_S | 1.12 GiB | 2.1 | 2.2 | 2.5 | 3.1 | 4.2 | 6.4 |
| I1-IQ2_M | 1.06 GiB | 2.0 | 2.2 | 2.4 | 3.0 | 4.1 | 6.4 |
| I1-IQ2_S | 0.99 GiB | 1.9 | 2.1 | 2.4 | 2.9 | 4.1 | 6.3 |
| I1-IQ2_XS | 0.96 GiB | 1.9 | 2.1 | 2.3 | 2.9 | 4.0 | 6.3 |
| I1-IQ2_XXS | 0.88 GiB | 1.8 | 2.0 | 2.3 | 2.8 | 3.9 | 6.2 |
| I1-IQ1_M | 0.79 GiB | 1.7 | 1.9 | 2.2 | 2.7 | 3.9 | 6.1 |
| I1-IQ1_S | 0.74 GiB | 1.7 | 1.8 | 2.1 | 2.7 | 3.8 | 6.0 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Run it
The best-fitting configuration above, as a command:
llama-cli -hf fgdrg/MATE-3B \ --ctx-size 131072 \ -ngl auto
Recent llama.cpp defaults to --fit on with -ngl auto, so it will size the offload for you. The question worth your attention is not how many layers to offload but what context and quantization you are willing to live with — which is what the grid above is for.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, most of this model's layers cache only a 32,768-token window rather than the full context.