Llama-4-Maverick-17B-128E-Instruct
meta-llama/Llama-4-Maverick-17B-128E-InstructLlama-4-Maverick-17B-128E-Instruct at Q4_K_M is exactly 242,767,153,440 bytes (226.09 GiB / 242.77 GB) — an effective 4.836 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| TQ1_02 shards | 92.59 GiB | 99,421,158,912 | 1.981 | — | GeorgyGUF |
| UD-TQ1_0 | 98.45 GiB | 105,711,920,832 | 2.106 | — | unsloth |
| UD-IQ1_S3 shards | 111.20 GiB | 119,399,114,784 | 2.379 | — | unsloth |
| UD-IQ1_S3 shards | 112.48 GiB | 120,769,472,544 | 2.406 | — | unsloth |
| UD-IQ1_M3 shards | 117.57 GiB | 126,239,967,232 | 2.515 | — | unsloth |
| UD-IQ1_M3 shards | 118.78 GiB | 127,539,546,144 | 2.541 | — | unsloth |
| UD-IQ2_XXS3 shards | 124.85 GiB | 134,055,135,232 | 2.671 | — | unsloth |
| UD-IQ2_XXS3 shards | 125.91 GiB | 135,196,116,992 | 2.693 | — | unsloth |
| UD-IQ2_M3 shards | 131.54 GiB | 141,235,914,784 | 2.814 | — | unsloth |
| Q2_K3 shards | 135.64 GiB | 145,645,120,544 | 2.901 | — | unsloth |
| Q2_K3 shards | 135.64 GiB | 145,645,120,544 | 2.901 | — | unsloth |
| Q2_K_L3 shards | 135.87 GiB | 145,887,578,144 | 2.906 | — | unsloth |
| Q2_K_L3 shards | 135.87 GiB | 145,887,578,144 | 2.906 | — | unsloth |
| UD-IQ3_XXS4 shards | 157.10 GiB | 168,681,080,960 | 3.360 | — | unsloth |
| UD-IQ3_XXS4 shards | 157.69 GiB | 169,314,814,080 | 3.373 | — | unsloth |
| Q3_K_S4 shards | 160.80 GiB | 172,655,990,400 | 3.439 | — | unsloth |
| Q3_K_S4 shards | 160.80 GiB | 172,655,990,400 | 3.439 | — | unsloth |
| Q3_K_M4 shards | 177.95 GiB | 191,068,984,960 | 3.806 | — | unsloth |
| Q3_K_M4 shards | 177.95 GiB | 191,068,984,960 | 3.806 | — | unsloth |
| IQ4_XS5 shards | 199.61 GiB | 214,324,857,120 | 4.270 | — | unsloth |
| IQ4_XS5 shards | 199.61 GiB | 214,324,857,120 | 4.270 | — | unsloth |
| UD-IQ4_XS5 shards | 205.52 GiB | 220,673,767,168 | 4.396 | — | unsloth |
| IQ4_NL5 shards | 210.26 GiB | 225,767,442,656 | 4.497 | — | unsloth |
| IQ4_NL5 shards | 210.26 GiB | 225,767,442,656 | 4.497 | — | unsloth |
| Q4_05 shards | 211.19 GiB | 226,766,211,296 | 4.517 | — | unsloth |
| Q4_K_S5 shards | 212.15 GiB | 227,799,058,688 | 4.538 | — | unsloth |
| Q4_K_S5 shards | 212.15 GiB | 227,799,058,688 | 4.538 | — | unsloth |
| Q4_K_M5 shards | 226.09 GiB | 242,767,153,440 | 4.836 | — | unsloth |
| Q4_K_M5 shards | 226.09 GiB | 242,767,153,440 | 4.836 | — | unsloth |
| Q4_16 shards | 233.50 GiB | 250,714,806,624 | 4.995 | — | unsloth |
| Q4_16 shards | 233.50 GiB | 250,714,806,624 | 4.995 | — | unsloth |
| Q5_K_S6 shards | 256.76 GiB | 275,693,627,776 | 5.492 | — | unsloth |
| Q5_K_S6 shards | 256.76 GiB | 275,693,627,776 | 5.492 | — | unsloth |
| Q5_K_M6 shards | 264.93 GiB | 284,467,259,776 | 5.667 | — | unsloth |
| Q5_K_M6 shards | 264.93 GiB | 284,467,259,776 | 5.667 | — | unsloth |
| UD-IQ2_M7 shards | 273.65 GiB | 293,827,825,664 | 5.853 | — | unsloth |
| Q6_K7 shards | 306.19 GiB | 328,773,622,752 | 6.550 | — | unsloth |
| Q6_K7 shards | 306.19 GiB | 328,773,622,752 | 6.550 | — | unsloth |
| Q8_09 shards | 396.57 GiB | 425,817,094,048 | 8.483 | — | unsloth |
| Q8_09 shards | 396.57 GiB | 425,817,094,048 | 8.483 | — | unsloth |
KV cache by context
This model declares a 8,192-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 210.38 GiB. The real file is 226.09 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does Llama-4-Maverick-17B-128E-Instruct need?
- Q4_K_M is exactly 242,767,153,440 bytes (226.09 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Is Llama-4-Maverick-17B-128E-Instruct a mixture-of-experts model?
- Yes — 128 experts, 1 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
- Which quantization of Llama-4-Maverick-17B-128E-Instruct should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.