Llama-4-Scout-17B-16E-Instruct
meta-llama/Llama-4-Scout-17B-16E-InstructLlama-4-Scout-17B-16E-Instruct at Q4_K_M is exactly 65,359,900,352 bytes (60.87 GiB / 65.36 GB) — an effective 4.813 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| IQ1_M | 24.51 GiB | 26,318,067,200 | 1.938 | — | bartowski |
| UD-TQ1_0 | 27.25 GiB | 29,261,483,520 | 2.155 | — | unsloth |
| IQ2_XXS | 28.09 GiB | 30,165,685,760 | 2.221 | — | bartowski |
| UD-IQ1_S | 30.24 GiB | 32,470,126,080 | 2.391 | — | unsloth |
| IQ2_XS | 30.68 GiB | 32,941,790,720 | 2.426 | — | bartowski |
| IQ2_S | 31.98 GiB | 34,336,604,160 | 2.528 | — | bartowski |
| UD-IQ1_M | 32.59 GiB | 34,988,428,800 | 2.576 | — | unsloth |
| IQ2_M | 34.56 GiB | 37,112,709,120 | 2.733 | — | bartowski |
| UD-IQ2_XXS | 34.83 GiB | 37,400,972,800 | 2.754 | — | unsloth |
| UD-IQ2_M | 36.39 GiB | 39,078,694,400 | 2.878 | — | unsloth |
| Q2_K | 36.85 GiB | 39,563,317,760 | 2.913 | — | unsloth |
| Q2_K_L | 37.07 GiB | 39,805,775,360 | 2.931 | — | unsloth |
| Q2_K | 40.03 GiB | 42,986,260,480 | 3.165 | — | bartowski |
| Q2_K_L | 40.97 GiB | 43,996,500,480 | 3.240 | — | bartowski |
| IQ3_XXS | 41.87 GiB | 44,955,402,240 | 3.310 | — | bartowski |
| UD-IQ3_XXS | 42.59 GiB | 45,725,847,040 | 3.367 | — | unsloth |
| Q3_K_S | 43.53 GiB | 46,740,372,480 | 3.442 | — | unsloth |
| IQ3_XS | 44.19 GiB | 47,452,090,880 | 3.494 | — | bartowski |
| Q3_K_S | 46.34 GiB | 49,754,370,560 | 3.664 | — | bartowski |
| IQ3_M2 shards | 46.87 GiB | 50,322,567,872 | 3.706 | — | bartowski |
| Q3_K_M2 shards | 48.20 GiB | 51,755,187,392 | 3.811 | — | unsloth |
| Q3_K_M2 shards | 50.59 GiB | 54,318,953,184 | 4.000 | — | bartowski |
| IQ4_XS2 shards | 53.69 GiB | 57,651,883,712 | 4.245 | — | unsloth |
| Q3_K_L2 shards | 53.83 GiB | 57,799,569,824 | 4.256 | — | lmstudio-community |
| Q3_K_L2 shards | 53.83 GiB | 57,799,570,112 | 4.256 | — | bartowski |
| IQ4_XS2 shards | 55.78 GiB | 59,892,341,984 | 4.410 | — | bartowski |
| IQ4_NL2 shards | 56.76 GiB | 60,947,033,792 | 4.488 | — | unsloth |
| Q4_02 shards | 56.98 GiB | 61,182,963,392 | 4.505 | — | unsloth |
| Q4_K_S2 shards | 57.23 GiB | 61,452,971,712 | 4.525 | — | unsloth |
| IQ4_NL2 shards | 58.67 GiB | 62,991,754,464 | 4.638 | — | bartowski |
| Q4_02 shards | 58.72 GiB | 63,054,669,024 | 4.643 | — | bartowski |
| Q4_K_M2 shards | 60.87 GiB | 65,359,900,352 | 4.813 | — | unsloth |
| Q4_K_M2 shards | 62.91 GiB | 67,546,178,464 | 4.974 | — | lmstudio-community |
| Q4_K_M2 shards | 62.91 GiB | 67,546,178,784 | 4.974 | — | bartowski |
| Q4_12 shards | 62.94 GiB | 67,586,260,672 | 4.977 | — | unsloth |
| Q4_K_L2 shards | 63.62 GiB | 68,313,961,152 | 5.030 | — | bartowski |
| Q4_12 shards | 64.35 GiB | 69,096,207,552 | 5.088 | — | bartowski |
| Q5_K_S2 shards | 69.16 GiB | 74,256,944,832 | 5.468 | — | unsloth |
| Q5_K_M2 shards | 71.29 GiB | 76,546,444,992 | 5.637 | — | unsloth |
| Q5_K_L2 shards | 73.87 GiB | 79,316,144,864 | 5.841 | — | bartowski |
KV cache by context
This model declares a 8,192-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 56.91 GiB. The real file is 60.87 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does Llama-4-Scout-17B-16E-Instruct need?
- Q4_K_M is exactly 65,359,900,352 bytes (60.87 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Is Llama-4-Scout-17B-16E-Instruct a mixture-of-experts model?
- Yes — 16 experts, 1 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
- Which quantization of Llama-4-Scout-17B-16E-Instruct should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.