gemma-3n-E2B-it
google/gemma-3n-E2B-itgemma-3n-E2B-it at Q4_K_M is exactly 2,787,806,304 bytes (2.60 GiB / 2.79 GB) — an effective 4.100 bits per weight, not the nominal 4. Its KV cache at 32K is 0.42 GiB, not the 1.88 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| Q2_K | 1.76 GiB | 1,887,882,336 | 2.777 | — | bartowski |
| UD-IQ2_XXS | 1.91 GiB | 2,053,426,528 | 3.020 | — | unsloth |
| IQ3_XS | 2.02 GiB | 2,170,211,424 | 3.192 | — | bartowski |
| UD-IQ2_M | 2.04 GiB | 2,188,660,064 | 3.219 | — | unsloth |
| Q3_K_S | 2.06 GiB | 2,209,582,176 | 3.250 | — | bartowski |
| Q2_K | 2.07 GiB | 2,221,329,760 | 3.267 | — | unsloth |
| Q2_K_L | 2.07 GiB | 2,221,329,760 | 3.267 | — | unsloth |
| IQ3_M | 2.08 GiB | 2,237,156,448 | 3.290 | — | bartowski |
| Q3_K_M | 2.14 GiB | 2,299,677,792 | 3.382 | — | bartowski |
| UD-IQ3_XXS | 2.16 GiB | 2,324,114,784 | 3.418 | — | unsloth |
| Q3_K_L | 2.22 GiB | 2,379,893,856 | 3.500 | — | bartowski |
| Q3_K_S | 2.23 GiB | 2,393,083,232 | 3.520 | — | unsloth |
| Q3_K_M | 2.31 GiB | 2,483,178,848 | 3.652 | — | unsloth |
| IQ4_XS | 2.43 GiB | 2,607,467,616 | 3.835 | — | bartowski |
| Q4_0 | 2.54 GiB | 2,726,612,064 | 4.010 | — | bartowski |
| IQ4_NL | 2.54 GiB | 2,727,398,496 | 4.011 | — | bartowski |
| Q4_K_S | 2.54 GiB | 2,730,282,080 | 4.016 | — | bartowski |
| Q4_K_M | 2.60 GiB | 2,787,806,304 | 4.100 | — | bartowski |
| IQ4_XS | 2.71 GiB | 2,909,785,440 | 4.279 | — | unsloth |
| Q4_1 | 2.76 GiB | 2,965,294,176 | 4.361 | — | bartowski |
| Q4_0 | 2.76 GiB | 2,965,687,648 | 4.362 | — | unsloth |
| IQ4_NL | 2.76 GiB | 2,966,474,080 | 4.363 | — | unsloth |
| Q4_K_S | 2.77 GiB | 2,969,357,664 | 4.367 | — | unsloth |
| Q4_K_M | 2.82 GiB | 3,026,881,888 | 4.452 | — | unsloth |
| Q4_1 | 2.87 GiB | 3,078,540,640 | 4.528 | — | unsloth |
| Q5_K_S | 2.99 GiB | 3,207,122,016 | 4.717 | — | bartowski |
| Q5_K_M | 3.02 GiB | 3,240,266,848 | 4.766 | — | bartowski |
| Q5_K_S | 3.04 GiB | 3,261,648,224 | 4.797 | — | unsloth |
| Q5_K_M | 3.07 GiB | 3,294,793,056 | 4.846 | — | unsloth |
| Q2_K_L | 3.26 GiB | 3,496,397,920 | 5.142 | — | bartowski |
| Q6_K | 3.47 GiB | 3,721,006,176 | 5.473 | — | bartowski |
| Q4_K_L | 3.65 GiB | 3,924,462,688 | 5.772 | — | bartowski |
| Q5_K_L | 3.84 GiB | 4,125,264,992 | 6.067 | — | bartowski |
| Q6_K | 3.92 GiB | 4,208,594,272 | 6.190 | — | unsloth |
| Q6_K_L | 4.04 GiB | 4,338,617,440 | 6.381 | — | bartowski |
| Q8_0 | 4.46 GiB | 4,788,112,064 | 7.042 | — | ggml-org |
| Q8_0 | 4.46 GiB | 4,788,112,480 | 7.042 | — | bartowski |
| Q8_0 | 4.46 GiB | 4,788,112,736 | 7.042 | — | unsloth |
| F16 | 8.31 GiB | 8,918,846,144 | 13.117 | — | ggml-org |
| F16 | 8.31 GiB | 8,918,846,560 | 13.117 | — | unsloth |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 0.09 GiB | 0.23 GiB | 2.50× | 6 / 24 / 0 |
| 8,192 | 0.14 GiB | 0.47 GiB | 3.33× | 6 / 24 / 0 |
| 16,384 | 0.23 GiB | 0.94 GiB | 4.00× | 6 / 24 / 0 |
| 32,768 | 0.42 GiB | 1.88 GiB | 4.44× | 6 / 24 / 0 |
| 65,536 | 0.80 GiB | 3.75 GiB | 4.71× | 6 / 24 / 0 |
| 131,072 | 1.55 GiB | 7.50 GiB | 4.85× | 6 / 24 / 0 |
24 of 30 layers cache only a 512-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.85 GiB. The real file is 2.60 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 1.88 GiB at 32K context where the real figure is 0.42 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does gemma-3n-E2B-it need?
- Q4_K_M is exactly 2,787,806,304 bytes (2.60 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is gemma-3n-E2B-it's KV cache?
- 0.42 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Which quantization of gemma-3n-E2B-it should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.