google · text

gemma-4-12B

google/gemma-4-12B

gemma-4-12B at Q4_K_M is exactly 7,740,990,944 bytes (7.21 GiB / 7.74 GB) — an effective 5.178 bits per weight, not the nominal 4. Its KV cache at 32K is 2.47 GiB, not the 12.00 GiB a flat formula predicts.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
12.0B
Architecture
gemma4
48 layers
Context
262,144
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S2.72 GiB2,924,755,4241.956Mungert
IQ1_M3.14 GiB3,372,361,1842.256Mungert
IQ2_XXS3.38 GiB3,631,146,4642.429Mungert
IQ2_XS3.68 GiB3,955,795,4242.646Mungert
IQ2_S3.93 GiB4,222,445,0242.824Mungert
IQ2_M4.16 GiB4,462,122,4642.985Mungert
Q2_K_S4.38 GiB4,706,285,0243.148Mungert
Q2_K_M4.57 GiB4,903,998,9443.280Mungert
IQ3_XXS4.67 GiB5,013,853,6643.354Mungert
IQ3_XS4.91 GiB5,272,393,1843.527Mungert
IQ3_M5.41 GiB5,812,327,9043.888Mungert
Q3_K_S5.66 GiB6,074,553,8244.063Mungert
Q3_K_M5.87 GiB6,302,250,4644.216Mungert
IQ4_XS6.18 GiB6,635,255,2644.438Mungert
IQ4_NL6.26 GiB6,716,356,0644.493Mungert
Q4_06.72 GiB7,219,672,5444.829Mungert
Q4_K_S6.78 GiB7,275,705,8244.867Mungert
Q4_17.01 GiB7,523,431,9045.032Mungert
Q4_K_M7.21 GiB7,740,990,9445.178Mungert
Q5_07.99 GiB8,582,165,9845.741Mungert
Q5_K_M8.37 GiB8,988,468,7046.013Mungert
Q5_18.63 GiB9,263,412,7046.196Mungert
Q6_K9.11 GiB9,786,003,3606.546sneedjak
Q6_K_M9.34 GiB10,029,815,2646.709Mungert
Q8_011.80 GiB12,669,628,0328.475ggml-org
Q8_011.80 GiB12,669,646,0168.475Mungert
BF1622.20 GiB23,832,047,23215.941ggml-org
BF1622.20 GiB23,832,065,21615.941Mungert

KV cache by context

computed per layer — this model uses sliding-window attention
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.72 GiB1.50 GiB2.09×8 / 40 / 0
8,1920.97 GiB3.00 GiB3.10×8 / 40 / 0
16,3841.47 GiB6.00 GiB4.09×8 / 40 / 0
32,7682.47 GiB12.00 GiB4.86×8 / 40 / 0
65,5364.47 GiB24.00 GiB5.37×8 / 40 / 0
131,0728.47 GiB48.00 GiB5.67×8 / 40 / 0

40 of 48 layers cache only a 1,024-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.27 GiB. The real file is 7.21 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 12.00 GiB at 32K context where the real figure is 2.47 GiB, because most of this model's layers cache a fixed window rather than the whole context.

Architecture

from config.json
Layers
48
Attention heads
16
KV heads
8
Head dim
256
Hidden size
3840
Vocab
262,144
Sliding window
1024
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does gemma-4-12B need?
Q4_K_M is exactly 7,740,990,944 bytes (7.21 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is gemma-4-12B's KV cache?
2.47 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of gemma-4-12B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.