google · vision language

gemma-3-27b-it-qat-q4_0-unquantized

google/gemma-3-27b-it-qat-q4_0-unquantized

gemma-3-27b-it-qat-q4_0-unquantized at Q4_K_M is exactly 16,546,688,736 bytes (15.41 GiB / 16.55 GB) — an effective 4.825 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
27.4B
Architecture
gemma3
Context
native (config.json)
License
gemma

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
UD-IQ1_S6.06 GiB6,506,339,4241.897unsloth
UD-IQ1_M6.51 GiB6,986,738,7842.038unsloth
IQ2_XXS7.16 GiB7,685,576,0642.241bartowski
UD-IQ2_XXS7.31 GiB7,850,253,4082.289unsloth
IQ2_XS7.86 GiB8,438,904,1922.461bartowski
IQ2_S8.18 GiB8,782,409,0882.561bartowski
IQ2_M8.84 GiB9,493,073,2802.768bartowski
UD-IQ2_M8.96 GiB9,624,505,4402.807unsloth
Q2_K9.78 GiB10,503,720,6723.063unsloth
Q2_K_L9.78 GiB10,503,720,6723.063unsloth
Q2_K9.78 GiB10,503,720,9603.063bartowski
IQ3_XXS9.98 GiB10,716,478,8483.125bartowski
UD-IQ3_XXS10.07 GiB10,810,020,9603.152unsloth
Q2_K_L10.10 GiB10,845,115,7763.163bartowski
IQ3_XS10.77 GiB11,562,233,8563.372bartowski
Q3_K_S11.33 GiB12,167,614,1763.548unsloth
Q3_K_S11.33 GiB12,167,614,4643.548bartowski
IQ3_M11.69 GiB12,547,074,0483.659bartowski
Q3_K_M12.51 GiB13,437,640,4163.919unsloth
Q3_K_M12.51 GiB13,437,640,7043.919bartowski
Q3_K_L13.54 GiB14,543,462,4004.241bartowski
IQ4_XS13.75 GiB14,767,447,7764.307unsloth
IQ4_XS13.75 GiB14,767,448,0644.307bartowski
IQ4_NL14.50 GiB15,567,396,5764.540unsloth
IQ4_NL14.50 GiB15,567,396,8644.540bartowski
Q4_014.55 GiB15,617,973,9844.555unsloth
Q4_014.55 GiB15,617,974,2724.555bartowski
Q4_K_S14.60 GiB15,674,056,4164.571unsloth
Q4_K_S14.60 GiB15,674,056,7044.571bartowski
Q4_K_M15.41 GiB16,546,688,7364.825unsloth
Q4_K_M15.41 GiB16,546,689,0244.825bartowski
Q4_K_L15.73 GiB16,888,083,8404.925bartowski
Q4_115.99 GiB17,167,294,1765.006unsloth
Q4_115.99 GiB17,167,294,4645.006bartowski
Q5_K_S17.48 GiB18,767,191,7765.473unsloth
Q5_K_S17.48 GiB18,767,192,0645.473bartowski
Q5_K_M17.95 GiB19,271,675,6165.620unsloth
Q5_K_M17.95 GiB19,271,675,9045.620bartowski
Q5_K_L18.27 GiB19,613,070,7205.720bartowski
Q6_K20.64 GiB22,166,974,1766.465unsloth

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 14.37 GiB. The real file is 15.41 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does gemma-3-27b-it-qat-q4_0-unquantized need?
Q4_K_M is exactly 16,546,688,736 bytes (15.41 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of gemma-3-27b-it-qat-q4_0-unquantized should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.