google · text

translategemma-4b-it

google/translategemma-4b-it

translategemma-4b-it at Q4_K_M is exactly 2,489,909,312 bytes (2.32 GiB / 2.49 GB) — an effective 4.007 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
5.0B
Architecture
gemma3
Context
native (config.json)
License
gemma

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K1.61 GiB1,729,180,1602.783mradermacher
Q3_K_S1.80 GiB1,937,379,8403.118mradermacher
Q3_K_M1.95 GiB2,098,475,5203.377mradermacher
Q3_K_L2.08 GiB2,236,100,6723.598bullerwins
Q3_K_L2.08 GiB2,236,101,1203.598mradermacher
IQ4_XS2.12 GiB2,279,641,6003.668mradermacher
Q4_K_S2.21 GiB2,377,945,1523.827bullerwins
Q4_K_S2.21 GiB2,377,945,6003.827mradermacher
Q4_K_M2.32 GiB2,489,909,3124.007bullerwins
Q4_K_M2.32 GiB2,489,909,7604.007mradermacher
Q5_K_S2.57 GiB2,764,607,5524.449bullerwins
Q5_K_S2.57 GiB2,764,608,0004.449mradermacher
Q5_K_M2.64 GiB2,829,713,4724.554bullerwins
Q5_K_M2.64 GiB2,829,713,9204.554mradermacher
Q6_K2.97 GiB3,190,755,3925.135bullerwins
Q6_K2.97 GiB3,190,755,8405.135mradermacher
Q8_03.85 GiB4,130,417,4726.647bullerwins
Q8_03.85 GiB4,130,417,9206.647mradermacher
F167.23 GiB7,767,819,52012.500mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.60 GiB. The real file is 2.32 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does translategemma-4b-it need?
Q4_K_M is exactly 2,489,909,312 bytes (2.32 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of translategemma-4b-it should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.