DavidAU · text

Gemma-3-4b-it-Uncensored-DBL-X

DavidAU/Gemma-3-4b-it-Uncensored-DBL-X

Gemma-3-4b-it-Uncensored-DBL-X at Q4_K_M is exactly 2,709,839,808 bytes (2.52 GiB / 2.71 GB) — an effective 4.635 bits per weight, not the nominal 4. Its KV cache at 32K is 0.94 GiB, not the 4.75 GiB a flat formula predicts.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
4.7B
Architecture
gemma3
38 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S1.13 GiB1,209,700,7042.069mradermacher
I1-IQ1_M1.20 GiB1,284,288,8642.196mradermacher
I1-IQ2_XXS1.31 GiB1,408,602,4642.409mradermacher
I1-IQ2_XS1.41 GiB1,514,279,2642.590mradermacher
I1-IQ2_S1.46 GiB1,563,062,6242.673mradermacher
I1-IQ2_M1.55 GiB1,662,513,5042.843mradermacher
I1-Q2_K_S1.64 GiB1,760,101,9843.010mradermacher
I1-IQ3_XXS1.71 GiB1,833,152,8643.135mradermacher
Q2_K1.74 GiB1,867,046,8483.193DavidAU
I1-Q2_K1.74 GiB1,867,048,5443.193mradermacher
I1-IQ3_XS1.88 GiB2,014,463,5843.445mradermacher
Q3_K_S1.96 GiB2,099,740,6083.591DavidAU
I1-Q3_K_S1.96 GiB2,099,742,3043.591mradermacher
I1-IQ3_S1.96 GiB2,099,742,3043.591mradermacher
I1-IQ3_M2.01 GiB2,153,358,9443.683mradermacher
Q3_K_M2.12 GiB2,278,940,6083.898DavidAU
I1-Q3_K_M2.12 GiB2,278,942,3043.898mradermacher
Q3_K_L2.27 GiB2,433,605,5684.162DavidAU
I1-Q3_K_L2.27 GiB2,433,607,2644.162mradermacher
I1-IQ4_XS2.29 GiB2,463,958,6244.214mradermacher
IQ4_XS2.31 GiB2,480,340,9284.242DavidAU
Q4_02.40 GiB2,576,023,4884.406DavidAU
I1-IQ4_NL2.40 GiB2,576,025,1844.406mradermacher
I1-Q4_02.41 GiB2,582,578,7844.417mradermacher
IQ4_NL2.41 GiB2,589,130,6884.428DavidAU
Q4_K_S2.41 GiB2,590,441,4084.430DavidAU
I1-Q4_K_S2.41 GiB2,590,443,1044.430mradermacher
Q4_K_M2.52 GiB2,709,839,8084.635DavidAU
I1-Q4_K_M2.52 GiB2,709,841,5044.635mradermacher
I1-Q4_12.61 GiB2,800,158,3044.789mradermacher
Q5_K_S2.82 GiB3,024,289,7285.172DavidAU
Q5_02.82 GiB3,024,289,7285.172DavidAU
I1-Q5_K_S2.82 GiB3,024,291,4245.172mradermacher
Q5_K_M2.88 GiB3,093,225,4085.290DavidAU
I1-Q5_K_M2.88 GiB3,093,227,1045.290mradermacher
Q6_K3.26 GiB3,500,572,6085.987DavidAU
I1-Q6_K3.26 GiB3,500,574,3045.987mradermacher
Q8_04.22 GiB4,531,657,4087.750DavidAU
F167.94 GiB8,522,953,40814.577DavidAU

KV cache by context

computed per layer — this model uses sliding-window attention
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.28 GiB0.59 GiB2.11×6 / 32 / 0
8,1920.38 GiB1.19 GiB3.17×6 / 32 / 0
16,3840.56 GiB2.38 GiB4.22×6 / 32 / 0
32,7680.94 GiB4.75 GiB5.07×6 / 32 / 0
65,5361.69 GiB9.50 GiB5.63×6 / 32 / 0
131,0723.19 GiB19.00 GiB5.96×6 / 32 / 0

32 of 38 layers cache only a 1,024-token window rather than the full context, on a period of 6. Figures assume the default configuration; --swa-full disables the saving entirely.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.45 GiB. The real file is 2.52 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 4.75 GiB at 32K context where the real figure is 0.94 GiB, because most of this model's layers cache a fixed window rather than the whole context.

Architecture

from config.json
Layers
38
Attention heads
8
KV heads
4
Head dim
256
Hidden size
2560
Vocab
262,208
Sliding window
1024
SWA period
6
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Gemma-3-4b-it-Uncensored-DBL-X need?
Q4_K_M is exactly 2,709,839,808 bytes (2.52 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Gemma-3-4b-it-Uncensored-DBL-X's KV cache?
0.94 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Gemma-3-4b-it-Uncensored-DBL-X should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.