Quantization format
Q3_K_M
Files labelled Q3_K_M average 4.127 effective bits per weight across 2461 real quantizations — not the nominal 3.4375. That is 20% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 2461 files measured
Nominal bpw
3.4375
from the block layout
Measured average
4.127
2461 files
Range
1.534–22.549
varies by architecture
File sizes
0.01 GiB+
up to 697.01 GiB
Note
Same block type as Q3_K_S; the suffix changes which tensors are promoted.
What files with this label actually contain
tensor types across 38 parsed files
F32
11849
Q3_K
6743
Q4_K
3589
Q8_0
1462
F16
700
Q5_K
452
IQ3_XXS
236
IQ4_NL
153
IQ4_XS
117
Q5_0
89
Q6_K
80
MXFP4
72
BF16
41
Q4_0
32
Q5_1
5
If Q3_K_M were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-27B | 12.65 GiB | 3.912 | +14% |
| Qwen3.6-35B-A3B | 15.94 GiB | 3.809 | +11% |
| Qwen3.5-9B | 4.35 GiB | 3.873 | +13% |
| gemma-4-26B-A4B-it | 12.13 GiB | 3.924 | +14% |
| gemma-4-12B-it | 5.30 GiB | 3.809 | +11% |
| Qwen3.5-4B | 2.14 GiB | 3.937 | +15% |
| Hy3 | 125.35 GiB | 3.604 | +5% |
| gemma-4-E4B-it | 3.78 GiB | 4.060 | +18% |
| Qwen3-Coder-30B-A3B-Instruct | 13.70 GiB | 3.855 | +12% |
| gemma-4-31B-it | 13.72 GiB | 3.770 | +10% |
| Qwen3-VL-30B-A3B-Instruct | 13.70 GiB | 3.788 | +10% |
| FLUX.2-klein-9B | 4.44 GiB | 4.203 | +22% |
| LTX-2.3 | 9.90 GiB | 9.246 | +169% |
| Llama-3.2-1B-Instruct | 0.64 GiB | 4.472 | +30% |
| gemma-4-E2B-it | 2.36 GiB | 3.961 | +15% |
| gpt-oss-20b | 10.72 GiB | 4.279 | +24% |
| Qwen3-8B | 3.84 GiB | 4.028 | +17% |
| Qwopus3.6-35B-A3B-v1 | 15.99 GiB | 3.820 | +11% |
| Qwen3.5-0.8B | 0.44 GiB | 4.306 | +25% |
| Qwen3-VL-8B-Instruct-abliterated-v1 | 3.84 GiB | 3.763 | +9% |
| UI-TARS-1.5-7B | 3.55 GiB | 3.674 | +7% |
| Qwen3-4B | 1.93 GiB | 4.128 | +20% |
| GLM-5.2 | 169.33 GiB | 1.931 | +-44% |
| Llama-3.1-8B-Instruct | 3.74 GiB | 4.004 | +16% |
| Ornith-1.0-35B | 15.11 GiB | 3.745 | +9% |
●From the filewhat these mean