Quantization format
Q2_K
Files labelled Q2_K average 3.337 effective bits per weight across 2331 real quantizations — not the nominal 2.625. That is 27% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 2331 files measured
Nominal bpw
2.625
from the block layout
Measured average
3.337
2331 files
Range
1.078–19.476
varies by architecture
File sizes
0.01 GiB+
up to 864.81 GiB
Note
Blogs usually cite 2.5625. The block layout gives 2.625.
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-35B-A3B | 12.58 GiB | 3.006 | +14% |
| gemma-4-26B-A4B-it | 10.20 GiB | 3.301 | +26% |
| gemma-4-12B-it | 4.50 GiB | 3.231 | +23% |
| DeepSeek-V4-Flash | 92.86 GiB | 2.742 | +4% |
| Qwen3.5-4B | 2.04 GiB | 3.763 | +43% |
| Hy3 | 99.36 GiB | 2.857 | +9% |
| Qwen3-Coder-30B-A3B-Instruct | 10.49 GiB | 2.950 | +12% |
| gemma-4-31B-it | 11.76 GiB | 3.231 | +23% |
| Qwen3-VL-30B-A3B-Instruct | 10.49 GiB | 2.899 | +10% |
| FLUX.2-klein-9B | 3.71 GiB | 3.508 | +34% |
| LTX-2.3 | 7.39 GiB | 6.903 | +163% |
| Llama-3.2-1B-Instruct | 0.54 GiB | 3.760 | +43% |
| gemma-4-E2B-it | 2.81 GiB | 4.716 | +80% |
| gpt-oss-20b | 10.68 GiB | 4.265 | +62% |
| Qwen3-8B | 3.06 GiB | 3.205 | +22% |
| Qwen3.5-0.8B | 0.43 GiB | 4.252 | +62% |
| Qwen3-VL-8B-Instruct-abliterated-v1 | 3.06 GiB | 2.995 | +14% |
| UI-TARS-1.5-7B | 2.81 GiB | 2.910 | +11% |
| Qwen3-4B | 1.55 GiB | 3.320 | +26% |
| Llama-3.1-8B-Instruct | 2.96 GiB | 3.167 | +21% |
| Ornith-1.0-35B | 11.75 GiB | 2.912 | +11% |
| Llama-3.2-3B-Instruct | 1.27 GiB | 3.396 | +29% |
| ThinkingCap-Qwen3.6-27B | 11.03 GiB | 3.462 | +32% |
| whisper-medium | 0.25 GiB | 2.795 | +6% |
| Wan2.2-I2V-A14B | 4.94 GiB | 2.968 | +13% |
●From the filewhat these mean