Quantization format
Q3_K_L
Files labelled Q3_K_L average 4.368 effective bits per weight across 2134 real quantizations — not the nominal 3.4375. That is 27% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 2134 files measured
Nominal bpw
3.4375
from the block layout
Measured average
4.368
2134 files
Range
1.647–23.759
varies by architecture
File sizes
0.02 GiB+
up to 454.05 GiB
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-35B-A3B | 16.56 GiB | 3.956 | +15% |
| gemma-4-26B-A4B-it | 12.29 GiB | 3.978 | +16% |
| gemma-4-12B-it | 6.12 GiB | 4.392 | +28% |
| Qwen3.5-4B | 2.48 GiB | 4.576 | +33% |
| Hy3 | 133.25 GiB | 3.831 | +11% |
| Qwen3-Coder-30B-A3B-Instruct | 13.58 GiB | 3.821 | +11% |
| gemma-4-31B-it | 15.66 GiB | 4.300 | +25% |
| Llama-3.2-1B-Instruct | 0.68 GiB | 4.742 | +38% |
| gemma-4-E2B-it | 3.06 GiB | 5.138 | +49% |
| Qwen3-8B | 4.13 GiB | 4.328 | +26% |
| Qwen3.5-0.8B | 0.49 GiB | 4.844 | +41% |
| Qwen3-VL-8B-Instruct-abliterated-v1 | 4.13 GiB | 4.044 | +18% |
| UI-TARS-1.5-7B | 3.81 GiB | 3.944 | +15% |
| Qwen3-4B | 2.09 GiB | 4.455 | +30% |
| Llama-3.1-8B-Instruct | 4.03 GiB | 4.306 | +25% |
| Ornith-1.0-35B | 15.73 GiB | 3.898 | +13% |
| Llama-3.2-3B-Instruct | 1.69 GiB | 4.520 | +32% |
| ThinkingCap-Qwen3.6-27B | 14.23 GiB | 4.468 | +30% |
| Wan2.1-T2V-1.3B | 1.15 GiB | 6.982 | +103% |
| Qwythos-9B-v2 | 4.89 GiB | 4.349 | +27% |
| Qwen3.5-122B-A10B | 57.37 GiB | 3.939 | +15% |
| Qwen3-Coder-Next | 35.60 GiB | 3.838 | +12% |
| Qwen2.5-32B-Instruct | 16.06 GiB | 4.211 | +23% |
| gemma-3-1b-it | 0.70 GiB | 6.013 | +75% |
| Qwen3.5-35B-A3B | 16.56 GiB | 3.956 | +15% |
●From the filewhat these mean