Quantization format
IQ3_XXS
Files labelled IQ3_XXS average 3.427 effective bits per weight across 542 real quantizations — not the nominal 3.0625. That is 12% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 542 files measured
Nominal bpw
3.0625
from the block layout
Measured average
3.427
542 files
Range
2.268–19.140
varies by architecture
File sizes
0.03 GiB+
up to 397.96 GiB
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-35B-A3B | 14.68 GiB | 3.507 | +15% |
| gemma-4-26B-A4B-it | 11.33 GiB | 3.665 | +20% |
| gemma-4-12B-it | 4.79 GiB | 3.442 | +12% |
| Qwen3.5-4B | 2.09 GiB | 3.852 | +26% |
| Hy3 | 109.28 GiB | 3.142 | +3% |
| gemma-4-31B-it | 12.09 GiB | 3.321 | +8% |
| gemma-4-E2B-it | 2.47 GiB | 4.135 | +35% |
| Qwen3-8B | 3.14 GiB | 3.291 | +7% |
| Qwen3.5-0.8B | 0.42 GiB | 4.161 | +36% |
| Ornith-1.0-35B | 13.85 GiB | 3.432 | +12% |
| ThinkingCap-Qwen3.6-27B | 11.76 GiB | 3.692 | +21% |
| Qwythos-9B-v2 | 4.11 GiB | 3.657 | +19% |
| Qwen3.5-122B-A10B | 50.74 GiB | 3.485 | +14% |
| Qwen3-Coder-Next | 29.55 GiB | 3.186 | +4% |
| Qwen3.5-35B-A3B | 14.68 GiB | 3.507 | +15% |
| Qwen3-0.6B | 0.32 GiB | 3.681 | +20% |
| Qwen3-VL-2B-Instruct | 0.70 GiB | 2.837 | +-7% |
| Jan-v3-4B-base-instruct | 1.71 GiB | 3.332 | +9% |
| Qwen2.5-VL-7B-Instruct | 2.90 GiB | 3.005 | +-2% |
| Ornith-1.0-9B | 3.98 GiB | 3.720 | +21% |
| gemma-3-4b-it | 1.57 GiB | 3.143 | +3% |
| Qwen3-30B-A3B | 11.38 GiB | 3.201 | +5% |
| Qwen3-1.7B | 0.83 GiB | 3.497 | +14% |
| Qwen2.5-Coder-32B-Instruct | 11.96 GiB | 3.135 | +2% |
| Qwen3.5-27B | 11.96 GiB | 3.697 | +21% |
●From the filewhat these mean