Quantization format
IQ2_XXS
Files labelled IQ2_XXS average 2.393 effective bits per weight across 342 real quantizations — not the nominal 2.0625. That is 16% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 342 files measured
Nominal bpw
2.0625
from the block layout
Measured average
2.393
342 files
Range
1.854–7.074
varies by architecture
File sizes
0.02 GiB+
up to 262.75 GiB
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-35B-A3B | 9.94 GiB | 2.376 | +15% |
| gemma-4-26B-A4B-it | 8.99 GiB | 2.910 | +41% |
| Hy3 | 76.47 GiB | 2.199 | +7% |
| gemma-4-31B-it | 10.09 GiB | 2.771 | +34% |
| Ornith-1.0-35B | 9.11 GiB | 2.257 | +9% |
| Qwen3.5-122B-A10B | 33.81 GiB | 2.322 | +13% |
| Qwen3-Coder-Next | 17.97 GiB | 1.938 | +-6% |
| Qwen2.5-32B-Instruct | 8.41 GiB | 2.204 | +7% |
| Qwen3.5-35B-A3B | 9.94 GiB | 2.376 | +15% |
| Qwen3-30B-A3B | 7.59 GiB | 2.135 | +3% |
| jina-embeddings-v5-text-small | 0.21 GiB | 3.080 | +49% |
| Qwen2.5-Coder-32B-Instruct | 8.41 GiB | 2.204 | +7% |
| Qwen3.5-27B | 8.95 GiB | 2.766 | +34% |
| Laguna-S-2.1 | 29.82 GiB | 2.179 | +6% |
| Qwen3-30B-A3B-Instruct-2507 | 7.05 GiB | 1.983 | +-4% |
| Llama-3.3-70B-Instruct | 17.79 GiB | 2.165 | +5% |
| DeepSeek-R1-Distill-Llama-70B | 17.79 GiB | 2.165 | +5% |
| Step-3.7-Flash | 51.22 GiB | 2.185 | +6% |
| Qwen2.5-72B-Instruct | 23.74 GiB | 2.805 | +36% |
| Llama-3.1-70B-Instruct | 17.79 GiB | 2.165 | +5% |
| jina-embeddings-v5-text-nano | 0.10 GiB | 3.985 | +93% |
| Mistral-Small-3.2-24B-Instruct-2506 | 6.10 GiB | 2.181 | +6% |
| Ornith-1.0-397B | 99.01 GiB | 2.143 | +4% |
| DeepSeek-R1-Distill-Qwen-32B | 8.41 GiB | 2.204 | +7% |
| Laguna-XS-2.1 | 8.76 GiB | 2.250 | +9% |
●From the filewhat these mean