Quantization format

Q4_1

Q4_1 has a nominal rate of bits per weight. We have no measured files carrying this label yet.

From the file· 878 files measured
Nominal bpw
from the block layout
Measured average
5.371
878 files
Range
1.431–30.248
varies by architecture
File sizes
0.02 GiB+
up to 598.99 GiB

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-27B16.07 GiB4.968
Qwen3.6-35B-A3B21.30 GiB5.089
Qwen3.5-9B5.44 GiB4.838
gemma-4-26B-A4B-it15.04 GiB4.866
gemma-4-12B-it6.89 GiB4.948
Qwen3.5-4B2.59 GiB4.780
Hy3174.90 GiB5.028
gemma-4-E4B-it4.73 GiB5.077
Qwen3-Coder-30B-A3B-Instruct17.87 GiB5.029
gemma-4-31B-it17.81 GiB4.891
Qwen3-VL-30B-A3B-Instruct17.87 GiB4.942
FLUX.2-klein-9B5.74 GiB5.429
LTX-2.312.81 GiB11.967
Llama-3.2-1B-Instruct0.77 GiB5.384
gemma-4-E2B-it2.94 GiB4.926
gpt-oss-20b10.78 GiB4.306
Qwen3-8B4.89 GiB5.126
Qwen3.5-0.8B0.50 GiB4.902
Qwen3-4B2.42 GiB5.164
Llama-3.1-8B-Instruct4.78 GiB5.111
Ornith-1.0-35B20.46 GiB5.072
Llama-3.2-3B-Instruct1.95 GiB5.213
ThinkingCap-Qwen3.6-27B16.60 GiB5.213
whisper-medium0.45 GiB5.050
Wan2.2-I2V-A14B8.62 GiB5.184
From the filewhat these mean

Other quantizations