Quantization format

Q5_K_M

Files labelled Q5_K_M average 5.919 effective bits per weight across 2776 real quantizations — not the nominal 5.5. That is 8% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.

From the file· 2776 files measured
Nominal bpw
5.5
from the block layout
Measured average
5.919
2776 files
Range
1.604–32.945
varies by architecture
File sizes
0.01 GiB+
up to 1038.96 GiB

What files with this label actually contain

tensor types across 37 parsed files
F32
13740
Q5_K
9310
Q6_K
2045
F16
940
Q8_0
936
Q5_1
93
MXFP4
72
BF16
41
Q4_0
16

If Q5_K_M were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-27B18.47 GiB5.712+4%
Qwen3.6-35B-A3B24.13 GiB5.766+5%
Qwen3.5-9B6.27 GiB5.577+1%
gemma-4-26B-A4B-it17.99 GiB5.822+6%
gemma-4-12B-it7.84 GiB5.628+2%
nemotron-3.5-asr-streaming-0.6b0.52 GiB7.018+28%
parakeet-unified-en-0.6b0.50 GiB6.997+27%
Qwen3.5-4B2.99 GiB5.516+0%
Qwythos-9B-Claude-Mythos-5-1M12.25 GiB11.182+103%
Hy3197.61 GiB5.681+3%
gemma-4-E4B-it5.11 GiB5.484+-0%
Qwen3-Coder-30B-A3B-Instruct20.23 GiB5.692+3%
gemma-4-31B-it20.17 GiB5.540+1%
Qwen3-VL-30B-A3B-Instruct20.23 GiB5.594+2%
FLUX.2-klein-9B6.54 GiB6.185+12%
cohere-transcribe-03-20261.65 GiB6.856+25%
LTX-2.314.97 GiB13.979+154%
Llama-3.2-1B-Instruct0.85 GiB5.901+7%
llama-3-youko-8b5.34 GiB5.711+4%
gemma-4-E2B-it3.41 GiB5.713+4%
gpt-oss-20b10.91 GiB4.357+-21%
Qwen3-8B5.45 GiB5.715+4%
Qwopus3.6-35B-A3B-v123.61 GiB5.640+3%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP39.03 GiB12.069+119%
Qwen3.5-0.8B0.55 GiB5.404+-2%
From the filewhat these mean

Other quantizations