Quantization format

Q8_0

Files labelled Q8_0 average 8.635 effective bits per weight across 3451 real quantizations — not the nominal 8.5. That is 2% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.

From the file· 3451 files measured
Nominal bpw
8.5
from the block layout
Measured average
8.635
3451 files
Range
1.210–31.059
varies by architecture
File sizes
0.01 GiB+
up to 1557.00 GiB

Note

Carries a per-block scale, so it is 8.5 and not 8.

What files with this label actually contain

tensor types across 37 parsed files
Q8_0
13156
F32
12790
F16
927
BF16
125
MXFP4
72

If Q8_0 were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-27B26.63 GiB8.235+-3%
embeddinggemma-300m0.31 GiB8.679+2%
Qwen3.6-35B-A3B35.21 GiB8.412+-1%
Qwen3.5-9B8.87 GiB7.896+-7%
gemma-4-26B-A4B-it25.02 GiB8.095+-5%
gemma-4-12B-it11.80 GiB8.475+-0%
nemotron-3.5-asr-streaming-0.6b0.70 GiB9.418+11%
parakeet-unified-en-0.6b0.68 GiB9.463+11%
Qwen3.5-4B4.30 GiB7.935+-7%
Qwythos-9B-Claude-Mythos-5-1M9.11 GiB8.320+-2%
Hy3295.84 GiB8.505+0%
gemma-4-E4B-it7.48 GiB8.035+-5%
Qwen3-Coder-30B-A3B-Instruct30.25 GiB8.511+0%
gemma-4-31B-it30.39 GiB8.349+-2%
Qwen3-VL-30B-A3B-Instruct30.25 GiB8.364+-2%
FLUX.2-klein-9B9.29 GiB8.793+3%
cohere-transcribe-03-20262.25 GiB9.335+10%
Qwen-AgentWorld-35B-A3B34.37 GiB8.518+0%
LTX-2.323.75 GiB22.180+161%
Llama-3.2-1B-Instruct1.23 GiB8.552+1%
llama-3-youko-8b7.95 GiB8.509+0%
gemma-4-E2B-it4.63 GiB7.757+-9%
gpt-oss-20b11.28 GiB4.503+-47%
Qwen3-8B8.11 GiB8.507+0%
Qwopus3.6-35B-A3B-v135.21 GiB8.412+-1%
From the filewhat these mean

Other quantizations