Quantization format

Q4_K

Q4_K has a nominal rate of bits per weight. We have no measured files carrying this label yet.

From the file· 164 files measured
Nominal bpw
from the block layout
Measured average
6.044
164 files
Range
3.201–28.888
varies by architecture
File sizes
0.02 GiB+
up to 165.75 GiB

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
embeddinggemma-300m0.28 GiB8.082
nemotron-3.5-asr-streaming-0.6b0.38 GiB5.123
DeepSeek-V4-Flash153.33 GiB4.527
Qwythos-9B-Claude-Mythos-5-1M5.38 GiB4.914
FLUX.2-klein-9B5.32 GiB5.038
cohere-transcribe-03-20261.41 GiB5.849
parakeet-tdt-0.6b-v30.39 GiB5.332
whisper-medium0.41 GiB4.655
Voxtral-Mini-4B-Realtime-26022.35 GiB4.559
whisper-large-v30.83 GiB4.609
whisper-large-v3-turbo0.44 GiB4.688
gemma-4-E2B-it-qat-q4_0-unquantized3.18 GiB5.354
GLM-4.7-Flash16.99 GiB4.675
Qwen3-ASR-1.7B1.39 GiB5.077
Z-Image-Turbo3.60 GiB5.023
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled20.22 GiB4.831
Qwen3-ASR-0.6B0.59 GiB5.382
Wan2.2-Animate-14B10.80 GiB5.372
parakeet-tdt-0.6b-v20.37 GiB5.139
Mistral-7B-Instruct-v0.34.07 GiB4.827
Qwen3-TTS-12Hz-0.6B-Base0.50 GiB4.661
gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking17.40 GiB4.780
canary-1b-v20.37 GiB3.201
FLUX.1-Fill-dev6.45 GiB4.657
Llama-3.1-8B4.58 GiB4.902
From the filewhat these mean

Other quantizations