Quantization format

Q3_K

Q3_K has a nominal rate of bits per weight. We have no measured files carrying this label yet.

From the file· 84 files measured
Nominal bpw
from the block layout
Measured average
4.493
84 files
Range
3.553–14.091
varies by architecture
File sizes
0.02 GiB+
up to 104.93 GiB

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
whisper-medium0.32 GiB3.601
whisper-large-v30.64 GiB3.553
whisper-large-v3-turbo0.34 GiB3.636
Z-Image-Turbo2.93 GiB4.086
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled15.99 GiB3.820
Wan2.2-Animate-14B8.14 GiB4.048
FLUX.1-Fill-dev4.99 GiB3.601
Llama-3.1-8B3.74 GiB4.004
WAN2.2-14B-Rapid-AllInOne8.03 GiB3.979
Codestral-22B-v0.110.02 GiB3.868
whisper-small0.11 GiB3.768
DeepSeek-Coder-V2-Lite-Instruct7.57 GiB4.139
Hermes-3-Llama-3.1-8B3.74 GiB4.004
Qwen2-0.5B-Instruct0.33 GiB5.756
Qwen2.5-Math-7B-Instruct3.55 GiB4.001
gemma-2-27b-it12.50 GiB3.945
MiniCPM-V-2_63.55 GiB3.760
s2-pro2.82 GiB5.312
whisper-base0.03 GiB4.088
glm-4-9b-chat4.72 GiB4.310
G4-Alice-v1.2-31B14.24 GiB3.911
Qwen2.5-Math-1.5B-Instruct0.77 GiB4.271
Meta-Llama-3.1-8B-Instruct-abliterated3.74 GiB4.004
Phi-3-mini-4k-instruct1.82 GiB4.094
openchat-3.6-8b-202405223.74 GiB4.004
From the filewhat these mean

Other quantizations