Quantization format

UD_IQ4_XS

UD_IQ4_XS has a nominal rate of bits per weight. We have no measured files carrying this label yet.

From the file· 36 files measured
Nominal bpw
from the block layout
Measured average
3.989
36 files
Range
3.594–4.947
varies by architecture
File sizes
1.02 GiB+
up to 461.08 GiB

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-35B-A3B16.51 GiB3.945
gemma-4-26B-A4B-it12.66 GiB4.098
DeepSeek-V4-Flash128.43 GiB3.792
Qwen-AgentWorld-35B-A3B16.56 GiB4.105
GLM-5.2340.22 GiB3.879
Ornith-1.0-35B16.56 GiB4.105
Kimi-K2.7-Code461.08 GiB3.741
Qwen3.5-122B-A10B56.09 GiB3.852
Qwen3-Coder-Next35.79 GiB3.859
Qwen3.5-35B-A3B16.29 GiB3.891
Laguna-S-2.153.61 GiB3.917
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1618.19 GiB4.734
Step-3.7-Flash88.79 GiB3.788
LFM2.5-8B-A1B3.97 GiB4.029
DeepSeek-R1-Distill-Qwen-1.5B1.02 GiB4.947
Ornith-1.0-397B178.91 GiB3.873
Qwen3.5-397B-A17B176.70 GiB3.763
KAT-Coder-V2.5-Dev16.96 GiB4.203
North-Mini-Code-1.014.18 GiB3.997
MiniMax-M2.7100.97 GiB3.792
MiniMax-M3193.31 GiB3.888
NVIDIA-Nemotron-3-Super-120B-A12B-BF1660.06 GiB4.173
Mistral-Small-4-119B-260354.13 GiB3.895
Huihui-gemma-4-26B-A4B-it-abliterated12.50 GiB4.044
MiMo-V2.5139.18 GiB3.847
From the filewhat these mean

Other quantizations