DavidAU · text

Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-Thinking

DavidAU/Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-Thinking

Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-Thinking at Q4_K_M is exactly 12,852,817,792 bytes (11.97 GiB / 12.85 GB) — an effective 4.835 bits per weight, not the nominal 4. Its KV cache at 32K is 1.50 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
21.3B
Architecture
qwen35
48 layers
Context
262,144
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S5.30 GiB5,687,925,9522.139mradermacher
I1-IQ1_M5.63 GiB6,048,870,5922.275mradermacher
I1-IQ2_XXS6.19 GiB6,650,444,9922.502mradermacher
I1-IQ2_XS6.65 GiB7,143,500,9922.687mradermacher
I1-IQ2_S6.87 GiB7,380,024,5122.776mradermacher
I1-IQ2_M7.32 GiB7,861,284,0322.957mradermacher
I1-Q2_K_S7.50 GiB8,054,016,1923.030mradermacher
Q2_K7.82 GiB8,401,520,5123.160mradermacher
I1-Q2_K7.82 GiB8,401,520,8323.160mradermacher
I1-IQ3_XXS8.15 GiB8,747,617,4723.290mradermacher
I1-IQ3_XS8.73 GiB9,375,401,1523.526mradermacher
Q3_K_S8.81 GiB9,455,518,5923.557mradermacher
I1-Q3_K_S8.81 GiB9,455,518,9123.557mradermacher
I1-IQ3_S9.05 GiB9,714,549,9523.654mradermacher
I1-IQ3_M9.16 GiB9,835,709,6323.700mradermacher
Q3_K_M9.67 GiB10,379,412,3523.904mradermacher
I1-Q3_K_M9.67 GiB10,379,412,6723.904mradermacher
Q3_K_L10.39 GiB11,158,635,3924.197mradermacher
I1-Q3_K_L10.39 GiB11,158,635,7124.197mradermacher
I1-IQ4_XS10.94 GiB11,744,215,2324.418mradermacher
IQ4_XS11.02 GiB11,827,773,3124.449mradermacher
I1-Q4_011.25 GiB12,083,343,5524.545mradermacher
Q4_K_S11.30 GiB12,137,082,7524.565mradermacher
I1-Q4_K_S11.30 GiB12,137,083,0724.565mradermacher
Q4_K_M11.97 GiB12,852,817,7924.835mradermacher
I1-Q4_K_M11.97 GiB12,852,818,1124.835mradermacher
I1-Q4_112.36 GiB13,270,814,9124.992mradermacher
Q5_K_S13.50 GiB14,491,709,3125.451mradermacher
I1-Q5_K_S13.50 GiB14,491,709,6325.451mradermacher
Q5_K_M13.88 GiB14,905,323,3925.607mradermacher
I1-Q5_K_M13.88 GiB14,905,323,7125.607mradermacher
Q6_K15.91 GiB17,086,110,5926.427mradermacher
I1-Q6_K15.91 GiB17,086,110,9126.427mradermacher
Q8_020.61 GiB22,124,994,4328.322mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.19 GiB0.75 GiB4.00×12 / 0 / 36
8,1920.38 GiB1.50 GiB4.00×12 / 0 / 36
16,3840.75 GiB3.00 GiB4.00×12 / 0 / 36
32,7681.50 GiB6.00 GiB4.00×12 / 0 / 36
65,5363.00 GiB12.00 GiB4.00×12 / 0 / 36
131,0726.00 GiB24.00 GiB4.00×12 / 0 / 36

36 of 48 layers use linear attention, which keeps a fixed-size recurrent state instead of a per-token cache. Those layers do not grow with context at all — treating them as ordinary attention, as a flat formula does, overstates this model's cache by roughly 4.0× at long context.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 11.14 GiB. The real file is 11.97 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
48
Attention heads
24
KV heads
4
Head dim
256
Hidden size
5120
Vocab
248,320
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-Thinking need?
Q4_K_M is exactly 12,852,817,792 bytes (11.97 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-Thinking's KV cache?
1.50 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-Thinking should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.