DavidAU · text

Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking

DavidAU/Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking

Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking at Q4_K_M is exactly 12,852,818,432 bytes (11.97 GiB / 12.85 GB) — an effective 4.835 bits per weight, not the nominal 4. Its KV cache at 32K is 1.50 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
21.3B
Architecture
qwen35
48 layers
Context
262,144
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S5.30 GiB5,687,926,6242.139mradermacher
I1-IQ1_M5.63 GiB6,048,871,2642.275mradermacher
I1-IQ2_XXS6.19 GiB6,650,445,6642.502mradermacher
I1-IQ2_XS6.65 GiB7,143,501,6642.687mradermacher
I1-IQ2_S6.87 GiB7,380,025,1842.776mradermacher
I1-IQ2_M7.32 GiB7,861,284,7042.957mradermacher
I1-Q2_K_S7.50 GiB8,054,016,8643.030mradermacher
Q2_K7.82 GiB8,401,521,1523.160mradermacher
I1-Q2_K7.82 GiB8,401,521,5043.160mradermacher
I1-IQ3_XXS8.15 GiB8,747,618,1443.290mradermacher
I1-IQ3_XS8.73 GiB9,375,401,8243.526mradermacher
Q3_K_S8.81 GiB9,455,519,2323.557mradermacher
I1-Q3_K_S8.81 GiB9,455,519,5843.557mradermacher
I1-IQ3_S9.05 GiB9,714,550,6243.654mradermacher
I1-IQ3_M9.16 GiB9,835,710,3043.700mradermacher
Q3_K_M9.67 GiB10,379,412,9923.904mradermacher
I1-Q3_K_M9.67 GiB10,379,413,3443.904mradermacher
Q3_K_L10.39 GiB11,158,636,0324.197mradermacher
I1-Q3_K_L10.39 GiB11,158,636,3844.197mradermacher
I1-IQ4_XS10.94 GiB11,744,215,9044.418mradermacher
IQ4_XS11.02 GiB11,827,773,9524.449mradermacher
I1-Q4_011.25 GiB12,083,344,2244.545mradermacher
Q4_K_S11.30 GiB12,137,083,3924.565mradermacher
I1-Q4_K_S11.30 GiB12,137,083,7444.565mradermacher
Q4_K_M11.97 GiB12,852,818,4324.835mradermacher
I1-Q4_K_M11.97 GiB12,852,818,7844.835mradermacher
I1-Q4_112.36 GiB13,270,815,5844.992mradermacher
Q5_K_S13.50 GiB14,491,709,9525.451mradermacher
I1-Q5_K_S13.50 GiB14,491,710,3045.451mradermacher
Q5_K_M13.88 GiB14,905,324,0325.607mradermacher
I1-Q5_K_M13.88 GiB14,905,324,3845.607mradermacher
Q6_K15.91 GiB17,086,111,2326.427mradermacher
I1-Q6_K15.91 GiB17,086,111,5846.427mradermacher
Q8_020.61 GiB22,124,995,0728.322mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.19 GiB0.75 GiB4.00×12 / 0 / 36
8,1920.38 GiB1.50 GiB4.00×12 / 0 / 36
16,3840.75 GiB3.00 GiB4.00×12 / 0 / 36
32,7681.50 GiB6.00 GiB4.00×12 / 0 / 36
65,5363.00 GiB12.00 GiB4.00×12 / 0 / 36
131,0726.00 GiB24.00 GiB4.00×12 / 0 / 36

36 of 48 layers use linear attention, which keeps a fixed-size recurrent state instead of a per-token cache. Those layers do not grow with context at all — treating them as ordinary attention, as a flat formula does, overstates this model's cache by roughly 4.0× at long context.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 11.14 GiB. The real file is 11.97 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
48
Attention heads
24
KV heads
4
Head dim
256
Hidden size
5120
Vocab
248,320
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking need?
Q4_K_M is exactly 12,852,818,432 bytes (11.97 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking's KV cache?
1.50 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.