hiebo · text

Qwen3.6-34B-80L-Fable-5-Heretic

hiebo/Qwen3.6-34B-80L-Fable-5-Heretic

Qwen3.6-34B-80L-Fable-5-Heretic at Q4_K_M is exactly 20,505,296,736 bytes (19.10 GiB / 20.51 GB) — an effective 4.905 bits per weight, not the nominal 4. Its KV cache at 32K is 2.50 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
33.4B
Architecture
qwen35
80 layers
Context
262,144
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K12.27 GiB13,174,736,7363.151mradermacher
I1-Q2_K12.27 GiB13,174,737,0243.151mradermacher
Q3_K_S13.85 GiB14,874,970,9763.558mradermacher
I1-Q3_K_S13.85 GiB14,874,971,2643.558mradermacher
I1-IQ3_S14.26 GiB15,307,385,9843.662mradermacher
I1-IQ3_M14.45 GiB15,513,496,7043.711mradermacher
Q3_K_M15.29 GiB16,422,767,4563.9281078mradermacher
I1-Q3_K_M15.29 GiB16,422,767,7443.928mradermacher
Q3_K_L16.53 GiB17,745,939,2964.245mradermacher
I1-Q3_K_L16.53 GiB17,745,939,5844.245mradermacher
I1-IQ4_XS17.37 GiB18,647,330,9444.460mradermacher
IQ4_XS17.50 GiB18,786,594,6564.4941078mradermacher
I1-Q4_017.88 GiB19,198,509,1844.592mradermacher
Q4_K_S17.95 GiB19,274,530,6564.610mradermacher
I1-Q4_K_S17.95 GiB19,274,530,9444.610mradermacher
Q4_K_M19.10 GiB20,505,296,7364.9051078mradermacher
I1-Q4_K_M19.10 GiB20,505,297,0244.905mradermacher
I1-Q4_119.70 GiB21,151,195,2645.059mradermacher
Q5_K_S21.57 GiB23,159,586,6565.540mradermacher
I1-Q5_K_S21.57 GiB23,159,586,9445.540mradermacher
Q5_K_M22.22 GiB23,861,477,2165.708mradermacher
I1-Q5_K_M22.22 GiB23,861,477,5045.708mradermacher
Q6_K25.54 GiB27,427,418,9766.5611078mradermacher
I1-Q6_K25.54 GiB27,427,419,2646.561mradermacher
Q8_033.08 GiB35,517,853,5368.4961078mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.31 GiB1.25 GiB4.00×20 / 0 / 60
8,1920.63 GiB2.50 GiB4.00×20 / 0 / 60
16,3841.25 GiB5.00 GiB4.00×20 / 0 / 60
32,7682.50 GiB10.00 GiB4.00×20 / 0 / 60
65,5365.00 GiB20.00 GiB4.00×20 / 0 / 60
131,07210.00 GiB40.00 GiB4.00×20 / 0 / 60

60 of 80 layers use linear attention, which keeps a fixed-size recurrent state instead of a per-token cache. Those layers do not grow with context at all — treating them as ordinary attention, as a flat formula does, overstates this model's cache by roughly 4.0× at long context.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 17.52 GiB. The real file is 19.10 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
80
Attention heads
24
KV heads
4
Head dim
256
Hidden size
5120
Vocab
248,320
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Qwen3.6-34B-80L-Fable-5-Heretic need?
Q4_K_M is exactly 20,505,296,736 bytes (19.10 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Qwen3.6-34B-80L-Fable-5-Heretic's KV cache?
2.50 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Qwen3.6-34B-80L-Fable-5-Heretic should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.