coder3101 · text

Skyfall-31B-v4.2-heretic

coder3101/Skyfall-31B-v4.2-heretic

Skyfall-31B-v4.2-heretic at Q4_K_M is exactly 18,978,620,704 bytes (17.68 GiB / 18.98 GB) — an effective 4.843 bits per weight, not the nominal 4. Its KV cache at 32K is 6.75 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
31.4B
Architecture
llama
54 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S6.39 GiB6,861,506,1121.751mradermacher
I1-IQ1_M6.99 GiB7,508,100,6721.916mradermacher
I1-IQ2_XXS8.00 GiB8,585,758,2722.191mradermacher
I1-IQ2_XS8.83 GiB9,483,273,7922.420mradermacher
I1-IQ2_S9.14 GiB9,812,919,8722.504mradermacher
I1-IQ2_M9.94 GiB10,675,045,9522.724mradermacher
I1-Q2_K_S10.18 GiB10,930,226,7522.789mradermacher
Q2_K10.92 GiB11,729,437,9842.993mradermacher
I1-Q2_K10.92 GiB11,729,438,2722.993mradermacher
I1-IQ3_XXS11.42 GiB12,263,638,5923.129mradermacher
I1-IQ3_XS12.17 GiB13,070,386,7523.335mradermacher
Q3_K_S12.80 GiB13,744,014,6243.507mradermacher
I1-Q3_K_S12.80 GiB13,744,014,9123.507mradermacher
I1-IQ3_S12.84 GiB13,781,616,1923.517mradermacher
I1-IQ3_M13.10 GiB14,065,714,7523.589mradermacher
Q3_K_M14.16 GiB15,199,487,2643.878mradermacher
I1-Q3_K_M14.16 GiB15,199,487,5523.878mradermacher
Q3_K_L15.32 GiB16,444,671,2644.196mradermacher
I1-Q3_K_L15.32 GiB16,444,671,5524.196mradermacher
I1-IQ4_XS15.74 GiB16,904,324,6724.313mradermacher
IQ4_XS15.89 GiB17,061,610,7844.353mradermacher
I1-Q4_016.65 GiB17,881,794,1124.563mradermacher
Q4_K_S16.71 GiB17,947,329,8244.579mradermacher
I1-Q4_K_S16.71 GiB17,947,330,1124.579mradermacher
Q4_K_M17.68 GiB18,978,620,7044.843mradermacher
I1-Q4_K_M17.68 GiB18,978,620,9924.843mradermacher
I1-Q4_118.38 GiB19,736,462,9125.036mradermacher
Q5_K_S20.17 GiB21,654,045,9845.525mradermacher
I1-Q5_K_S20.17 GiB21,654,046,2725.525mradermacher
Q5_K_M20.72 GiB22,251,488,5445.678mradermacher
I1-Q5_K_M20.72 GiB22,251,488,8325.678mradermacher
Q6_K23.96 GiB25,728,910,6246.565mradermacher
I1-Q6_K23.96 GiB25,728,910,9126.565mradermacher
Q8_031.03 GiB33,322,075,4248.502mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.84 GiB0.84 GiB54 / 0 / 0
8,1921.69 GiB1.69 GiB54 / 0 / 0
16,3843.38 GiB3.38 GiB54 / 0 / 0
32,7686.75 GiB6.75 GiB54 / 0 / 0
65,53613.50 GiB13.50 GiB54 / 0 / 0
131,07227.00 GiB27.00 GiB54 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.42 GiB. The real file is 17.68 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
54
Attention heads
32
KV heads
8
Head dim
128
Hidden size
5120
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Skyfall-31B-v4.2-heretic need?
Q4_K_M is exactly 18,978,620,704 bytes (17.68 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Skyfall-31B-v4.2-heretic's KV cache?
6.75 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Skyfall-31B-v4.2-heretic should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.