TheDrummer · text

Skyfall-31B-v4.2

TheDrummer/Skyfall-31B-v4.2

Skyfall-31B-v4.2 at Q4_K_M is exactly 19,296,142,016 bytes (17.97 GiB / 19.30 GB) — an effective 4.924 bits per weight, not the nominal 4. Its KV cache at 32K is 6.75 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
31.4B
Architecture
llama
54 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S6.39 GiB6,861,505,9841.751mradermacher
I1-IQ1_M6.99 GiB7,508,100,5441.916mradermacher
I1-IQ2_XXS8.00 GiB8,585,758,1442.191mradermacher
I1-IQ2_XS8.83 GiB9,483,273,6642.420mradermacher
IQ2_XXS9.01 GiB9,677,423,2962.469bartowski
I1-IQ2_S9.14 GiB9,812,919,7442.504mradermacher
IQ2_XS9.75 GiB10,467,787,4562.671bartowski
I1-IQ2_M9.94 GiB10,675,045,8242.724mradermacher
IQ2_S10.10 GiB10,848,551,6162.768bartowski
I1-Q2_K_S10.18 GiB10,930,226,6242.789mradermacher
IQ2_M10.81 GiB11,603,526,3362.961bartowski
I1-Q2_K10.92 GiB11,729,438,1442.993mradermacher
Q2_K11.17 GiB11,997,807,2963.061bartowski
I1-IQ3_XXS11.42 GiB12,263,638,4643.129mradermacher
IQ3_XXS11.73 GiB12,590,662,3363.213bartowski
Q2_K_L11.78 GiB12,653,167,2963.229bartowski
I1-IQ3_XS12.17 GiB13,070,386,6243.335mradermacher
IQ3_XS12.43 GiB13,349,733,0563.406bartowski
I1-Q3_K_S12.80 GiB13,744,014,7843.507mradermacher
I1-IQ3_S12.84 GiB13,781,616,0643.517mradermacher
Q3_K_S13.00 GiB13,957,006,0163.561bartowski
I1-IQ3_M13.10 GiB14,065,714,6243.589mradermacher
IQ3_M13.33 GiB14,314,095,2963.652bartowski
I1-Q3_K_M14.16 GiB15,199,487,4243.878mradermacher
Q3_K_M14.37 GiB15,428,207,2963.937bartowski
I1-Q3_K_L15.32 GiB16,444,671,4244.196mradermacher
Q3_K_L15.38 GiB16,516,104,8964.214bartowski
I1-IQ4_XS15.74 GiB16,904,324,5444.313mradermacher
IQ4_XS15.93 GiB17,099,539,1364.363bartowski
I1-Q4_016.65 GiB17,881,793,9844.563mradermacher
I1-Q4_K_S16.71 GiB17,947,329,9844.579mradermacher
Q4_016.78 GiB18,022,367,9364.599bartowski
IQ4_NL16.79 GiB18,032,444,0964.601bartowski
Q4_K_S16.86 GiB18,102,321,8564.619bartowski
I1-Q4_K_M17.68 GiB18,978,620,8644.843mradermacher
Q4_K_M17.97 GiB19,296,142,0164.924bartowski
I1-Q4_118.38 GiB19,736,462,7845.036mradermacher
Q4_K_L18.43 GiB19,794,215,6165.051bartowski
Q4_118.48 GiB19,842,958,0165.063bartowski
I1-Q5_K_S20.17 GiB21,654,046,1445.525mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.84 GiB0.84 GiB54 / 0 / 0
8,1921.69 GiB1.69 GiB54 / 0 / 0
16,3843.38 GiB3.38 GiB54 / 0 / 0
32,7686.75 GiB6.75 GiB54 / 0 / 0
65,53613.50 GiB13.50 GiB54 / 0 / 0
131,07227.00 GiB27.00 GiB54 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.42 GiB. The real file is 17.97 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
54
Attention heads
32
KV heads
8
Head dim
128
Hidden size
5120
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Skyfall-31B-v4.2 need?
Q4_K_M is exactly 19,296,142,016 bytes (17.97 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Skyfall-31B-v4.2's KV cache?
6.75 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Skyfall-31B-v4.2 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.