TheDrummer · text

Valkyrie-49B-v2.1

TheDrummer/Valkyrie-49B-v2.1

Valkyrie-49B-v2.1 at Q4_K_M is exactly 30,215,578,688 bytes (28.14 GiB / 30.22 GB) — an effective 4.847 bits per weight, not the nominal 4. Its KV cache at 32K is 80.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
49.9B
Architecture
deci
80 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S10.27 GiB11,029,865,9521.770mradermacher
IQ1_M11.19 GiB12,015,854,6561.928bartowski
I1-IQ1_M11.19 GiB12,015,855,0721.928mradermacher
IQ2_XXS12.72 GiB13,659,169,8562.191bartowski
I1-IQ2_XXS12.72 GiB13,659,170,2722.191mradermacher
IQ2_XS14.04 GiB15,076,582,4642.419bartowski
I1-IQ2_XS14.04 GiB15,076,582,8802.419mradermacher
IQ2_S14.76 GiB15,848,481,8562.542bartowski
I1-IQ2_S14.76 GiB15,848,482,2722.542mradermacher
IQ2_M15.98 GiB17,163,134,0162.753bartowski
I1-IQ2_M15.98 GiB17,163,134,4322.753mradermacher
I1-Q2_K_S16.30 GiB17,507,198,4322.809mradermacher
Q2_K17.45 GiB18,739,799,1043.006bartowski
I1-Q2_K17.45 GiB18,739,799,5203.006mradermacher
IQ3_XXS18.18 GiB19,519,022,1443.131bartowski
I1-IQ3_XXS18.18 GiB19,519,022,5603.131mradermacher
Q2_K_L18.41 GiB19,765,847,1043.171bartowski
IQ3_XS19.47 GiB20,908,008,5123.354bartowski
I1-IQ3_XS19.47 GiB20,908,008,9283.354mradermacher
Q3_K_S20.45 GiB21,955,339,3283.522bartowski
I1-IQ3_S20.45 GiB21,955,339,7443.522mradermacher
I1-Q3_K_S20.45 GiB21,955,339,7443.522mradermacher
IQ3_M21.10 GiB22,657,229,8883.635bartowski
I1-IQ3_M21.10 GiB22,657,230,3043.635mradermacher
Q3_K_M22.64 GiB24,311,227,4563.900bartowski
I1-Q3_K_M22.64 GiB24,311,227,8723.900mradermacher
Q3_K_L24.47 GiB26,272,064,5764.215bartowski
I1-Q3_K_L24.47 GiB26,272,064,9924.215mradermacher
IQ4_XS25.03 GiB26,871,407,6804.311bartowski
I1-IQ4_XS25.03 GiB26,871,408,0964.311mradermacher
IQ4_NL26.43 GiB28,384,044,0964.553bartowski
Q4_026.50 GiB28,457,444,4164.565bartowski
I1-Q4_026.50 GiB28,457,444,8324.565mradermacher
Q4_K_S26.67 GiB28,633,605,1844.594bartowski
I1-Q4_K_S26.67 GiB28,633,605,6004.594mradermacher
Q4_K_M28.14 GiB30,215,578,6884.847bartowski
I1-Q4_K_M28.14 GiB30,215,579,1044.847mradermacher
Q4_K_L28.87 GiB30,995,375,1684.973bartowski
Q4_129.23 GiB31,383,626,8165.035bartowski
I1-Q4_129.23 GiB31,383,627,2325.035mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,09610.00 GiB10.00 GiB80 / 0 / 0
8,19220.00 GiB20.00 GiB80 / 0 / 0
16,38440.00 GiB40.00 GiB80 / 0 / 0
32,76880.00 GiB80.00 GiB80 / 0 / 0
65,536160.00 GiB160.00 GiB80 / 0 / 0
131,072320.00 GiB320.00 GiB80 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 26.12 GiB. The real file is 28.14 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
80
Attention heads
64
KV heads
64
Head dim
128
Hidden size
8192
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Valkyrie-49B-v2.1 need?
Q4_K_M is exactly 30,215,578,688 bytes (28.14 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Valkyrie-49B-v2.1's KV cache?
80.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Valkyrie-49B-v2.1 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.