RekaAI · text

reka-flash-3

RekaAI/reka-flash-3

reka-flash-3 at Q4_K_M is exactly 13,610,362,336 bytes (12.68 GiB / 13.61 GB) — an effective 5.208 bits per weight, not the nominal 4. Its KV cache at 32K is 4.13 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
20.9B
Architecture
llama
44 layers
Context
32,768
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS6.88 GiB7,385,212,6722.826bartowski
IQ2_XS7.29 GiB7,827,482,3682.995bartowski
IQ2_S7.57 GiB8,123,672,3203.109bartowski
IQ2_M7.93 GiB8,514,037,5043.258bartowski
Q2_K8.04 GiB8,630,896,2243.303unsloth
Q2_K8.04 GiB8,630,896,3843.303bartowski
Q2_K_L8.17 GiB8,775,403,1043.358unsloth
IQ3_XXS8.55 GiB9,177,982,7203.512bartowski
Q2_K_L8.60 GiB9,233,008,3843.533bartowski
IQ3_XS8.85 GiB9,501,144,8323.636bartowski
Q3_K_S9.25 GiB9,934,628,6083.802bartowski
IQ3_M9.55 GiB10,258,245,3763.926bartowski
Q3_K_M10.12 GiB10,863,011,4244.157unsloth
Q3_K_M10.12 GiB10,863,011,5844.157bartowski
Q3_K_L10.63 GiB11,412,284,8964.367lmstudio-community
Q3_K_L10.63 GiB11,412,285,1844.367bartowski
IQ4_XS10.70 GiB11,488,151,2964.396bartowski
IQ4_NL11.13 GiB11,949,688,5764.573bartowski
Q4_011.14 GiB11,961,460,4804.577bartowski
Q4_K_S11.76 GiB12,627,764,9924.832bartowski
Q4_112.29 GiB13,191,759,6165.048bartowski
Q4_K_M12.68 GiB13,610,362,3365.208lmstudio-community
Q4_K_M12.68 GiB13,610,362,4645.208unsloth
Q4_K_M12.68 GiB13,610,362,6245.208bartowski
Q4_K_L13.10 GiB14,067,967,7445.383bartowski
Q5_K_S13.78 GiB14,791,755,5205.660bartowski
Q5_K_M14.56 GiB15,635,474,0165.983unsloth
Q5_K_M14.56 GiB15,635,474,1765.983bartowski
Q5_K_L14.92 GiB16,016,008,9606.129bartowski
Q6_K17.17 GiB18,440,725,9847.057lmstudio-community
Q6_K17.17 GiB18,440,726,1127.057unsloth
Q6_K17.17 GiB18,440,726,2727.057bartowski
Q6_K_L17.45 GiB18,739,373,8247.171bartowski
Q8_020.69 GiB22,217,246,1768.502lmstudio-community
Q8_020.69 GiB22,217,246,3048.502unsloth
Q8_020.69 GiB22,217,246,4648.502bartowski
BF1638.94 GiB41,815,623,13616.002bartowski
BF1638.94 GiB41,815,623,26416.002unsloth

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.52 GiB0.52 GiB44 / 0 / 0
8,1921.03 GiB1.03 GiB44 / 0 / 0
16,3842.06 GiB2.06 GiB44 / 0 / 0
32,7684.13 GiB4.13 GiB44 / 0 / 0
65,5368.25 GiB8.25 GiB44 / 0 / 0
131,07216.50 GiB16.50 GiB44 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 10.95 GiB. The real file is 12.68 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
44
Attention heads
64
KV heads
8
Head dim
96
Hidden size
6144
Vocab
100,352
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does reka-flash-3 need?
Q4_K_M is exactly 13,610,362,336 bytes (12.68 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is reka-flash-3's KV cache?
4.13 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of reka-flash-3 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.