RekaAI · text

reka-flash-3.1

RekaAI/reka-flash-3.1

reka-flash-3.1 at Q4_K_M is exactly 13,610,362,432 bytes (12.68 GiB / 13.61 GB) — an effective 5.208 bits per weight, not the nominal 4. Its KV cache at 32K is 4.13 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
20.9B
Architecture
llama
44 layers
Context
98,304
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S6.15 GiB6,604,482,7842.527mradermacher
I1-IQ1_M6.42 GiB6,897,256,6722.639mradermacher
I1-IQ2_XXS6.88 GiB7,385,213,1522.826mradermacher
I1-IQ2_XS7.29 GiB7,827,482,8482.995mradermacher
IQ2_S7.57 GiB8,123,672,4163.109bartowski
I1-IQ2_S7.57 GiB8,123,672,8003.109mradermacher
IQ2_M7.93 GiB8,514,037,6003.258bartowski
I1-IQ2_M7.93 GiB8,514,037,9843.258mradermacher
I1-Q2_K_S7.95 GiB8,537,655,5203.267mradermacher
Q2_K8.04 GiB8,630,896,4803.303bartowski
I1-Q2_K8.04 GiB8,630,896,8643.303mradermacher
IQ3_XXS8.55 GiB9,177,982,8163.512bartowski
I1-IQ3_XXS8.55 GiB9,177,983,2003.512mradermacher
Q2_K_L8.60 GiB9,233,008,4803.533bartowski
IQ3_XS8.85 GiB9,501,144,9283.636bartowski
I1-IQ3_XS8.85 GiB9,501,145,3123.636mradermacher
Q3_K_S9.25 GiB9,934,628,7043.802bartowski
I1-Q3_K_S9.25 GiB9,934,629,0883.802mradermacher
I1-IQ3_S9.28 GiB9,962,203,3603.812mradermacher
IQ3_M9.55 GiB10,258,245,4723.926bartowski
I1-IQ3_M9.55 GiB10,258,245,8563.926mradermacher
Q3_K_M10.12 GiB10,863,011,6804.157bartowski
I1-Q3_K_M10.12 GiB10,863,012,0644.157mradermacher
Q3_K_L10.63 GiB11,412,284,9924.367lmstudio-community
Q3_K_L10.63 GiB11,412,285,2804.367bartowski
I1-Q3_K_L10.63 GiB11,412,285,6644.367mradermacher
IQ4_XS10.70 GiB11,488,151,3924.396bartowski
I1-IQ4_XS10.70 GiB11,488,151,7764.396mradermacher
IQ4_NL11.13 GiB11,949,688,6724.573bartowski
Q4_011.14 GiB11,961,460,5764.577bartowski
I1-Q4_011.14 GiB11,961,460,9604.577mradermacher
Q4_K_S11.76 GiB12,627,765,0884.832bartowski
I1-Q4_K_S11.76 GiB12,627,765,4724.832mradermacher
Q4_112.29 GiB13,191,759,7125.048bartowski
I1-Q4_112.29 GiB13,191,760,0965.048mradermacher
Q4_K_M12.68 GiB13,610,362,4325.208lmstudio-community
Q4_K_M12.68 GiB13,610,362,7205.208bartowski
I1-Q4_K_M12.68 GiB13,610,363,1045.208mradermacher
Q4_K_L13.10 GiB14,067,967,8405.383bartowski
Q5_K_S13.78 GiB14,791,755,6165.660bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.52 GiB0.52 GiB44 / 0 / 0
8,1921.03 GiB1.03 GiB44 / 0 / 0
16,3842.06 GiB2.06 GiB44 / 0 / 0
32,7684.13 GiB4.13 GiB44 / 0 / 0
65,5368.25 GiB8.25 GiB44 / 0 / 0
131,07216.50 GiB16.50 GiB44 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 10.95 GiB. The real file is 12.68 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
44
Attention heads
64
KV heads
8
Head dim
96
Hidden size
6144
Vocab
100,352
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does reka-flash-3.1 need?
Q4_K_M is exactly 13,610,362,432 bytes (12.68 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is reka-flash-3.1's KV cache?
4.13 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of reka-flash-3.1 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.