LiquidAI · text · mixture of experts

LFM2-24B-A2B

LiquidAI/LFM2-24B-A2B

LFM2-24B-A2B at Q4_K_M is exactly 14,415,473,952 bytes (13.43 GiB / 14.42 GB) — an effective 4.837 bits per weight, not the nominal 4. Its KV cache at 32K is 0.63 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
23.8B
total, not active
Architecture
lfm2moe
40 layers
Context
128,000
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS5.35 GiB5,746,134,9441.928bartowski
IQ2_XS6.17 GiB6,621,958,0482.222bartowski
IQ2_S6.45 GiB6,929,190,8162.325bartowski
IQ2_M7.21 GiB7,738,953,6322.597bartowski
Q2_K7.75 GiB8,320,749,4722.792bartowski
Q2_K_L7.78 GiB8,353,255,3282.803bartowski
IQ3_XXS8.75 GiB9,390,100,3843.151bartowski
IQ3_XS9.11 GiB9,780,662,1763.282bartowski
Q3_K_S9.64 GiB10,348,859,2963.472bartowski
IQ3_M10.09 GiB10,836,561,8243.636bartowski
Q3_K_M10.10 GiB10,842,591,1363.638bartowski
Q3_K_L10.50 GiB11,279,191,9683.784bartowski
IQ4_XS11.87 GiB12,749,918,1124.278bartowski
Q4_012.54 GiB13,467,405,6004.519LiquidAI
IQ4_NL12.56 GiB13,488,705,4404.526bartowski
Q4_012.77 GiB13,707,399,0724.599bartowski
Q4_K_S12.99 GiB13,947,719,5844.680bartowski
Q4_K_M13.43 GiB14,415,473,9524.837LiquidAI
Q4_K_M13.44 GiB14,435,422,1124.843bartowski
Q4_K_L13.47 GiB14,467,927,9684.854bartowski
Q4_113.93 GiB14,958,088,0965.019bartowski
Q5_K_S15.31 GiB16,443,854,7525.517bartowski
Q5_K_M15.76 GiB16,918,818,0805.677LiquidAI
Q5_K_M15.77 GiB16,931,557,2805.681bartowski
Q5_K_L15.80 GiB16,964,063,1365.692bartowski
Q6_K18.23 GiB19,578,621,2166.569LiquidAI
Q6_K18.24 GiB19,583,700,8966.571bartowski
Q6_K_L18.27 GiB19,616,206,7526.582bartowski
Q8_023.61 GiB25,351,965,9848.506LiquidAI
Q8_023.61 GiB25,351,966,6248.506bartowski
BF1644.42 GiB47,700,397,34416.004LiquidAI
F1644.42 GiB47,700,397,34416.004LiquidAI
BF1644.42 GiB47,700,397,69616.004bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.08 GiB0.31 GiB4.00×10 / 0 / 30
8,1920.16 GiB0.63 GiB4.00×10 / 0 / 30
16,3840.31 GiB1.25 GiB4.00×10 / 0 / 30
32,7680.63 GiB2.50 GiB4.00×10 / 0 / 30
65,5361.25 GiB5.00 GiB4.00×10 / 0 / 30
131,0722.50 GiB10.00 GiB4.00×10 / 0 / 30

30 of 40 layers use linear attention, which keeps a fixed-size recurrent state instead of a per-token cache. Those layers do not grow with context at all — treating them as ordinary attention, as a flat formula does, overstates this model's cache by roughly 4.0× at long context.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 12.49 GiB. The real file is 13.43 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
40
Attention heads
32
KV heads
8
Head dim
64
Hidden size
2048
Vocab
65,536
Sliding window
none
SWA period
MLA
no
Experts
64
Experts per token
4
use_sliding_window

Questions people ask

How much VRAM does LFM2-24B-A2B need?
Q4_K_M is exactly 14,415,473,952 bytes (13.43 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is LFM2-24B-A2B's KV cache?
0.63 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Is LFM2-24B-A2B a mixture-of-experts model?
Yes — 64 experts, 4 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
Which quantization of LFM2-24B-A2B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.