baidu · text

ERNIE-4.5-21B-A3B-PT

baidu/ERNIE-4.5-21B-A3B-PT

ERNIE-4.5-21B-A3B-PT at Q4_K_M is exactly 13,331,015,456 bytes (12.42 GiB / 13.33 GB) — an effective 4.859 bits per weight, not the nominal 4. Its KV cache at 32K is 1.75 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
21.9B
Architecture
ernie4_5-moe
28 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XS5.91 GiB6,346,976,3522.313bartowski
IQ2_S5.93 GiB6,369,913,9522.322bartowski
UD-TQ1_06.06 GiB6,504,202,0162.371unsloth
UD-IQ1_S6.58 GiB7,066,500,8962.576unsloth
IQ2_M6.67 GiB7,164,537,9522.611bartowski
UD-IQ1_M6.74 GiB7,235,276,5762.637unsloth
UD-IQ2_XXS7.22 GiB7,757,506,3362.828unsloth
UD-IQ2_M7.47 GiB8,025,599,7762.925unsloth
Q2_K7.54 GiB8,091,872,3522.949bartowski
Q2_K_L7.59 GiB8,152,596,2562.971unsloth
Q2_K7.59 GiB8,152,596,2562.971unsloth
Q2_K_L7.60 GiB8,155,995,2322.973bartowski
IQ3_XXS8.38 GiB8,993,094,7523.278bartowski
IQ3_XS8.71 GiB9,350,859,8723.408bartowski
UD-IQ3_XXS8.89 GiB9,540,402,9763.477unsloth
Q3_K_S8.98 GiB9,639,608,0963.514unsloth
Q3_K_S9.17 GiB9,850,039,3923.590bartowski
IQ3_M9.59 GiB10,293,595,2323.752bartowski
Q3_K_M9.59 GiB10,297,855,0723.753bartowski
Q3_K_M9.80 GiB10,524,815,1363.836unsloth
Q3_K_L9.93 GiB10,663,750,4643.887lmstudio-community
Q3_K_L9.93 GiB10,663,750,7523.887bartowski
IQ4_XS11.01 GiB11,823,021,8564.309unsloth
IQ4_XS11.14 GiB11,958,004,8324.359bartowski
IQ4_NL11.62 GiB12,475,596,5764.547unsloth
Q4_011.64 GiB12,501,360,4164.557unsloth
Q4_K_S11.68 GiB12,536,094,4964.569unsloth
IQ4_NL11.74 GiB12,603,698,2724.594bartowski
Q4_011.90 GiB12,782,611,5524.659bartowski
Q4_K_S12.12 GiB13,012,642,9124.743bartowski
Q4_K_M12.42 GiB13,331,015,4564.859unsloth
Q4_K_M12.57 GiB13,499,267,9044.920lmstudio-community
Q4_K_M12.57 GiB13,499,268,1924.920bartowski
Q4_K_L12.63 GiB13,563,391,0724.944bartowski
Q4_112.83 GiB13,778,452,2565.022unsloth
Q4_112.93 GiB13,881,322,5925.060bartowski
Q5_K_S14.10 GiB15,142,297,3765.519unsloth
Q5_K_S14.18 GiB15,230,340,1925.551bartowski
Q5_K_M14.51 GiB15,585,330,9765.681unsloth
Q5_K_M14.67 GiB15,751,453,7925.741bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.22 GiB0.22 GiB28 / 0 / 0
8,1920.44 GiB0.44 GiB28 / 0 / 0
16,3840.88 GiB0.88 GiB28 / 0 / 0
32,7681.75 GiB1.75 GiB28 / 0 / 0
65,5363.50 GiB3.50 GiB28 / 0 / 0
131,0727.00 GiB7.00 GiB28 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 11.50 GiB. The real file is 12.42 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
28
Attention heads
20
KV heads
4
Head dim
128
Hidden size
2560
Vocab
103,424
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does ERNIE-4.5-21B-A3B-PT need?
Q4_K_M is exactly 13,331,015,456 bytes (12.42 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is ERNIE-4.5-21B-A3B-PT's KV cache?
1.75 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of ERNIE-4.5-21B-A3B-PT should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.