LGAI-EXAONE · vision language

EXAONE-4.5-33B

LGAI-EXAONE/EXAONE-4.5-33B

EXAONE-4.5-33B at Q4_K_M is exactly 20,047,839,424 bytes (18.67 GiB / 20.05 GB) — an effective 4.669 bits per weight, not the nominal 4. Its KV cache at 32K is 2.84 GiB, not the 8.00 GiB a flat formula predicts.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
34.4B
Architecture
exaone4
64 layers
Context
262,144
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K11.57 GiB12,424,016,8642.893mradermacher
I1-Q2_K11.57 GiB12,424,017,0882.893mradermacher
Q3_K_S13.53 GiB14,523,319,2643.382mradermacher
I1-Q3_K_S13.53 GiB14,523,319,4883.382mradermacher
I1-IQ3_S13.57 GiB14,568,580,2883.393mradermacher
I1-IQ3_M13.92 GiB14,943,896,7683.480mradermacher
Q3_K_M14.97 GiB16,077,044,7043.744mradermacher
I1-Q3_K_M14.97 GiB16,077,044,9283.744mradermacher
Q3_K_L16.21 GiB17,400,708,0644.053mradermacher
I1-Q3_K_L16.21 GiB17,400,708,2884.053mradermacher
I1-IQ4_XS16.63 GiB17,854,647,4884.158mradermacher
IQ4_XS16.79 GiB18,029,955,2644.199LGAI-EXAONE
IQ4_XS16.79 GiB18,029,956,0644.199mradermacher
I1-Q4_017.58 GiB18,880,163,0084.397mradermacher
Q4_K_S17.65 GiB18,952,907,7444.414mradermacher
I1-Q4_K_S17.65 GiB18,952,907,9684.414mradermacher
Q4_K_M18.67 GiB20,047,839,4244.669LGAI-EXAONE
Q4_K_M18.67 GiB20,047,840,2244.669mradermacher
I1-Q4_K_M18.67 GiB20,047,840,4484.669mradermacher
I1-Q4_119.40 GiB20,827,319,4884.851mradermacher
Q5_K_S21.28 GiB22,844,599,2645.320mradermacher
I1-Q5_K_S21.28 GiB22,844,599,4885.320mradermacher
Q5_K_M21.87 GiB23,482,253,5045.469LGAI-EXAONE
Q5_K_M21.87 GiB23,482,254,3045.469mradermacher
I1-Q5_K_M21.87 GiB23,482,254,5285.469mradermacher
Q6_K25.27 GiB27,131,318,4646.319LGAI-EXAONE
Q6_K25.27 GiB27,131,319,2646.319mradermacher
I1-Q6_K25.27 GiB27,131,319,4886.319mradermacher
Q8_032.73 GiB35,138,742,4648.184LGAI-EXAONE
Q8_032.73 GiB35,138,743,2648.184mradermacher
BF1661.59 GiB66,135,222,46415.403LGAI-EXAONE

KV cache by context

computed per layer — this model uses sliding-window attention
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.00 GiB1.00 GiB16 / 48 / 0
8,1921.34 GiB2.00 GiB1.49×16 / 48 / 0
16,3841.84 GiB4.00 GiB2.17×16 / 48 / 0
32,7682.84 GiB8.00 GiB2.81×16 / 48 / 0
65,5364.84 GiB16.00 GiB3.30×16 / 48 / 0
131,0728.84 GiB32.00 GiB3.62×16 / 48 / 0

48 of 64 layers cache only a 4,096-token window rather than the full context, on a period of 4. Figures assume the default configuration; --swa-full disables the saving entirely.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 17.99 GiB. The real file is 18.67 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 8.00 GiB at 32K context where the real figure is 2.84 GiB, because most of this model's layers cache a fixed window rather than the whole context.

Architecture

from config.json
Layers
64
Attention heads
40
KV heads
8
Head dim
128
Hidden size
5120
Vocab
153,600
Sliding window
4096
SWA period
4
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does EXAONE-4.5-33B need?
Q4_K_M is exactly 20,047,839,424 bytes (18.67 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is EXAONE-4.5-33B's KV cache?
2.84 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of EXAONE-4.5-33B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.