ibm-granite · text

granite-4.1-8b

ibm-granite/granite-4.1-8b

granite-4.1-8b at Q4_K_M is exactly 5,347,914,400 bytes (4.98 GiB / 5.35 GB) — an effective 4.866 bits per weight, not the nominal 4. Its KV cache at 32K is 5.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.8B
Architecture
40 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
UD-IQ2_M3.05 GiB3,272,078,7202.978unsloth
Q2_K3.18 GiB3,412,308,6403.105ibm-granite
UD-IQ3_XXS3.36 GiB3,606,148,4803.281unsloth
Q3_K_S3.67 GiB3,942,953,6323.588ibm-granite
Q3_K_S3.67 GiB3,942,954,3683.588unsloth
Q3_K_M4.05 GiB4,347,048,6083.956ibm-granite
Q3_K_M4.05 GiB4,347,049,3443.956unsloth
Q3_K_L4.38 GiB4,699,894,4324.277ibm-granite
IQ4_XS4.50 GiB4,833,129,8564.398unsloth
Q4_04.71 GiB5,055,951,5204.601ibm-granite
Q4_04.72 GiB5,072,336,2564.616unsloth
IQ4_NL4.73 GiB5,076,923,7764.620unsloth
Q4_K_S4.74 GiB5,090,816,6724.632ibm-granite
Q4_K_S4.74 GiB5,090,817,4084.632unsloth
Q4_K_M4.98 GiB5,347,914,4004.866ibm-granite
Q4_K_M4.98 GiB5,347,915,1364.866unsloth
Q4_15.20 GiB5,579,715,2325.077ibm-granite
Q4_15.20 GiB5,579,715,9685.077unsloth
Q5_05.68 GiB6,103,478,9445.554ibm-granite
Q5_K_S5.68 GiB6,103,478,9445.554ibm-granite
Q5_K_S5.68 GiB6,103,479,6805.554unsloth
Q5_K_M5.82 GiB6,253,884,0645.691ibm-granite
Q5_K_M5.82 GiB6,253,884,8005.691unsloth
Q5_16.17 GiB6,627,242,6566.030ibm-granite
Q6_K6.72 GiB7,216,476,8326.567ibm-granite
Q6_K6.72 GiB7,216,477,5686.567unsloth
Q8_08.70 GiB9,345,610,4008.504ibm-granite
Q8_08.70 GiB9,345,611,1368.504unsloth
BF1616.38 GiB17,587,417,76016.004ibm-granite
BF1616.38 GiB17,587,418,24016.004unsloth

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.63 GiB0.63 GiB40 / 0 / 0
8,1921.25 GiB1.25 GiB40 / 0 / 0
16,3842.50 GiB2.50 GiB40 / 0 / 0
32,7685.00 GiB5.00 GiB40 / 0 / 0
65,53610.00 GiB10.00 GiB40 / 0 / 0
131,07220.00 GiB20.00 GiB40 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.61 GiB. The real file is 4.98 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
40
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
100,352
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does granite-4.1-8b need?
Q4_K_M is exactly 5,347,914,400 bytes (4.98 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is granite-4.1-8b's KV cache?
5.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of granite-4.1-8b should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.