ibm-granite · text

granite-3.1-8b-instruct

ibm-granite/granite-3.1-8b-instruct

granite-3.1-8b-instruct at Q4_K_M is exactly 4,899,475,200 bytes (4.56 GiB / 4.90 GB) — an effective 4.797 bits per weight, not the nominal 4. Its KV cache at 32K is 5.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.2B
Architecture
granite
40 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
TQ1_01.72 GiB1,849,232,9601.811Mungert
TQ2_02.07 GiB2,222,788,1602.176Mungert
IQ2_XXS2.29 GiB2,463,434,9122.412Mungert
IQ2_XS2.50 GiB2,689,534,1122.633Mungert
IQ2_S2.60 GiB2,788,493,4722.730Mungert
IQ2_M2.64 GiB2,836,824,9602.777bartowski
IQ2_M2.73 GiB2,935,294,1122.874Mungert
Q2_K_S2.82 GiB3,029,403,8082.966Mungert
Q2_K_M2.85 GiB3,056,144,1282.992Mungert
Q2_K2.89 GiB3,103,590,8803.039bartowski
Q2_K_L2.94 GiB3,152,352,6403.086bartowski
IQ3_XXS3.10 GiB3,323,398,3043.254Mungert
IQ3_XS3.19 GiB3,424,979,1043.353Mungert
IQ3_XS3.19 GiB3,427,994,0803.356bartowski
Q3_K_S3.35 GiB3,592,489,4403.517bartowski
IQ3_M3.48 GiB3,738,716,6403.660bartowski
IQ3_S3.56 GiB3,823,569,0563.744Mungert
IQ3_M3.56 GiB3,823,569,0563.744Mungert
Q3_K_S3.67 GiB3,938,912,4163.857Mungert
Q3_K_M3.69 GiB3,965,652,7363.883Mungert
Q3_K_M3.72 GiB3,996,584,4163.913bartowski
Q3_K_L4.05 GiB4,349,429,9524.258lmstudio-community
Q3_K_L4.05 GiB4,349,430,2404.258bartowski
IQ4_XS4.12 GiB4,428,073,4404.335bartowski
IQ4_XS4.12 GiB4,428,074,7524.335Mungert
Q4_04.28 GiB4,598,989,4724.503Mungert
Q4_04.35 GiB4,667,279,8404.570bartowski
IQ4_NL4.35 GiB4,671,867,3604.574bartowski
IQ4_NL4.35 GiB4,671,868,6724.574Mungert
Q4_K_S4.36 GiB4,685,760,9924.588bartowski
Q4_K_S4.41 GiB4,730,851,0724.632Mungert
Q4_K_M4.56 GiB4,899,475,2004.797Mungert
Q4_K_M4.60 GiB4,942,858,4324.840lmstudio-community
Q4_K_M4.60 GiB4,942,858,7204.840bartowski
Q4_K_L4.65 GiB4,991,620,4804.887bartowski
Q4_14.76 GiB5,109,646,7525.003Mungert
Q5_05.23 GiB5,620,304,0325.503Mungert
Q5_K_S5.26 GiB5,647,043,0405.529bartowski
Q5_K_S5.33 GiB5,725,032,1925.605Mungert
Q5_K_M5.40 GiB5,797,448,1605.676bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.63 GiB0.63 GiB40 / 0 / 0
8,1921.25 GiB1.25 GiB40 / 0 / 0
16,3842.50 GiB2.50 GiB40 / 0 / 0
32,7685.00 GiB5.00 GiB40 / 0 / 0
65,53610.00 GiB10.00 GiB40 / 0 / 0
131,07220.00 GiB20.00 GiB40 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.28 GiB. The real file is 4.56 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
40
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
49,155
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does granite-3.1-8b-instruct need?
Q4_K_M is exactly 4,899,475,200 bytes (4.56 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is granite-3.1-8b-instruct's KV cache?
5.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of granite-3.1-8b-instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.