internlm · text

internlm3-8b-instruct

internlm/internlm3-8b-instruct

internlm3-8b-instruct at Q4_K_M is exactly 5,358,623,936 bytes (4.99 GiB / 5.36 GB) — an effective 4.869 bits per weight, not the nominal 4. Its KV cache at 32K is 1.50 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.8B
Architecture
llama
48 layers
Context
32,768
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M2.98 GiB3,203,374,6562.911bartowski
Q2_K3.21 GiB3,450,641,6003.135internlm
Q2_K3.21 GiB3,450,641,9843.135bartowski
IQ3_XS3.56 GiB3,818,282,5603.470bartowski
Q2_K_L3.69 GiB3,964,689,9843.603bartowski
Q3_K_S3.72 GiB3,993,263,6803.628bartowski
IQ3_M3.86 GiB4,140,326,4643.762bartowski
Q3_K_M4.09 GiB4,390,280,3843.989internlm
Q3_K_M4.09 GiB4,390,280,7683.989bartowski
Q3_K_L4.41 GiB4,732,902,6884.301lmstudio-community
Q3_K_L4.41 GiB4,732,902,9764.301bartowski
IQ4_XS4.51 GiB4,841,807,4244.399bartowski
Q4_04.74 GiB5,092,613,3124.627internlm
IQ4_NL4.75 GiB5,098,905,1524.633bartowski
Q4_04.76 GiB5,108,342,3364.642bartowski
Q4_K_S4.77 GiB5,124,595,2644.657bartowski
Q4_K_M4.99 GiB5,358,623,9364.869internlm
Q4_K_M4.99 GiB5,358,624,0324.869lmstudio-community
Q4_K_M4.99 GiB5,358,624,3204.869bartowski
Q4_15.22 GiB5,609,954,8805.098bartowski
Q4_K_L5.35 GiB5,749,300,8005.224bartowski
Q5_05.71 GiB6,127,295,6805.568internlm
Q5_K_S5.71 GiB6,127,296,0645.568bartowski
Q5_K_M5.83 GiB6,264,331,4565.692internlm
Q5_K_M5.83 GiB6,264,331,8405.692bartowski
Q5_K_L6.14 GiB6,589,210,1765.987bartowski
Q6_K6.73 GiB7,226,645,6966.566internlm
Q6_K6.73 GiB7,226,645,7926.566lmstudio-community
Q6_K6.73 GiB7,226,646,0806.566bartowski
Q6_K_L6.97 GiB7,481,613,8886.798bartowski
Q8_08.72 GiB9,358,826,6888.504internlm
Q8_08.72 GiB9,358,826,7848.504lmstudio-community
Q8_08.72 GiB9,358,827,0728.504bartowski
F1616.40 GiB17,612,430,91216.004bartowski
F3232.80 GiB35,220,118,81632.003bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.19 GiB0.19 GiB48 / 0 / 0
8,1920.38 GiB0.38 GiB48 / 0 / 0
16,3840.75 GiB0.75 GiB48 / 0 / 0
32,7681.50 GiB1.50 GiB48 / 0 / 0
65,5363.00 GiB3.00 GiB48 / 0 / 0
131,0726.00 GiB6.00 GiB48 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.61 GiB. The real file is 4.99 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
48
Attention heads
32
KV heads
2
Head dim
128
Hidden size
4096
Vocab
128,512
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does internlm3-8b-instruct need?
Q4_K_M is exactly 5,358,623,936 bytes (4.99 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is internlm3-8b-instruct's KV cache?
1.50 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of internlm3-8b-instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.