PrimeIntellect · text

INTELLECT-1-Instruct

PrimeIntellect/INTELLECT-1-Instruct

INTELLECT-1-Instruct at Q4_K_M is exactly 6,229,006,784 bytes (5.80 GiB / 6.23 GB) — an effective 4.880 bits per weight, not the nominal 4. Its KV cache at 32K is 5.25 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
10.2B
Architecture
llama
42 layers
Context
8,192
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S2.31 GiB2,479,635,5841.943mradermacher
I1-IQ1_M2.48 GiB2,666,806,4002.089mradermacher
I1-IQ2_XXS2.77 GiB2,978,757,7602.334mradermacher
IQ2_XS3.03 GiB3,250,338,5282.546bartowski
I1-IQ2_XS3.03 GiB3,250,338,9442.546mradermacher
IQ2_S3.20 GiB3,432,602,3362.689bartowski
I1-IQ2_S3.20 GiB3,432,602,7522.689mradermacher
IQ2_M3.43 GiB3,682,163,4242.885bartowski
I1-IQ2_M3.43 GiB3,682,163,8402.885mradermacher
I1-Q2_K_S3.47 GiB3,728,399,4882.921mradermacher
Q2_K3.71 GiB3,981,630,1763.119bartowski
I1-Q2_K3.71 GiB3,981,630,5923.119mradermacher
IQ3_XXS3.83 GiB4,112,472,8003.222bartowski
I1-IQ3_XXS3.83 GiB4,112,473,2163.222mradermacher
IQ3_XS4.11 GiB4,413,455,0723.458bartowski
I1-IQ3_XS4.11 GiB4,413,455,4883.458mradermacher
Q2_K_L4.19 GiB4,494,654,1763.521bartowski
Q3_K_S4.29 GiB4,602,002,1443.605bartowski
I1-Q3_K_S4.29 GiB4,602,002,5603.605mradermacher
I1-IQ3_S4.31 GiB4,625,398,9123.624mradermacher
IQ3_M4.43 GiB4,757,977,8243.728bartowski
I1-IQ3_M4.43 GiB4,757,978,2403.728mradermacher
Q3_K_M4.71 GiB5,062,261,4723.966bartowski
I1-Q3_K_M4.71 GiB5,062,261,8883.966mradermacher
Q3_K_L5.09 GiB5,464,914,3684.281lmstudio-community
Q3_K_L5.09 GiB5,464,914,6564.281bartowski
I1-Q3_K_L5.09 GiB5,464,915,0724.281mradermacher
IQ4_XS5.23 GiB5,613,230,8164.398bartowski
I1-IQ4_XS5.23 GiB5,613,231,2324.398mradermacher
Q4_05.50 GiB5,906,733,7924.628bartowski
I1-Q4_05.50 GiB5,906,734,2084.628mradermacher
IQ4_NL5.50 GiB5,910,403,8084.630bartowski
I1-IQ4_NL5.50 GiB5,910,404,2244.630mradermacher
Q4_K_S5.52 GiB5,927,181,0244.644bartowski
I1-Q4_K_S5.52 GiB5,927,181,4404.644mradermacher
Q4_K_M5.80 GiB6,229,006,7844.880lmstudio-community
Q4_K_M5.80 GiB6,229,007,0724.880bartowski
I1-Q4_K_M5.80 GiB6,229,007,4884.880mradermacher
I1-Q4_16.05 GiB6,493,740,1605.088mradermacher
Q4_K_L6.16 GiB6,618,905,3125.186bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.66 GiB0.66 GiB42 / 0 / 0
8,1921.31 GiB1.31 GiB42 / 0 / 0
16,3842.63 GiB2.63 GiB42 / 0 / 0
32,7685.25 GiB5.25 GiB42 / 0 / 0
65,53610.50 GiB10.50 GiB42 / 0 / 0
131,07221.00 GiB21.00 GiB42 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 5.35 GiB. The real file is 5.80 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
42
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does INTELLECT-1-Instruct need?
Q4_K_M is exactly 6,229,006,784 bytes (5.80 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is INTELLECT-1-Instruct's KV cache?
5.25 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of INTELLECT-1-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.