utter-project · text

EuroLLM-9B-Instruct

utter-project/EuroLLM-9B-Instruct

EuroLLM-9B-Instruct at Q4_K_M is exactly 5,582,838,208 bytes (5.20 GiB / 5.58 GB) — an effective 4.880 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
9.2B
Architecture
llama
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M3.10 GiB3,333,085,9202.913bartowski
Q2_K3.35 GiB3,593,067,2323.141bartowski
IQ3_XS3.70 GiB3,977,599,7123.477bartowski
Q2_K_L3.82 GiB4,105,067,2323.588bartowski
Q3_K_S3.86 GiB4,141,767,3923.620bartowski
IQ3_M4.00 GiB4,292,172,5123.752bartowski
Q3_K_M4.24 GiB4,553,136,8643.980bartowski
Q3_K_L4.58 GiB4,913,846,7204.295lmstudio-community
Q3_K_L4.58 GiB4,913,847,0084.295bartowski
IQ4_XS4.70 GiB5,045,541,6004.410bartowski
Q4_04.94 GiB5,303,360,2244.636bartowski
IQ4_NL4.94 GiB5,309,651,6804.641bartowski
Q4_K_S4.96 GiB5,321,186,0164.651bartowski
Q4_K_M5.20 GiB5,582,838,2084.880lmstudio-community
Q4_K_M5.20 GiB5,582,838,4964.880bartowski
Q4_K_L5.56 GiB5,971,958,4965.220bartowski
Q5_K_S5.93 GiB6,366,092,0005.565bartowski
Q5_K_M6.07 GiB6,518,168,2885.697bartowski
Q5_K_L6.37 GiB6,841,752,2885.980bartowski
Q6_K7.00 GiB7,511,955,9046.566lmstudio-community
Q6_K7.00 GiB7,511,956,1926.566bartowski
Q6_K_L7.23 GiB7,765,908,1926.788bartowski
Q8_09.06 GiB9,728,448,9608.504lmstudio-community
Q8_09.06 GiB9,728,449,2488.504bartowski
F1617.05 GiB18,308,422,08016.003bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.79 GiB. The real file is 5.20 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does EuroLLM-9B-Instruct need?
Q4_K_M is exactly 5,582,838,208 bytes (5.20 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of EuroLLM-9B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.