mistralai · text

Mistral-Large-Instruct-2411

mistralai/Mistral-Large-Instruct-2411

Mistral-Large-Instruct-2411 at Q4_K_M is exactly 73,219,623,136 bytes (68.19 GiB / 73.22 GB) — an effective 4.777 bits per weight, not the nominal 4. Its KV cache at 32K is 11.00 GiB.

From the file· summed from 2 file(s)From the file· KV per layer
Parameters
123B
Architecture
llama
88 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_M26.44 GiB28,386,314,5281.852bartowski
IQ2_XXS30.20 GiB32,430,541,0882.116bartowski
IQ2_XS33.60 GiB36,081,158,4322.354bartowski
IQ2_M38.76 GiB41,619,605,7922.716bartowski
Q2_K42.09 GiB45,196,298,2722.949MaziyarPanahi
Q2_K42.09 GiB45,196,298,5282.949bartowski
Q2_K_L42.46 GiB45,589,514,5282.975bartowski
IQ3_XXS43.78 GiB47,009,024,2883.067bartowski
Q3_K_S2 shards49.22 GiB52,849,854,9763.448bartowski
IQ3_M2 shards51.48 GiB55,276,390,8803.607bartowski
Q3_K_M2 shards55.04 GiB59,102,775,8083.856bartowski
Q3_K_L2 shards60.12 GiB64,554,322,1444.212lmstudio-community
Q3_K_L2 shards60.12 GiB64,554,322,4324.212bartowski
IQ4_XS2 shards60.94 GiB65,434,339,8404.269bartowski
Q4_02 shards64.56 GiB69,322,459,6484.523bartowski
Q4_K_S2 shards64.79 GiB69,570,972,1604.539bartowski
Q4_K_M2 shards68.19 GiB73,219,623,1364.777lmstudio-community
Q4_K_M2 shards68.19 GiB73,219,623,4244.777bartowski
Q5_K_S3 shards78.56 GiB84,355,893,8565.504bartowski
Q5_K_M3 shards80.55 GiB86,488,304,2245.643bartowski
Q6_K3 shards93.68 GiB100,586,277,1846.563lmstudio-community
Q6_K3 shards93.68 GiB100,586,277,4726.563bartowski
Q8_04 shards121.33 GiB130,280,376,8008.501lmstudio-community
Q8_04 shards121.33 GiB130,280,376,8008.501bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.38 GiB1.38 GiB88 / 0 / 0
8,1922.75 GiB2.75 GiB88 / 0 / 0
16,3845.50 GiB5.50 GiB88 / 0 / 0
32,76811.00 GiB11.00 GiB88 / 0 / 0
65,53622.00 GiB22.00 GiB88 / 0 / 0
131,07244.00 GiB44.00 GiB88 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 64.23 GiB. The real file is 68.19 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
88
Attention heads
96
KV heads
8
Head dim
128
Hidden size
12288
Vocab
32,768
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Mistral-Large-Instruct-2411 need?
Q4_K_M is exactly 73,219,623,136 bytes (68.19 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Mistral-Large-Instruct-2411's KV cache?
11.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Mistral-Large-Instruct-2411 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.