mistralai · text

Mistral-Large-3-675B-Instruct-2512

mistralai/Mistral-Large-3-675B-Instruct-2512

Mistral-Large-3-675B-Instruct-2512 at Q4_K_M is exactly 406,989,520,384 bytes (379.04 GiB / 406.99 GB) — an effective 4.835 bits per weight, not the nominal 4.

From the file· summed from 9 file(s)
Parameters
673B
Architecture
deepseek2
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S4 shards132.50 GiB142,267,169,7601.690bartowski
IQ1_M4 shards138.29 GiB148,492,663,7761.764bartowski
IQ2_XXS5 shards153.81 GiB165,151,554,6561.962bartowski
UD-TQ1_0157.03 GiB168,604,875,2002.003unsloth
IQ2_XS5 shards176.52 GiB189,541,563,5202.252bartowski
IQ2_S5 shards177.66 GiB190,766,333,0562.266bartowski
UD-IQ1_S4 shards177.70 GiB190,801,132,3522.267unsloth
UD-IQ1_M5 shards192.28 GiB206,460,042,2402.453unsloth
IQ2_M6 shards201.36 GiB216,213,306,5922.568bartowski
UD-IQ2_XXS5 shards209.37 GiB224,807,238,6562.671unsloth
UD-IQ2_M5 shards219.21 GiB235,376,884,6722.796unsloth
Q2_K7 shards223.36 GiB239,826,713,9522.849bartowski
Q2_K_L7 shards224.21 GiB240,744,217,9522.860bartowski
Q2_K5 shards230.14 GiB247,110,065,1522.936unsloth
Q2_K_L6 shards230.34 GiB247,330,266,2082.938unsloth
IQ3_XXS7 shards251.10 GiB269,619,903,8083.203bartowski
IQ3_XS8 shards259.96 GiB279,130,750,4323.316bartowski
UD-IQ3_XXS6 shards260.04 GiB279,215,127,6163.317unsloth
Q3_K_S6 shards271.81 GiB291,854,199,8723.467unsloth
Q3_K_S8 shards275.05 GiB295,337,541,0243.509bartowski
IQ3_M8 shards288.58 GiB309,858,876,8643.681bartowski
Q3_K_M8 shards288.62 GiB309,902,917,0563.682bartowski
Q3_K_L9 shards299.59 GiB321,679,081,0243.821bartowski
Q3_K_M7 shards299.74 GiB321,845,570,7203.823unsloth
IQ4_XS8 shards335.23 GiB359,954,316,7364.276unsloth
IQ4_XS10 shards337.09 GiB361,952,919,2004.300bartowski
IQ4_NL8 shards354.59 GiB380,738,813,2484.523unsloth
Q4_08 shards355.49 GiB381,700,357,4404.535unsloth
IQ4_NL10 shards356.18 GiB382,449,958,6244.543bartowski
Q4_K_S8 shards356.38 GiB382,661,901,6324.546unsloth
Q4_010 shards361.71 GiB388,388,044,4804.614bartowski
Q4_K_S11 shards368.91 GiB396,117,098,3364.706bartowski
Q4_K_M9 shards379.04 GiB406,989,520,3844.835unsloth
Q4_K_M11 shards383.10 GiB411,354,713,9524.887bartowski
Q4_19 shards393.36 GiB422,366,526,9125.018unsloth
Q4_111 shards393.95 GiB422,996,295,4245.025bartowski
Q5_K_S10 shards432.55 GiB464,446,570,0805.518unsloth
Q5_K_M10 shards445.14 GiB477,969,661,5365.678unsloth
Q5_K_M13 shards447.75 GiB480,762,839,1685.711bartowski
Q6_K12 shards515.32 GiB553,317,879,6166.573unsloth

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 352.78 GiB. The real file is 379.04 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Mistral-Large-3-675B-Instruct-2512 need?
Q4_K_M is exactly 406,989,520,384 bytes (379.04 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Mistral-Large-3-675B-Instruct-2512 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.