DavidAU · text

L3.2-8X4B-MOE-V2-Dark-Champion-Inst-21B-uncen-ablit

DavidAU/L3.2-8X4B-MOE-V2-Dark-Champion-Inst-21B-uncen-ablit

L3.2-8X4B-MOE-V2-Dark-Champion-Inst-21B-uncen-ablit at Q4_K_M is exactly 12,850,124,384 bytes (11.97 GiB / 12.85 GB) — an effective 4.914 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
20.9B
Architecture
llama
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S4.35 GiB4,673,622,5281.787mradermacher
I1-IQ1_M4.76 GiB5,114,810,8801.956mradermacher
I1-IQ2_XXS5.45 GiB5,850,124,8002.237mradermacher
I1-IQ2_XS6.00 GiB6,438,375,9362.462mradermacher
I1-IQ2_S6.11 GiB6,560,180,7362.509mradermacher
I1-IQ2_M6.66 GiB7,148,431,8722.733mradermacher
I1-Q2_K_S6.90 GiB7,406,897,6642.832mradermacher
Q2_K7.43 GiB7,980,992,0963.052DavidAU
Q2_K7.43 GiB7,980,992,0963.052themrmmdhosen
I1-Q2_K7.43 GiB7,980,993,0243.052mradermacher
I1-IQ3_XXS7.79 GiB8,368,974,3363.200mradermacher
I1-IQ3_XS8.28 GiB8,893,161,9843.401mradermacher
Q3_K_S8.72 GiB9,360,301,6643.579themrmmdhosen
Q3_K_S8.72 GiB9,360,301,6643.579DavidAU
I1-IQ3_S8.72 GiB9,360,302,5923.579mradermacher
I1-Q3_K_S8.72 GiB9,360,302,5923.579mradermacher
I1-IQ3_M9.12 GiB9,788,121,6003.743mradermacher
Q3_K_M9.56 GiB10,266,271,3283.926DavidAU
Q3_K_M9.56 GiB10,266,271,3283.926themrmmdhosen
I1-Q3_K_M9.56 GiB10,266,272,2563.926mradermacher
I1-Q3_K_L10.19 GiB10,943,390,2084.184mradermacher
I1-IQ4_XS10.61 GiB11,393,923,5844.357mradermacher
IQ4_XS10.73 GiB11,519,751,7764.405themrmmdhosen
IQ4_XS10.73 GiB11,519,751,7764.405DavidAU
I1-Q4_011.21 GiB12,032,236,0324.601mradermacher
Q4_K_S11.29 GiB12,120,315,4884.635themrmmdhosen
Q4_K_S11.29 GiB12,120,315,4884.635DavidAU
I1-Q4_K_S11.29 GiB12,120,316,4164.635mradermacher
Q4_K_M11.97 GiB12,850,124,3844.914themrmmdhosen
Q4_K_M11.97 GiB12,850,124,3844.914DavidAU
I1-Q4_K_M11.97 GiB12,850,125,3124.914mradermacher
I1-Q4_112.34 GiB13,252,237,8245.067mradermacher
Q5_K_S13.53 GiB14,522,570,3365.553DavidAU
Q5_K_S13.53 GiB14,522,570,3365.553themrmmdhosen
I1-Q5_K_S13.53 GiB14,522,571,2645.553mradermacher
Q5_K_M13.92 GiB14,950,389,3445.717DavidAU
Q5_K_M13.92 GiB14,950,389,3445.717themrmmdhosen
I1-Q5_K_M13.92 GiB14,950,390,2725.717mradermacher
Q6_K16.04 GiB17,222,028,8966.585DavidAU
Q6_K16.04 GiB17,222,028,8966.585themrmmdhosen

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 10.96 GiB. The real file is 11.97 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does L3.2-8X4B-MOE-V2-Dark-Champion-Inst-21B-uncen-ablit need?
Q4_K_M is exactly 12,850,124,384 bytes (11.97 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of L3.2-8X4B-MOE-V2-Dark-Champion-Inst-21B-uncen-ablit should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.