huihui-ai · text

Llama-3.1-Nemotron-70B-Instruct-HF-abliterated

huihui-ai/Llama-3.1-Nemotron-70B-Instruct-HF-abliterated

Llama-3.1-Nemotron-70B-Instruct-HF-abliterated at Q4_K_M is exactly 42,520,398,816 bytes (39.60 GiB / 42.52 GB) — an effective 4.821 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
70.6B
Architecture
llama
Context
native (config.json)
License
llama3.1

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S14.29 GiB15,343,488,4161.740mradermacher
IQ1_M15.60 GiB16,751,201,2481.899bartowski
I1-IQ1_M15.60 GiB16,751,201,6961.899mradermacher
IQ2_XXS17.79 GiB19,097,390,0482.165bartowski
I1-IQ2_XXS17.79 GiB19,097,390,4962.165mradermacher
IQ2_XS19.69 GiB21,142,113,2482.397bartowski
I1-IQ2_XS19.69 GiB21,142,113,6962.397mradermacher
I1-IQ2_S20.71 GiB22,242,348,4482.522mradermacher
IQ2_M22.46 GiB24,119,299,0402.735bartowski
I1-IQ2_M22.46 GiB24,119,299,4882.735mradermacher
Q2_K24.56 GiB26,375,113,6962.991bartowski
I1-Q2_K24.56 GiB26,375,114,1442.991mradermacher
Q2_K_L25.52 GiB27,401,161,6963.107bartowski
IQ3_XXS25.58 GiB27,469,499,3603.115bartowski
I1-IQ3_XXS25.58 GiB27,469,499,8083.115mradermacher
I1-IQ3_XS27.29 GiB29,307,735,4563.323mradermacher
Q3_K_S28.79 GiB30,912,056,2883.505bartowski
I1-IQ3_S28.79 GiB30,912,056,7363.505mradermacher
I1-Q3_K_S28.79 GiB30,912,056,7363.505mradermacher
IQ3_M29.74 GiB31,937,039,3283.621bartowski
I1-IQ3_M29.74 GiB31,937,039,7763.621mradermacher
Q3_K_M31.91 GiB34,267,499,4883.886bartowski
I1-Q3_K_M31.91 GiB34,267,499,9363.886mradermacher
Q3_K_L34.59 GiB37,140,597,7284.211bartowski
I1-Q3_K_L34.59 GiB37,140,598,1764.211mradermacher
IQ4_XS35.30 GiB37,902,666,7204.298bartowski
I1-IQ4_XS35.30 GiB37,902,667,1684.298mradermacher
Q4_037.36 GiB40,116,538,3364.549bartowski
I1-Q4_037.36 GiB40,116,538,7844.549mradermacher
I1-Q4_K_S37.58 GiB40,347,225,5044.575mradermacher
Q4_K_M39.60 GiB42,520,398,8164.821bartowski
I1-Q4_K_M39.60 GiB42,520,399,2644.821mradermacher
I1-Q5_K_S45.32 GiB48,657,452,4485.517mradermacher
Q5_K_M2 shards46.52 GiB49,949,822,1445.664bartowski
I1-Q5_K_M46.52 GiB49,949,822,3685.664mradermacher
Q6_K2 shards53.91 GiB57,888,148,6726.564bartowski
Q8_02 shards69.83 GiB74,975,055,0088.501bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 36.96 GiB. The real file is 39.60 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Llama-3.1-Nemotron-70B-Instruct-HF-abliterated need?
Q4_K_M is exactly 42,520,398,816 bytes (39.60 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Llama-3.1-Nemotron-70B-Instruct-HF-abliterated should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.