NousResearch · text

Hermes-4.3-36B

NousResearch/Hermes-4.3-36B

Hermes-4.3-36B at Q4_K_M is exactly 21,762,145,888 bytes (20.27 GiB / 21.76 GB) — an effective 4.816 bits per weight, not the nominal 4. Its KV cache at 32K is 8.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
36.2B
Architecture
seed_oss
64 layers
Context
524,288
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS9.23 GiB9,910,923,1042.193bartowski
IQ2_XS10.19 GiB10,945,081,1842.422bartowski
IQ2_S10.82 GiB11,612,626,7842.570bartowski
IQ2_M11.68 GiB12,541,927,2642.775bartowski
Q2_K12.67 GiB13,604,183,6483.010MaziyarPanahi
Q2_K12.67 GiB13,604,183,9043.010bartowski
IQ3_XXS13.15 GiB14,116,757,3443.124bartowski
Q2_K_L13.39 GiB14,379,863,9043.182bartowski
IQ3_XS14.05 GiB15,089,946,4643.339bartowski
Q3_K_S14.77 GiB15,855,406,9443.509bartowski
IQ3_M15.36 GiB16,496,021,3443.651bartowski
Q3_K_M16.41 GiB17,620,946,5283.899MaziyarPanahi
Q3_K_M16.41 GiB17,620,946,7843.899bartowski
Q3_K_L17.83 GiB19,142,692,4484.236MaziyarPanahi
Q3_K_L17.83 GiB19,142,692,7044.236bartowski
IQ4_XS18.16 GiB19,498,614,6244.315bartowski
IQ4_NL19.18 GiB20,592,983,9044.557bartowski
Q4_019.21 GiB20,621,819,7444.564bartowski
Q4_K_S19.27 GiB20,695,220,0644.580bartowski
Q4_K_M20.27 GiB21,762,145,8884.816MaziyarPanahi
Q4_K_M20.27 GiB21,762,146,1444.816bartowski
Q4_K_L20.82 GiB22,351,662,9444.946bartowski
Q4_121.20 GiB22,760,750,9445.037bartowski
Q5_K_S23.26 GiB24,970,461,0245.526bartowski
Q5_K_M23.84 GiB25,594,363,4885.664MaziyarPanahi
Q5_K_M23.84 GiB25,594,363,7445.664bartowski
Q5_K_L24.29 GiB26,084,593,5045.772bartowski
Q6_K27.63 GiB29,666,094,6886.565MaziyarPanahi
Q6_K27.63 GiB29,666,094,9446.565bartowski
Q6_K_L27.99 GiB30,050,832,2246.650bartowski
Q8_035.78 GiB38,421,090,1448.502bartowski
BF162 shards67.35 GiB72,311,394,08016.002bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.00 GiB1.00 GiB64 / 0 / 0
8,1922.00 GiB2.00 GiB64 / 0 / 0
16,3844.00 GiB4.00 GiB64 / 0 / 0
32,7688.00 GiB8.00 GiB64 / 0 / 0
65,53616.00 GiB16.00 GiB64 / 0 / 0
131,07232.00 GiB32.00 GiB64 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 18.94 GiB. The real file is 20.27 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
64
Attention heads
80
KV heads
8
Head dim
128
Hidden size
5120
Vocab
155,136
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Hermes-4.3-36B need?
Q4_K_M is exactly 21,762,145,888 bytes (20.27 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Hermes-4.3-36B's KV cache?
8.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Hermes-4.3-36B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.