Hermes-Trismegistus-Mistral-7B
teknium/Hermes-Trismegistus-Mistral-7BHermes-Trismegistus-Mistral-7B at Q4_K_M is exactly 4,368,450,304 bytes (4.07 GiB / 4.37 GB) — an effective 4.826 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| Q2_K | 2.87 GiB | 3,083,107,200 | 3.406 | — | orangejuicesmith |
| Q2_K | 2.87 GiB | 3,083,107,200 | 3.406 | — | TheBloke |
| Q3_K_S | 2.95 GiB | 3,164,577,472 | 3.496 | — | TheBloke |
| Q3_K_S | 2.95 GiB | 3,164,577,472 | 3.496 | — | orangejuicesmith |
| Q3_K_M | 3.28 GiB | 3,518,996,160 | 3.888 | — | orangejuicesmith |
| Q3_K_M | 3.28 GiB | 3,518,996,160 | 3.888 | — | TheBloke |
| Q3_K_L | 3.56 GiB | 3,822,034,624 | 4.222 | — | TheBloke |
| Q3_K_L | 3.56 GiB | 3,822,034,624 | 4.222 | — | orangejuicesmith |
| Q4_0 | 3.83 GiB | 4,108,927,744 | 4.539 | — | orangejuicesmith |
| Q4_0 | 3.83 GiB | 4,108,927,744 | 4.539 | — | TheBloke |
| Q4_K_S | 3.86 GiB | 4,140,385,024 | 4.574 | — | orangejuicesmith |
| Q4_K_S | 3.86 GiB | 4,140,385,024 | 4.574 | — | TheBloke |
| Q4_K_M | 4.07 GiB | 4,368,450,304 | 4.826 | — | TheBloke |
| Q4_K_M | 4.07 GiB | 4,368,450,304 | 4.826 | — | orangejuicesmith |
| Q5_K_S | 4.65 GiB | 4,997,728,000 | 5.521 | — | orangejuicesmith |
| Q5_0 | 4.65 GiB | 4,997,728,000 | 5.521 | — | orangejuicesmith |
| Q5_K_S | 4.65 GiB | 4,997,728,000 | 5.521 | — | TheBloke |
| Q5_0 | 4.65 GiB | 4,997,728,000 | 5.521 | — | TheBloke |
| Q5_K_M | 4.78 GiB | 5,131,421,440 | 5.669 | — | TheBloke |
| Q5_K_M | 4.78 GiB | 5,131,421,440 | 5.669 | — | orangejuicesmith |
| Q6_K | 5.53 GiB | 5,942,078,272 | 6.564 | — | TheBloke |
| Q6_K | 5.53 GiB | 5,942,078,272 | 6.564 | — | orangejuicesmith |
| Q8_0 | 7.17 GiB | 7,695,874,752 | 8.502 | — | TheBloke |
| Q8_0 | 7.17 GiB | 7,695,874,752 | 8.502 | — | orangejuicesmith |
KV cache by context
This model declares a 4,096-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 3.79 GiB. The real file is 4.07 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does Hermes-Trismegistus-Mistral-7B need?
- Q4_K_M is exactly 4,368,450,304 bytes (4.07 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of Hermes-Trismegistus-Mistral-7B should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.