Trinity-Large-TrueBase
arcee-ai/Trinity-Large-TrueBaseTrinity-Large-TrueBase at Q4_K_M is exactly 241,632,208,032 bytes (225.04 GiB / 241.63 GB) — an effective 4.849 bits per weight, not the nominal 4. Its KV cache at 32K is 2.67 GiB, not the 7.50 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| I1-IQ1_S | 73.49 GiB | 78,905,750,240 | 1.583 | — | mradermacher |
| IQ1_S3 shards | 76.06 GiB | 81,666,716,288 | 1.639 | — | bartowski |
| IQ1_M3 shards | 79.29 GiB | 85,135,766,112 | 1.708 | — | bartowski |
| I1-IQ1_M | 82.07 GiB | 88,126,026,464 | 1.769 | — | mradermacher |
| IQ2_XXS3 shards | 88.58 GiB | 95,107,625,568 | 1.909 | — | bartowski |
| I1-IQ2_XXS | 96.39 GiB | 103,493,153,504 | 2.077 | — | mradermacher |
| IQ2_XS3 shards | 102.27 GiB | 109,811,937,920 | 2.204 | — | bartowski |
| IQ2_S3 shards | 102.49 GiB | 110,050,460,288 | 2.208 | — | bartowski |
| I1-IQ2_XS | 107.87 GiB | 115,822,244,576 | 2.324 | — | mradermacher |
| I1-IQ2_S | 108.32 GiB | 116,312,326,880 | 2.334 | — | mradermacher |
| IQ2_M4 shards | 116.50 GiB | 125,094,511,328 | 2.510 | — | bartowski |
| I1-IQ2_M | 119.77 GiB | 128,606,028,512 | 2.581 | — | mradermacher |
| I1-Q2_K_S | 122.88 GiB | 131,936,813,792 | 2.648 | — | mradermacher |
| Q2_K4 shards | 129.44 GiB | 138,989,781,728 | 2.789 | — | bartowski |
| Q2_K_L4 shards | 130.00 GiB | 139,590,357,728 | 2.801 | — | bartowski |
| I1-Q2_K | 134.81 GiB | 144,754,868,960 | 2.905 | — | mradermacher |
| I1-IQ3_XXS | 142.48 GiB | 152,986,993,376 | 3.070 | — | mradermacher |
| IQ3_XXS4 shards | 146.20 GiB | 156,981,428,960 | 3.150 | — | bartowski |
| I1-IQ3_XS | 150.34 GiB | 161,421,759,200 | 3.240 | — | mradermacher |
| IQ3_XS5 shards | 151.27 GiB | 162,428,146,560 | 3.260 | — | bartowski |
| I1-Q3_K_S | 159.90 GiB | 171,690,595,040 | 3.446 | — | mradermacher |
| I1-IQ3_S | 159.92 GiB | 171,715,662,560 | 3.446 | — | mradermacher |
| I1-IQ3_M | 160.39 GiB | 172,218,266,336 | 3.456 | — | mradermacher |
| Q3_K_S5 shards | 160.91 GiB | 172,772,184,928 | 3.467 | — | bartowski |
| IQ3_M5 shards | 168.61 GiB | 181,045,941,056 | 3.633 | — | bartowski |
| Q3_K_M5 shards | 168.74 GiB | 181,184,746,304 | 3.636 | — | bartowski |
| Q3_K_L5 shards | 175.82 GiB | 188,781,482,848 | 3.789 | — | bartowski |
| I1-Q3_K_M | 176.30 GiB | 189,305,443,040 | 3.799 | — | mradermacher |
| IQ4_XS6 shards | 198.19 GiB | 212,807,902,144 | 4.271 | — | bartowski |
| IQ4_NL6 shards | 209.71 GiB | 225,174,496,160 | 4.519 | — | bartowski |
| Q4_06 shards | 213.32 GiB | 229,048,460,224 | 4.597 | — | bartowski |
| Q4_K_S6 shards | 217.12 GiB | 233,131,615,168 | 4.679 | — | bartowski |
| Q4_K_M7 shards | 225.04 GiB | 241,632,208,032 | 4.849 | — | bartowski |
| Q4_K_L7 shards | 225.46 GiB | 242,088,645,792 | 4.858 | — | bartowski |
| Q4_17 shards | 232.76 GiB | 249,924,199,488 | 5.016 | — | bartowski |
| Q5_K_S8 shards | 255.88 GiB | 274,746,156,256 | 5.514 | — | bartowski |
| Q5_K_M8 shards | 263.87 GiB | 283,323,229,408 | 5.686 | — | bartowski |
| Q6_K9 shards | 305.07 GiB | 327,566,357,792 | 6.574 | — | bartowski |
| Q8_011 shards | 394.59 GiB | 423,684,396,160 | 8.503 | — | bartowski |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 0.94 GiB | 0.94 GiB | — | 15 / 45 / 0 |
| 8,192 | 1.26 GiB | 1.88 GiB | 1.49× | 15 / 45 / 0 |
| 16,384 | 1.73 GiB | 3.75 GiB | 2.17× | 15 / 45 / 0 |
| 32,768 | 2.67 GiB | 7.50 GiB | 2.81× | 15 / 45 / 0 |
| 65,536 | 4.54 GiB | 15.00 GiB | 3.30× | 15 / 45 / 0 |
| 131,072 | 8.29 GiB | 30.00 GiB | 3.62× | 15 / 45 / 0 |
45 of 60 layers cache only a 4,096-token window rather than the full context, on a period of 4. Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 208.83 GiB. The real file is 225.04 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 7.50 GiB at 32K context where the real figure is 2.67 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does Trinity-Large-TrueBase need?
- Q4_K_M is exactly 241,632,208,032 bytes (225.04 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is Trinity-Large-TrueBase's KV cache?
- 2.67 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Is Trinity-Large-TrueBase a mixture-of-experts model?
- Yes — 256 experts, 4 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
- Which quantization of Trinity-Large-TrueBase should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.