Olmo-3-32B-Think
allenai/Olmo-3-32B-ThinkOlmo-3-32B-Think at Q4_K_M is exactly 19,482,033,568 bytes (18.14 GiB / 19.48 GB) — an effective 4.835 bits per weight, not the nominal 4. Its KV cache at 32K is 2.84 GiB, not the 8.00 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| UD-IQ1_S | 6.75 GiB | 7,245,908,416 | 1.798 | — | unsloth |
| UD-IQ1_M | 7.33 GiB | 7,865,223,616 | 1.952 | — | unsloth |
| UD-IQ2_XXS | 8.30 GiB | 8,914,618,816 | 2.212 | — | unsloth |
| IQ2_XS | 9.02 GiB | 9,685,606,272 | 2.404 | — | bartowski |
| IQ2_S | 9.40 GiB | 10,088,696,128 | 2.504 | — | bartowski |
| IQ2_M | 10.21 GiB | 10,965,567,808 | 2.721 | — | bartowski |
| UD-IQ2_M | 10.29 GiB | 11,047,569,856 | 2.742 | — | unsloth |
| Q2_K | 11.18 GiB | 12,005,939,968 | 2.980 | — | bartowski |
| Q2_K | 11.18 GiB | 12,005,940,096 | 2.980 | — | unsloth |
| Q2_K_L | 11.29 GiB | 12,126,273,696 | 3.010 | — | unsloth |
| Q2_K_L | 11.65 GiB | 12,507,329,952 | 3.104 | — | bartowski |
| IQ3_XXS | 11.68 GiB | 12,540,397,888 | 3.112 | — | bartowski |
| UD-IQ3_XXS | 11.78 GiB | 12,650,375,616 | 3.140 | — | unsloth |
| IQ3_XS | 12.45 GiB | 13,371,425,984 | 3.319 | — | bartowski |
| Q3_K_S | 13.09 GiB | 14,058,243,264 | 3.489 | — | bartowski |
| Q3_K_S | 13.09 GiB | 14,058,243,392 | 3.489 | — | unsloth |
| IQ3_M | 13.48 GiB | 14,476,035,264 | 3.593 | — | bartowski |
| Q3_K_M | 14.53 GiB | 15,600,960,704 | 3.872 | — | bartowski |
| Q3_K_M | 14.53 GiB | 15,600,960,832 | 3.872 | — | unsloth |
| Q3_K_L | 15.75 GiB | 16,912,991,424 | 4.198 | — | bartowski |
| IQ4_XS | 16.14 GiB | 17,332,137,568 | 4.302 | — | bartowski |
| IQ4_XS | 16.16 GiB | 17,348,182,176 | 4.306 | — | unsloth |
| IQ4_NL | 17.06 GiB | 18,312,871,968 | 4.545 | — | bartowski |
| IQ4_NL | 17.06 GiB | 18,312,872,096 | 4.545 | — | unsloth |
| Q4_0 | 17.08 GiB | 18,341,707,808 | 4.552 | — | bartowski |
| Q4_0 | 17.08 GiB | 18,341,707,936 | 4.552 | — | unsloth |
| Q4_K_S | 17.15 GiB | 18,415,108,128 | 4.570 | — | bartowski |
| Q4_K_S | 17.15 GiB | 18,415,108,256 | 4.570 | — | unsloth |
| Q4_K_M | 18.14 GiB | 19,482,033,568 | 4.835 | — | lmstudio-community |
| Q4_K_M | 18.14 GiB | 19,482,034,208 | 4.835 | — | bartowski |
| Q4_K_M | 18.14 GiB | 19,482,034,336 | 4.835 | — | unsloth |
| Q4_K_L | 18.50 GiB | 19,863,090,592 | 4.930 | — | bartowski |
| Q4_1 | 18.86 GiB | 20,253,369,248 | 5.027 | — | bartowski |
| Q4_1 | 18.86 GiB | 20,253,369,376 | 5.027 | — | unsloth |
| Q5_K_S | 20.71 GiB | 22,235,809,568 | 5.519 | — | bartowski |
| Q5_K_S | 20.71 GiB | 22,235,809,696 | 5.519 | — | unsloth |
| Q5_K_M | 21.29 GiB | 22,859,712,288 | 5.673 | — | bartowski |
| Q5_K_M | 21.29 GiB | 22,859,712,416 | 5.673 | — | unsloth |
| Q5_K_L | 21.58 GiB | 23,176,590,752 | 5.752 | — | bartowski |
| Q6_K | 24.63 GiB | 26,448,494,624 | 6.564 | — | lmstudio-community |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 1.00 GiB | 1.00 GiB | — | 16 / 48 / 0 |
| 8,192 | 1.34 GiB | 2.00 GiB | 1.49× | 16 / 48 / 0 |
| 16,384 | 1.84 GiB | 4.00 GiB | 2.17× | 16 / 48 / 0 |
| 32,768 | 2.84 GiB | 8.00 GiB | 2.81× | 16 / 48 / 0 |
| 65,536 | 4.84 GiB | 16.00 GiB | 3.30× | 16 / 48 / 0 |
| 131,072 | 8.84 GiB | 32.00 GiB | 3.62× | 16 / 48 / 0 |
48 of 64 layers cache only a 4,096-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.89 GiB. The real file is 18.14 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 8.00 GiB at 32K context where the real figure is 2.84 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does Olmo-3-32B-Think need?
- Q4_K_M is exactly 19,482,033,568 bytes (18.14 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is Olmo-3-32B-Think's KV cache?
- 2.84 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Which quantization of Olmo-3-32B-Think should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.