Step-3.5-Flash-REAP-121B-A11B
cerebras/Step-3.5-Flash-REAP-121B-A11BStep-3.5-Flash-REAP-121B-A11B at I1-IQ1_S is exactly 24,414,167,744 bytes (22.74 GiB / 24.41 GB) — an effective 1.615 bits per weight, not the nominal 1. Its KV cache at 32K is 13.03 GiB, not the 45.00 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| I1-IQ1_S | 22.74 GiB | 24,414,167,744 | 1.615 | — | mradermacher |
| I1-IQ1_M | 25.28 GiB | 27,146,462,912 | 1.795 | — | mradermacher |
| I1-IQ2_XXS | 29.52 GiB | 31,700,288,192 | 2.096 | — | mradermacher |
| I1-IQ2_XS | 32.98 GiB | 35,407,835,840 | 2.342 | — | mradermacher |
| I1-IQ2_S | 33.40 GiB | 35,858,358,976 | 2.371 | — | mradermacher |
| I1-IQ2_M | 36.79 GiB | 39,501,419,200 | 2.612 | — | mradermacher |
| I1-Q2_K_S | 37.79 GiB | 40,572,452,544 | 2.683 | — | mradermacher |
| I1-Q2_K | 41.19 GiB | 44,227,395,264 | 2.925 | — | mradermacher |
| I1-IQ3_XXS | 43.40 GiB | 46,597,518,016 | 3.082 | — | mradermacher |
| I1-IQ3_XS | 45.92 GiB | 49,308,838,592 | 3.261 | — | mradermacher |
| I1-Q3_K_S | 48.71 GiB | 52,297,181,888 | 3.459 | — | mradermacher |
| I1-IQ3_S | 48.73 GiB | 52,322,249,408 | 3.460 | — | mradermacher |
| I1-IQ3_M | 49.23 GiB | 52,857,023,168 | 3.496 | — | mradermacher |
| I1-Q3_K_M | 53.75 GiB | 57,715,993,280 | 3.817 | — | mradermacher |
| I1-Q3_K_L | 58.48 GiB | 62,791,625,408 | 4.153 | — | mradermacher |
| I1-IQ4_XS | 60.12 GiB | 64,555,701,952 | 4.269 | — | mradermacher |
| I1-Q4_0 | 63.71 GiB | 68,411,672,256 | 4.524 | — | mradermacher |
| I1-Q4_K_S | 63.83 GiB | 68,536,452,800 | 4.533 | — | mradermacher |
| I1-Q4_K_M | 67.82 GiB | 72,817,116,864 | 4.816 | — | mradermacher |
| I1-Q4_1 | 70.61 GiB | 75,814,545,088 | 5.014 | — | mradermacher |
| I1-Q5_K_S | 77.62 GiB | 83,340,101,312 | 5.512 | — | mradermacher |
| I1-Q5_K_M | 79.79 GiB | 85,672,773,312 | 5.666 | — | mradermacher |
| I1-Q6_K | 92.51 GiB | 99,331,908,288 | 6.569 | — | mradermacher |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 2.53 GiB | 5.63 GiB | 2.22× | 12 / 33 / 0 |
| 8,192 | 4.03 GiB | 11.25 GiB | 2.79× | 12 / 33 / 0 |
| 16,384 | 7.03 GiB | 22.50 GiB | 3.20× | 12 / 33 / 0 |
| 32,768 | 13.03 GiB | 45.00 GiB | 3.45× | 12 / 33 / 0 |
| 65,536 | 25.03 GiB | 90.00 GiB | 3.60× | 12 / 33 / 0 |
| 131,072 | 49.03 GiB | 180.00 GiB | 3.67× | 12 / 33 / 0 |
33 of 45 layers cache only a 512-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts I1-IQ1_S at roughly 63.37 GiB. The real file is 22.74 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 45.00 GiB at 32K context where the real figure is 13.03 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does Step-3.5-Flash-REAP-121B-A11B need?
- I1-IQ1_S is exactly 24,414,167,744 bytes (22.74 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is Step-3.5-Flash-REAP-121B-A11B's KV cache?
- 13.03 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Which quantization of Step-3.5-Flash-REAP-121B-A11B should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.