Will it fit?
Every quantization of a model against your card, at any context length. Weights are the summed bytes of the real published files; the cache is computed layer by layer from the architecture. Nothing here multiplies a parameter count by a constant.
Context
KV cache
Qwen3.6-35B-A3B fits a GeForce RTX 4090 at UD-Q4_K_S — 21.35 GiB of 22.32 GiB usable at 32K context. Expect roughly 155 tokens/sec (modeled, ±37%).
| Quant | Weights● | Exact bytes● | KV● | Total◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|
| BF16 | 66.19 GiB | 71,065,942,560 | 0.63 GiB | 67.61 GiB | -45.29 GiB | — |
| Q8_0 | 35.21 GiB | 37,801,097,504 | 0.63 GiB | 36.63 GiB | -14.31 GiB | — |
| UD-Q6_K | 27.95 GiB | 30,011,242,784 | 0.63 GiB | 29.38 GiB | -7.06 GiB | — |
| Q6_K | 26.56 GiB | 28,514,152,288 | 0.63 GiB | 27.99 GiB | -5.67 GiB | — |
| UD-Q5_K_M | 25.23 GiB | 27,087,812,896 | 0.63 GiB | 26.66 GiB | -4.34 GiB | — |
| UD-Q5_K_S | 23.78 GiB | 25,538,017,568 | 0.63 GiB | 25.21 GiB | -2.89 GiB | — |
| UD-Q4_K_M | 21.11 GiB | 22,663,387,424 | 0.63 GiB | 22.54 GiB | -0.22 GiB | — |
| UD-Q4_K_S | 19.92 GiB | 21,388,319,008 | 0.63 GiB | 21.35 GiB | 0.97 GiB | 155±37% |
| Q4_K_M | 19.71 GiB | 21,166,757,728 | 0.63 GiB | 21.14 GiB | 1.18 GiB | 156±37% |
| UD-IQ4_NL | 17.26 GiB | 18,536,192,288 | 0.63 GiB | 18.69 GiB | 3.63 GiB | 170±37% |
| UD-IQ4_XS | 16.96 GiB | 18,209,036,576 | 0.63 GiB | 18.39 GiB | 3.93 GiB | 172±37% |
| UD-Q3_K_M | 15.93 GiB | 17,104,402,720 | 0.63 GiB | 17.36 GiB | 4.96 GiB | 179±37% |
| UD-Q3_K_S | 14.30 GiB | 15,359,196,128 | 0.63 GiB | 15.73 GiB | 6.59 GiB | 191±37% |
| UD-IQ3_S | 14.29 GiB | 15,346,432,288 | 0.63 GiB | 15.72 GiB | 6.60 GiB | 191±37% |
| UD-IQ3_XXS | 13.10 GiB | 14,069,266,720 | 0.63 GiB | 14.53 GiB | 7.79 GiB | 200±37% |
| UD-IQ2_M | 11.07 GiB | 11,882,969,376 | 0.63 GiB | 12.50 GiB | 9.82 GiB | 219±37% |
| UD-IQ2_XXS | 11.01 GiB | 11,819,399,456 | 0.63 GiB | 12.44 GiB | 9.88 GiB | 220±37% |
| UD-IQ1_M | 10.59 GiB | 11,366,414,624 | 0.63 GiB | 12.02 GiB | 10.30 GiB | 224±37% |
Weights are exact. The cache is computed per layer. The remaining term — the runtime's working buffers — is modeled, which is why a total carries a band when its parts do not. Speed is a bandwidth roofline, and the ± figure is our measured error against real benchmark runs, published on the methodology page.