Can I run Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness on a GeForce RTX 3050?
Not at these settings. No indexed quantization of Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness fits GeForce RTX 3050 at any context we compute, with f16 KV. The smallest shipped quantization is 4.88 GiB in weights alone, against 5.58 GiB usable. CPU offload can still run it, slowly.
From the file· weights summed from filesFrom the file· KV computed per layerPredicted· speed and compute buffer
Every quantization at every context
total memory required; green fits, red does not
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, the KV cache grows linearly with context and, past about 32K, becomes larger than the weights for many models.