grug-35b Q4_1
This file is exactly 21,973,769,152 bytes — 20.46 GiB / 21.97 GB at an effective 5.007 bits per weight. The nominal rate for Q4_1 is lower; mixed-precision tensors make the real figure higher, always.
Get it
llama-cli -hf bartowski/ProCreations_grug-35b-GGUF:Q4_1
Downloads and runs in one step, resolving the quantization by name.
hf download bartowski/ProCreations_grug-35b-GGUF ProCreations_grug-35b-Q4_1.gguf
Straight from the Hugging Face CDN — we host nothing and earn nothing from this. Verify what you received against the exact byte count above; a size mismatch is the usual cause of a file that will not load.
Fit by accelerator
| Accelerator | Memory○ | Total @4K◐ | Total @32K◐ | Fits◐ | tok/s @4K◐ |
|---|---|---|---|---|---|
| Arc A310 4GB | 4 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc A350M 4GB | 4 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc A370M 4GB | 4 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc A530M 4GB | 4 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc Pro A30M 4GB | 4 GB | 21.35 GiB | 21.89 GiB | no | — |
| Radeon Pro W6400 | 4 GB | 21.45 GiB | 21.99 GiB | no | — |
| Radeon RX 6400 | 4 GB | 21.45 GiB | 21.99 GiB | no | — |
| Radeon RX 6500 XT | 4 GB | 21.45 GiB | 21.99 GiB | no | — |
| RTX A400 | 4 GB | 21.55 GiB | 22.09 GiB | no | — |
| Arc A380 6GB | 6 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc Pro A40 6GB | 6 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc Pro A50 6GB | 6 GB | 21.35 GiB | 21.89 GiB | no | — |
| GeForce RTX 2060 | 6 GB | 21.35 GiB | 21.89 GiB | no | — |
| GeForce RTX 3050 | 6 GB | 21.35 GiB | 21.89 GiB | no | — |
| GeForce RTX 3060 OEM | 6 GB | 21.35 GiB | 21.89 GiB | no | — |
| RTX A2000 | 6 GB | 21.55 GiB | 22.09 GiB | no | — |
| Apple M1 | 8 GB | 21.10 GiB | 21.64 GiB | no | — |
| Apple M2 | 8 GB | 21.10 GiB | 21.64 GiB | no | — |
| Apple M3 | 8 GB | 21.10 GiB | 21.64 GiB | no | — |
| Apple M4 | 8 GB | 21.10 GiB | 21.64 GiB | no | — |
| Arc A530M 8GB | 8 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc A550M 8GB | 8 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc A570M 8GB | 8 GB | 21.35 GiB | 21.89 GiB | no | — |
| Arc A580 8GB | 8 GB | 21.35 GiB | 21.89 GiB | no | — |
Memory at context
| Context | Weights● | KV (f16)● | KV (q8_0)● | Working buffer◐ | Total (f16)◐ |
|---|---|---|---|---|---|
| 4,096 | 20.46 GiB | 0.08 GiB | 0.04 GiB | 0.30 GiB | 20.85 GiB |
| 8,192 | 20.46 GiB | 0.16 GiB | 0.08 GiB | 0.30 GiB | 20.93 GiB |
| 16,384 | 20.46 GiB | 0.31 GiB | 0.17 GiB | 0.30 GiB | 21.08 GiB |
| 32,768 | 20.46 GiB | 0.63 GiB | 0.33 GiB | 0.30 GiB | 21.39 GiB |
| 65,536 | 20.46 GiB | 1.25 GiB | 0.66 GiB | 0.30 GiB | 22.02 GiB |
| 131,072 | 20.46 GiB | 2.50 GiB | 1.33 GiB | 0.30 GiB | 23.27 GiB |
Quantizing the KV cache to q8_0 is a roughly 2× lever on the dominant term at long context, and it is the single most useful setting most local users never touch. Totals here exclude the allocator reserve your driver takes, which is hardware-specific — the per-accelerator table above includes it.