EsDrac-v1-7B Q4_K_S
This file is exactly 4,457,770,112 bytes — 4.15 GiB / 4.46 GB at an effective 4.683 bits per weight. The nominal rate for Q4_K_S is lower; mixed-precision tensors make the real figure higher, always.
Get it
llama-cli -hf mradermacher/EsDrac-v1-7B-GGUF:Q4_K_S
Downloads and runs in one step, resolving the quantization by name.
hf download mradermacher/EsDrac-v1-7B-GGUF EsDrac-v1-7B.Q4_K_S.gguf
Straight from the Hugging Face CDN — we host nothing and earn nothing from this. Verify what you received against the exact byte count above; a size mismatch is the usual cause of a file that will not load.
Fit by accelerator
| Accelerator | Memory○ | Total @4K◐ | Total @32K◐ | Fits◐ | tok/s @4K◐ |
|---|---|---|---|---|---|
| Arc A310 4GB | 4 GB | 5.22 GiB | 6.76 GiB | no | — |
| Arc A350M 4GB | 4 GB | 5.22 GiB | 6.76 GiB | no | — |
| Arc A370M 4GB | 4 GB | 5.22 GiB | 6.76 GiB | no | — |
| Arc A530M 4GB | 4 GB | 5.22 GiB | 6.76 GiB | no | — |
| Arc Pro A30M 4GB | 4 GB | 5.22 GiB | 6.76 GiB | no | — |
| Radeon Pro W6400 | 4 GB | 5.32 GiB | 6.86 GiB | no | — |
| Radeon RX 6400 | 4 GB | 5.32 GiB | 6.86 GiB | no | — |
| Radeon RX 6500 XT | 4 GB | 5.32 GiB | 6.86 GiB | no | — |
| RTX A400 | 4 GB | 5.42 GiB | 6.96 GiB | no | — |
| Arc A380 6GB | 6 GB | 5.22 GiB | 6.76 GiB | yes4K only | 23±30% |
| Arc Pro A40 6GB | 6 GB | 5.22 GiB | 6.76 GiB | yes4K only | 23±30% |
| Arc Pro A50 6GB | 6 GB | 5.22 GiB | 6.76 GiB | yes4K only | 23±30% |
| GeForce RTX 2060 | 6 GB | 5.22 GiB | 6.76 GiB | yes4K only | 54±12.9% |
| GeForce RTX 3050 | 6 GB | 5.22 GiB | 6.76 GiB | yes4K only | 28±12.9% |
| GeForce RTX 3060 OEM | 6 GB | 5.22 GiB | 6.76 GiB | yes4K only | 54±12.9% |
| RTX A2000 | 6 GB | 5.42 GiB | 6.96 GiB | yes4K only | 37±22% |
| Apple M1 | 8 GB | 4.97 GiB | 6.51 GiB | yes4K only | 12±8.3% |
| Apple M2 | 8 GB | 4.97 GiB | 6.51 GiB | yes4K only | 18±8.3% |
| Apple M3 | 8 GB | 4.97 GiB | 6.51 GiB | yes4K only | 18±8.3% |
| Apple M4 | 8 GB | 4.97 GiB | 6.51 GiB | yes4K only | 21±8.3% |
| Arc A530M 8GB | 8 GB | 5.22 GiB | 6.76 GiB | yes | 27±30% |
| Arc A550M 8GB | 8 GB | 5.22 GiB | 6.76 GiB | yes | 27±30% |
| Arc A570M 8GB | 8 GB | 5.22 GiB | 6.76 GiB | yes | 27±30% |
| Arc A580 8GB | 8 GB | 5.22 GiB | 6.76 GiB | yes | 58±30% |
Memory at context
| Context | Weights● | KV (f16)● | KV (q8_0)● | Working buffer◐ | Total (f16)◐ |
|---|---|---|---|---|---|
| 4,096 | 4.15 GiB | 0.22 GiB | 0.12 GiB | 0.35 GiB | 4.72 GiB |
| 8,192 | 4.15 GiB | 0.44 GiB | 0.23 GiB | 0.35 GiB | 4.94 GiB |
| 16,384 | 4.15 GiB | 0.88 GiB | 0.46 GiB | 0.35 GiB | 5.38 GiB |
| 32,768 | 4.15 GiB | 1.75 GiB | 0.93 GiB | 0.35 GiB | 6.26 GiB |
| 65,536 | 4.15 GiB | 3.50 GiB | 1.86 GiB | 0.35 GiB | 8.01 GiB |
| 131,072 | 4.15 GiB | 7.00 GiB | 3.72 GiB | 0.35 GiB | 11.51 GiB |
Quantizing the KV cache to q8_0 is a roughly 2× lever on the dominant term at long context, and it is the single most useful setting most local users never touch. Totals here exclude the allocator reserve your driver takes, which is hardware-specific — the per-accelerator table above includes it.