Best local AI models for 12GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 12GB card gives you about 11.16 GiB to work with after driver overhead. 14 indexed models fit at 32K context — the largest being Wan2.1-VACE-14B at 17.3B parameters in Q4_K_S.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 12GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| Wan2.2-Animate-14B | video generation | Q4_K_S | 17.3B | 10.70 GiB | 0.46 GiB |
| Wan2.1-I2V-14B-480P | video generation | Q4_1 | 16.4B | 11.15 GiB | 0.01 GiB |
| Bernini-R | video generation | Q5_1 | 14.3B | 11.10 GiB | 0.06 GiB |
| Wan2.2-Distill-Models | video generation | Q5_1 | 14.3B | 11.10 GiB | 0.06 GiB |
| Wan2.2-TI2V-5B | video generation | Q8_0 | 5.0B | 5.87 GiB | 5.29 GiB |
| Wan2.1-T2V-14B | video generation | Q5_0 | 14.3B | 10.88 GiB | 0.28 GiB |
| Wan2.1-I2V-14B-720P | video generation | Q4_1 | 16.4B | 11.15 GiB | 0.01 GiB |
| Wan2.1-VACE-14B | video generation | Q4_K_S | 17.3B | 10.66 GiB | 0.50 GiB |
| Wan2.2-S2V-14B | video generation | IQ3_XXS | 16.3B | 10.82 GiB | 0.34 GiB |
| Wan2.1-FLF2V-14B-720P | video generation | Q4_1 | 16.4B | 11.16 GiB | 0.00 GiB |
| JoyAI-Echo | video generation | Q5_1 | 12.2B | 10.08 GiB | 1.08 GiB |
| HunyuanVideo-1.5 | video generation | Q8_0 | 8.3B | 9.22 GiB | 1.94 GiB |
| Wan2.2-TI2V-5B-Turbo | video generation | Q8_0 | 5.0B | 5.87 GiB | 5.29 GiB |
| SkyReels-V2-DF-14B-540P | video generation | Q5_1 | 14.3B | 11.10 GiB | 0.06 GiB |
This page models a generic 12GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.