Best local AI models for 128GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 128GB card gives you about 119.04 GiB to work with after driver overhead. 16 indexed models fit at 32K context — the largest being lingbot-world-v2-14b-causal-fast at 18.5B parameters in Q8_0.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 128GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| Wan2.2-Animate-14B | video generation | Q8_0 | 17.3B | 18.27 GiB | 100.77 GiB |
| Wan2.1-I2V-14B-480P | video generation | BF16 | 16.4B | 31.83 GiB | 87.21 GiB |
| Bernini-R | video generation | Q8_0 | 14.3B | 29.56 GiB | 89.48 GiB |
| Wan2.2-Distill-Models | video generation | Q8_0 | 14.3B | 15.19 GiB | 103.85 GiB |
| Wan2.2-TI2V-5B | video generation | Q8_0 | 5.0B | 5.87 GiB | 113.17 GiB |
| Wan2.1-T2V-14B | video generation | BF16 | 14.3B | 27.89 GiB | 91.15 GiB |
| Wan-Dancer-14B | video generation | Q6_K | 17.2B | 27.75 GiB | 91.29 GiB |
| Wan2.1-I2V-14B-720P | video generation | BF16 | 16.4B | 31.83 GiB | 87.21 GiB |
| Wan2.1-VACE-14B | video generation | BF16 | 17.3B | 33.14 GiB | 85.90 GiB |
| Wan2.2-S2V-14B | video generation | BF16 | 16.3B | 31.37 GiB | 87.67 GiB |
| Wan2.1-FLF2V-14B-720P | video generation | BF16 | 16.4B | 31.84 GiB | 87.20 GiB |
| JoyAI-Echo | video generation | Q6_K | 12.2B | 19.07 GiB | 99.97 GiB |
| HunyuanVideo-1.5 | video generation | Q8_0 | 8.3B | 9.22 GiB | 109.82 GiB |
| Wan2.2-TI2V-5B-Turbo | video generation | Q8_0 | 5.0B | 5.87 GiB | 113.17 GiB |
| SkyReels-V2-DF-14B-540P | video generation | BF16 | 14.3B | 27.46 GiB | 91.58 GiB |
| lingbot-world-v2-14b-causal-fast | video generation | Q8_0 | 18.5B | 19.40 GiB | 99.64 GiB |
This page models a generic 128GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.