Best local AI models for 4GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 4GB card gives you about 3.72 GiB to work with after driver overhead. 17 indexed models fit at 32K context — the largest being s2-pro at 4.6B parameters in Q3_K.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 4GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Ace-Step1.5speech synthesisF32160M2.71 GiB1.01 GiB
Qwen3-TTS-12Hz-0.6B-Basespeech synthesisQ4_K915M1.34 GiB2.38 GiB
OmniVoicespeech synthesisF16613M2.37 GiB1.35 GiB
VieNeu-TTS-0.3Bspeech synthesisQ8_0244M1.94 GiB1.78 GiB
neutts-airspeech synthesisQ8_0748M1.90 GiB1.82 GiB
s2-proMoEspeech synthesisQ3_K4.6B3.64 GiB0.08 GiB
Fun-CosyVoice3-0.5B-2512speech synthesisQ4_K1.70 GiB2.02 GiB
VoxCPM2speech synthesisQ8_02.3B3.48 GiB0.24 GiB
Kokoro-82Mspeech synthesisF1682M1.00 GiB2.72 GiB
VieNeu-TTSspeech synthesisQ4_0553M1.54 GiB2.18 GiB
csm-1bspeech synthesisQ8_01.6B3.67 GiB0.05 GiB
VibeVoice-Realtime-0.5Bspeech synthesisF161.0B2.74 GiB0.98 GiB
Qwen3-TTS-12Hz-1.7B-Basespeech synthesisQ8_01.9B2.77 GiB0.95 GiB
Qwen3-TTS-12Hz-1.7B-VoiceDesignspeech synthesisQ8_01.9B2.75 GiB0.97 GiB
VibeVoice-1.5Bspeech synthesisQ4_K2.7B2.61 GiB1.11 GiB
VoxCPM-0.5Bspeech synthesisQ4_K728M3.29 GiB0.43 GiB
Qwen3-TTS-12Hz-1.7B-CustomVoicespeech synthesisQ8_01.9B2.75 GiB0.97 GiB
Spec sheetPredictedwhat these mean

This page models a generic 4GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.