NVIDIA · workstation

RTX A400

RTX A400 has 4 GB of VRAM at 96 GB/s — about 3.72 GiB usable after driver and compositor overhead. 266 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
96 GB/s
64-bit bus
Tensor FP16
11 TF
dense
TDP
50 W
$135 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 187vision language 20audio tts 15video 2embedding 13audio asr 29

What fits at 128K context

largest quantization that fits, per model · 266 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
LFM2-2.6BQ5_02.6B1.65 GiB1.06 GiB3.72 GiB0.00 GiB20±22%
LFM2.5-Audio-1.5B-JPF161.5B2.67 GiB0.00 GiB3.72 GiB0.00 GiB20±22%
umt5-xxlQ3_K_S5.7B2.66 GiB0.00 GiB3.71 GiB0.01 GiB21±22%
SEX_ROLEPLAY-3.2-1BI1-IQ2_M1.5B0.59 GiB2.13 GiB3.71 GiB0.01 GiB20±22%
Llama-3.2-1B-Instruct-abliteratedI1-IQ2_M1.5B0.59 GiB2.13 GiB3.71 GiB0.01 GiB20±22%
Novaciano-3.2-1BI1-IQ2_M1.5B0.59 GiB2.13 GiB3.71 GiB0.01 GiB20±22%
Imp-RPG.System-1BI1-IQ2_M1.5B0.59 GiB2.13 GiB3.71 GiB0.01 GiB20±22%
FableForge-1.5BQ3_K_M1.5B0.85 GiB1.86 GiB3.71 GiB0.01 GiB20±22%
LFM2-1.2BQ3_K_S1.2B0.58 GiB2.13 GiB3.71 GiB0.01 GiB20±22%
Qwen2.5-Omni-7BUD-IQ2_M10.7B2.66 GiB0.00 GiB3.70 GiB0.02 GiB21±22%
Dolphin3.0-Llama3.2-1BIQ3_XS1.2B0.58 GiB2.13 GiB3.70 GiB0.02 GiB20±22%
OneLLM-Doey-ChatQA-V1-Llama-3.2-1BI1-IQ3_XS1.2B0.58 GiB2.13 GiB3.70 GiB0.02 GiB20±22%
Llama-3.2-1B-Instruct-hereticI1-IQ3_XS1.2B0.58 GiB2.13 GiB3.70 GiB0.02 GiB20±22%
Llama-3.2-1B-Instruct-UncensoredI1-IQ3_XS1.2B0.58 GiB2.13 GiB3.70 GiB0.02 GiB20±22%
Qwen2.5-1.5B-Instruct-abliteratedIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
Qwen2.5-1.5B-VibeThinker-heretic-uncensored-abliteratedIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
NEXUS-Coder-OBLITERATEDIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
NEXUS-Coder-AbliteratedIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
Qwen2.5-1.5B-hereticIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
ShellWhisperer-1.5BIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
Qwen2.5-1.5BIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
PiCo-1BIQ4_XS1.5B0.84 GiB1.86 GiB3.70 GiB0.02 GiB20±22%
granite-4.0-h-microQ5_K_L3.2B2.16 GiB0.53 GiB3.69 GiB0.03 GiB20±22%
Dolphin3.0-Qwen2.5-1.5BIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Qwen2.5-Coder-1.5B-Unsensored-DPOI1-IQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Qwen2.5-Math-1.5B-InstructIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Qwen2.5-Coder-1.5B-InstructIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Qwen2.5-1.5B-InstructIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Fourier-Qwen2-VL-2B-0.67I1-IQ4_XS2.2B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Qwen2-VL-2B-InstructIQ4_XS2.2B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
Qwen2-1.5B-InstructIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
NEXUS-SecurityIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
NEXUS-ScienceIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
NEXUS-LegalIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
NEXUS-CoderIQ4_XS1.5B0.83 GiB1.86 GiB3.69 GiB0.03 GiB20±22%
LFM2-2.6B-ExpQ4_K_L2.6B1.62 GiB1.06 GiB3.69 GiB0.03 GiB20±22%
moondream2F161.9B2.64 GiB0.00 GiB3.69 GiB0.03 GiB21±22%
medgemma-1.5-4b-itUD-IQ2_XXS4.3B1.25 GiB1.42 GiB3.69 GiB0.03 GiB21±22%
medgemma-4b-itUD-IQ2_XXS4.3B1.25 GiB1.42 GiB3.69 GiB0.03 GiB21±22%
gemma-2bQ4_K_S2.5B1.45 GiB1.20 GiB3.68 GiB0.04 GiB21±22%
gemma-4-E2B-itUD-IQ3_XXS5.1B2.21 GiB0.48 GiB3.68 GiB0.04 GiB20±22%
LFM2-700MQ6_K742M0.57 GiB2.13 GiB3.68 GiB0.04 GiB20±22%
DeepSeek-R1-Distill-Qwen-1.5BUD-IQ3_XXS1.8B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
NEXUS-MedicalQ3_K_L1.5B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
NEXUS-FinanceQ3_K_L1.5B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
gemma-4-E2B-itUD-IQ3_XXS5.1B2.21 GiB0.48 GiB3.68 GiB0.04 GiB20±22%
Surogate-3.5-2BI1-Q6_K2.8B1.88 GiB0.80 GiB3.68 GiB0.04 GiB20±22%
VoxCPM2Q8_02.3B2.63 GiB0.00 GiB3.68 GiB0.04 GiB21±22%
DeepScaleR-1.5B-PreviewIQ3_M1.8B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
Qwen2.5-Coder-1.5B-Instruct-abliteratedIQ3_M1.8B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
DeepSeek-R1-Distill-Qwen-1.5B-uncensoredI1-IQ3_M1.8B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
Nemotron-Research-Reasoning-Qwen-1.5BIQ3_M1.8B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
dots.ocrI1-IQ3_M3.0B0.82 GiB1.86 GiB3.68 GiB0.04 GiB20±22%
gemma-4-E2B-it-ultra-uncensored-hereticIQ3_XXS5.1B2.20 GiB0.48 GiB3.67 GiB0.05 GiB20±22%
EXAONE-Deep-7.8BIQ2_M7.8B2.63 GiB0.00 GiB3.67 GiB0.05 GiB21±22%
EXAONE-3.5-7.8B-InstructIQ2_M7.8B2.63 GiB0.00 GiB3.67 GiB0.05 GiB21±22%
Qwen3.5-2B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGQ8_02.2B1.87 GiB0.80 GiB3.67 GiB0.05 GiB21±22%
Miss-MARTHA-hot-POCKET-edition-2B-OMNIQ8_02.3B1.87 GiB0.80 GiB3.67 GiB0.05 GiB21±22%
Huihui-Qwen3.5-2B-abliteratedQ8_02.3B1.87 GiB0.80 GiB3.67 GiB0.05 GiB21±22%
Noema-2BQ8_01.9B1.87 GiB0.80 GiB3.67 GiB0.05 GiB21±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A400 run?
266 of 2118 indexed open-weight models fit a RTX A400 at 131,072 context with q8_0 KV cache, the largest being LFM2-2.6B at Q5_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A400 actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A400 fast for local AI?
Its memory bandwidth is 96 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.