NVIDIA · workstation

RTX A2000

RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 925 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 762audio asr 38vision language 79audio tts 19embedding 25video 2

What fits at 64K context

largest quantization that fits, per model · 925 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
SmolLM2-1.7B-Instruct-UncensoredQ5_K_M1.8B1.21 GiB3.38 GiB5.58 GiB0.00 GiB36±22%
Yi-6B-ChatI1-Q4_K_M6.1B3.42 GiB1.13 GiB5.57 GiB0.01 GiB36±22%
Yi-1.5-6B-ChatQ4_K_M6.1B3.42 GiB1.13 GiB5.57 GiB0.01 GiB36±22%
legitus-instruct-v1I1-IQ2_XXS8.1B2.25 GiB2.25 GiB5.57 GiB0.01 GiB36±22%
Apertus-8B-Instruct-2509I1-IQ2_XXS8.1B2.25 GiB2.25 GiB5.57 GiB0.01 GiB36±22%
Jan-v3-4B-base-instructQ2_K_L4.4B2.03 GiB2.53 GiB5.57 GiB0.01 GiB36±22%
Jan-code-4bQ2_K_L4.4B2.03 GiB2.53 GiB5.57 GiB0.01 GiB36±22%
granite-speech-4.1-2b-narQ4_K2.3B3.18 GiB1.41 GiB5.57 GiB0.01 GiB35±22%
Crow-9B-HERETIC-4.6I1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSOREDI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCTI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
NaNovel-9BI1-IQ3_S9.7B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-Unredacted-MAXI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-abliteratedI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Ken3.5-9BI1-IQ3_S9.7B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-BaseIQ3_S9.7B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-gemini-3.1-opus-4.6-reasoningI1-IQ3_S9.4B3.97 GiB0.56 GiB5.57 GiB0.01 GiB36±22%
Darwin-4B-ChimeraQ8_04.0B3.99 GiB0.57 GiB5.57 GiB0.01 GiB36±22%
granite-3.1-8b-instructTQ1_08.2B1.72 GiB2.81 GiB5.57 GiB0.01 GiB36±22%
Vero-Qwen35-9B-BaseI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Vero-Qwen35-9BI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Morphos-9BI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwable-9B-Claude-Fable-5-hereticI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Holo-3.1-9BI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwable-9B-Claude-Fable-5I1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-imabari-v2I1-Q3_K_S9.7B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-abliterated-v2-MAXI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwable-9B-Claude-Fable-5-StraTAI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-RpRMax-v1I1-Q3_K_S9.7B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
AdQWENistrator-9BI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
cajal-9b-v2-fullI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-ultra-uncensored-hereticQ3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Holo-3.1-9B-CoderI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
PlutoI1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Holo-3.1-9B-abliterated-rdoI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-Uncensored-cyber-v3Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Miss_MARTHA-9B-Qwen3.5-OmniI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Huihui-Qwen3.5-9B-abliteratedQ3_K_S9.7B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
qwen3.5-9b-nsfw-captioning-v5I1-Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-DS-v4-Flash-v3.0Q3_K_S9.4B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwen3.5-9B-DeepSeek-V4-FlashI1-Q3_K_S9.7B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Huihui-Qwen3.5-9B-Claude-4.6-Opus-abliteratedI1-Q3_K_S9.7B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Qwopus3.5-9B-v3.5Q3_K_S9.7B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
Katarau-9B-ru-RP-nsfwI1-Q3_K_S9.0B3.97 GiB0.56 GiB5.56 GiB0.02 GiB36±22%
internlm3-8b-instructQ2_K_L8.8B3.69 GiB0.84 GiB5.56 GiB0.02 GiB36±22%
Olmo-3-7B-InstructUD-IQ1_S7.3B1.81 GiB2.72 GiB5.56 GiB0.02 GiB36±22%
Olmo-3-7B-ThinkUD-IQ1_S7.3B1.81 GiB2.72 GiB5.56 GiB0.02 GiB36±22%
Ministral-3-3B-Instruct-2512-BF16Q6_K_L4.3B2.72 GiB1.83 GiB5.56 GiB0.02 GiB36±22%
Nemotron-Mini-4B-InstructIQ4_XS4.2B2.29 GiB2.25 GiB5.56 GiB0.02 GiB36±22%
Hubble-4B-v1Q3_K_L4.5B2.30 GiB2.25 GiB5.56 GiB0.02 GiB36±22%
Aura-4BI1-Q3_K_L4.5B2.30 GiB2.25 GiB5.56 GiB0.02 GiB36±22%
magnum-v2-4bI1-Q3_K_L4.5B2.30 GiB2.25 GiB5.56 GiB0.02 GiB36±22%
Impish_LLAMA_4BQ3_K_L4.5B2.30 GiB2.25 GiB5.56 GiB0.02 GiB36±22%
Llama-3.1-Minitron-4B-Width-BaseQ3_K_L4.5B2.30 GiB2.25 GiB5.56 GiB0.02 GiB36±22%
AfriqueGemma-12BI1-IQ1_M12.2B3.26 GiB1.26 GiB5.56 GiB0.02 GiB36±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A2000 run?
925 of 2118 indexed open-weight models fit a RTX A2000 at 65,536 context with q4_0 KV cache, the largest being SmolLM2-1.7B-Instruct-Uncensored at Q5_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A2000 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A2000 fast for local AI?
Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.