NVIDIA · workstation

RTX A2000

RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1089 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 920embedding 26vision language 84audio asr 38audio tts 19video 2

What fits at 16K context

largest quantization that fits, per model · 1089 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
LFM2-8B-A1BMoEQ4_08.3B4.48 GiB0.10 GiB5.58 GiB0.00 GiB100±37%
Aya-Medikal-V2I1-IQ3_XS8.0B3.47 GiB1.06 GiB5.58 GiB0.00 GiB36±22%
legitus-instruct-v1I1-IQ3_S8.1B3.44 GiB1.06 GiB5.57 GiB0.01 GiB36±22%
Apertus-8B-Instruct-2509I1-IQ3_S8.1B3.44 GiB1.06 GiB5.57 GiB0.01 GiB36±22%
Nemotron-3-Embed-8B-BF16IQ3_S8.0B3.40 GiB1.13 GiB5.57 GiB0.01 GiB36±22%
FrickFritz-4BQ8_04.7B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
Newton-bot-3-VLM-mini-4BQ8_04.7B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
qwen3.5-4b-agentic-coder-v4Q8_04.7B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
Myth-4BQ8_04.3B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-4B-UncensoredQ8_04.7B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
JOSIE-2-4B-PreviewQ8_04.7B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
Surogate-3.5-4BQ8_05.3B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
Qwopus3.5-4B-v3Q8_04.7B4.29 GiB0.27 GiB5.57 GiB0.01 GiB36±22%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-IQ2_S12.2B3.74 GiB0.78 GiB5.57 GiB0.01 GiB36±22%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-IQ2_S12.2B3.74 GiB0.78 GiB5.57 GiB0.01 GiB36±22%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-IQ2_S12.2B3.74 GiB0.78 GiB5.57 GiB0.01 GiB36±22%
Floppa-12B-Gemma3-UncensoredI1-IQ2_S12.2B3.74 GiB0.78 GiB5.57 GiB0.01 GiB36±22%
gemma-3-12b-it-hereticI1-IQ2_S12.2B3.74 GiB0.78 GiB5.57 GiB0.01 GiB36±22%
gemma-3-12b-it-abliteratedIQ2_S12.2B3.74 GiB0.78 GiB5.57 GiB0.01 GiB36±22%
Phi-3.5-mini-instructIQ3_XXS3.8B1.37 GiB3.19 GiB5.57 GiB0.01 GiB36±22%
Yi-6B-ChatI1-Q5_K_M6.1B4.01 GiB0.53 GiB5.57 GiB0.01 GiB36±22%
Yi-1.5-6B-ChatQ5_K_M6.1B4.01 GiB0.53 GiB5.57 GiB0.01 GiB36±22%
dolphin-2.9.2-Phi-3-MediumKV unresolvedIQ1_S14.0B2.84 GiB1.66 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-itIQ4_XS8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
Gemma-4-E4B-it-Minecraft-MT-en-zh-v0.1I1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-Queen-it-qat-q4_0-unquantizedI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
Gemma-4-E4B-Luchador-RudoI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
supergemma4-e4b-abliteratedI1-IQ3_M7.5B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
Gemma-4-E4B-AbliteratedI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-The-DECKARD-Claude-Opus-Expresso-Universe-HERETIC-UNCENSORED-ThinkingI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-The-DECKARD-Expresso-Universe-HERETIC-UNCENSORED-ThinkingI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-hereticI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-Claude-Opus-4.5-HERETIC-UNCENSORED-ThinkingI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
Huihui-gemma-4-E4B-it-abliteratedI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-Uncensored-MAXI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
Darkidol-Gemma-4-E4B-itI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-abliteratedI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-itIQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
OpenMedResearch-Gemma-4E4NI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
Reasoning-Medical0.1-E4B-sftI1-IQ3_M8.0B4.39 GiB0.15 GiB5.56 GiB0.02 GiB36±22%
granite-8b-code-instruct-4kI1-IQ3_S8.1B3.32 GiB1.20 GiB5.56 GiB0.02 GiB36±22%
granite-8b-code-base-4kI1-IQ3_S8.1B3.32 GiB1.20 GiB5.56 GiB0.02 GiB36±22%
granite-3.3-8b-instructIQ3_XS8.2B3.19 GiB1.33 GiB5.55 GiB0.03 GiB36±22%
granite-3.2-8b-instructIQ3_XS8.2B3.19 GiB1.33 GiB5.55 GiB0.03 GiB36±22%
granite-3.1-8b-instructIQ3_XS8.2B3.19 GiB1.33 GiB5.55 GiB0.03 GiB36±22%
Qwen2.5-3B-Instruct-abliteratedI1-Q5_K_S3.1B4.24 GiB0.30 GiB5.55 GiB0.03 GiB36±22%
Marco-Nano-InstructMoEI1-IQ3_M8.0B3.65 GiB0.93 GiB5.55 GiB0.03 GiB72±37%
Fara1.5-9BIQ3_XS9.4B4.25 GiB0.27 GiB5.55 GiB0.03 GiB36±22%
QwenPaw-Flash-9BIQ3_XS9.4B4.25 GiB0.27 GiB5.55 GiB0.03 GiB36±22%
grug-9bIQ3_XS9.4B4.25 GiB0.27 GiB5.55 GiB0.03 GiB36±22%
OmniCoder-9BIQ3_XS9.4B4.25 GiB0.27 GiB5.55 GiB0.03 GiB36±22%
Ornith-1.0-9BIQ3_XS9.2B4.25 GiB0.27 GiB5.55 GiB0.03 GiB36±22%
Qwen3.5-9B-NeoIQ3_XS9.7B4.25 GiB0.27 GiB5.55 GiB0.03 GiB36±22%
Assistant_Pepe_8BQ2_K_L3.44 GiB1.06 GiB5.55 GiB0.03 GiB36±22%
Gemma-4-E4B-LuchadorQ3_K_S8.0B4.38 GiB0.15 GiB5.55 GiB0.03 GiB36±22%
Llama-3.1-Tulu-3-8BQ2_K_L8.0B3.44 GiB1.06 GiB5.54 GiB0.04 GiB36±22%
Llama-3-Groq-8B-Tool-UseQ2_K_L8.0B3.44 GiB1.06 GiB5.54 GiB0.04 GiB36±22%
Dolphin3.0-Llama3.1-8BQ2_K_L8.0B3.44 GiB1.06 GiB5.54 GiB0.04 GiB36±22%
dolphin-2.9.4-llama3.1-8bQ2_K_L8.0B3.44 GiB1.06 GiB5.54 GiB0.04 GiB36±22%
dolphin-2.9-llama3-8bQ2_K_L8.0B3.44 GiB1.06 GiB5.54 GiB0.04 GiB36±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A2000 run?
1089 of 2118 indexed open-weight models fit a RTX A2000 at 16,384 context with q8_0 KV cache, the largest being LFM2-8B-A1B at Q4_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A2000 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A2000 fast for local AI?
Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.