NVIDIA · workstation

RTX A2000

RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1211 of 2118 indexed models fit at 8K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1035image 1vision language 90embedding 26audio asr 38audio tts 19video 2

What fits at 8K context

largest quantization that fits, per model · 1211 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
gemma-4-E4B-itQ4_K_S8.0B4.51 GiB0.05 GiB5.58 GiB0.00 GiB36±22%
gemma-4-E4B-itQ4_K_S8.0B4.51 GiB0.05 GiB5.58 GiB0.00 GiB36±22%
LFM2-8B-A1BMoEQ4_K_S8.3B4.56 GiB0.03 GiB5.58 GiB0.00 GiB106±37%
Qwen3-14BUD-IQ2_XXS14.8B4.16 GiB0.35 GiB5.57 GiB0.01 GiB36±22%
deepseek-math-7b-instructQ3_K_L6.9B3.49 GiB1.05 GiB5.57 GiB0.01 GiB36±22%
deepseek-llm-7b-chatQ3_K_L6.9B3.49 GiB1.05 GiB5.57 GiB0.01 GiB36±22%
Janus-Pro-7BI1-Q3_K_L7.4B3.49 GiB1.05 GiB5.57 GiB0.01 GiB36±22%
deepseek-coder-7b-instruct-v1.5I1-Q3_K_L6.9B3.49 GiB1.05 GiB5.57 GiB0.01 GiB36±22%
Falcon3-7B-InstructQ4_K_M7.5B4.26 GiB0.25 GiB5.57 GiB0.01 GiB36±22%
Mistral-7B-Instruct-v0.3-ParasiteI1-Q4_17.2B4.24 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
Mistral-7B-Instruct-v0.3-JbliteratedI1-Q4_17.2B4.24 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
Mistral-7B-Instruct-v0.3Q4_17.2B4.24 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
Mistral-7B-v0.3-Chinese-ChatQ4_17.2B4.24 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
dolphin-2.8-mistral-7b-v02Q4_17.2B4.24 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
Kunoichi-DPO-v2-7BKV unresolvedQ4_17.2B4.24 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
Tini-Cybersec-8B-A1BMoEQ4_08.5B4.54 GiB0.03 GiB5.56 GiB0.02 GiB107±37%
gemma-4-E4B-uncensoredI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-qat-heretic_decensoredI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma4-e4b-mahou-nsfwI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-mentalchat16kI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma4-E4B-it-abliteratedI1-Q3_K_M7.9B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-OBLITERATEDI1-Q3_K_M8.0B4.49 GiB0.05 GiB5.56 GiB0.02 GiB36±22%
legitus-instruct-v1IQ4_XS8.1B4.21 GiB0.28 GiB5.56 GiB0.02 GiB36±22%
LFM2.5-Queen-Opus-4.7-8B-A1BMoEI1-Q4_K_S8.5B4.53 GiB0.03 GiB5.55 GiB0.03 GiB107±37%
LFM2.5-8B-A1B-KO-SFTMoEI1-Q4_K_S8.5B4.53 GiB0.03 GiB5.55 GiB0.03 GiB107±37%
LFM2.5-8B-A1B-hereticMoEI1-Q4_K_S8.5B4.53 GiB0.03 GiB5.55 GiB0.03 GiB107±37%
LFM2.5-8B-A1B-SOMPOA-heresyMoEI1-Q4_K_S8.5B4.53 GiB0.03 GiB5.55 GiB0.03 GiB107±37%
Huihui-LFM2.5-8B-A1B-abliteratedMoEI1-Q4_K_S8.5B4.53 GiB0.03 GiB5.55 GiB0.03 GiB107±37%
Supertron2.1-8B-A1BMoEI1-Q4_K_S8.5B4.53 GiB0.03 GiB5.55 GiB0.03 GiB107±37%
OLMo-2-1124-7B-InstructQ3_K_M7.3B3.40 GiB1.13 GiB5.55 GiB0.03 GiB36±22%
Anubis-Mini-8B-v1IQ4_XS8.0B4.23 GiB0.28 GiB5.55 GiB0.03 GiB36±22%
Qwen3-VL-8B-Instruct-HereticI1-IQ1_M8.8B4.20 GiB0.32 GiB5.55 GiB0.03 GiB36±22%
GrammarCoder-7B-BaseI1-Q4_K_M7.6B4.37 GiB0.12 GiB5.55 GiB0.03 GiB36±22%
gemma-4-E4B-it-hereticQ4_K_S8.0B4.48 GiB0.05 GiB5.55 GiB0.03 GiB36±22%
Swallow-7b-NVE-instruct-hfIQ4_XS6.7B3.40 GiB1.13 GiB5.55 GiB0.03 GiB36±22%
Aya-Medikal-V2I1-Q3_K_L8.0B4.22 GiB0.28 GiB5.54 GiB0.04 GiB36±22%
DeepHat-V1-7B-Heretic-AbliteratedI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
ShizhenGPT-7B-VLI1-Q4_K_M8.3B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
DeepHat-V1-7BQ4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
HuatuoGPT-o1-7BI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
MathSmith-DS-Qwen-7B-LongCoTI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
AstraGPTCoder-7BI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-7B-Instruct-Ghidra-v2I1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
EsDrac-v1-7BI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Hemlock-Apothecary-7B-GRPO-e3I1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
openhands-lm-7b-v0.1I1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Hemlock2-Coder-7B-GRPOI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
shellwhiz-7bI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-7B-Instruct-abliteratedI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-7B-Instruct-OBLITERATED-advancedI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Qwen-STEM-Specialist-7BI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
VulnLLM-R-7BI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Garnet-OCR-7B-0422I1-Q4_K_M8.3B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
UwU-7B-InstructI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Video-R1-7BI1-Q4_K_M8.3B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
HARC-Qwen2.5-7B-InstructI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-7B-AbliteratedI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
Bozdogan-7BI1-Q4_K_M7.6B4.36 GiB0.12 GiB5.54 GiB0.04 GiB36±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A2000 run?
1211 of 2118 indexed open-weight models fit a RTX A2000 at 8,192 context with q4_0 KV cache, the largest being gemma-4-E4B-it at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A2000 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A2000 fast for local AI?
Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.