NVIDIA · workstation

RTX A2000

RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1163 of 2118 indexed models fit at 16K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 989vision language 88embedding 26audio asr 38image 1audio tts 19video 2

What fits at 16K context

largest quantization that fits, per model · 1163 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
llama-3.2-3b-instructQ2_K3.2B4.08 GiB0.49 GiB5.58 GiB0.00 GiB36±22%
LFM2.5-Queen-Opus-4.7-8B-A1BMoEI1-Q4_K_S8.5B4.53 GiB0.05 GiB5.58 GiB0.00 GiB104±37%
LFM2.5-8B-A1B-KO-SFTMoEI1-Q4_K_S8.5B4.53 GiB0.05 GiB5.58 GiB0.00 GiB104±37%
LFM2.5-8B-A1B-hereticMoEI1-Q4_K_S8.5B4.53 GiB0.05 GiB5.58 GiB0.00 GiB104±37%
LFM2.5-8B-A1B-SOMPOA-heresyMoEI1-Q4_K_S8.5B4.53 GiB0.05 GiB5.58 GiB0.00 GiB104±37%
Huihui-LFM2.5-8B-A1B-abliteratedMoEI1-Q4_K_S8.5B4.53 GiB0.05 GiB5.58 GiB0.00 GiB104±37%
Supertron2.1-8B-A1BMoEI1-Q4_K_S8.5B4.53 GiB0.05 GiB5.58 GiB0.00 GiB104±37%
Falcon3-7B-InstructQ4_07.5B4.02 GiB0.49 GiB5.58 GiB0.00 GiB36±22%
zeta-2Q3_K_M8.3B3.97 GiB0.56 GiB5.58 GiB0.00 GiB36±22%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseIQ4_XS8.1B4.42 GiB0.12 GiB5.58 GiB0.00 GiB36±22%
Fara1.5-9BIQ3_M9.4B4.40 GiB0.14 GiB5.57 GiB0.01 GiB36±22%
QwenPaw-Flash-9BIQ3_M9.4B4.40 GiB0.14 GiB5.57 GiB0.01 GiB36±22%
grug-9bIQ3_M9.4B4.40 GiB0.14 GiB5.57 GiB0.01 GiB36±22%
OmniCoder-9BIQ3_M9.4B4.40 GiB0.14 GiB5.57 GiB0.01 GiB36±22%
Ornith-1.0-9BIQ3_M9.2B4.40 GiB0.14 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9B-NeoIQ3_M9.7B4.40 GiB0.14 GiB5.57 GiB0.01 GiB36±22%
Tini-Cybersec-8B-A1BMoEIQ4_NL8.5B4.52 GiB0.05 GiB5.57 GiB0.01 GiB104±37%
gemma-4-E4B-it-hereticQ4_08.0B4.48 GiB0.08 GiB5.57 GiB0.01 GiB36±22%
Marco-Nano-InstructMoEI1-IQ4_XS8.0B4.10 GiB0.49 GiB5.57 GiB0.01 GiB102±37%
Falcon3-10B-InstructI1-IQ3_XXS10.3B3.80 GiB0.70 GiB5.57 GiB0.01 GiB36±22%
LocateAnything-3BQ4_K3.8B4.39 GiB0.16 GiB5.56 GiB0.02 GiB36±22%
LFM2.5-8B-A1BMoEQ4_08.5B4.51 GiB0.05 GiB5.56 GiB0.02 GiB104±37%
Teuken-7B-instruct-research-v0.4I1-Q4_K_S7.5B4.38 GiB0.14 GiB5.56 GiB0.02 GiB36±22%
Qwythos-9B-v2IQ3_XS9.7B4.37 GiB0.14 GiB5.55 GiB0.03 GiB36±22%
Tess-4-9BIQ3_XS9.7B4.37 GiB0.14 GiB5.55 GiB0.03 GiB36±22%
Qwen3-14BUD-IQ1_M14.8B3.79 GiB0.70 GiB5.55 GiB0.03 GiB36±22%
SambaLingo-Japanese-ChatI1-Q2_K_S6.9B2.27 GiB2.25 GiB5.55 GiB0.03 GiB36±22%
Qwen3-VL-Embedding-8BQ3_K_L8.1B3.88 GiB0.63 GiB5.54 GiB0.04 GiB36±22%
qwen-indic-v1I1-Q3_K_L7.6B3.88 GiB0.63 GiB5.54 GiB0.04 GiB36±22%
Qwen3-Embedding-8BQ3_K_L7.6B3.88 GiB0.63 GiB5.54 GiB0.04 GiB36±22%
Ling-mini-2.0MoEIQ2_S16.3B4.38 GiB0.18 GiB5.54 GiB0.04 GiB152±37%
NVIDIA-Nemotron-3-Nano-4B-BF16Q6_K4.0B3.78 GiB0.74 GiB5.54 GiB0.04 GiB36±22%
Aya-Medikal-V2I1-Q3_K_M8.0B3.93 GiB0.56 GiB5.54 GiB0.04 GiB36±22%
granite-speech-4.1-2b-narBF162.3B4.20 GiB0.35 GiB5.54 GiB0.04 GiB36±22%
Qwen3-16B-A3BMoEIQ2_XXS16.0B4.12 GiB0.42 GiB5.54 GiB0.04 GiB78±37%
Gemma-4-E4B-LuchadorIQ3_M8.0B4.44 GiB0.08 GiB5.54 GiB0.04 GiB36±22%
Ministral-3-14B-Instruct-2512UD-IQ2_XXS13.9B3.78 GiB0.70 GiB5.54 GiB0.04 GiB36±22%
Ministral-3-14B-Reasoning-2512UD-IQ2_XXS13.9B3.78 GiB0.70 GiB5.54 GiB0.04 GiB36±22%
Nanbeige4.2-3B-hereticQ8_04.2B4.13 GiB0.39 GiB5.54 GiB0.04 GiB36±22%
Nanbeige4.2-3BQ8_04.2B4.13 GiB0.39 GiB5.54 GiB0.04 GiB36±22%
Rocinante-XL-16B-v1I1-IQ1_S16.1B3.54 GiB0.95 GiB5.54 GiB0.04 GiB36±22%
SuperGemma-4-12b-abliteratedI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-hereticI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-hereticI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12B-it-uncensored-hereticI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
Grug-12BI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
Aura-Medium-v1-BF16I1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12B-it-Esper4I1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12B-it-GuardpointI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
Gemma-4-12B-it-AEON-Abliterated-K4-BF16I1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12B-it-Tachibana-AgentI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12b-marvin-gutenberg-rp-v2I1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12b-crownelius-writerI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliteratedI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma-4-12b-asterion-agenticI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
Huihui-gemma-4-12B-agentic-fable5-abliteratedI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
g4-12b-it-trismegistusI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
gemma4-12b-it-asimovI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
FabGemmaI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliteratedI1-IQ2_M12.0B4.07 GiB0.41 GiB5.53 GiB0.05 GiB36±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A2000 run?
1163 of 2118 indexed open-weight models fit a RTX A2000 at 16,384 context with q4_0 KV cache, the largest being llama-3.2-3b-instruct at Q2_K. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A2000 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A2000 fast for local AI?
Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.
RTX A2000 — what AI models can it run locally? — ossmodeldb