RTX A2000
RTX A2000 has 12 GB of VRAM at 288 GB/s — about 11.16 GiB usable after driver and compositor overhead. 653 of 2118 indexed models fit at 128K context with f16 KV.
What fits at 128K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| granite-3.1-3b-a800m-instructMoE | Q5_K_M | 3.3B | 2.19 GiB | 8.00 GiB | 11.16 GiB | 0.00 GiB | 11±37% |
| Qwen3.5-9B | Q5_K_S | 9.7B | 6.11 GiB | 4.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| GLM-4.6V-Flash | Q4_0 | 10.3B | 5.10 GiB | 5.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| GLM-Z1-9B-0414 | Q4_0 | 9.4B | 5.10 GiB | 5.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| glm4.1v-9b-base-sft | I1-Q4_0 | 10.3B | 5.10 GiB | 5.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| GLM-4-9B-0414 | Q4_0 | 9.4B | 5.10 GiB | 5.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| GLM-4.1V-9B-Thinking | Q4_0 | 10.3B | 5.10 GiB | 5.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| Teuken-7B-instruct-research-v0.4 | I1-Q6_K | 7.5B | 6.10 GiB | 4.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| Falcon3-1B-Instruct | Q5_K_M | 1.7B | 1.13 GiB | 9.00 GiB | 11.13 GiB | 0.03 GiB | 16±22% |
| DeepScaleR-1.5B-Preview | F32 | 1.8B | 6.63 GiB | 3.50 GiB | 11.13 GiB | 0.03 GiB | 16±22% |
| DeepSeek-R1-Distill-Qwen-1.5B | F32 | 1.8B | 6.63 GiB | 3.50 GiB | 11.13 GiB | 0.03 GiB | 16±22% |
| Aurora-Code-1MoE | I1-IQ2_XXS | 34.7B | 7.62 GiB | 2.50 GiB | 11.12 GiB | 0.04 GiB | 29±37% |
| VibeVoice-1.5B | F32 | 2.7B | 10.07 GiB | 0.00 GiB | 11.12 GiB | 0.04 GiB | 16±22% |
| granite-34b-code-base-8k | I1-IQ2_S | 33.7B | 10.04 GiB | 0.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Fara1.5-9B | Q5_K_S | 9.4B | 6.08 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| QwenPaw-Flash-9B | Q5_K_S | 9.4B | 6.08 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| grug-9b | Q5_K_S | 9.4B | 6.08 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| OmniCoder-9B | Q5_K_S | 9.4B | 6.08 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Ornith-1.0-9B | Q5_K_S | 9.2B | 6.08 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Neo | Q5_K_S | 9.7B | 6.08 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| internlm3-8b-instruct | Q3_K_M | 8.8B | 4.09 GiB | 6.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Muse-Glimmer-30B | IQ2_XXS | 29.8B | 8.31 GiB | 1.72 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Crow-9B-HERETIC-4.6 | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSORED | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCT | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| NaNovel-9B | I1-Q5_K_M | 9.7B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Unredacted-MAX | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-abliterated | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Miss_MARTHA-9B-Qwen3.5-Omni | I1-Q5_K_M | 9.0B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Ken3.5-9B | I1-Q5_K_M | 9.7B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Huihui-Qwen3.5-9B-abliterated | Q5_K_M | 9.7B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-abliterated | Q5_K_M | 9.0B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-Base | Q5_K | 9.7B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-heretic-v2 | Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning | I1-Q5_K_M | 9.4B | 6.07 GiB | 4.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Nanbeige4.1-3B | IQ4_XS | 3.9B | 2.09 GiB | 8.00 GiB | 11.10 GiB | 0.06 GiB | 16±22% |
| Wan2.2-Distill-Models | Q5_K_M | 14.3B | 10.06 GiB | 0.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| SkyReels-V2-DF-14B-540P | Q5_K_M | 14.3B | 10.06 GiB | 0.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| nomic-embed-code | Q3_K_S | 7.1B | 3.03 GiB | 7.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| SmolLM3-3B | UD-IQ2_M | 3.1B | 1.07 GiB | 9.00 GiB | 11.08 GiB | 0.08 GiB | 16±22% |
| Wan2.1-T2V-14B | Q5_0 | 14.3B | 10.05 GiB | 0.00 GiB | 11.08 GiB | 0.08 GiB | 16±22% |
| Qwen3.5-9B-Coder | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Qwythos-9B-Claude-Mythos-5-1M-MTP | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Qwen3.5-9B-Fable-5-v1 | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Qwythos-9B-v2 | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| PINQWEN-3.5-9B-1M-BF16 | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Openprose-2-Flash | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Qwen3.5-9B-Nikusui-v1 | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Ornstein-3.5-9B-V1.5 | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Ornith-1.0-9B-heretic-MTP | I1-Q5_K_S | 9.4B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Tess-4-9B | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| dotwebs-1 | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| lift | Q5_0 | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Hemlock-Qwopus3.5-9B-Coder | I1-Q5_K_S | 9.7B | 6.03 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Vero-Qwen35-9B-Base | I1-Q5_K_M | 9.4B | 6.02 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Vero-Qwen35-9B | I1-Q5_K_M | 9.4B | 6.02 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
| Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-Thinking | I1-Q5_K_M | 9.4B | 6.02 GiB | 4.00 GiB | 11.06 GiB | 0.10 GiB | 16±22% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Image generation | 4.98 it/s | 3.58–6.36 | 66 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.
Questions people ask
- What AI models can a RTX A2000 run?
- 653 of 2118 indexed open-weight models fit a RTX A2000 at 131,072 context with f16 KV cache, the largest being granite-3.1-3b-a800m-instruct at Q5_K_M. That covers text, vision-language, image, video and speech models.
- How much usable memory does a RTX A2000 actually have?
- Its nameplate is 12 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a RTX A2000 fast for local AI?
- Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.