RTX A1000
RTX A1000 has 8 GB of VRAM at 192 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1346 of 2118 indexed models fit at 8K context with f16 KV.
What fits at 8K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Ling-liteMoE | IQ2_S | 16.8B | 6.01 GiB | 0.44 GiB | 7.44 GiB | 0.00 GiB | 46±37% |
| L3-Dark-Planet-8B | Q5_K_S | 8.0B | 5.40 GiB | 1.00 GiB | 7.44 GiB | 0.00 GiB | 17±22% |
| deepseek-coder-6.7B-kexer | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Magicoder-S-DS-6.7B | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| deepseek-coder-6.7b-base | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| WizardLM-7B-Uncensored | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Llama-2-7B-32K-Instruct | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Luna-AI-Llama2-Uncensored | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Swallow-7b-NVE-instruct-hf | I1-IQ3_XXS | 6.7B | 2.41 GiB | 4.00 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Rocinante-XL-16B-v1 | I1-IQ2_XS | 16.1B | 4.70 GiB | 1.69 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Bonsai-8B-unpacked | Q4_K_L | 8.2B | 5.27 GiB | 1.13 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| gemma-4-12B | IQ3_M | 12.0B | 5.41 GiB | 0.97 GiB | 7.43 GiB | 0.01 GiB | 17±22% |
| Apriel-1.6-15b-Thinker | I1-Q2_K_S | 14.9B | 4.88 GiB | 1.50 GiB | 7.42 GiB | 0.02 GiB | 17±22% |
| Gemma-The-Writer-N-Restless-Quill-10B-Uncensored | Q3_K_S | 10.0B | 4.14 GiB | 2.25 GiB | 7.42 GiB | 0.02 GiB | 17±22% |
| Ministral-3-14B-Instruct-2512 | UD-IQ3_XXS | 13.9B | 5.12 GiB | 1.25 GiB | 7.42 GiB | 0.02 GiB | 17±22% |
| Ministral-3-14B-Reasoning-2512 | UD-IQ3_XXS | 13.9B | 5.12 GiB | 1.25 GiB | 7.42 GiB | 0.02 GiB | 17±22% |
| NVIDIA-Nemotron-Nano-9B-v2 | IQ2_S | 8.9B | 4.62 GiB | 1.75 GiB | 7.42 GiB | 0.02 GiB | 17±22% |
| Phi-3-medium-128k-instruct | Q2_K | 14.0B | 4.79 GiB | 1.56 GiB | 7.41 GiB | 0.03 GiB | 17±22% |
| Phi-3-medium-4k-instruct | I1-Q2_K | 14.0B | 4.79 GiB | 1.56 GiB | 7.41 GiB | 0.03 GiB | 17±22% |
| zeta-2.1 | I1-Q5_K_S | 8.3B | 5.37 GiB | 1.00 GiB | 7.41 GiB | 0.03 GiB | 17±22% |
| Kimi-VL-A3B-InstructMoE | I1-Q2_K_S | 16.4B | 6.15 GiB | 0.24 GiB | 7.40 GiB | 0.04 GiB | 53±37% |
| Moonlight-16B-A3B-InstructMoE | Q2_K_S | 16.0B | 6.15 GiB | 0.24 GiB | 7.40 GiB | 0.04 GiB | 53±37% |
| Gemma-4-E4B-Luchador | Q5_K_L | 8.0B | 6.21 GiB | 0.18 GiB | 7.40 GiB | 0.04 GiB | 17±22% |
| Nemotron-3-Embed-8B-BF16 | Q5_K_M | 8.0B | 5.30 GiB | 1.06 GiB | 7.40 GiB | 0.04 GiB | 17±22% |
| Ling-mini-2.0MoE | IQ3_XXS | 16.3B | 6.10 GiB | 0.31 GiB | 7.40 GiB | 0.04 GiB | 71±37% |
| Wan2.2-Animate-14B | Q2_K | 17.3B | 6.36 GiB | 0.00 GiB | 7.40 GiB | 0.04 GiB | 17±22% |
| granite-4.0-h-3b-ar | F16 | 3.4B | 6.33 GiB | 0.06 GiB | 7.40 GiB | 0.04 GiB | 17±22% |
| Qwen3.5-9B | Q5_K_S | 9.7B | 6.11 GiB | 0.25 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| Olmo-3-7B-Instruct | Q3_K_L | 7.3B | 3.68 GiB | 2.69 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| Olmo-3-7B-Think | I1-Q3_K_L | 7.3B | 3.68 GiB | 2.69 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| Falcon3-10B-Instruct | Q3_K_L | 10.3B | 5.08 GiB | 1.25 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| ERNIE-4.5-21B-A3B-Thinking | IQ2_S | 21.8B | 5.93 GiB | 0.44 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| ERNIE-4.5-21B-A3B-PT | IQ2_S | 21.9B | 5.93 GiB | 0.44 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| NVIDIA-Nemotron-Nano-12B-v2 | Q2_K | 12.3B | 4.38 GiB | 1.94 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| Teuken-7B-instruct-research-v0.4 | I1-Q6_K | 7.5B | 6.10 GiB | 0.25 GiB | 7.39 GiB | 0.05 GiB | 17±22% |
| Assistant_Pepe_8B | Q5_K_M | — | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| MathCoder2-CodeLlama-7B | Q2_K | 6.7B | 2.36 GiB | 4.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| llava-v1.5-7b | Q2_K | 6.7B | 2.36 GiB | 4.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Phi-4-reasoning-plus | IQ2_M | 14.7B | 4.76 GiB | 1.56 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Phi-4-reasoning | IQ2_M | 14.7B | 4.76 GiB | 1.56 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| phi-4 | IQ2_M | 14.7B | 4.76 GiB | 1.56 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Kimi-VL-A3B-Thinking-2506MoE | Q2_K | 16.4B | 6.13 GiB | 0.24 GiB | 7.38 GiB | 0.06 GiB | 53±37% |
| Foundation-Sec-8B-Instruct | I1-Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Foundation-Sec-8B-Instruct-heretic | I1-Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| dolphin-2.9-llama3-8b | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| saiga_llama3_8b | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Meta-Llama-3-8B-Instruct | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| MiniCPM-Llama3-V-2_5 | Q5_K | 8.5B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| llama-3-8b-bnb-4bit | Q5_K_M | 8.2B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Llama3-ChatQA-1.5-8B | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Meta-Llama-3-8B | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Llama-3.1-Tulu-3-8B | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Llama-3-Groq-8B-Tool-Use | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| llama3.1-heretic-unsensored | I1-Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Dolphin3.0-Llama3.1-8B | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| dolphin-2.9.4-llama3.1-8b | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Dolphin3.0-Llama3.1-8B-abliterated | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| LLAMA-3_8B_Unaligned_BETA | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Llama-3.1-8B | Q5_K | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
| Meta-Llama-3-8B | Q5_K_M | 8.0B | 5.34 GiB | 1.00 GiB | 7.38 GiB | 0.06 GiB | 17±22% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Image generation | 3.75 it/s | 3.59–4.05 | 7 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.
Questions people ask
- What AI models can a RTX A1000 run?
- 1346 of 2118 indexed open-weight models fit a RTX A1000 at 8,192 context with f16 KV cache, the largest being Ling-lite at IQ2_S. That covers text, vision-language, image, video and speech models.
- How much usable memory does a RTX A1000 actually have?
- Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a RTX A1000 fast for local AI?
- Its memory bandwidth is 192 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.