NVIDIA · workstation

RTX A1000

RTX A1000 has 8 GB of VRAM at 192 GB/s — about 7.44 GiB usable after driver and compositor overhead. 917 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
192 GB/s
128-bit bus
Tensor FP16
27 TF
dense
TDP
50 W
$365 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 84text 743audio asr 38audio tts 20video 7embedding 24image 1

What fits at 64K context

largest quantization that fits, per model · 917 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwythos-9B-v2Q4_K_S9.7B5.34 GiB1.06 GiB7.44 GiB0.00 GiB17±22%
Tess-4-9BQ4_K_S9.7B5.34 GiB1.06 GiB7.44 GiB0.00 GiB17±22%
SciPhi-Self-RAG-Mistral-7B-32kKV unresolvedI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ3_K_M12.1B5.58 GiB0.80 GiB7.44 GiB0.00 GiB17±22%
dolphin-2.2.1-mistral-7bKV unresolvedI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
OpenChat-3.5-7B-Qwen-v2.0KV unresolvedI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
openchat-3.5-0106KV unresolvedI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Mistral-7B-Instruct-v0.1KV unresolvedI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Mistral-7B-Instruct-v0.2I1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
ContextualKunoichi_KTO-7BI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
xLAM-7b-rI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Ninja-v1-RP-WIPKV unresolvedI1-IQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Kunoichi-DPO-v2-7BKV unresolvedIQ2_S7.2B2.15 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Marco-Nano-InstructMoEI1-IQ2_XXS8.0B2.75 GiB3.72 GiB7.44 GiB0.00 GiB16±37%
Phi-4-mini-reasoningQ4_K_S3.8B2.18 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Phi-4-mini-instructQ4_K_S3.8B2.18 GiB4.25 GiB7.44 GiB0.00 GiB17±22%
Qwen3.6-28BMoEI1-IQ1_S28.2B5.77 GiB0.66 GiB7.44 GiB0.00 GiB50±37%
Qwen3.5-28BMoEI1-IQ1_S28.7B5.77 GiB0.66 GiB7.44 GiB0.00 GiB50±37%
Jan-v1-4BQ2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
Jan-nano-128kQ2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
Qwen3-4B-Instruct-2507Q2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
Qwen3-4B-Thinking-2507Q2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
Jan-nanoQ2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
Qwen3-4B-abliteratedQ2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
Qwen3-4B-Instruct-2507-hereticQ2_K_L4.0B1.64 GiB4.78 GiB7.43 GiB0.01 GiB17±22%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-IQ2_M12.2B4.01 GiB2.37 GiB7.43 GiB0.01 GiB17±22%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-IQ2_M12.2B4.01 GiB2.37 GiB7.43 GiB0.01 GiB17±22%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-IQ2_M12.2B4.01 GiB2.37 GiB7.43 GiB0.01 GiB17±22%
Floppa-12B-Gemma3-UncensoredI1-IQ2_M12.2B4.01 GiB2.37 GiB7.43 GiB0.01 GiB17±22%
gemma-3-12b-it-hereticI1-IQ2_M12.2B4.01 GiB2.37 GiB7.43 GiB0.01 GiB17±22%
gemma-3-12b-it-abliteratedIQ2_M12.2B4.01 GiB2.37 GiB7.43 GiB0.01 GiB17±22%
Llama-3.1-8B-InstructUD-IQ1_M8.0B2.13 GiB4.25 GiB7.42 GiB0.02 GiB17±22%
Llama-3.1-Nemotron-Nano-8B-v1UD-IQ1_M8.0B2.13 GiB4.25 GiB7.42 GiB0.02 GiB17±22%
DeepSeek-R1-Distill-Llama-8BUD-IQ1_M8.0B2.13 GiB4.25 GiB7.42 GiB0.02 GiB17±22%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseI1-Q5_K_M8.1B5.45 GiB0.93 GiB7.42 GiB0.02 GiB17±22%
Gemma-4-E4B-LuchadorQ6_K8.0B5.90 GiB0.50 GiB7.41 GiB0.03 GiB17±22%
Voxtral-Mini-3B-2507Q4_14.7B2.42 GiB3.98 GiB7.41 GiB0.03 GiB17±22%
orpheus-3b-0.1-pretrainedQ5_13.8B2.68 GiB3.72 GiB7.41 GiB0.03 GiB17±22%
salamandra-7b-instruct-2606I1-IQ1_S7.8B2.13 GiB4.25 GiB7.41 GiB0.03 GiB17±22%
rnj-1-instructUD-IQ1_M8.3B2.11 GiB4.25 GiB7.41 GiB0.03 GiB17±22%
zeta-2.1I1-IQ1_M8.3B2.12 GiB4.25 GiB7.41 GiB0.03 GiB17±22%
Hubble-4B-v1Q3_K_M4.5B2.14 GiB4.25 GiB7.40 GiB0.04 GiB17±22%
Aura-4BI1-Q3_K_M4.5B2.14 GiB4.25 GiB7.40 GiB0.04 GiB17±22%
magnum-v2-4bI1-Q3_K_M4.5B2.14 GiB4.25 GiB7.40 GiB0.04 GiB17±22%
Impish_LLAMA_4BQ3_K_M4.5B2.14 GiB4.25 GiB7.40 GiB0.04 GiB17±22%
Llama-3.1-Minitron-4B-Width-BaseQ3_K_M4.5B2.14 GiB4.25 GiB7.40 GiB0.04 GiB17±22%
Luna-7B-A4BMoEI1-IQ1_M6.7B1.61 GiB4.78 GiB7.40 GiB0.04 GiB11±37%
Kimi-VL-A3B-InstructMoEI1-IQ2_XXS16.4B5.38 GiB1.01 GiB7.40 GiB0.04 GiB34±37%
Moonlight-16B-A3B-InstructMoEIQ2_XXS16.0B5.38 GiB1.01 GiB7.40 GiB0.04 GiB34±37%
Wan2.2-Animate-14BQ2_K17.3B6.36 GiB0.00 GiB7.40 GiB0.04 GiB17±22%
Apertus-8B-Instruct-2509UD-IQ1_S8.1B2.08 GiB4.25 GiB7.39 GiB0.05 GiB17±22%
internlm3-8b-instructQ4_K_S8.8B4.77 GiB1.59 GiB7.39 GiB0.05 GiB17±22%
OLMoE-1B-7B-0924-InstructMoEI1-IQ2_M6.9B2.17 GiB4.25 GiB7.39 GiB0.05 GiB14±37%
Qwen2.5-Omni-3BBF165.5B6.33 GiB0.00 GiB7.38 GiB0.06 GiB17±22%
gemma-3n-E4B-itQ6_K7.8B5.84 GiB0.49 GiB7.37 GiB0.07 GiB17±22%
Teuken-7B-instruct-research-v0.4I1-Q5_K_M7.5B5.27 GiB1.06 GiB7.37 GiB0.07 GiB17±22%
Llama-3.2-3B-Instruct-uncensoredQ5_K_L3.6B2.64 GiB3.72 GiB7.37 GiB0.07 GiB17±22%
CycleGRPO-4BI1-Q2_K_S4.8B1.58 GiB4.78 GiB7.37 GiB0.07 GiB17±22%
Nemotron-3-Embed-8B-BF16IQ1_S8.0B1.81 GiB4.52 GiB7.37 GiB0.07 GiB17±22%
Wan2.1-T2V-1.3BQ6_K1.4B6.35 GiB0.00 GiB7.36 GiB0.08 GiB17±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.75 it/s3.594.057
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a RTX A1000 run?
917 of 2118 indexed open-weight models fit a RTX A1000 at 65,536 context with q8_0 KV cache, the largest being Qwythos-9B-v2 at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A1000 actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A1000 fast for local AI?
Its memory bandwidth is 192 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.