NVIDIA · workstation

RTX A1000

RTX A1000 has 8 GB of VRAM at 192 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1170 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
192 GB/s
128-bit bus
Tensor FP16
27 TF
dense
TDP
50 W
$365 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 94text 984embedding 26video 7audio asr 38image 1audio tts 20

What fits at 32K context

largest quantization that fits, per model · 1170 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
LocateAnything-3BQ8_03.8B5.83 GiB0.60 GiB7.44 GiB0.00 GiB17±22%
Vero-Qwen35-9B-BaseI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Vero-Qwen35-9BI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Crow-9B-HERETIC-4.6I1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSOREDI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCTI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Morphos-9BI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwable-9B-Claude-Fable-5-hereticI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Holo-3.1-9BI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwable-9B-Claude-Fable-5I1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-imabari-v2I1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-abliterated-v2-MAXI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwable-9B-Claude-Fable-5-StraTAI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
NaNovel-9BI1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Unredacted-MAXI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-RpRMax-v1I1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-abliteratedI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
AdQWENistrator-9BI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
cajal-9b-v2-fullI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-ultra-uncensored-hereticQ5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Holo-3.1-9B-CoderI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
PlutoI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Holo-3.1-9B-abliterated-rdoI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-Uncensored-cyber-v3Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Miss_MARTHA-9B-Qwen3.5-OmniI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Ken3.5-9BI1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Huihui-Qwen3.5-9B-abliteratedQ5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-BaseI1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
qwen3.5-9b-nsfw-captioning-v5I1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-abliteratedQ5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-DS-v4-Flash-v3.0Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
OmniCoder-9BQ5_09.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-gemini-3.1-opus-4.6-reasoningI1-Q5_K_S9.4B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwen3.5-9B-DeepSeek-V4-FlashI1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Huihui-Qwen3.5-9B-Claude-4.6-Opus-abliteratedI1-Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Qwopus3.5-9B-v3.5Q5_K_S9.7B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
Katarau-9B-ru-RP-nsfwI1-Q5_K_S9.0B5.87 GiB0.53 GiB7.44 GiB0.00 GiB17±22%
G9v3-3BBF163.0B5.57 GiB0.86 GiB7.43 GiB0.01 GiB17±22%
Trinity-MiniMoEIQ2_XXS26.1B6.11 GiB0.33 GiB7.43 GiB0.01 GiB58±37%
rocket-3BQ2_K2.8B1.12 GiB5.31 GiB7.43 GiB0.01 GiB17±22%
stable-code-3bI1-IQ3_XS2.8B1.11 GiB5.31 GiB7.42 GiB0.02 GiB17±22%
granite-8b-code-instruct-4kI1-Q3_K_L8.1B3.99 GiB2.39 GiB7.42 GiB0.02 GiB17±22%
granite-8b-code-base-4kI1-Q3_K_L8.1B3.99 GiB2.39 GiB7.42 GiB0.02 GiB17±22%
Smilodon-9B-v1I1-IQ2_M10.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
bella-bartender-v2I1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
Gemma-2-9B-It-SPPO-Iter3I1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
Gemma-SEA-LION-v3-9B-ITI1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
G2-Darkest-Writer-9B-v1I1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
Tiger-Gemma-9B-v3I1-IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
gemma-2-9b-it-abliteratedIQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
gemma-2-9b-itIQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
Tiger-Gemma-9B-v1IQ2_M9.2B3.20 GiB3.18 GiB7.42 GiB0.02 GiB17±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.75 it/s3.594.057
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a RTX A1000 run?
1170 of 2118 indexed open-weight models fit a RTX A1000 at 32,768 context with q8_0 KV cache, the largest being LocateAnything-3B at Q8_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A1000 actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A1000 fast for local AI?
Its memory bandwidth is 192 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.