NVIDIA · datacenter

A100 80GB

A100 80GB has 80 GB of VRAM at 2039 GB/s — about 74.40 GiB usable after driver and compositor overhead. 2071 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
80 GB
HBM2e
Bandwidth
2039 GB/s
5120-bit bus
Tensor FP16
312 TF
dense
TDP
400 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 186text 1781image 2audio asr 39audio tts 21video 16embedding 26

What fits at 32K context

largest quantization that fits, per model · 2071 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3.5-122B-A10BMoEUD-Q4_K_M125B72.89 GiB0.40 GiB74.31 GiB0.09 GiB92±37%
Step-3.7-FlashQ2_K201B66.34 GiB6.92 GiB74.29 GiB0.11 GiB16±22%
GLM-4.7-REAP-218B-A32BMoEUD-IQ2_XXS218B67.13 GiB6.11 GiB74.28 GiB0.12 GiB42±37%
OYM-Qimi-122B-A10B-K2.6MoEI1-Q4_1125B72.82 GiB0.40 GiB74.25 GiB0.15 GiB92±37%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ4_XS139B69.15 GiB4.12 GiB74.25 GiB0.15 GiB56±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ4_XS139B69.15 GiB4.12 GiB74.25 GiB0.15 GiB56±37%
dots.llm1.instMoEQ2_K_L143B56.67 GiB16.47 GiB74.16 GiB0.24 GiB29±37%
MiniMax-M2.5MoEUD-IQ2_XXS229B69.03 GiB4.12 GiB74.13 GiB0.27 GiB63±37%
MiniMax-M2.1MoEUD-IQ2_XXS229B68.98 GiB4.12 GiB74.08 GiB0.32 GiB63±37%
MiniMax-M2MoEUD-IQ2_XXS229B68.92 GiB4.12 GiB74.02 GiB0.38 GiB63±37%
c4ai-command-r-plus-08-2024Q5_K_M104B68.57 GiB4.25 GiB74.00 GiB0.40 GiB16±22%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEQ4_1124B71.37 GiB1.46 GiB73.83 GiB0.57 GiB74±37%
Llama-3_1-Nemotron-51B-InstructQ4_151.5B30.18 GiB42.50 GiB73.82 GiB0.58 GiB16±22%
DeepSeek-Coder-V2-Instruct-0724MoEIQ2_M236B71.64 GiB1.12 GiB73.80 GiB0.60 GiB80±37%
DeepSeek-V2.5MoEIQ2_M236B71.64 GiB1.12 GiB73.80 GiB0.60 GiB80±37%
DeepSeek-Coder-V2-InstructMoEIQ2_M236B71.64 GiB1.12 GiB73.80 GiB0.60 GiB80±37%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedQ5_K_S109B69.16 GiB3.19 GiB73.37 GiB1.03 GiB57±37%
Qwen3.5-REAP-262B-A17BMoEIQ2_XS262B71.81 GiB0.50 GiB73.36 GiB1.04 GiB93±37%
Devstral-2-123B-Instruct-2512Q4_K_S125B66.36 GiB5.84 GiB73.36 GiB1.04 GiB16±22%
Mistral-Medium-3.5-128BI1-Q4_K_S128B66.36 GiB5.84 GiB73.36 GiB1.04 GiB16±22%
XORTRON-NXTXPRTXXLI1-Q4_K_S128B66.36 GiB5.84 GiB73.36 GiB1.04 GiB16±22%
step-3.5-flashQ2_K_L199B65.26 GiB6.92 GiB73.21 GiB1.19 GiB16±22%
command-a-plus-05-2026-bf16MoEIQ2_M219B71.32 GiB0.76 GiB73.08 GiB1.32 GiB70±37%
GLM-4.5-Air-DerestrictedMoEQ4_K_L110B68.88 GiB3.05 GiB72.97 GiB1.43 GiB58±37%
Llama-3_3-Nemotron-Super-49B-v1_5Q4_149.9B29.23 GiB42.50 GiB72.87 GiB1.53 GiB16±22%
Valkyrie-49B-v2.1I1-Q4_149.9B29.23 GiB42.50 GiB72.87 GiB1.53 GiB16±22%
Llama-3_3-Nemotron-Super-49B-v1Q4_149.9B29.23 GiB42.50 GiB72.87 GiB1.53 GiB16±22%
Qwen3.5-122B-A10B-hereticMoEI1-Q4_1123B71.35 GiB0.40 GiB72.78 GiB1.62 GiB93±37%
Seed-OSS-36B-InstructBF1636.2B67.35 GiB4.25 GiB72.69 GiB1.71 GiB16±22%
Hermes-4.3-36BBF1636.2B67.35 GiB4.25 GiB72.69 GiB1.71 GiB16±22%
GLM-4.5-AirMoEQ4_K_M110B68.45 GiB3.05 GiB72.54 GiB1.86 GiB58±37%
Mixtral-8x22B-Instruct-v0.1MoEQ3_K_L141B67.60 GiB3.72 GiB72.38 GiB2.02 GiB28±37%
Mixtral-8x22B-v0.1MoEQ3_K_L141B67.60 GiB3.72 GiB72.37 GiB2.03 GiB28±37%
Mixtral-8x22B-v0.1MoEQ3_K_L141B67.60 GiB3.72 GiB72.37 GiB2.03 GiB28±37%
Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliteratedMoEQ4_K_M123B70.63 GiB0.40 GiB72.06 GiB2.34 GiB94±37%
Behemoth-X-123B-v2Q4_K_S123B64.79 GiB5.84 GiB71.79 GiB2.61 GiB17±22%
Mistral-Large-Instruct-2411Q4_K_S123B64.79 GiB5.84 GiB71.79 GiB2.61 GiB17±22%
Step-3.5-Flash-REAP-121B-A11BI1-Q4_K_S121B63.83 GiB6.92 GiB71.78 GiB2.62 GiB16±22%
MiMo-V2-FlashMoEKV unresolvedIQ2_XXS310B68.47 GiB1.99 GiB71.51 GiB2.89 GiB79±37%
CalmeRys-78B-Orpo-v0.1Q6_K78.0B64.27 GiB5.71 GiB71.12 GiB3.28 GiB17±22%
calme-2.3-rys-78bQ6_K78.0B64.27 GiB5.71 GiB71.12 GiB3.28 GiB17±22%
GLM-4.6VMoEQ4_K_L108B66.89 GiB3.05 GiB70.97 GiB3.43 GiB59±37%
Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16MoEQ8_035.1B69.57 GiB0.33 GiB70.91 GiB3.49 GiB97±37%
Laguna-S-2.1MoEQ4_1118B68.96 GiB0.87 GiB70.85 GiB3.55 GiB84±37%
MiniMax-M2.7MoEUD-IQ2_M229B65.32 GiB4.12 GiB70.42 GiB3.98 GiB65±37%
Qwen2.5-Coder-32B-InstructQ8_032.8B64.86 GiB4.25 GiB70.21 GiB4.19 GiB17±22%
Mistral-Small-4-119B-2603MoEUD-Q4_K_M119B68.70 GiB0.37 GiB70.10 GiB4.30 GiB97±37%
MiMo-V2.5MoEKV unresolvedIQ1_M311B67.01 GiB1.99 GiB70.05 GiB4.35 GiB80±37%
gpt-oss-120b-Uncensored-xCloudMoEI1-Q4_1117B68.42 GiB0.61 GiB70.02 GiB4.38 GiB94±37%
gpt-oss-120b-abliteratedMoEI1-Q4_1117B68.42 GiB0.61 GiB70.02 GiB4.38 GiB94±37%
HunyuanImage-2.1Q6_K17.5B68.97 GiB0.00 GiB70.02 GiB4.38 GiB17±22%
GLM-4.5VMoEI1-Q4_K_M108B65.61 GiB3.05 GiB69.69 GiB4.71 GiB60±37%
Qwen3-235B-A22B-abliteratedMoEI1-IQ2_S235B65.40 GiB3.12 GiB69.56 GiB4.84 GiB59±37%
grok-2MoEIQ2_XXS270B63.81 GiB4.25 GiB69.20 GiB5.20 GiB29±37%
Qwen3-VL-235B-A22B-ThinkingMoEUD-IQ1_M236B64.90 GiB3.12 GiB69.06 GiB5.34 GiB60±37%
Qwen3-VL-235B-A22B-InstructMoEUD-IQ1_M236B64.83 GiB3.12 GiB68.98 GiB5.42 GiB60±37%
gpt-oss-20b-hereticMoEIQ4_NL20.9B67.58 GiB0.41 GiB68.97 GiB5.43 GiB52±37%
HarmonicHarlequin_v5-20BQ8_033.3B32.97 GiB34.53 GiB68.55 GiB5.85 GiB17±22%
Qwen3.5-88BMoEI1-Q6_K87.7B67.08 GiB0.40 GiB68.50 GiB5.90 GiB88±37%
MiniMax-M2.7-BF16-ultra-uncensored-hereticMoEI1-IQ2_S229B63.36 GiB4.12 GiB68.47 GiB5.93 GiB66±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation32.79 it/s18.5843.5581
Prompt processing4666.46 tok/s3574.565059.4918
Text generation179.67 tok/s169.96187.4916
Benchmarked· n=81

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a A100 80GB run?
2071 of 2118 indexed open-weight models fit a A100 80GB at 32,768 context with q8_0 KV cache, the largest being Qwen3.5-122B-A10B at UD-Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a A100 80GB actually have?
Its nameplate is 80 GB, but about 74.40 GiB is available to a model once driver and compositor overhead is accounted for.
Is a A100 80GB fast for local AI?
Its memory bandwidth is 2039 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.