NVIDIA · datacenter

A100 40GB

A100 40GB has 40 GB of VRAM at 1555 GB/s — about 37.20 GiB usable after driver and compositor overhead. 1924 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
40 GB
HBM2
Bandwidth
1555 GB/s
5120-bit bus
Tensor FP16
312 TF
dense
TDP
400 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1643vision language 177image 2audio asr 39audio tts 21video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 1924 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
spoomplesmaxx-v2.1-30BI1-Q5_K_M28.9B19.09 GiB17.00 GiB37.20 GiB0.00 GiB25±22%
Huihui-granite-4.1-30b-abliteratedI1-Q5_K_M28.9B19.09 GiB17.00 GiB37.20 GiB0.00 GiB25±22%
granite-4.1-30b-hereticI1-Q5_K_M28.9B19.09 GiB17.00 GiB37.20 GiB0.00 GiB25±22%
granite-4.1-30bQ5_K_M28.9B19.09 GiB17.00 GiB37.20 GiB0.00 GiB25±22%
WizardLM-7B-UncensoredI1-Q2_K_S6.7B2.16 GiB34.00 GiB37.18 GiB0.02 GiB24±22%
Llama-2-7B-32K-InstructI1-Q2_K_S6.7B2.16 GiB34.00 GiB37.18 GiB0.02 GiB24±22%
Phi-3.5-MoE-instructMoEKV unresolvedQ5_K_M41.9B27.68 GiB8.50 GiB37.18 GiB0.02 GiB37±37%
SambaLingo-Japanese-ChatI1-IQ2_S6.9B2.15 GiB34.00 GiB37.18 GiB0.02 GiB24±22%
Phi-3-mini-4k-instructKV unresolvedQ6_K3.8B10.67 GiB25.50 GiB37.17 GiB0.03 GiB24±22%
DeepSeek-R1-Distill-Llama-70BUD-IQ1_S70.6B14.79 GiB21.25 GiB37.17 GiB0.03 GiB25±22%
Hy-MT2-30B-A3BMoEQ8_030.1B29.79 GiB6.38 GiB37.16 GiB0.04 GiB50±37%
Hermes-4-70BUD-IQ1_S70.6B14.77 GiB21.25 GiB37.14 GiB0.06 GiB25±22%
Llama-3.3-70B-InstructUD-IQ1_S70.6B14.77 GiB21.25 GiB37.14 GiB0.06 GiB25±22%
CallerQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Dumpling-Qwen2.5-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
OREAL-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
openhands-lm-32b-v0.1Q4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
LongWriter-Zero-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
OpenCodeReasoning-Nemotron-32B-IOIQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-Coder-32B-Instruct-abliteratedQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
OlympicCoder-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
OpenCodeReasoning-Nemotron-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
OpenThinker-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
QwQ-32B-ArliAI-RpR-v4Q4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-Coder-32B-InstructQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
QwQ-32B-abliteratedQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
OpenThinker2-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
INTELLECT-2Q4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-32B-InstructQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
QwQ-32B-PreviewQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-Coder-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-32b-RP-InkQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
TinyR1-32B-PreviewQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
deepseek-r1-qwen-2.5-32B-ablatedQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Rombos-LLM-V2.5-Qwen-32bQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
DeepSeek-R1-Distill-Qwen-32B-abliteratedQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-32B-ArliAI-RPMax-v1.3Q4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
DeepSeek-R1-Distill-Qwen-32BQ4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
Qwen2.5-VL-32B-InstructQ4_K_L33.5B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
EVA-Qwen2.5-32B-v0.2Q4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
EVA-Qwen2.5-32B-v0.1Q4_K_L32.8B19.03 GiB17.00 GiB37.13 GiB0.07 GiB25±22%
cogito-v1-preview-qwen-32BQ4_K_L32.8B19.02 GiB17.00 GiB37.12 GiB0.08 GiB25±22%
QwQ-32B-Snowdrop-v0Q4_K_L32.8B19.02 GiB17.00 GiB37.12 GiB0.08 GiB25±22%
Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoEI1-Q5_K_S42.4B27.23 GiB8.90 GiB37.12 GiB0.08 GiB41±37%
llm-jp-4-32b-a3b-thinkingMoEQ8_032.1B31.84 GiB4.25 GiB37.09 GiB0.11 GiB61±37%
deepseek-coder-6.7B-kexerI1-IQ2_S6.7B2.05 GiB34.00 GiB37.07 GiB0.13 GiB25±22%
Magicoder-S-DS-6.7BI1-IQ2_S6.7B2.05 GiB34.00 GiB37.07 GiB0.13 GiB25±22%
deepseek-coder-6.7b-baseI1-IQ2_S6.7B2.05 GiB34.00 GiB37.07 GiB0.13 GiB25±22%
Luna-AI-Llama2-UncensoredI1-IQ2_S6.7B2.05 GiB34.00 GiB37.07 GiB0.13 GiB25±22%
Swallow-7b-NVE-instruct-hfI1-IQ2_S6.7B2.05 GiB34.00 GiB37.07 GiB0.13 GiB25±22%
OpenBuddy-R1-0528-Distill-Qwen3-32B-Preview0-QATQ4_K_L32.8B18.94 GiB17.00 GiB37.04 GiB0.16 GiB25±22%
KAT-DevQ4_K_L32.8B18.94 GiB17.00 GiB37.03 GiB0.17 GiB25±22%
Qwen3-VL-32B-InstructQ4_K_L33.4B18.94 GiB17.00 GiB37.03 GiB0.17 GiB25±22%
DeepSWE-PreviewQ4_K_L32.8B18.94 GiB17.00 GiB37.03 GiB0.17 GiB25±22%
internlm3-8b-instructF328.8B32.80 GiB3.19 GiB37.01 GiB0.19 GiB25±22%
Janus-Pro-7BI1-Q4_17.4B4.10 GiB31.88 GiB37.00 GiB0.20 GiB25±22%
deepseek-math-7b-instructQ4_16.9B4.10 GiB31.88 GiB37.00 GiB0.20 GiB25±22%
Mistral-Small-4-119B-2603MoEIQ2_XS119B34.48 GiB1.49 GiB37.00 GiB0.20 GiB105±37%
Qwythos-9B-Claude-Mythos-5-1MBF169.4B33.83 GiB2.13 GiB36.99 GiB0.21 GiB25±22%
Qwythos-9B-v2BF169.7B33.83 GiB2.13 GiB36.99 GiB0.21 GiB25±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a A100 40GB run?
1924 of 2118 indexed open-weight models fit a A100 40GB at 131,072 context with q8_0 KV cache, the largest being spoomplesmaxx-v2.1-30B at I1-Q5_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a A100 40GB actually have?
Its nameplate is 40 GB, but about 37.20 GiB is available to a model once driver and compositor overhead is accounted for.
Is a A100 40GB fast for local AI?
Its memory bandwidth is 1555 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.