NVIDIA · workstation

RTX PRO 5000 Blackwell

RTX PRO 5000 Blackwell has 72 GB of VRAM at 1344 GB/s — about 66.96 GiB usable after driver and compositor overhead. 2061 of 2118 indexed models fit at 128K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
72 GB
GDDR7
Bandwidth
1344 GB/s
384-bit bus
Tensor FP16
295 TF
dense
TDP
300 W
$4569 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1771vision language 186image 2audio asr 39audio tts 21video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 2061 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Mistral-Small-4-119B-2603MoEQ4_K_S119B65.08 GiB0.79 GiB66.91 GiB0.05 GiB64±37%
c4ai-command-r-08-2024F1632.3B60.17 GiB5.63 GiB66.91 GiB0.05 GiB12±22%
GLM-4.5-Air-DerestrictedMoEQ4_0110B59.38 GiB6.47 GiB66.88 GiB0.08 GiB32±37%
GLM-4.5-AirMoEQ4_0110B59.38 GiB6.47 GiB66.88 GiB0.08 GiB32±37%
Qwen3.5-REAP-212B-A17BMoEIQ2_M212B64.76 GiB1.05 GiB66.87 GiB0.09 GiB59±37%
command-r-35b-writer-v2I1-Q4_135.0B20.75 GiB45.00 GiB66.86 GiB0.10 GiB12±22%
Gemma-4-Novelist-Eclipse-31BBF1632.7B59.82 GiB5.95 GiB66.86 GiB0.10 GiB12±22%
Gemma-4-31B-StyleTuneBF1632.7B59.82 GiB5.95 GiB66.86 GiB0.10 GiB12±22%
Llama-3.3-70B-InstructQ6_K_L70.6B54.39 GiB11.25 GiB66.76 GiB0.20 GiB12±22%
Anubis-70B-v1.2Q6_K_L70.6B54.39 GiB11.25 GiB66.76 GiB0.20 GiB12±22%
Tess-R1-Limerick-Llama-3.1-70BQ6_K_L70.6B54.39 GiB11.25 GiB66.76 GiB0.20 GiB12±22%
Infinity-Instruct-7M-Gen-Llama3_1-70BQ6_K_L70.6B54.39 GiB11.25 GiB66.76 GiB0.20 GiB12±22%
Athene-70BQ6_K_L70.6B54.39 GiB11.25 GiB66.76 GiB0.20 GiB12±22%
Qwen3.5-122B-A10B-hereticMoEI1-Q4_K_S123B64.86 GiB0.84 GiB66.73 GiB0.23 GiB64±37%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEIQ4_XS124B62.59 GiB3.09 GiB66.68 GiB0.28 GiB46±37%
Laguna-S-2.1MoEUD-Q4_K_S118B63.88 GiB1.73 GiB66.63 GiB0.33 GiB54±37%
Qwen3.5-REAP-262B-A17BMoEIQ2_XXS262B64.50 GiB1.05 GiB66.61 GiB0.35 GiB63±37%
Mistral-Medium-3.5-128BQ3_K_S128B53.05 GiB12.38 GiB66.58 GiB0.38 GiB12±22%
MiMo-V2-FlashMoEKV unresolvedIQ1_M310B61.31 GiB4.22 GiB66.57 GiB0.39 GiB44±37%
step-3.5-flashIQ2_S199B51.70 GiB13.79 GiB66.52 GiB0.44 GiB12±22%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ3_M139B56.81 GiB8.72 GiB66.52 GiB0.44 GiB29±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ3_M139B56.81 GiB8.72 GiB66.52 GiB0.44 GiB29±37%
Phi-3-mini-4k-instructKV unresolvedF323.8B52.01 GiB13.50 GiB66.51 GiB0.45 GiB12±22%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedQ4_0109B58.72 GiB6.75 GiB66.50 GiB0.46 GiB32±37%
Devstral-2-123B-Instruct-2512IQ3_M125B52.89 GiB12.38 GiB66.43 GiB0.53 GiB12±22%
XORTRON-NXTXPRTXXLI1-IQ3_M128B52.89 GiB12.38 GiB66.43 GiB0.53 GiB12±22%
Apertus-70B-Instruct-2509Q6_K70.6B53.95 GiB11.25 GiB66.38 GiB0.58 GiB12±22%
MiniMax-M2.7MoEIQ2_XXS229B56.67 GiB8.72 GiB66.37 GiB0.59 GiB31±37%
Wizard-Vicuna-30B-UncensoredI1-IQ2_M32.5B10.43 GiB54.84 GiB66.34 GiB0.62 GiB12±22%
archangel_sft-kto_llama30bI1-IQ2_M32.5B10.43 GiB54.84 GiB66.34 GiB0.62 GiB12±22%
Qwen3.6-35B-A3B-uncensored-hereticMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Darwin-35B-A3B-OpusMoEBF1636.0B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Carnice-MoE-35B-A3BMoEF1636.0B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.6-35B-A3B-hereticMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Aurora-Code-1MoEBF1634.7B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
grug-35b-v2MoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
grug-35bMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
WorldSim-Opus-3.6-35B-A3BMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.6-35B-A3B-abliterated-MAXMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Huihui-Nex-N2-mini-abliteratedMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.6-35B-A3B-AnkoMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
KAT-Coder-V2.5-DevMoEBF1634.7B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Ornith-1.0-35B-uncensored-hereticMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.6-35B-A3B-abliterated-v4MoEBF1634.7B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.5-35B-A3B-ultra-uncensored-hereticMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Nex-N2-mini-ultra-uncensored-hereticMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Agents-A1MoEF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.6-35B-A3B-java-v1MoEBF1634.7B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Nex-N2-miniMoEBF1635.1B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Qwen3.5-35B-A3B-BaseMoEBF1636.0B64.61 GiB0.70 GiB66.32 GiB0.64 GiB66±37%
Meta-Llama-3-70B-InstructQ6_K70.6B53.92 GiB11.25 GiB66.30 GiB0.66 GiB12±22%
Qwen3-VL-235B-A22B-ThinkingMoEUD-IQ1_S236B58.65 GiB6.61 GiB66.29 GiB0.67 GiB32±37%
calme-2.4-llama3-70bQ6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
calme-2.2-llama3-70bQ6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
L3.3-70B-Magnum-v4-SEQ6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
L3.3-Electra-R1-70bQ6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
Latxa-Llama-3.1-70B-Instruct-v2I1-Q6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
Llama-3.3_70_b_uncensored_continuedI1-Q6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
grok-oss-Revenant-70BI1-Q6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
Hermes-4-70BQ6_K70.6B53.91 GiB11.25 GiB66.29 GiB0.67 GiB12±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX PRO 5000 Blackwell run?
2061 of 2118 indexed open-weight models fit a RTX PRO 5000 Blackwell at 131,072 context with q4_0 KV cache, the largest being Mistral-Small-4-119B-2603 at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX PRO 5000 Blackwell actually have?
Its nameplate is 72 GB, but about 66.96 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX PRO 5000 Blackwell fast for local AI?
Its memory bandwidth is 1344 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.