NVIDIA · workstation

RTX PRO 6000 Blackwell Workstation Edition

RTX PRO 6000 Blackwell Workstation Edition has 96 GB of VRAM at 1792 GB/s — about 89.28 GiB usable after driver and compositor overhead. 2069 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
96 GB
GDDR7
Bandwidth
1792 GB/s
512-bit bus
Tensor FP16
504 TF
dense
TDP
600 W
$8565 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 189text 1776audio tts 21image 2audio asr 39video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 2069 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
GLM-4.6VMoEQ5_K_L108B75.96 GiB12.22 GiB89.21 GiB0.07 GiB27±37%
Behemoth-X-123B-v2Q4_0123B64.56 GiB23.38 GiB89.09 GiB0.19 GiB12±22%
Mistral-Large-Instruct-2411Q4_0123B64.56 GiB23.38 GiB89.09 GiB0.19 GiB12±22%
Hunyuan-A13B-InstructMoEQ8_080.4B79.58 GiB8.50 GiB89.08 GiB0.20 GiB12±22%
Mistral-Medium-3.5-128BIQ4_XS128B64.39 GiB23.38 GiB88.93 GiB0.35 GiB12±22%
Noromaid-20b-v0.1.1I1-IQ2_XS20.0B5.53 GiB82.34 GiB88.92 GiB0.36 GiB12±22%
MiMo-V2.5MoEKV unresolvedIQ2_XXS311B79.83 GiB7.97 GiB88.85 GiB0.43 GiB38±37%
Delphi-25B-SimpleRL-MathI1-Q5_K_M25.0B16.52 GiB71.12 GiB88.72 GiB0.56 GiB12±22%
MiniMax-M2.7MoEIQ2_M229B71.17 GiB16.47 GiB88.62 GiB0.66 GiB25±37%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEQ5_K_S124B81.55 GiB5.84 GiB88.39 GiB0.89 GiB41±37%
MiMo-V2-FlashMoEKV unresolvedIQ2_S310B79.34 GiB7.97 GiB88.35 GiB0.93 GiB38±37%
GLM-4.5VMoEI1-Q5_K_M108B75.10 GiB12.22 GiB88.35 GiB0.93 GiB27±37%
CalmeRys-78B-Orpo-v0.1Q6_K78.0B64.27 GiB22.84 GiB88.25 GiB1.03 GiB12±22%
calme-2.3-rys-78bQ6_K78.0B64.27 GiB22.84 GiB88.25 GiB1.03 GiB12±22%
Ornith-1.0-397BMoEIQ1_M397B85.09 GiB1.99 GiB88.13 GiB1.15 GiB65±37%
GLM-4.7-REAP-218B-A32BMoEUD-IQ1_M218B62.63 GiB24.44 GiB88.11 GiB1.17 GiB18±37%
Qwen3.5-122B-A10BMoEQ5_K_M125B85.23 GiB1.59 GiB87.86 GiB1.42 GiB62±37%
internlm2-math-plus-20bF3219.9B73.99 GiB12.75 GiB87.80 GiB1.48 GiB12±22%
Qwen3-235B-A22B-abliteratedMoEI1-Q2_K_S235B74.28 GiB12.48 GiB87.80 GiB1.48 GiB27±37%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedQ5_K_L109B73.87 GiB12.75 GiB87.65 GiB1.63 GiB27±37%
XORTRON-NXTXPRTXXLIQ4_XS128B63.03 GiB23.38 GiB87.56 GiB1.72 GiB12±22%
Phi-3.5-MoE-instructMoEKV unresolvedF1641.9B78.00 GiB8.50 GiB87.51 GiB1.77 GiB26±37%
Trinity-Large-TrueBaseMoEI1-IQ1_M399B82.07 GiB4.40 GiB87.50 GiB1.78 GiB54±37%
Step-3.5-Flash-REAP-121B-A11BI1-IQ4_XS121B60.12 GiB26.05 GiB87.20 GiB2.08 GiB12±22%
MiniMax-M2.7-BF16-ultra-uncensored-hereticMoEI1-IQ2_M229B69.70 GiB16.47 GiB87.15 GiB2.13 GiB26±37%
MiniMax-M2.1MoEI1-IQ2_M229B69.70 GiB16.47 GiB87.15 GiB2.13 GiB26±37%
MiniMax-M2.5MoEI1-IQ2_M229B69.70 GiB16.47 GiB87.15 GiB2.13 GiB26±37%
Mixtral-8x22B-Instruct-v0.1MoEIQ4_XS141B71.12 GiB14.88 GiB87.05 GiB2.23 GiB16±37%
Devstral-2-123B-Instruct-2512IQ4_XS125B62.51 GiB23.38 GiB87.05 GiB2.23 GiB12±22%
Mixtral-8x22B-v0.1MoEIQ4_XS141B71.11 GiB14.88 GiB87.05 GiB2.23 GiB16±37%
Mixtral-8x22B-v0.1MoEIQ4_XS141B71.11 GiB14.88 GiB87.05 GiB2.23 GiB16±37%
DeepSeek-Coder-V2-InstructMoEQ2_K_L236B81.44 GiB4.48 GiB86.96 GiB2.32 GiB47±37%
ERNIE-4.5-300B-A47B-PTUD-TQ1_0300B71.48 GiB14.34 GiB86.95 GiB2.33 GiB12±22%
GPT-NeoX-20B-ErebusI1-Q6_K20.6B15.72 GiB70.13 GiB86.93 GiB2.35 GiB12±22%
c4ai-command-r-plus-08-2024Q5_K_M104B68.57 GiB17.00 GiB86.75 GiB2.53 GiB12±22%
Qwen3-235B-A22B-Thinking-2507MoEIQ2_M235B73.16 GiB12.48 GiB86.68 GiB2.60 GiB27±37%
Qwen3-235B-A22B-Instruct-2507MoEIQ2_M235B73.16 GiB12.48 GiB86.68 GiB2.60 GiB27±37%
step-3.5-flashIQ2_M199B59.59 GiB26.05 GiB86.66 GiB2.62 GiB12±22%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ4_XS139B69.15 GiB16.47 GiB86.60 GiB2.68 GiB24±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ4_XS139B69.15 GiB16.47 GiB86.60 GiB2.68 GiB24±37%
GLM-4.5-Air-DerestrictedMoEQ5_K_S110B73.16 GiB12.22 GiB86.40 GiB2.88 GiB28±37%
GLM-4.5-AirMoEQ5_K_S110B73.16 GiB12.22 GiB86.40 GiB2.88 GiB28±37%
MiniMax-M2MoEUD-IQ2_XXS229B68.92 GiB16.47 GiB86.37 GiB2.91 GiB26±37%
Qwen3-Coder-NextMoEQ8_079.7B78.99 GiB6.38 GiB86.35 GiB2.93 GiB46±37%
Qwen3-Next-80B-A3B-ThinkingMoEQ8_081.3B78.99 GiB6.38 GiB86.35 GiB2.93 GiB46±37%
Qwen3-Next-80B-A3B-InstructMoEQ8_081.3B78.99 GiB6.38 GiB86.35 GiB2.93 GiB46±37%
Ace-Step1.5BF16160M82.03 GiB3.21 GiB86.23 GiB3.05 GiB12±22%
Laguna-S-2.1MoEUD-Q5_K_M118B81.83 GiB3.26 GiB86.11 GiB3.17 GiB51±37%
DeepSeek-Coder-V2-Instruct-0724MoEQ2_K_L236B80.52 GiB4.48 GiB86.04 GiB3.24 GiB47±37%
DeepSeek-V2.5MoEQ2_K_L236B80.52 GiB4.48 GiB86.04 GiB3.24 GiB47±37%
GLM-4.6-REAP-268B-A32BMoEUD-TQ1_0269B60.36 GiB24.44 GiB85.84 GiB3.44 GiB18±37%
DeepSeek-V4-Flash-0731MoEUD-IQ2_M304B84.68 GiB0.03 GiB85.76 GiB3.52 GiB80±37%
DeepSeek-V4-FlashMoEUD-IQ2_M291B84.68 GiB0.03 GiB85.76 GiB3.52 GiB80±37%
Qwen3.5-397B-A17BMoEIQ1_S403B82.64 GiB1.99 GiB85.68 GiB3.60 GiB67±37%
Mistral-Small-4-119B-2603MoEUD-Q5_K_M119B83.04 GiB1.49 GiB85.56 GiB3.72 GiB64±37%
Seed-OSS-36B-InstructBF1636.2B67.35 GiB17.00 GiB85.44 GiB3.84 GiB12±22%
Hermes-4.3-36BBF1636.2B67.35 GiB17.00 GiB85.44 GiB3.84 GiB12±22%
gpt-oss-120b-Uncensored-xCloudMoEI1-Q5_K_S117B81.92 GiB2.40 GiB85.32 GiB3.96 GiB59±37%
gpt-oss-120b-abliteratedMoEI1-Q5_K_S117B81.92 GiB2.40 GiB85.32 GiB3.96 GiB59±37%
OYM-Qimi-122B-A10B-K2.6MoEI1-Q5_K_M125B82.62 GiB1.59 GiB85.24 GiB4.04 GiB63±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing14316.57 tok/s9546.2516645.0038
Text generation267.03 tok/s256.42275.7724
Benchmarked· n=38

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-15013.

Questions people ask

What AI models can a RTX PRO 6000 Blackwell Workstation Edition run?
2069 of 2118 indexed open-weight models fit a RTX PRO 6000 Blackwell Workstation Edition at 131,072 context with q8_0 KV cache, the largest being GLM-4.6V at Q5_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX PRO 6000 Blackwell Workstation Edition actually have?
Its nameplate is 96 GB, but about 89.28 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX PRO 6000 Blackwell Workstation Edition fast for local AI?
Its memory bandwidth is 1792 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.