NVIDIA · workstation

RTX 4000 SFF Ada Generation

RTX 4000 SFF Ada Generation has 20 GB of VRAM at 280 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1594 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6
Bandwidth
280 GB/s
160-bit bus
Tensor FP16
77 TF
dense
TDP
70 W
$1250 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1335vision language 156video 16audio asr 39embedding 26image 1audio tts 21

What fits at 128K context

largest quantization that fits, per model · 1594 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nexa-AI-4x4B-InstructMoEI1-Q5_K_M12.1B8.03 GiB9.56 GiB18.60 GiB0.00 GiB7±37%
Huihui-Qwen3-Coder-Next-abliteratedMoEIQ1_M79.7B16.01 GiB1.59 GiB18.60 GiB0.00 GiB32±37%
Aurora-Code-1MoEI1-Q4_K_S34.7B16.26 GiB1.33 GiB18.59 GiB0.01 GiB32±37%
Goetia-26B-A4B-v1.4MoEI1-Q4_K_S26.0B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
G4-Moonlight-Dusk-26B-A4BMoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
Chimera-X-26B-A4BMoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
Gemma-4-26B-A4B-StyleTuneMoEI1-Q4_K_S26.5B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-Q4_K_S25.8B14.79 GiB2.81 GiB18.59 GiB0.01 GiB9±22%
DeepSeek-V2-Lite-ChatMoEQ8_015.7B15.56 GiB2.02 GiB18.58 GiB0.02 GiB21±37%
DeepSeek-Coder-V2-Lite-InstructMoEQ8_015.7B15.56 GiB2.02 GiB18.58 GiB0.02 GiB21±37%
DeepSeek-Coder-V2-Lite-BaseMoEQ8_015.7B15.56 GiB2.02 GiB18.58 GiB0.02 GiB21±37%
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-IQ4_XS13.9B6.90 GiB10.63 GiB18.58 GiB0.02 GiB9±22%
Ministral-3-14B-Instruct-2512-BF16IQ4_XS13.9B6.90 GiB10.63 GiB18.58 GiB0.02 GiB9±22%
Ministral-3-14B-Reasoning-2512-UncensoredI1-IQ4_XS13.9B6.90 GiB10.63 GiB18.58 GiB0.02 GiB9±22%
granite-8b-code-instruct-4kQ8_08.1B7.98 GiB9.56 GiB18.58 GiB0.02 GiB9±22%
granite-8b-code-base-4kQ8_08.1B7.98 GiB9.56 GiB18.58 GiB0.02 GiB9±22%
DeepSeek-V2-Lite-Chat-Uncensored-Unbiased-ReasonerMoEQ8_015.7B15.55 GiB2.02 GiB18.57 GiB0.03 GiB21±37%
DeepSeek-V2-Lite-Chat-UncensoredMoEQ8_015.7B15.55 GiB2.02 GiB18.57 GiB0.03 GiB21±37%
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16UD-IQ3_S33.0B17.53 GiB0.00 GiB18.57 GiB0.03 GiB9±22%
Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-IQ2_XS39.5B11.13 GiB6.38 GiB18.57 GiB0.03 GiB9±22%
Skywork-R1V3-38BQ4_K_S38.4B17.49 GiB0.00 GiB18.56 GiB0.04 GiB9±22%
gemma-4-26B-A4B-itMoEQ4_K_S26.5B14.76 GiB2.81 GiB18.56 GiB0.04 GiB9±22%
Qwen3-Coder-REAP-25B-A3BMoEQ3_K_M24.9B11.18 GiB6.38 GiB18.55 GiB0.05 GiB12±37%
GLM-4-32B-0414-Korean-CultureI1-IQ3_S32.6B13.40 GiB4.05 GiB18.54 GiB0.06 GiB9±22%
Falcon3-10B-InstructQ5_K_M10.3B6.84 GiB10.63 GiB18.53 GiB0.07 GiB9±22%
GLM-Z1-32B-0414Q3_K_S32.6B13.38 GiB4.05 GiB18.52 GiB0.08 GiB9±22%
GLM-4-32B-0414Q3_K_S32.6B13.38 GiB4.05 GiB18.52 GiB0.08 GiB9±22%
phi-4IQ2_XS14.7B4.18 GiB13.28 GiB18.52 GiB0.08 GiB9±22%
Salience-1.5-ProMoEIQ3_M36.0B16.18 GiB1.33 GiB18.51 GiB0.09 GiB33±37%
Qwable-v1MoEIQ3_M36.0B16.18 GiB1.33 GiB18.51 GiB0.09 GiB33±37%
T-SearchMoEIQ3_M36.0B16.18 GiB1.33 GiB18.51 GiB0.09 GiB33±37%
Marco-Mini-InstructMoEI1-Q4_117.3B10.10 GiB7.44 GiB18.51 GiB0.09 GiB11±37%
gpt-oss-20b-hereticMoEQ5_K_L20.9B15.91 GiB1.60 GiB18.50 GiB0.10 GiB21±37%
NousCoder-14BQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
spoomplesmaxx-mini-14BI1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
vanilla-cn-roleplay-0.2I1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Claria-14bI1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
qwen3-14b-code-reasoning-conversationalQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
NTX-2.1-ProI1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Qwen3-14B-UncensoredI1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Qwen3-14B-Claude-4.5-Opus-High-Reasoning-DistillQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Qwen3-14BQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
FrogMini-14B-2510I1-Q3_K_M6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Qwen3-14B-abliteratedQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Josiefied-Qwen3-14B-abliterated-v3Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Hermes-4-14BQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Slava-Qwen3-14B-SerbianI1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Qwen3-14B-BaseQ3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Huihui-Qwen3-14B-abliterated-v2I1-Q3_K_M14.8B6.82 GiB10.63 GiB18.50 GiB0.10 GiB9±22%
Muse-Glimmer-30BQ4_029.8B16.50 GiB0.91 GiB18.50 GiB0.10 GiB9±22%
Qwen3.6-27B-Fable-5-ExperimentalQ3_K_M27.8B13.18 GiB4.25 GiB18.49 GiB0.11 GiB9±22%
Ministral-3-8B-Instruct-2512Q8_08.9B8.41 GiB9.03 GiB18.48 GiB0.12 GiB9±22%
Ministral-3-8B-Reasoning-2512Q8_08.9B8.41 GiB9.03 GiB18.48 GiB0.12 GiB9±22%
Ministral-3-8B-Instruct-2512-BF16-abliteratedQ8_08.9B8.41 GiB9.03 GiB18.48 GiB0.12 GiB9±22%
Ministral-3-8B-Instruct-2512-BF16Q8_08.9B8.41 GiB9.03 GiB18.48 GiB0.12 GiB9±22%
Amaretto-8BQ8_08.9B8.41 GiB9.03 GiB18.48 GiB0.12 GiB9±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation10.30 it/s7.6410.699
Benchmarked· n=9

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a RTX 4000 SFF Ada Generation run?
1594 of 2118 indexed open-weight models fit a RTX 4000 SFF Ada Generation at 131,072 context with q8_0 KV cache, the largest being Nexa-AI-4x4B-Instruct at I1-Q5_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX 4000 SFF Ada Generation actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX 4000 SFF Ada Generation fast for local AI?
Its memory bandwidth is 280 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.