NVIDIA · workstation

RTX 4500 Ada Generation

RTX 4500 Ada Generation has 24 GB of VRAM at 432 GB/s — about 22.32 GiB usable after driver and compositor overhead. 1960 of 2118 indexed models fit at 16K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
24 GB
GDDR6
Bandwidth
432 GB/s
192-bit bus
Tensor FP16
159 TF
dense
TDP
210 W
$2250 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1683vision language 173image 2video 16audio tts 21audio asr 39embedding 26

What fits at 16K context

largest quantization that fits, per model · 1960 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Delphi-25B-SimpleRL-MathI1-Q5_K_M25.0B16.52 GiB4.71 GiB22.31 GiB0.01 GiB12±22%
Skyfall-31B-v4.2Q5_K_S31.4B20.23 GiB0.95 GiB22.30 GiB0.02 GiB12±22%
command-r-35b-writer-v2I1-IQ3_M35.0B15.55 GiB5.63 GiB22.28 GiB0.04 GiB12±22%
medgemma-27b-itI1-Q6_K28.8B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
gemma-3-27b-it-abliterated-refined-visionI1-Q6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
gemma-3-27b-it-abliteratedQ6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
Nidum-Gemma-3-27B-it-UncensoredI1-Q6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
gemma-3-27b-itQ6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
AtomicGPT-gemma3-27bI1-Q6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
Unbound-v1.12.0-27BI1-Q6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
Mira-v1.12-Ties-27BI1-Q6_K27.4B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
Medgamma27BI1-Q6_K27.0B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
medgemma-27b-text-itQ6_K27.0B20.64 GiB0.52 GiB22.25 GiB0.07 GiB12±22%
Kimi-Linear-48B-A3B-InstructMoEIQ3_M49.1B21.10 GiB0.13 GiB22.24 GiB0.08 GiB12±22%
Pantheon-Reasoning-27BI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-PreservedI1-Q6_K27.4B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-Fable-5-ExperimentalI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwable-5-27B-CoderI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16Q6_K27.4B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
EVE-27b-XENO-HAT-DeepSeek-V4-FlashI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
EVE-27B-XENO-HATI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Godoter-27BI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Reasoning-Medical-27BI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwopus3.6-27B-v2-abliteratedI1-Q6_K27.4B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-BF16I1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Reasoning-Medical0.1-27BI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Huihui-ThinkingCap-Qwen3.6-27B-abliteratedI1-Q6_K27.4B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Semancer-27BI1-Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-Uncensored-CyberQ6_K27.4B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwen3.6-27B-Omnimerge-v4Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwopus3.6-27B-v2Q6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Darwin-28B-CoderI1-Q6_K26.9B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Qwopus3.6-27B-CoderQ6_K27.8B20.89 GiB0.28 GiB22.23 GiB0.09 GiB12±22%
Skyfall-31B-v4.2-hereticI1-Q5_K_S31.4B20.17 GiB0.95 GiB22.23 GiB0.09 GiB12±22%
Maenad-70BI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
DeepSeek-R1-Distill-Llama-70B-Uncensored-v2-Unbiased-ReasonerI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Rombos-LLM-70b-Llama-3.3I1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
L3.3-Electra-R1-70bI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
L3.3-70B-Magnum-v4-SEIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Latxa-Llama-3.1-70B-Instruct-v2I1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Llama-3.3_70_b_uncensored_continuedI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Llama-3.3-70B-Instruct-abliteratedI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
grok-oss-Revenant-70BI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Llama-3.1-Nemotron-70B-Instruct-HFI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
L3.3-70B-Euryale-v2.3I1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Hermes-3-Llama-3.1-70BIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Hermes-4-70B-hereticI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Llama-3.3-70B-InstructIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Llama-3.1-70BIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Anubis-70B-v1.2IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Hermes-4-70BIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Golem-70B-v1bI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
DeepSeek-R1-Distill-Llama-70B-abliteratedI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
DeepSeek-R1-Distill-Llama-70B-hereticI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
llama-3-firefunction-v2IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
DeepSeek-R1-Distill-Llama-70BIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Legion-V2.1-LLaMa-70BI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Assistant_Pepe_70BI1-IQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
Tess-R1-Limerick-Llama-3.1-70BIQ2_XS70.6B19.69 GiB1.41 GiB22.22 GiB0.10 GiB12±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX 4500 Ada Generation run?
1960 of 2118 indexed open-weight models fit a RTX 4500 Ada Generation at 16,384 context with q4_0 KV cache, the largest being Delphi-25B-SimpleRL-Math at I1-Q5_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX 4500 Ada Generation actually have?
Its nameplate is 24 GB, but about 22.32 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX 4500 Ada Generation fast for local AI?
Its memory bandwidth is 432 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.