GPU comparison for local inference

GeForce RTX 5090 vs GeForce RTX 4080 Super

GeForce RTX 5090 holds more — 32 GB against 16 GB, which is what decides whether a model runs at all. GeForce RTX 5090 has 2.43× the memory bandwidth, which is what decides how fast it generates. Of the models indexed here, 266 fit only on GeForce RTX 5090 and 0 fit only on GeForce RTX 4080 Super.

Spec sheet· specsFrom the file· fitPredicted· speed

Specifications

GeForce RTX 5090GeForce RTX 4080 Super
Memory32 GB16 GB
Usable to GPU32 GB16 GB
Bandwidth1792 GB/s736 GB/s
Memory typeGDDR7GDDR6X
Bus width512-bit256-bit
Tensor FP16 (dense)419 TFLOPS209 TFLOPS
TDP575 W320 W
MSRP at launch$1999$999
Models that fit19541688
Spec sheetPredictedwhat these mean

Same model, both cards

modeled tokens per second at 32K context
ModelQuantParamsGeForce RTX 5090GeForce RTX 4080 SuperDifference
Qwen3-Coder-30B-A3B-InstructMoEQ6_K30.5B11969+72%
Qwen3.6-27BQ6_K27.8B5438+41%
Qwen3.6-35B-A3BMoEUD-Q6_K36.0B204152+34%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPIQ3_M27.8B4438+17%
Qwen3.8-27BQ6_K27.8B5438+41%
Qwen3.5-9BBF169.7B6951+34%
gemma-4-26B-A4B-itMoEQ8_026.5B4638+23%
gemma-4-12B-itBF1612.0B5143+20%
nemotron-3.5-asr-streaming-0.6bF32638M393195+102%
Qwen3.5-4BBF164.7B13157+129%
gemma-4-12B-it-qat-q4_0-unquantizedQ4_012.0B13258+129%
Qwythos-9B-Claude-Mythos-5-1MQ8_09.4B6640+67%

Token rates are modeled from memory bandwidth, so on models both cards can hold the ratio tracks bandwidth closely. That is the honest shape of the answer: for inference, capacity decides what you can run and bandwidth decides how fast. Neither is teraflops.

How to read this

We earn nothing from either of these cards. If both hold the models you care about, the faster one wins on bandwidth alone. If one holds a model the other cannot, that difference usually matters far more than any percentage of token rate — a model that does not fit runs 5–20× slower once it starts spilling to system memory, not slightly slower.