nvidia · text

Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 at Q4_K_M is exactly 24,515,129,536 bytes (22.83 GiB / 24.52 GB) — an effective 5.940 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
33.0B
Architecture
nemotron_h_moe
null layers
Context
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
UD-IQ2_M17.23 GiB18,499,970,4964.483unsloth
UD-IQ2_XXS17.23 GiB18,499,970,4964.483unsloth
UD-IQ3_S17.53 GiB18,820,767,1684.560unsloth
UD-IQ3_XXS18.12 GiB19,459,349,9524.715unsloth
UD-IQ4_XS18.19 GiB19,536,678,3364.734unsloth
UD-IQ4_NL18.19 GiB19,536,678,3364.734unsloth
UD-Q3_K_M18.19 GiB19,536,678,3364.734unsloth
UD-Q4_K_S21.47 GiB23,048,883,6485.585unsloth
UD-Q4_K_M22.25 GiB23,887,023,5525.788unsloth
Q4_K_M22.83 GiB24,515,129,5365.940401lmstudio-community
UD-Q5_K_S23.10 GiB24,804,986,3046.011unsloth
UD-Q5_K_M27.00 GiB28,995,685,8247.026unsloth
Q6_K31.21 GiB33,508,166,8488.119401lmstudio-community
Q8_031.28 GiB33,585,495,2328.138401lmstudio-community
UD-Q6_K31.28 GiB33,585,499,5848.138unsloth
Q8_031.28 GiB33,585,499,5848.138401unsloth
F1658.84 GiB63,181,504,83215.309401mudler

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 17.30 GiB. The real file is 22.83 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
KV heads
Head dim
Hidden size
Vocab
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 need?
Q4_K_M is exactly 24,515,129,536 bytes (22.83 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.