nvidia · text

NVIDIA-Nemotron-Nano-12B-v2

nvidia/NVIDIA-Nemotron-Nano-12B-v2

NVIDIA-Nemotron-Nano-12B-v2 at Q4_K_M is exactly 7,494,497,504 bytes (6.98 GiB / 7.49 GB) — an effective 4.870 bits per weight, not the nominal 4. Its KV cache at 32K is 7.75 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
12.3B
Architecture
nemotron_h
62 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_S3.79 GiB4,066,452,9922.643bartowski
IQ2_M4.08 GiB4,380,288,5122.847bartowski
Q2_K4.38 GiB4,703,360,2243.057MaziyarPanahi
Q2_K4.38 GiB4,703,360,5123.057bartowski
IQ3_XXS4.62 GiB4,960,282,1123.224bartowski
Q2_K_L4.99 GiB5,358,720,5123.482bartowski
IQ3_XS5.08 GiB5,455,775,2323.546bartowski
Q3_K_S5.19 GiB5,567,841,7923.618bartowski
IQ3_M5.30 GiB5,690,394,1123.698bartowski
Q3_K_M5.61 GiB6,023,480,5443.914MaziyarPanahi
Q3_K_M5.61 GiB6,023,480,8323.914bartowski
Q3_K_L5.94 GiB6,373,442,7844.142MaziyarPanahi
Q3_K_L5.94 GiB6,373,443,0724.142bartowski
IQ4_XS6.29 GiB6,749,681,1524.386bartowski
IQ4_NL6.62 GiB7,113,324,0324.623bartowski
Q4_06.67 GiB7,159,199,2324.653bartowski
Q4_K_S6.71 GiB7,207,695,8724.684bartowski
Q4_K_M6.98 GiB7,494,497,5044.870MaziyarPanahi
Q4_K_M6.98 GiB7,494,497,7924.870bartowski
Q4_17.30 GiB7,840,609,7925.095bartowski
Q4_K_L7.44 GiB7,992,571,3925.194bartowski
Q5_K_S7.98 GiB8,567,895,5525.568bartowski
Q5_K_M8.16 GiB8,764,257,5045.696MaziyarPanahi
Q5_K_M8.16 GiB8,764,257,7925.696bartowski
Q5_K_L8.55 GiB9,178,445,3125.965bartowski
Q6_K9.42 GiB10,113,377,5046.572MaziyarPanahi
Q6_K9.42 GiB10,113,377,7926.572bartowski
Q6_K_L9.72 GiB10,438,436,3526.784bartowski
Q8_012.19 GiB13,094,139,3928.510bartowski
BF1622.94 GiB24,632,571,10416.008bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.97 GiB0.97 GiB62 / 0 / 0
8,1921.94 GiB1.94 GiB62 / 0 / 0
16,3843.88 GiB3.88 GiB62 / 0 / 0
32,7687.75 GiB7.75 GiB62 / 0 / 0
65,53615.50 GiB15.50 GiB62 / 0 / 0
131,07231.00 GiB31.00 GiB62 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.45 GiB. The real file is 6.98 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
62
Attention heads
40
KV heads
8
Head dim
128
Hidden size
5120
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does NVIDIA-Nemotron-Nano-12B-v2 need?
Q4_K_M is exactly 7,494,497,504 bytes (6.98 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is NVIDIA-Nemotron-Nano-12B-v2's KV cache?
7.75 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of NVIDIA-Nemotron-Nano-12B-v2 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.