Model comparison

Qwen3-Embedding-8B vs Nemotron-3-Embed-8B-BF16

At Q4_K_M, Qwen3-Embedding-8B is the smaller download — 4,676,804,928 bytes against 4,896,389,984.

From the file· summed bytes, KV per layer

Side by side

Qwen3-Embedding-8BNemotron-3-Embed-8B-BF16
Parameters7.6B8.0B
Architectureqwen3mistral3
Layers3634
Native context40,960262,144
Mixture of expertsnono
Quantizations published1736
Smallest quantization2.87 GiB1.81 GiB
Q4_K_M4.36 GiB4.56 GiB
Licenceapache-2.0openmdw-1.1

KV cache by context

the term that decides long-context viability
ContextQwen3-Embedding-8BNemotron-3-Embed-8B-BF16Ratio
4,0960.56 GiB0.53 GiB1.06×
8,1921.13 GiB1.06 GiB1.06×
16,3842.25 GiB2.13 GiB1.06×
32,7684.50 GiB4.25 GiB1.06×
65,5369.00 GiB8.50 GiB1.06×
131,07218.00 GiB17.00 GiB1.06×