darkc0de · vision language

XORTRON-NXTXPRTXXL

darkc0de/XORTRON-NXTXPRTXXL

XORTRON-NXTXPRTXXL at Q4_K_M is exactly 74,897,136,096 bytes (69.75 GiB / 74.90 GB) — an effective 4.692 bits per weight, not the nominal 4. Its KV cache at 32K is 11.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
128B
Architecture
llama
88 layers
Context
262,144
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S25.33 GiB27,193,744,0961.704mradermacher
I1-IQ1_M27.59 GiB29,620,280,0321.856mradermacher
I1-IQ2_XXS31.35 GiB33,664,506,5922.109mradermacher
I1-IQ2_XS34.75 GiB37,315,123,9362.338mradermacher
I1-IQ2_S37.01 GiB39,740,873,4402.490mradermacher
I1-IQ2_M40.02 GiB42,976,254,6882.692mradermacher
I1-Q2_K_S40.05 GiB43,000,634,0802.694mradermacher
Q2_K43.39 GiB46,590,695,9042.919mradermacher
I1-Q2_K43.39 GiB46,590,696,1602.919mradermacher
I1-IQ3_XXS45.04 GiB48,365,673,1843.030mradermacher
I1-IQ3_XS48.11 GiB51,659,250,4003.236mradermacher
Q3_K_S50.63 GiB54,366,935,5203.406mradermacher
I1-Q3_K_S50.63 GiB54,366,935,7763.406mradermacher
I1-IQ3_S50.77 GiB54,513,998,5603.415mradermacher
I1-IQ3_M52.89 GiB56,793,471,7123.558mradermacher
Q3_K_M56.46 GiB60,619,856,3523.797mradermacher
I1-Q3_K_M56.46 GiB60,619,856,6083.797mradermacher
Q3_K_L61.53 GiB66,071,402,9764.139mradermacher
I1-Q3_K_L61.53 GiB66,071,403,2324.139mradermacher
I1-IQ4_XS62.47 GiB67,074,104,0324.202mradermacher
IQ4_XS63.03 GiB67,679,656,4164.240mradermacher
I1-Q4_066.12 GiB70,999,972,5764.448mradermacher
Q4_K_S66.36 GiB71,248,484,8324.463mradermacher
I1-Q4_K_S66.36 GiB71,248,485,0884.463mradermacher
Q4_K_M69.75 GiB74,897,136,0964.692mradermacher
I1-Q4_K_M69.75 GiB74,897,136,3524.692mradermacher
I1-Q4_173.08 GiB78,471,076,5764.916mradermacher
Q5_K_S80.27 GiB86,184,401,3765.399mradermacher
I1-Q5_K_S80.27 GiB86,184,401,6325.399mradermacher
Q5_K_M82.25 GiB88,316,811,7445.533mradermacher
I1-Q5_K_M82.25 GiB88,316,812,0005.533mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.38 GiB1.38 GiB88 / 0 / 0
8,1922.75 GiB2.75 GiB88 / 0 / 0
16,3845.50 GiB5.50 GiB88 / 0 / 0
32,76811.00 GiB11.00 GiB88 / 0 / 0
65,53622.00 GiB22.00 GiB88 / 0 / 0
131,07244.00 GiB44.00 GiB88 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 66.90 GiB. The real file is 69.75 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
88
Attention heads
96
KV heads
8
Head dim
128
Hidden size
12288
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does XORTRON-NXTXPRTXXL need?
Q4_K_M is exactly 74,897,136,096 bytes (69.75 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is XORTRON-NXTXPRTXXL's KV cache?
11.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of XORTRON-NXTXPRTXXL should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.