upstage · text

solar-pro-preview-instruct

upstage/solar-pro-preview-instruct

solar-pro-preview-instruct at Q4_K_M is exactly 13,310,219,072 bytes (12.40 GiB / 13.31 GB) — an effective 4.809 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
22.1B
Architecture
llama
64 layers
Context
4,096
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S4.46 GiB4,786,663,2321.730MaziyarPanahi
IQ1_M4.87 GiB5,231,488,8321.890MaziyarPanahi
IQ2_XS6.16 GiB6,618,394,4322.392MaziyarPanahi
Q2_K7.65 GiB8,213,924,6722.968MaziyarPanahi
IQ3_XS8.50 GiB9,125,197,6323.297MaziyarPanahi
Q3_K_S8.92 GiB9,580,672,8323.462MaziyarPanahi
Q3_K_M9.95 GiB10,686,592,8323.861579MaziyarPanahi
Q3_K_L10.84 GiB11,635,226,4324.204MaziyarPanahi
IQ4_XS11.06 GiB11,878,032,1924.292579MaziyarPanahi
Q4_K_S11.73 GiB12,594,238,2724.551MaziyarPanahi
Q4_K_M12.40 GiB13,310,219,0724.809579MaziyarPanahi
Q5_K_S14.20 GiB15,246,070,5925.509MaziyarPanahi
Q5_K_M14.59 GiB15,663,862,5925.660579MaziyarPanahi
Q6_K16.92 GiB18,164,608,8326.564579MaziyarPanahi
Q8_021.91 GiB23,526,487,8728.501579MaziyarPanahi

KV cache by context

unresolved

This model declares a 2,047-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.

A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 11.60 GiB. The real file is 12.40 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
64
Attention heads
40
KV heads
10
Head dim
128
Hidden size
5120
Vocab
32,128
Sliding window
2047
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does solar-pro-preview-instruct need?
Q4_K_M is exactly 13,310,219,072 bytes (12.40 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of solar-pro-preview-instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.