FreedomIntelligence · text

ShizhenGPT-32B-VL

FreedomIntelligence/ShizhenGPT-32B-VL

ShizhenGPT-32B-VL at Q4_K_M is exactly 19,851,337,408 bytes (18.49 GiB / 19.85 GB) — an effective 4.747 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
33.5B
Architecture
qwen2vl
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S6.77 GiB7,274,508,2241.740mradermacher
I1-IQ1_M7.39 GiB7,932,161,9841.897mradermacher
I1-IQ2_XXS8.41 GiB9,028,251,5842.159mradermacher
I1-IQ2_XS9.27 GiB9,957,552,0642.381mradermacher
I1-IQ2_S9.67 GiB10,387,570,6242.484mradermacher
I1-IQ2_M10.49 GiB11,264,442,3042.694mradermacher
I1-Q2_K_S10.70 GiB11,488,001,9842.747mradermacher
Q2_K11.47 GiB12,313,099,9682.945mradermacher
I1-Q2_K11.47 GiB12,313,100,2242.945mradermacher
I1-IQ3_XXS11.96 GiB12,839,272,3843.070mradermacher
I1-IQ3_XS12.76 GiB13,705,514,9443.278mradermacher
Q3_K_S13.40 GiB14,392,331,9683.442mradermacher
I1-Q3_K_S13.40 GiB14,392,332,2243.442mradermacher
I1-IQ3_S13.45 GiB14,436,896,7043.453mradermacher
I1-IQ3_M13.79 GiB14,810,124,2243.542mradermacher
Q3_K_M14.84 GiB15,935,049,4083.811mradermacher
I1-Q3_K_M14.84 GiB15,935,049,6643.811mradermacher
Q3_K_L16.06 GiB17,247,080,1284.125mradermacher
I1-Q3_K_L16.06 GiB17,247,080,3844.125mradermacher
I1-IQ4_XS16.48 GiB17,693,155,2644.231mradermacher
IQ4_XS16.64 GiB17,870,102,2084.274mradermacher
I1-Q4_017.43 GiB18,711,011,2644.475mradermacher
Q4_K_S17.49 GiB18,784,411,3284.492mradermacher
I1-Q4_K_S17.49 GiB18,784,411,5844.492mradermacher
Q4_K_M18.49 GiB19,851,337,4084.747mradermacher
I1-Q4_K_M18.49 GiB19,851,337,6644.747mradermacher
I1-Q4_119.22 GiB20,639,244,2244.936mradermacher
Q5_K_S21.08 GiB22,638,255,8085.414mradermacher
I1-Q5_K_S21.08 GiB22,638,256,0645.414mradermacher
Q5_K_M21.66 GiB23,262,158,5285.563mradermacher
I1-Q5_K_M21.66 GiB23,262,158,7845.563mradermacher
Q6_K25.04 GiB26,886,155,9686.430mradermacher
I1-Q6_K25.04 GiB26,886,156,2246.430mradermacher
Q8_032.43 GiB34,820,886,2088.327mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 17.52 GiB. The real file is 18.49 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does ShizhenGPT-32B-VL need?
Q4_K_M is exactly 19,851,337,408 bytes (18.49 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of ShizhenGPT-32B-VL should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.