prithivMLmods · text

Qwen2.5-VL-7B-Abliterated-Caption-it

prithivMLmods/Qwen2.5-VL-7B-Abliterated-Caption-it

Qwen2.5-VL-7B-Abliterated-Caption-it at Q4_K_M is exactly 1,929,902,080 bytes (1.80 GiB / 1.93 GB) — an effective 1.862 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
8.3B
Architecture
qwen2vl
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K1.19 GiB1,274,755,0721.230prithivMLmods
Q3_K_S1.35 GiB1,454,356,4801.403prithivMLmods
Q3_K_M1.48 GiB1,590,474,7521.534prithivMLmods
Q3_K_L1.59 GiB1,707,390,9761.647prithivMLmods
IQ4_XS1.63 GiB1,753,184,2561.691prithivMLmods
Q4_K_S1.71 GiB1,834,383,3601.770prithivMLmods
I1-IQ1_S1.77 GiB1,903,667,6481.837mradermacher
Q4_K_M1.80 GiB1,929,902,0801.862prithivMLmods
I1-IQ1_M1.90 GiB2,042,196,4161.970mradermacher
Q5_K_S2.02 GiB2,169,665,5362.093prithivMLmods
Q5_K_M2.07 GiB2,224,814,0802.146prithivMLmods
I1-IQ2_XXS2.12 GiB2,273,077,6962.193mradermacher
I1-IQ2_XS2.30 GiB2,469,022,1442.382mradermacher
Q6_K2.36 GiB2,538,158,0802.449prithivMLmods
I1-IQ2_S2.42 GiB2,595,637,6962.504mradermacher
I1-IQ2_M2.59 GiB2,780,342,7202.682mradermacher
I1-Q2_K_S2.64 GiB2,834,074,0482.734mradermacher
Q2_K2.81 GiB3,015,940,2562.910mradermacher
Q2_K2.81 GiB3,015,940,2562.910prithivMLmods
I1-Q2_K2.81 GiB3,015,940,5442.910mradermacher
I1-IQ3_XXS2.90 GiB3,114,514,8803.005mradermacher
Q8_03.06 GiB3,285,475,3283.170prithivMLmods
I1-IQ3_XS3.12 GiB3,346,256,3203.228mradermacher
Q3_K_S3.25 GiB3,492,368,5443.369prithivMLmods
Q3_K_S3.25 GiB3,492,368,5443.369mradermacher
I1-Q3_K_S3.25 GiB3,492,368,8323.369mradermacher
I1-IQ3_S3.26 GiB3,499,192,7683.376mradermacher
I1-IQ3_M3.33 GiB3,574,012,3523.448mradermacher
Q3_K_M3.55 GiB3,808,391,3283.674prithivMLmods
Q3_K_M3.55 GiB3,808,391,3283.674mradermacher
I1-Q3_K_M3.55 GiB3,808,391,6163.674mradermacher
Q3_K_L3.81 GiB4,088,459,4243.944prithivMLmods
Q3_K_L3.81 GiB4,088,459,4243.944mradermacher
I1-Q3_K_L3.81 GiB4,088,459,7123.944mradermacher
I1-IQ4_XS3.93 GiB4,218,472,8964.070mradermacher
IQ4_XS3.96 GiB4,250,298,5284.101mradermacher
IQ4_XS3.96 GiB4,250,298,5284.101prithivMLmods
I1-IQ4_NL4.13 GiB4,437,813,6964.282mradermacher
I1-Q4_04.14 GiB4,444,121,5364.287mradermacher
Q4_K_S4.15 GiB4,457,769,1204.301mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.34 GiB. The real file is 1.80 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Qwen2.5-VL-7B-Abliterated-Caption-it need?
Q4_K_M is exactly 1,929,902,080 bytes (1.80 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Qwen2.5-VL-7B-Abliterated-Caption-it should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.