DavidAU · text

Qwen3-VL-12B-Instruct-Brainstorm20x

DavidAU/Qwen3-VL-12B-Instruct-Brainstorm20x

Qwen3-VL-12B-Instruct-Brainstorm20x at Q4_K_M is exactly 7,216,982,464 bytes (6.72 GiB / 7.22 GB) — an effective 4.870 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
11.9B
Architecture
qwen3vl
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S2.70 GiB2,894,961,3761.953mradermacher
I1-IQ1_M2.90 GiB3,109,559,0082.098mradermacher
I1-IQ2_XXS3.23 GiB3,467,221,7282.339mradermacher
I1-IQ2_XS3.52 GiB3,782,187,7442.552mradermacher
I1-IQ2_S3.73 GiB4,005,825,2482.703mradermacher
I1-IQ2_M4.00 GiB4,291,955,4242.896mradermacher
I1-Q2_K_S4.03 GiB4,329,327,3282.921mradermacher
Q2_K4.32 GiB4,633,414,0803.126mradermacher
I1-Q2_K4.32 GiB4,633,414,3683.126mradermacher
I1-IQ3_XXS4.45 GiB4,777,970,4003.224mradermacher
IQ2_M4.76 GiB5,108,762,0483.447DavidAU
I1-IQ3_XS4.77 GiB5,123,816,1603.457mradermacher
Q3_K_S4.98 GiB5,345,425,8563.607mradermacher
I1-Q3_K_S4.98 GiB5,345,426,1443.607mradermacher
I1-IQ3_S5.01 GiB5,376,064,2243.627mradermacher
I1-IQ3_M5.16 GiB5,538,724,5763.737mradermacher
Q3_K_M5.48 GiB5,886,196,1603.972mradermacher
I1-Q3_K_M5.48 GiB5,886,196,4483.972mradermacher
IQ3_M5.84 GiB6,272,878,0164.232DavidAU
Q3_K_L5.92 GiB6,356,482,4964.289mradermacher
I1-Q3_K_L5.92 GiB6,356,482,7844.289mradermacher
I1-IQ4_XS6.07 GiB6,522,415,8404.401mradermacher
IQ4_XS6.12 GiB6,569,601,4724.433mradermacher
I1-Q4_06.39 GiB6,856,305,3764.626mradermacher
I1-IQ4_NL6.39 GiB6,866,266,8484.633mradermacher
Q4_K_S6.40 GiB6,877,276,6084.640mradermacher
I1-Q4_K_S6.40 GiB6,877,276,8964.640mradermacher
Q4_K_M6.72 GiB7,216,982,4644.870mradermacher
I1-Q4_K_M6.72 GiB7,216,982,7524.870mradermacher
IQ4_XS6.76 GiB7,256,569,2804.896DavidAU
I1-Q4_17.02 GiB7,539,550,9445.087mradermacher
Q4_K_S7.09 GiB7,611,430,3365.136DavidAU
Q4_K_M7.41 GiB7,951,136,1925.365DavidAU
Q5_K_S7.68 GiB8,241,670,5925.561mradermacher
I1-Q5_K_S7.68 GiB8,241,670,8805.561mradermacher
Q5_K_M7.86 GiB8,437,197,2485.693mradermacher
I1-Q5_K_M7.86 GiB8,437,197,5365.693mradermacher
Q5_K_S8.36 GiB8,975,824,3206.056DavidAU
Q5_K_M8.54 GiB9,171,350,9766.188DavidAU
Q6_K9.07 GiB9,733,675,4566.567mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.21 GiB. The real file is 6.72 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Qwen3-VL-12B-Instruct-Brainstorm20x need?
Q4_K_M is exactly 7,216,982,464 bytes (6.72 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Qwen3-VL-12B-Instruct-Brainstorm20x should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.