Nimbz · text

Prosopon-31B

Nimbz/Prosopon-31B

Prosopon-31B at Q4_K_M is exactly 19,479,785,536 bytes (18.14 GiB / 19.48 GB) — an effective 4.768 bits per weight, not the nominal 4. Its KV cache at 32K is 6.17 GiB, not the 30.00 GiB a flat formula predicts.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
32.7B
Architecture
gemma4
60 layers
Context
262,144
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S7.10 GiB7,618,913,7281.865mradermacher
I1-IQ1_M7.63 GiB8,188,296,6402.004mradermacher
I1-IQ2_XXS8.51 GiB9,137,268,1602.237mradermacher
I1-IQ2_XS9.31 GiB9,992,783,2962.446mradermacher
I1-IQ2_S10.02 GiB10,763,443,6482.635mradermacher
I1-Q2_K_S10.65 GiB11,439,013,3122.800mradermacher
I1-IQ2_M10.73 GiB11,522,620,8642.821mradermacher
Q2_K11.53 GiB12,378,737,8563.030mradermacher
I1-Q2_K11.53 GiB12,378,738,1123.030mradermacher
I1-IQ3_XXS11.81 GiB12,683,062,7203.105mradermacher
I1-IQ3_XS12.74 GiB13,677,923,7763.348mradermacher
Q3_K_S13.38 GiB14,366,911,6803.517mradermacher
I1-Q3_K_S13.38 GiB14,366,911,9363.517mradermacher
I1-IQ3_S13.38 GiB14,366,911,9363.517mradermacher
I1-IQ3_M14.00 GiB15,030,052,2883.679mradermacher
Q3_K_M14.80 GiB15,892,663,4883.890mradermacher
I1-Q3_K_M14.80 GiB15,892,663,7443.890mradermacher
Q3_K_L16.05 GiB17,233,824,9604.218mradermacher
I1-Q3_K_L16.05 GiB17,233,825,2164.218mradermacher
I1-IQ4_XS16.28 GiB17,484,475,8404.280mradermacher
IQ4_XS16.40 GiB17,610,919,1044.311mradermacher
I1-Q4_017.22 GiB18,494,303,6804.527mradermacher
Q4_K_S17.28 GiB18,555,890,8804.542mradermacher
I1-Q4_K_S17.28 GiB18,555,891,1364.542mradermacher
Q4_K_M18.14 GiB19,479,785,5364.768Nimbz
Q4_K_M18.14 GiB19,479,788,7364.768mradermacher
I1-Q4_K_M18.14 GiB19,479,788,9924.768mradermacher
I1-Q4_118.96 GiB20,362,227,1364.984mradermacher
Q5_K_S20.75 GiB22,280,724,5445.454Nimbz
Q5_K_S20.75 GiB22,280,727,7445.454mradermacher
I1-Q5_K_S20.75 GiB22,280,728,0005.454mradermacher
Q5_K_M21.25 GiB22,814,453,8245.585Nimbz
Q5_K_M21.25 GiB22,814,457,0245.585mradermacher
I1-Q5_K_M21.25 GiB22,814,457,2805.585mradermacher
Q6_K24.55 GiB26,357,538,8806.452Nimbz
Q6_K24.55 GiB26,357,542,0806.452mradermacher
I1-Q6_K24.55 GiB26,357,542,3366.452mradermacher
Q8_031.79 GiB34,133,041,2168.355Nimbz
Q8_031.79 GiB34,133,044,4168.355mradermacher

KV cache by context

computed per layer — this model uses sliding-window attention
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.80 GiB3.75 GiB2.09×10 / 50 / 0
8,1922.42 GiB7.50 GiB3.10×10 / 50 / 0
16,3843.67 GiB15.00 GiB4.09×10 / 50 / 0
32,7686.17 GiB30.00 GiB4.86×10 / 50 / 0
65,53611.17 GiB60.00 GiB5.37×10 / 50 / 0
131,07221.17 GiB120.00 GiB5.67×10 / 50 / 0

50 of 60 layers cache only a 1,024-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 17.12 GiB. The real file is 18.14 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 30.00 GiB at 32K context where the real figure is 6.17 GiB, because most of this model's layers cache a fixed window rather than the whole context.

Architecture

from config.json
Layers
60
Attention heads
32
KV heads
16
Head dim
256
Hidden size
5376
Vocab
262,144
Sliding window
1024
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Prosopon-31B need?
Q4_K_M is exactly 19,479,785,536 bytes (18.14 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Prosopon-31B's KV cache?
6.17 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Prosopon-31B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.
Prosopon-31B — VRAM requirements, exact quant sizes — ossmodeldb