dphn · text

dolphincoder-starcoder2-15b

dphn/dolphincoder-starcoder2-15b

dolphincoder-starcoder2-15b at Q4_K_M is exactly 9,860,206,880 bytes (9.18 GiB / 9.86 GB) — an effective 4.943 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
16.0B
Architecture
starcoder2
40 layers
Context
16,384
native (config.json)
License
bigcode-openrail-m

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S3.31 GiB3,558,621,3121.784mradermacher
I1-IQ1_M3.60 GiB3,862,380,6721.936mradermacher
I1-IQ2_XXS4.07 GiB4,368,646,2722.190mradermacher
I1-IQ2_XS4.49 GiB4,820,844,6722.417mradermacher
I1-IQ2_S4.79 GiB5,140,530,5282.577mradermacher
I1-IQ2_M5.16 GiB5,545,543,0082.780mradermacher
Q2_K5.77 GiB6,192,973,2803.105mradermacher
I1-Q2_K5.77 GiB6,192,973,5363.105mradermacher
I1-IQ3_XXS5.79 GiB6,217,942,3683.117mradermacher
IQ3_XS6.25 GiB6,714,182,3363.366mradermacher
I1-IQ3_XS6.25 GiB6,714,182,5923.366mradermacher
Q3_K_S6.51 GiB6,986,484,4163.502mradermacher
I1-Q3_K_S6.51 GiB6,986,484,6723.502mradermacher
IQ3_S6.52 GiB7,003,196,0963.511mradermacher
I1-IQ3_S6.52 GiB7,003,196,3523.511mradermacher
IQ3_M6.80 GiB7,304,006,3363.662mradermacher
I1-IQ3_M6.80 GiB7,304,006,5923.662mradermacher
Q3_K_M7.49 GiB8,044,432,0644.033mradermacher
I1-Q3_K_M7.49 GiB8,044,432,3204.033mradermacher
I1-IQ4_XS8.01 GiB8,595,919,0084.309mradermacher
IQ4_XS8.12 GiB8,713,883,5524.368mradermacher
Q3_K_L8.35 GiB8,965,343,9364.495mradermacher
I1-Q3_K_L8.35 GiB8,965,344,1924.495mradermacher
I1-Q4_08.49 GiB9,112,605,2164.568mradermacher
Q4_K_S8.53 GiB9,161,363,7444.593mradermacher
I1-Q4_K_S8.53 GiB9,161,364,0004.593mradermacher
Q4_K_M9.18 GiB9,860,206,8804.943mradermacher
I1-Q4_K_M9.18 GiB9,860,207,1364.943mradermacher
Q5_K_S10.27 GiB11,022,063,3925.526mradermacher
I1-Q5_K_S10.27 GiB11,022,063,6485.526mradermacher
Q5_K_M10.65 GiB11,431,499,5525.731mradermacher
I1-Q5_K_M10.65 GiB11,431,499,8085.731mradermacher
Q6_K12.20 GiB13,100,998,0166.568mradermacher
I1-Q6_K12.20 GiB13,100,998,2726.568mradermacher
Q8_015.80 GiB16,965,137,6008.505mradermacher

KV cache by context

unresolved

This model declares a 4,096-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.

A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 8.36 GiB. The real file is 9.18 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
40
Attention heads
48
KV heads
4
Head dim
128
Hidden size
6144
Vocab
49,154
Sliding window
4096
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does dolphincoder-starcoder2-15b need?
Q4_K_M is exactly 9,860,206,880 bytes (9.18 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of dolphincoder-starcoder2-15b should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.