circlestone-labs · image

Anima

circlestone-labs/Anima

Anima at Q4_K_M is exactly 1,384,407,168 bytes (1.29 GiB / 1.38 GB) — an effective 5.296 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
2.1B
Architecture
cosmos
Context
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q3_K_M0.95 GiB1,022,137,4723.910Abiray
Q3_K_M0.95 GiB1,022,137,4723.910Abiray
Q3_K_L1.07 GiB1,146,410,1124.386Bedovyy
Q4_K_S1.18 GiB1,269,080,1924.855Abiray
Q4_K_S1.18 GiB1,269,080,1924.855Abiray
Q4_01.19 GiB1,272,635,5204.869Bedovyy
Q4_K_S1.22 GiB1,306,321,0244.998Bedovyy
Q4_11.28 GiB1,370,046,5925.242Bedovyy
Q4_K_M1.29 GiB1,384,407,1685.296Abiray
Q4_K_M1.29 GiB1,384,407,1685.296Abiray
Q5_01.43 GiB1,531,682,9445.860Bedovyy
Q5_K_M1.45 GiB1,555,636,3525.952Abiray
Q5_K_M1.45 GiB1,555,636,3525.952Abiray
Q5_11.52 GiB1,629,094,0166.233Bedovyy
Q6_K1.62 GiB1,737,567,3606.648Abiray
Q6_K1.62 GiB1,737,567,3606.648Abiray
Q8_02.09 GiB2,239,471,7448.568Abiray
Q8_02.09 GiB2,239,471,7448.568Abiray
Q4_K_M3 shards3.97 GiB4,264,944,00016.317Bedovyy
Q5_K_S3 shards4.19 GiB4,498,710,91217.211Bedovyy
Q5_K_M3 shards4.45 GiB4,778,631,55218.282Bedovyy
Q6_K3 shards4.96 GiB5,324,424,57620.370Bedovyy
Q8_03 shards6.36 GiB6,830,137,72826.131Bedovyy

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.10 GiB. The real file is 1.29 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Anima need?
Q4_K_M is exactly 1,384,407,168 bytes (1.29 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Anima should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.