CohereLabs · text

c4ai-command-r-v01

CohereLabs/c4ai-command-r-v01

c4ai-command-r-v01 at Q4_K_M is exactly 21,527,051,232 bytes (20.05 GiB / 21.53 GB) — an effective 4.923 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
35.0B
Architecture
command-r
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S7.94 GiB8,523,398,1121.949dranger003
IQ1_M8.52 GiB9,146,645,4722.092dranger003
IQ2_XXS9.49 GiB10,185,391,0722.329dranger003
IQ2_XS10.34 GiB11,100,273,6322.539dranger003
IQ2_S11.03 GiB11,844,107,2322.709dranger003
IQ2_M11.80 GiB12,675,103,7122.899dranger003
Q2_K_S11.86 GiB12,738,673,6322.913dranger003
Q2_K12.87 GiB13,817,395,8723.160DavidAU
Q2_K12.87 GiB13,817,396,1923.160dranger003
IQ3_XXS12.88 GiB13,832,469,4723.163dranger003
IQ3_XS14.05 GiB15,091,416,0323.451dranger003
IQ3_S14.77 GiB15,862,119,3923.628dranger003
Q3_K_S14.77 GiB15,862,119,3923.628dranger003
IQ3_M15.55 GiB16,697,703,3923.819dranger003
Q3_K_M16.41 GiB17,618,484,1924.029dranger003
Q3_K_L17.83 GiB19,149,405,1524.379dranger003
IQ4_XS17.88 GiB19,201,833,9524.391dranger003
IQ4_NL18.84 GiB20,229,438,4324.626dranger003
Q4_K_S18.98 GiB20,378,336,2244.660dranger003
Q4_K_M20.05 GiB21,527,051,2324.923dranger003
IQ2_XS2 shards20.68 GiB22,200,547,2005.077DavidAU
Q5_K_S22.67 GiB24,339,856,3525.566dranger003
Q5_K_M23.29 GiB25,008,323,5525.719dranger003
Q6_K26.74 GiB28,707,175,1046.565dranger003
IQ3_XS2 shards28.11 GiB30,182,832,0006.903DavidAU
IQ3_M2 shards31.10 GiB33,395,406,7207.637DavidAU
Q8_034.63 GiB37,179,013,8248.503dranger003
Q4_K_S2 shards37.96 GiB40,756,672,3849.321DavidAU
Q5_K_S2 shards45.34 GiB48,679,712,64011.133DavidAU
Q6_K2 shards53.47 GiB57,414,350,72013.130DavidAU
F162 shards65.17 GiB69,973,228,38416.003dranger003

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 18.33 GiB. The real file is 20.05 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does c4ai-command-r-v01 need?
Q4_K_M is exactly 21,527,051,232 bytes (20.05 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of c4ai-command-r-v01 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.