CohereLabs · text

aya-expanse-8b

CohereLabs/aya-expanse-8b

aya-expanse-8b at Q4_K_M is exactly 5,056,982,464 bytes (4.71 GiB / 5.06 GB) — an effective 5.039 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
8.0B
Architecture
command-r
Context
native (config.json)
License
cc-by-nc-4.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M2.87 GiB3,084,807,8723.074bartowski
Q2_K3.20 GiB3,438,505,6643.426bartowski
Q2_K_L3.44 GiB3,692,457,6643.680bartowski
IQ3_XS3.47 GiB3,724,766,9123.712bartowski
Q3_K_S3.60 GiB3,870,518,9763.857bartowski
IQ3_M3.72 GiB3,990,843,0723.977bartowski
Q3_K_M3.93 GiB4,224,937,6644.210bartowski
Q3_K_L4.22 GiB4,527,975,8724.512lmstudio-community
Q3_K_L4.22 GiB4,527,976,1284.512bartowski
IQ4_XS4.28 GiB4,600,327,8724.584bartowski
Q4_04.48 GiB4,812,140,2244.795bartowski
Q4_K_S4.50 GiB4,828,917,4404.812bartowski
Q4_K_M4.71 GiB5,056,982,4645.039lmstudio-community
Q4_K_M4.71 GiB5,056,982,7205.039bartowski
Q4_K_L4.95 GiB5,310,934,7205.292bartowski
Q5_K_S5.28 GiB5,669,875,3925.650bartowski
Q5_K_M5.40 GiB5,803,568,8325.783bartowski
Q5_K_L5.64 GiB6,057,520,8326.036bartowski
Q6_K6.14 GiB6,596,816,3206.574lmstudio-community
Q6_K6.14 GiB6,596,816,5766.574bartowski
Q6_K_L6.38 GiB6,850,768,5766.827bartowski
Q8_07.95 GiB8,541,072,8328.511lmstudio-community
Q8_07.95 GiB8,541,073,0888.511bartowski
F1614.96 GiB16,067,227,07216.011bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.21 GiB. The real file is 4.71 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does aya-expanse-8b need?
Q4_K_M is exactly 5,056,982,464 bytes (4.71 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of aya-expanse-8b should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.