functionary-medium-v3.2
meetkai/functionary-medium-v3.2functionary-medium-v3.2 at Q4_K_M is exactly 42,520,394,592 bytes (39.60 GiB / 42.52 GB) — an effective 4.821 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| IQ1_M | 15.60 GiB | 16,751,197,024 | 1.899 | — | bartowski |
| IQ2_XXS | 17.79 GiB | 19,097,385,824 | 2.165 | — | bartowski |
| IQ2_XS | 19.69 GiB | 21,142,109,024 | 2.397 | — | bartowski |
| IQ2_M | 22.46 GiB | 24,119,294,816 | 2.735 | — | bartowski |
| Q2_K | 24.56 GiB | 26,375,109,472 | 2.991 | — | bartowski |
| Q2_K_L | 25.52 GiB | 27,401,157,472 | 3.107 | — | bartowski |
| IQ3_XXS | 25.58 GiB | 27,469,495,136 | 3.115 | — | bartowski |
| Q3_K_S | 28.79 GiB | 30,912,052,064 | 3.505 | — | bartowski |
| IQ3_M | 29.74 GiB | 31,937,035,104 | 3.621 | — | bartowski |
| Q3_K_M | 31.91 GiB | 34,267,495,264 | 3.886 | — | bartowski |
| Q3_K_L | 34.59 GiB | 37,140,593,504 | 4.211 | — | bartowski |
| IQ4_XS | 35.30 GiB | 37,902,662,496 | 4.298 | — | bartowski |
| Q4_K_S | 37.58 GiB | 40,347,220,832 | 4.575 | — | bartowski |
| Q4_K_M | 39.60 GiB | 42,520,394,592 | 4.821 | — | bartowski |
| Q4_K_L | 40.33 GiB | 43,300,191,072 | 4.910 | — | bartowski |
| Q5_K_M2 shards | 46.52 GiB | 49,949,817,888 | 5.664 | — | bartowski |
| Q6_K2 shards | 53.91 GiB | 57,888,144,448 | 6.564 | — | bartowski |
| Q8_02 shards | 69.83 GiB | 74,975,050,784 | 8.501 | — | bartowski |
KV cache by context
This model declares a 8,192-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 36.96 GiB. The real file is 39.60 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does functionary-medium-v3.2 need?
- Q4_K_M is exactly 42,520,394,592 bytes (39.60 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of functionary-medium-v3.2 should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.