Voxtral-Mini-4B-Realtime-2602
mistralai/Voxtral-Mini-4B-Realtime-2602Voxtral-Mini-4B-Realtime-2602 at Q4_K_M is exactly 2,830,493,984 bytes (2.64 GiB / 2.83 GB) — an effective 5.112 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| Q4_K | 2.35 GiB | 2,524,137,216 | 4.559 | — | cstr |
| Q4_K_M | 2.64 GiB | 2,830,493,984 | 5.112 | 714 | handy-computer |
| Q5_K_M | 3.06 GiB | 3,281,439,008 | 5.926 | 714 | handy-computer |
| Q6_K | 3.41 GiB | 3,661,018,912 | 6.612 | 714 | handy-computer |
| Q8_0 | 4.41 GiB | 4,731,791,648 | 8.546 | 714 | handy-computer |
| Q8_0 | 4.41 GiB | 4,733,486,848 | 8.549 | 714 | cstr |
| BF16 | 8.26 GiB | 8,868,301,088 | 16.016 | — | handy-computer |
| F16 | 8.27 GiB | 8,879,114,528 | 16.036 | 714 | handy-computer |
KV cache by context
This model declares a 8,192-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Measured
| Metric | Value | What it means |
|---|---|---|
| rtf | 105.1 | |
| RTFx | 105.1 | Higher is better — audio seconds processed per second of compute. |
| Word error rate | 13.34% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 2.60% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 15.91% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 8.86% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 2.23% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 8.00% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 4.92% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 8.80% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 6.42% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 1.61% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 11.46% | Lower is better — the share of words transcribed incorrectly. |
RTFx measured by the Open ASR Leaderboard on a single datacenter GPU at a large batch size. It ranks models against each other; it says nothing about throughput on consumer hardware. We reproduce these figures with attribution; they are not ours and we have not verified the runs. Source: open-asr-leaderboard-english-short-latest.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.32 GiB. The real file is 2.64 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does Voxtral-Mini-4B-Realtime-2602 need?
- Q4_K_M is exactly 2,830,493,984 bytes (2.64 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of Voxtral-Mini-4B-Realtime-2602 should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.