canary-180m-flash
nvidia/canary-180m-flashcanary-180m-flash at Q4_K_M is exactly 139,223,744 bytes (0.13 GiB / 0.14 GB) — an effective 5.891 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| Q4_K_M | 0.13 GiB | 139,223,744 | 5.891 | 789 | handy-computer |
| Q5_K_M | 0.15 GiB | 158,704,320 | 6.715 | 789 | handy-computer |
| Q6_K | 0.16 GiB | 176,291,520 | 7.459 | 789 | handy-computer |
| Q8_0 | 0.20 GiB | 218,447,552 | 9.242 | 789 | handy-computer |
| F16 | 0.36 GiB | 381,632,192 | 16.147 | 789 | handy-computer |
| F32 | 0.70 GiB | 756,498,112 | 32.007 | — | handy-computer |
Measured
| Metric | Value | What it means |
|---|---|---|
| rtf | 2484.2 | |
| RTFx | 2484.2 | Higher is better — audio seconds processed per second of compute. |
| Word error rate | 6.11% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 12.10% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 12.50% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 8.96% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 2.04% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 3.41% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 3.57% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 13.66% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 1.52% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 8.87% | Lower is better — the share of words transcribed incorrectly. |
| Word error rate | 6.29% | Lower is better — the share of words transcribed incorrectly. |
RTFx measured by the Open ASR Leaderboard on a single datacenter GPU at a large batch size. It ranks models against each other; it says nothing about throughput on consumer hardware. We reproduce these figures with attribution; they are not ours and we have not verified the runs. Source: open-asr-leaderboard-english-short-latest.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 0.10 GiB. The real file is 0.13 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.
Questions people ask
- How much VRAM does canary-180m-flash need?
- Q4_K_M is exactly 139,223,744 bytes (0.13 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of canary-180m-flash should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.