Exact memory for every model you can run yourself.
We don't estimate weights — we sum the actual published file bytes. We don't guess KV cache — we compute it per layer, including sliding-window, latent and hybrid architectures. Speed is modeled, and we publish how wrong that model is against real measurements.
From the file· file bytes, summedPredicted· speed, with error barsSpec sheet· hardware specswhat do these mean?
Base models
2382
Quantizations
67146
Accelerators
239
Bytes indexed
1350.6 TB
Will it fit?
Pick a model and a card, move the context slider, watch what fits. Exact file bytes, not estimates.
I have this card
Pick your GPU and see everything that fits, across every modality, at any context length.
I want this model
Every shipped quantization with its real file size, and which hardware holds it.
I have this much VRAM
What actually runs in 8, 12, 24, 48 or 128 GB — ranked by what fits, not by guesswork.
Where each number comes from
strongest first — every figure on the site is tagged
- From the file
- A fact, not a calculation. File sizes are the bytes of the actual published files added up; cache sizes come from the model's own architecture. If this is wrong, the source is wrong.
- Benchmarked
- Somebody ran it and wrote down the result. Not us — these come from public benchmark threads, reproduced with a link to the original. We aggregate them and have not rerun them.
- Spec sheet
- What the manufacturer publishes. Memory bandwidth, capacity, throughput. Theoretical peak, which no real workload reaches.
- Predicted
- Our prediction, never an observation. Always carries a ± figure, and that figure is how wrong we have actually been against real benchmarks — not a guess at our own accuracy.
- Rough guess
- A fallback we could not avoid. Rare, and flagged rather than quietly mixed in with everything else.
Most downloaded
| Model | Modality | Params | Quants | Smallest | Downloads |
|---|---|---|---|---|---|
| Qwen3.6-35B-A3BMoE | vision language | 36.0B | 34 | 9.36 GiB | 4.4M |
| embeddinggemma-300m | text | 303M | 9 | 0.26 GiB | 4.1M |
| Qwen3.5-9B | vision language | 9.7B | 34 | 2.97 GiB | 2.4M |
| gemma-4-26B-A4B-itMoE | vision language | 26.5B | 43 | 8.99 GiB | 2.2M |
| nemotron-3.5-asr-streaming-0.6b | audio asr | 638M | 8 | 0.38 GiB | 1.8M |
| Qwen3-Coder-30B-A3B-InstructMoE | text | 30.5B | 46 | 7.46 GiB | 1.8M |
| DeepSeek-V4-FlashMoE | text | 158B | 10 | 76.87 GiB | 1.8M |
| parakeet-unified-en-0.6b | audio asr | 618M | 6 | 0.44 GiB | 1.6M |
| Qwen3.5-4B | vision language | 4.7B | 34 | 1.42 GiB | 1.6M |
| Hy3MoE | text | 299B | 37 | 59.65 GiB | 1.4M |
| Qwythos-9B-Claude-Mythos-5-1M | vision language | 9.4B | 16 | 5.38 GiB | 1.4M |
| gemma-4-12B-it-qat-q4_0-unquantized | text | 12.0B | 2 | 6.50 GiB | 1.1M |