Status
What is in the dataset right now, and what it cannot yet tell you.
Base models
2,493
Quantizations
67,146
Accelerators
239
Fit counts
478
Measured rows
57,390
Coverage
last ingest 2026-07-28 11:16
| Field | Count | Coverage |
|---|---|---|
| architecture resolved | 2,167 | 87% |
| from an ungated mirror | 50 | 2% |
| effective bpw computed | 66,655 | 99% |
| tensor histogram parsed | 256 | 10% |
Measured data
117 accelerators covered
Reproduced from third-party benchmark publishers with attribution, never relicensed. None of it was measured by us, and we say so wherever it appears. It exists so that our own modeled figures can eventually be scored against something real rather than merely banded.
By modality
audio asr 52image 47video 40embedding 27audio tts 25vision language 216text 2086
Known limitations
- No first-party measured performance
- We now reproduce third-party measured benchmarks with attribution — image generation throughput per GPU, and speech recognition accuracy and speed. But every figure WE compute is still modeled from memory bandwidth with a published error band. Nothing on this site has been measured by us.
- Measured data is not redistributable
- The sources we reproduce publish no licence of their own. We display and attribute them, and exclude them from our CC-BY bulk dump, which therefore contains only data we derived ourselves.
- Image throughput is community-submitted, not controlled
- The image-generation figures per GPU aggregate thousands of community runs across different models, resolutions and step counts. Read the spread, not the median alone — it is not one controlled configuration.
- Video has no throughput data at all
- Component sizes are exact. Seconds per clip on consumer hardware is unpublished anywhere credible, so we show nothing rather than a guess.
- The catalog is a slice, not the whole Hub
- We index the most-downloaded base models. The full deduplicated set is roughly 1,500–3,000.
- Gated models may lack architecture
- Exact file sizes are always available, even for gated repositories. Architecture sometimes has to come from an ungated mirror, which we flag with reduced confidence, or is unavailable entirely.
- Compute buffer is modeled with constants we chose
- There is no closed form for the runtime's working buffers. Ours scales with the terms that should drive it — feed-forward width, micro-batch, and context without flash attention — but its constants are not fitted against measured allocations. It is the widest error band on the site and the reason a total carries one when its parts do not.