# ossmodeldb — full reference # Exact memory and modeled speed for locally-run open-weight models, across every modality. # Data licensed CC BY 4.0. Attribute as "ossmodeldb.org". # 2493 base models · 67146 quantizations · 239 accelerators · 179 pipeline components ## THE CORE CLAIM Every comparable site estimates a model's size as parameters times a bits-per-weight constant. That is wrong before it starts, for two independent reasons: 1. A quantization label names a MIXTURE, not a uniform precision. Normalization and bias tensors stay at full precision and attention/output tensors are commonly promoted, so a real file always exceeds the nominal rate. We sum published file bytes instead of estimating. 2. The KV cache is a per-layer sum. Three architecture families break the flat context x layers x heads formula, each by multiples rather than percentages. ## MEASURED EFFECTIVE BITS PER WEIGHT # What files actually weigh, versus what their label implies. # quant files mean min max UD-TQ1_0 50 2.067 1.810 2.771 I1-IQ1_S 957 2.068 1.351 5.114 IQ1_S 192 2.075 1.561 7.254 IQ1_M 254 2.170 1.660 7.918 UD-IQ1_S 156 2.172 1.387 4.457 I1-IQ1_M 956 2.209 1.416 5.149 UD-IQ1_M 177 2.323 1.467 4.479 IQ2_XXS 340 2.393 1.854 7.074 I1-IQ2_XXS 957 2.445 1.483 5.207 UD-IQ2_XXS 196 2.601 1.596 5.374 IQ2_XS 445 2.614 1.608 9.993 I1-IQ2_XS 957 2.647 1.546 5.613 IQ2_S 449 2.712 2.145 9.233 I1-IQ2_S 957 2.762 1.686 5.853 TQ2_0 19 2.808 2.069 4.495 UD-IQ2_M 203 2.947 1.874 5.853 I1-IQ2_M 957 2.950 1.740 6.260 IQ2_M 738 2.975 1.178 10.009 I1-Q2_K_S 829 3.030 1.756 5.955 Q2_K_S 86 3.184 2.701 14.149 I1-Q2_K 1036 3.223 1.823 6.874 TQ1_0 14 3.256 1.811 10.411 I1-IQ3_XXS 959 3.278 1.822 6.998 UD-IQ3_XXS 210 3.310 2.075 5.506 Q2_K 2322 3.338 1.078 19.476 UD-IQ3_S 31 3.344 2.903 5.937 Q2_K_L 759 3.376 2.095 7.075 IQ3_XXS 536 3.427 2.268 19.140 Q2_K_M 7 3.492 2.895 4.210 I1-IQ3_XS 960 3.518 1.996 7.563 IQ3_XS 782 3.561 1.297 13.839 I1-Q3_K_S 1035 3.680 2.040 7.887 I1-IQ3_S 1035 3.698 2.045 7.900 Q3_K_S 2232 3.742 1.403 21.256 I1-IQ3_M 1033 3.796 2.094 8.066 IQ3_M 898 3.860 2.453 13.190 IQ3_S 234 3.867 2.884 20.569 I1-Q3_K_M 1034 4.030 1.061 8.593 UD-IQ4_XS 34 4.088 3.741 6.979 Q3_K_M 2452 4.126 1.534 22.549 UD-IQ4_NL 28 4.140 3.843 6.979 I1-Q3_K_L 1033 4.321 1.151 9.199 Q3_K_L 2128 4.369 1.647 23.759 I1-IQ4_XS 1033 4.413 1.080 9.445 Q3_K 83 4.496 3.553 14.091 UD-Q3_K_M 36 4.524 3.504 11.747 IQ4_XS 1960 4.545 1.206 27.605 I1-Q4_0 1032 4.625 1.131 9.934 I1-Q4_K_S 1033 4.672 1.131 9.965 I1-IQ4_NL 528 4.683 1.131 9.321 UD-Q3_K_S 10 4.737 3.259 9.909 IQ4_NL 831 4.766 2.217 32.274 Q4_K_S 2350 4.788 1.303 27.527 Q4_0 1367 4.912 1.303 32.453 I1-Q4_K_M 1032 4.925 1.238 10.460 I1-Q4_1 896 5.051 1.232 10.222 Q4_K_M 3293 5.130 1.346 28.901 Q4_K_L 559 5.201 3.776 8.749 MXFP4 11 5.207 4.211 8.501 UD-Q4_K_S 28 5.285 4.521 12.367 Q4_1 872 5.372 1.431 30.248 I1-Q5_K_S 1015 5.536 1.333 11.804 Q5_K_S 2309 5.648 1.558 24.887 I1-Q5_K_M 1017 5.691 1.389 12.090 Q5_K_M 2766 5.912 1.604 32.945 UD-Q4_K_M 32 5.947 4.841 14.358 Q5_K_L 521 5.963 4.279 8.803 Q5_0 503 6.106 1.558 31.893 Q4_K 162 6.143 3.201 28.888 UD-Q5_K_S 27 6.268 5.485 14.200 Q5_K 108 6.374 4.174 21.264 Q6_K_M 10 6.415 3.644 7.657 I1-Q6_K 989 6.541 1.548 13.822 UD-Q6_K 27 6.733 6.366 8.138 UD-Q2_K 5 6.751 3.951 9.479 Q6_K 2983 6.797 1.038 29.692 Q6_K_L 520 6.811 4.606 9.239 NVFP4 30 6.849 4.463 10.637 Q5_1 268 6.942 1.685 31.474 UD-Q5_K_M 30 7.094 5.820 15.896 Q8_0 3438 8.648 1.210 31.059 Q2_0 7 8.699 2.132 18.098 BF16 936 15.490 1.033 32.883 F16 1165 15.676 1.014 32.396 F32 157 31.325 5.061 32.414 ## THE THREE ARCHITECTURES A FLAT KV FORMULA GETS WRONG - SLIDING-WINDOW ATTENTION. Most layers cache a fixed window, not the context. Gemma-3-27B has 10 full and 52 windowed layers of 62; its cache at 32K is 3.109 GiB rather than 15.50 GiB. The period comes from the model file when present and a per-architecture default otherwise. A period of 0 means EVERY layer is windowed, not none. - MULTI-HEAD LATENT ATTENTION (MLA). No value cache is allocated at all; the key cache stores a latent of kv_lora_rank + qk_rope_head_dim elements. DeepSeek-V3 caches 2.145 GiB at 32K. These models still declare a large num_key_value_heads, which is the trap. - HYBRID LINEAR ATTENTION. Some layers keep a fixed-size recurrent state that does not grow with context. Qwen3.6-27B has only 16 full-attention layers of 64; its cache at 128K is 8.00 GiB, not 32.00 GiB. ## SLIDING-WINDOW MODELS IN THE INDEX embeddinggemma-300m 24 layers, window 512, period 6 gemma-4-26B-A4B-it 30 layers, window 1024, period from layer map DeepSeek-V4-Flash 43 layers, window 128, period from layer map gemma-4-12B-it-qat-q4_0-unquantized 48 layers, window 1024, period from layer map gemma-4-26B-A4B-it-qat-q4_0-unquantized 30 layers, window 1024, period from layer map gpt-oss-20b 24 layers, window 128, period from layer map gemma-4-e4b-it 42 layers, window 512, period from layer map gemma-4-31B-it-qat-q4_0-unquantized 60 layers, window 1024, period from layer map gemma-4-E4B-it-qat-q4_0-unquantized 42 layers, window 512, period from layer map Voxtral-Mini-4B-Realtime-2602 26 layers, window 8192, period from layer map gemma-3-1b-it 26 layers, window 512, period 6 gemma-3-4b-it 34 layers, window 1024, period 6 Laguna-S-2.1 48 layers, window 512, period 4 gemma-4-26B-A4B-it-ultra-uncensored-heretic 30 layers, window 1024, period from layer map gemma-4-E2B-it-qat-q4_0-unquantized 35 layers, window 512, period from layer map gemma-3-12b-it 48 layers, window 1024, period 6 Phi-3.5-mini-instruct 32 layers, window 262144, period 1 gemma-2-2b-it 26 layers, window 4096, period 2 embeddinggemma-300m-qat-q8_0-unquantized 24 layers, window 512, period 6 gemma-4-12b-heretic-abliterated 48 layers, window 1024, period from layer map gpt-oss-120b 36 layers, window 128, period from layer map Phi-4-mini-instruct 32 layers, window 262144, period 1 gemma-4-31b-it 60 layers, window 1024, period from layer map gemma-4-e2b-it 35 layers, window 512, period from layer map gemma-3-27b-it 62 layers, window 1024, period 6 ## MIXTURE OF EXPERTS Resident size and per-token bytes are different questions. gpt-oss-120b holds 128 experts but routes to 4 per token, so roughly 3.1% of expert bytes are read per token while 100% must be resident. Collapsing these into one "VRAM required" figure is wrong in both directions, and it is why MoE models are the best candidates for CPU offload. ## DIFFUSION MODELS ARE PIPELINES A published parameter count for an image or video model describes the denoiser alone. FLUX.2 Klein: denoiser 16.91 GiB + text encoder 15.26 GiB + VAE 0.16 GiB = 32.32 GiB. The text encoder is 47% of the pipeline and is routinely held on the CPU, which is the largest single VRAM lever in image generation. Repositories often also ship a single-file copy of the denoiser at the root; counting both roughly doubles every figure. ## SPEED Token generation is bound by memory bandwidth, not arithmetic throughput: tokens/sec = 1 / (bytes_read_per_token / (bandwidth * efficiency) + fixed_overhead) Fitted per backend from public benchmark sets: Metal efficiency 0.928 with 6.16 ms fixed overhead; consumer CUDA 0.789 with 0.74 ms. MEASURED ACCURACY of those predictions, scored against 635 independent runs harvested from public benchmark threads (median / mean / 90th percentile absolute error): Apple Silicon n=116 4.3% / 8.3% / 16% consumer NVIDIA n=158 6.6% / 12.9% / 38% datacenter NVIDIA n=146 14.4% / 22.0% / 52% AMD n=215 13.6% / 26.5% / 62% The bands shown on the site are these measured figures, not the residual of the fit that produced eta and t0. The tail is longest on AMD and datacenter cards, where configuration varies most. Assuming a CONSTANT fraction of peak bandwidth — the common approach — is 16-24% MAPE with up to 139% max error, because the achieved fraction falls as bandwidth rises and a fixed per-token cost comes to dominate. Prefill is compute-bound instead, so quantizing weights does NOT speed up prompt processing and is usually slightly slower there, while decode at Q4 is roughly 2.3x faster. ## ACCELERATORS BY BANDWIDTH GeForce RTX 5090 D 32 GB 1792 GB/s NVIDIA GeForce RTX 5090 32 GB 1792 GB/s NVIDIA GeForce RTX 5090 D V2 24 GB 1344 GB/s NVIDIA Radeon VII 16 GB 1024 GB/s AMD GeForce RTX 3090 Ti 24 GB 1008 GB/s NVIDIA GeForce RTX 4090 24 GB 1008 GB/s NVIDIA GeForce RTX 4090 D 24 GB 1008 GB/s NVIDIA Radeon RX 7900 XTX 24 GB 960 GB/s AMD GeForce RTX 5080 16 GB 960 GB/s NVIDIA GeForce RTX 3090 24 GB 936 GB/s NVIDIA GeForce RTX 3080 12 GB 912 GB/s NVIDIA GeForce RTX 3080 Ti 12 GB 912 GB/s NVIDIA GeForce RTX 5090 Laptop 24 GB 896 GB/s NVIDIA GeForce RTX 5070 Ti 16 GB 896 GB/s NVIDIA GeForce RTX 5080 Laptop 16 GB 896 GB/s NVIDIA Titan V 32 GB 868 GB/s NVIDIA Apple M3 Ultra 256 GB 819 GB/s Apple Apple M3 Ultra 96 GB 819 GB/s Apple Apple M3 Ultra 512 GB 819 GB/s Apple Apple M1 Ultra 64 GB 819 GB/s Apple Apple M2 Ultra 64 GB 819 GB/s Apple Apple M2 Ultra 192 GB 819 GB/s Apple Apple M2 Ultra 128 GB 819 GB/s Apple Apple M1 Ultra 128 GB 819 GB/s Apple Radeon RX 7900 XT 20 GB 800 GB/s AMD GeForce RTX 3080 Ti 20 GB 760 GB/s NVIDIA GeForce RTX 3080 10 GB 760 GB/s NVIDIA GeForce RTX 4080 Super 16 GB 736 GB/s NVIDIA GeForce RTX 4080 16 GB 717 GB/s NVIDIA GeForce RTX 4070 Ti Super 16 GB 672 GB/s NVIDIA Titan RTX 24 GB 672 GB/s NVIDIA GeForce RTX 5070 12 GB 672 GB/s NVIDIA GeForce RTX 5070 Ti Laptop 12 GB 672 GB/s NVIDIA Titan V 12 GB 651 GB/s NVIDIA Radeon RX 9070 XT 16 GB 640 GB/s AMD Radeon RX 9070 16 GB 640 GB/s AMD Radeon RX 7800 XT 16 GB 624 GB/s AMD Radeon RX 7700 16 GB 624 GB/s AMD GeForce RTX 2080 Ti 11 GB 616 GB/s NVIDIA Apple M5 Max 48 GB 614 GB/s Apple Apple M5 Max 128 GB 614 GB/s Apple Apple M5 Max 64 GB 614 GB/s Apple GeForce RTX 3060 Ti GDDR6X 8 GB 608 GB/s NVIDIA GeForce RTX 3070 Ti 8 GB 608 GB/s NVIDIA Radeon RX 7900 GRE 16 GB 576 GB/s AMD Radeon RX 6950 XT 16 GB 576 GB/s AMD GeForce RTX 4090 Laptop 16 GB 576 GB/s NVIDIA Arc A770 16GB 16 GB 560 GB/s Intel Titan Xp 12 GB 548 GB/s NVIDIA Apple M4 Max 128 GB 546 GB/s Apple Apple M4 Max 64 GB 546 GB/s Apple Apple M4 Max 48 GB 546 GB/s Apple Radeon RX 6800 16 GB 512 GB/s AMD Arc A580 8GB 8 GB 512 GB/s Intel Arc A770 8GB 8 GB 512 GB/s Intel Radeon RX 6900 XT 16 GB 512 GB/s AMD Arc A750 8GB 8 GB 512 GB/s Intel GeForce RTX 3080 Ti Laptop 16 GB 512 GB/s NVIDIA Radeon RX 6800 XT 16 GB 512 GB/s AMD Arc A770M 16GB 16 GB 512 GB/s Intel ## MOST-DOWNLOADED MODELS Qwen3.6-35B-A3B 36.0B params smallest 9.36 GiB across 34 quants apache-2.0 embeddinggemma-300m 0.3B params smallest 0.26 GiB across 9 quants license unknown Qwen3.5-9B 9.7B params smallest 2.97 GiB across 34 quants apache-2.0 gemma-4-26B-A4B-it 26.5B params smallest 8.99 GiB across 43 quants apache-2.0 nemotron-3.5-asr-streaming-0.6b 0.6B params smallest 0.38 GiB across 8 quants other Qwen3-Coder-30B-A3B-Instruct 30.5B params smallest 7.46 GiB across 46 quants apache-2.0 DeepSeek-V4-Flash 158.1B params smallest 76.87 GiB across 10 quants mit parakeet-unified-en-0.6b 0.6B params smallest 0.44 GiB across 6 quants cc-by-4.0 Qwen3.5-4B 4.7B params smallest 1.42 GiB across 34 quants apache-2.0 Hy3 298.8B params smallest 59.65 GiB across 37 quants apache-2.0 Qwythos-9B-Claude-Mythos-5-1M 9.4B params smallest 5.38 GiB across 16 quants apache-2.0 gemma-4-12B-it-qat-q4_0-unquantized 12.0B params smallest 6.50 GiB across 2 quants apache-2.0 HyperCLOVAX-SEED-Text-Instruct-1.5B 1.6B params smallest 1.06 GiB across 1 quants other Qwen3-VL-30B-A3B-Instruct 31.1B params smallest 7.60 GiB across 26 quants apache-2.0 FLUX.2-klein-9B 9.1B params smallest 3.71 GiB across 25 quants other cohere-transcribe-03-2026 2.1B params smallest 1.41 GiB across 11 quants apache-2.0 gemma-4-26B-A4B-it-qat-q4_0-unquantized 26.5B params smallest 13.45 GiB across 1 quants apache-2.0 LTX-2.3 9.2B params smallest 7.39 GiB across 92 quants license unknown Llama-3.2-1B-Instruct 1.2B params smallest 0.39 GiB across 39 quants license unknown Qwen-AgentWorld-35B-A3B 34.7B params smallest 10.71 GiB across 15 quants apache-2.0 llama-3-youko-8b 8.0B params smallest 5.34 GiB across 2 quants llama3 Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP 27.8B params smallest 22.16 GiB across 10 quants apache-2.0 gpt-oss-20b 21.5B params smallest 10.68 GiB across 15 quants apache-2.0 Qwen3-8B 8.2B params smallest 2.12 GiB across 28 quants apache-2.0 Qwopus3.6-35B-A3B-v1 36.0B params smallest 6.97 GiB across 29 quants apache-2.0 Qwen3.5-0.8B 0.9B params smallest 0.31 GiB across 42 quants apache-2.0 Qwen3-VL-8B-Instruct-abliterated-v1 8.8B params smallest 1.97 GiB across 48 quants apache-2.0 gemma-4-e4b-it 8.0B params smallest 3.30 GiB across 36 quants gemma gemma-4-31B-it-qat-q4_0-unquantized 32.7B params smallest 16.44 GiB across 1 quants apache-2.0 Qwen3-4B 4.0B params smallest 1.01 GiB across 29 quants apache-2.0 UI-TARS-1.5-7B 8.3B params smallest 2.81 GiB across 14 quants license unknown Llama-3.1-8B-Instruct 8.0B params smallest 2.02 GiB across 45 quants llama3.1 GLM-5.2 753.3B params smallest 169.33 GiB across 20 quants mit Ornith-1.0-35B 34.7B params smallest 9.11 GiB across 43 quants mit parakeet-tdt-0.6b-v3 0.6B params smallest 0.39 GiB across 10 quants cc-by-4.0 ThinkingCap-Qwen3.6-27B 27.4B params smallest 9.30 GiB across 28 quants license unknown Llama-3.2-3B-Instruct 3.2B params smallest 0.85 GiB across 39 quants llama3.2 Qwopus3.6-27B-Coder 27.8B params smallest 8.89 GiB across 14 quants apache-2.0 whisper-medium 0.8B params smallest 0.25 GiB across 19 quants apache-2.0 Qwythos-9B-v2 9.7B params smallest 3.64 GiB across 43 quants apache-2.0 gemma-4-E4B-it-qat-q4_0-unquantized 7.9B params smallest 4.80 GiB across 1 quants apache-2.0 Wan2.2-I2V-A14B 14.3B params smallest 4.94 GiB across 30 quants apache-2.0 Voxtral-Mini-4B-Realtime-2602 4.4B params smallest 2.35 GiB across 8 quants apache-2.0 Qwen3.5-122B-A10B 125.1B params smallest 26.92 GiB across 58 quants apache-2.0 Wan2.1-T2V-1.3B 1.4B params smallest 0.61 GiB across 31 quants license unknown Qwen3-Coder-Next 79.7B params smallest 15.44 GiB across 61 quants apache-2.0 Kimi-K2.7-Code 1058.6B params smallest 283.04 GiB across 19 quants other Bielik-11B-v3.0-Instruct 11.2B params smallest 6.26 GiB across 6 quants apache-2.0 LFM2.5-1.2B-Instruct 1.2B params smallest 0.45 GiB across 23 quants other gemma-3-1b-it 1.0B params smallest 0.52 GiB across 28 quants gemma Qwen2.5-7B-Instruct 7.6B params smallest 2.59 GiB across 32 quants apache-2.0 Qwen3-14B 14.8B params smallest 3.56 GiB across 29 quants license unknown Qwen3-0.6B 0.8B params smallest 0.20 GiB across 44 quants license unknown Qwen3.5-35B-A3B 36.0B params smallest 8.77 GiB across 47 quants apache-2.0 Qwen2.5-Coder-7B-Instruct 7.6B params smallest 2.59 GiB across 32 quants apache-2.0 Qwen2.5-32B-Instruct 32.8B params smallest 8.41 GiB across 35 quants apache-2.0 Qwen3-VL-4B-Instruct 4.4B params smallest 1.01 GiB across 26 quants apache-2.0 Qwen2.5-1.5B-Instruct 1.5B params smallest 0.56 GiB across 32 quants apache-2.0 Jan-v3-4B-base-instruct 4.4B params smallest 1.56 GiB across 49 quants apache-2.0 gemma-3-4b-it 4.3B params smallest 1.43 GiB across 28 quants license unknown Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking 39.5B params smallest 13.76 GiB across 30 quants apache-2.0 whisper-large-v3 1.5B params smallest 0.49 GiB across 18 quants apache-2.0 MiniCPM5-1B-Claude-Opus-Fable5-Thinking 1.1B params smallest 0.43 GiB across 15 quants apache-2.0 Qwen3-30B-A3B 30.5B params smallest 7.59 GiB across 51 quants license unknown Qwen2.5-VL-7B-Instruct 8.3B params smallest 1.93 GiB across 50 quants apache-2.0 Qwen3-VL-2B-Instruct 2.1B params smallest 0.50 GiB across 46 quants apache-2.0 Ornith-1.0-9B 9.2B params smallest 2.26 GiB across 48 quants mit Qwen3-1.7B 2.0B params smallest 0.50 GiB across 25 quants license unknown LTX-2 18.9B params smallest 7.48 GiB across 28 quants license unknown Laguna-S-2.1 117.6B params smallest 23.15 GiB across 46 quants openmdw-1.1 canary-180m-flash 0.2B params smallest 0.13 GiB across 6 quants cc-by-4.0 Agents-A1 35.1B params smallest 19.71 GiB across 3 quants apache-2.0 gemma-4-26B-A4B-it-ultra-uncensored-heretic 25.8B params smallest 7.72 GiB across 43 quants apache-2.0 Qwen2.5-Coder-32B-Instruct 32.8B params smallest 8.41 GiB across 36 quants apache-2.0 Qwen3.5-27B 27.8B params smallest 7.98 GiB across 58 quants apache-2.0 jina-embeddings-v5-text-small 0.6B params smallest 0.19 GiB across 56 quants cc-by-nc-4.0 whisper-large-v3-turbo 0.8B params smallest 0.27 GiB across 31 quants apache-2.0 Qwen2.5-Coder-14B-Instruct 14.8B params smallest 4.38 GiB across 35 quants apache-2.0 gemma-4-E2B-it-qat-q4_0-unquantized 5.1B params smallest 3.12 GiB across 1 quants apache-2.0 Qwen3-4B-Instruct-2507 4.0B params smallest 1.01 GiB across 47 quants license unknown ## KV CACHE DTYPES f16 16 bits per element q8_0 8.5 bits per element (not 8 — carries a per-block scale) q4_0 4.5 bits per element (not 4) Quantizing the cache is roughly a 2x lever on the dominant term at long context. ## MACHINE INTERFACE - https://ossmodeldb.org/api/v1 — discovery document, keyless, CORS * - https://ossmodeldb.org/api/v1/models?modality=&limit= - https://ossmodeldb.org/api/v1/fit?model={slug}&hw={slug}&ctx={n}&kv={f16|q8_0|q4_0} - https://ossmodeldb.org/api/gguf?url={hf-gguf-url} — parse any Hub GGUF header live over HTTP Range - https://ossmodeldb.org/data — full dataset, JSONL or JSON ## GUIDES - https://ossmodeldb.org/guides/multi-gpu-for-local-ai — llama.cpp's default layer split runs your GPUs one at a time, so two 3090s give you 48GB at roughly one card's token rate. Here is the arithmetic. - https://ossmodeldb.org/guides/vision-language-models-locally — VLMs need a separate mmproj file alongside the main GGUF, and one image can burn 4096 context tokens — 576 MiB of KV cache on Qwen3-VL-8B. - https://ossmodeldb.org/guides/apple-silicon-for-local-ai — Unified memory buys capacity no consumer GPU can match, but prefill is 3-4x slower and the same chip name ships at two bandwidths 33% apart. - https://ossmodeldb.org/guides/choosing-a-quantization — Q4_K_M is a good default and Q5_K_M or Q6_K if it fits — but the per-model data needed to do better than that is not published anywhere. - https://ossmodeldb.org/guides/choosing-a-gpu — Capacity decides what you can run; memory bandwidth decides how fast. Teraflops barely matter for generation, and your current card may already be enough. - https://ossmodeldb.org/guides/choosing-a-runtime — Most of these are layers over llama.cpp, not rivals to it. Pick a wrapper for ergonomics; switch to vLLM only when you serve concurrent requests. - https://ossmodeldb.org/guides/what-quantization-costs-you — On the one model with public measurements, Q4_K_M costs 2.4% perplexity. The step from Q4 to Q3 costs more than everything above it combined. - https://ossmodeldb.org/guides/local-embeddings-and-rag — Retrieval models run from 91 MiB to 1.1 GiB and work fine on CPU. Your real costs are vector storage, the reranker pass, and MTEB's doubled memory field. - https://ossmodeldb.org/guides/moe-and-cpu-offload — MoE models offload to system RAM well because all experts must be resident but only a few are read per token. Here is the arithmetic and the llama.cpp flags. - https://ossmodeldb.org/guides/local-speech-tts-and-asr — Voice models are 60 MB to 3 GB, not 30 GB. Almost any laptop runs them, but the RTF figures you've read were measured on an H200 at batch 64. - https://ossmodeldb.org/guides/running-your-first-model — Pick a runtime, pick a GGUF file that actually fits your VRAM, and know how to tell whether it fit. - https://ossmodeldb.org/guides/when-it-does-not-fit — Hard OOM, mid-context OOM, and the silent spill that runs 5-20x slow are different problems with different fixes. Diagnose first, then work down the ladder. - https://ossmodeldb.org/guides/local-image-and-video-generation — A diffusion model is a pipeline of four models, not one file, and the headline parameter count covers only the denoiser. Here is the real memory math. - https://ossmodeldb.org/guides/kv-cache-and-context — KV cache is sized by num_key_value_heads and allocated up front. For Gemma-3, DeepSeek, and Qwen3.5/3.6 the flat formula overstates it 4x to 71x. ## DEFINITIONS - KV cache: Stored attention keys and values for tokens already processed, so they aren't recomputed each step. - GGUF: The single-file model format used by llama.cpp, carrying weights plus all metadata needed to run them. - Quantization: Storing weights at reduced precision to shrink a model, trading some quality for memory. - Sliding-window attention: Layers that attend only to a fixed recent window rather than the whole context. - MLA (multi-head latent attention): An attention variant that caches a compressed latent instead of full keys and values. - GQA (grouped-query attention): Multiple query heads share a smaller number of key/value heads, shrinking the cache. - Mixture of experts (MoE): Only a few of many feed-forward experts run per token, but all must be held in memory. - Active parameters: The subset of an MoE model's parameters involved in producing any single token. - Memory bandwidth: How fast a processor can read from memory — the main determinant of token generation speed. - Prefill (prompt processing): Processing the input prompt before generation starts. Compute-bound, unlike generation. - KV cache quantization: Storing the attention cache at reduced precision, roughly halving its size for little quality cost. - CPU offload: Keeping some layers in system RAM when a model does not fit entirely in VRAM. - Importance matrix (imatrix): Calibration data guiding which weights to preserve at higher precision during quantization. - Roofline model: Predicting performance from whichever resource saturates first — here, memory bandwidth. - Context window: The maximum number of tokens a model can attend to at once, including both your input and its output. - Tokens per second: How fast a model generates. Bound by memory bandwidth, not by arithmetic throughput. - Flash attention: An attention implementation that avoids materializing the full attention matrix, saving memory and time. - mmproj (multimodal projector): The separate file that lets a vision-language model actually see images. - Tensor parallelism: Splitting each layer across several GPUs so they compute together, rather than taking turns. - Sharding: Splitting one large model file into numbered parts. You need all of them. - Unified memory: One memory pool shared by CPU and GPU, as on Apple Silicon and some AMD systems. ## WHAT WE DO NOT CLAIM - All speed figures WE compute are MODELED and carry the measured error band above. Figures labelled measured were measured by third parties and are reproduced with per-row attribution; we have not verified those runs and we aggregate them rather than quoting any single one. Nothing on this site has been measured by us. - Image, video and speech models have exact memory but no throughput data. No public source has measured iterations per second or real-time factor on consumer hardware, and published blog figures contradict each other by several times over on identical hardware. - Where a model's sliding-window layout cannot be resolved, we publish no KV total rather than a confidently wrong one. - Comparison multipliers against named competitors are only published when we have queried that competitor and quoted what it actually returned. ## METHOD AND CORRECTIONS - https://ossmodeldb.org/methodology — how each figure is produced - https://ossmodeldb.org/corrections — every error we have found in our own data, and what changed - https://ossmodeldb.org/status — coverage, gaps and known limitations