Best local AI models for 128GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 128GB card gives you about 119.04 GiB to work with after driver overhead. 26 indexed models fit at 32K context — the largest being Nemotron-3-Embed-8B-BF16 at 8.0B parameters in BF16.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 128GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| jina-embeddings-v5-text-small | embeddings | F16 | 596M | 5.39 GiB | 113.65 GiB |
| embeddinggemma-300m-qat-q8_0-unquantized | embeddings | Q8_0 | 303M | 1.22 GiB | 117.82 GiB |
| KaLM-embedding-multilingual-mini-instruct-v2.5 | embeddings | Q8_0 | 494M | 1.65 GiB | 117.39 GiB |
| nomic-embed-text-v1.5 | embeddings | F32 | 137M | 2.40 GiB | 116.64 GiB |
| jina-embeddings-v5-text-nano | embeddings | F16 | 212M | 2.30 GiB | 116.74 GiB |
| all-MiniLM-L6-v2 | embeddings | F32 | 23M | 1.13 GiB | 117.91 GiB |
| mxbai-embed-xsmall-v1 | embeddings | F32 | 24M | 1.13 GiB | 117.91 GiB |
| nomic-embed-text-v2-moeMoE | embeddings | F32 | 475M | 2.58 GiB | 116.46 GiB |
| bge-m3 | embeddings | Q8_0 | 567M | 4.37 GiB | 114.67 GiB |
| snowflake-arctic-embed-l-v2.0 | embeddings | F32 | 568M | 5.90 GiB | 113.14 GiB |
| jina-reranker-v1-tiny-en | embeddings | F16 | 33M | 1.01 GiB | 118.03 GiB |
| Qwen3.5-9B-DFlash | embeddings | BF16 | 1.3B | 3.47 GiB | 115.57 GiB |
| Qwen3-Embedding-0.6B | embeddings | F16 | 596M | 5.39 GiB | 113.65 GiB |
| gte-small | embeddings | Q8_0 | 33M | 1.36 GiB | 117.68 GiB |
| Qwen3-Embedding-8B | embeddings | F16 | 7.6B | 19.43 GiB | 99.61 GiB |
| nomic-embed-text-v1 | embeddings | F32 | 137M | 2.40 GiB | 116.64 GiB |
| LFM2.5-Embedding-350M | embeddings | F16 | 354M | 1.82 GiB | 117.22 GiB |
| Qwen3-Embedding-4B | embeddings | F16 | 4.0B | 12.81 GiB | 106.23 GiB |
| Nemotron-3-Embed-8B-BF16 | embeddings | BF16 | 8.0B | 19.91 GiB | 99.13 GiB |
| nomic-embed-code | embeddings | F32 | 7.1B | 28.95 GiB | 90.09 GiB |
| LCO-Embedding-Omni-3B-2605 | embeddings | Q8_0 | 4.7B | 5.30 GiB | 113.74 GiB |
| qwen-indic-v1 | embeddings | F16 | 7.6B | 19.43 GiB | 99.61 GiB |
| Octen-Embedding-4B | embeddings | F16 | 4.0B | 12.81 GiB | 106.23 GiB |
| bge-reranker-v2-m3 | embeddings | Q8_0 | 568M | 4.37 GiB | 114.67 GiB |
| LFM2.5-ColBERT-350M | embeddings | F16 | 353M | 1.82 GiB | 117.22 GiB |
| gte-large | embeddings | Q8_0 | 335M | 4.11 GiB | 114.93 GiB |
This page models a generic 128GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.