Best local AI models for 6GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 6GB card gives you about 5.58 GiB to work with after driver overhead. 21 indexed models fit at 32K context — the largest being nomic-embed-code at 7.1B parameters in IQ3_XS.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 6GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| jina-embeddings-v5-text-small | embeddings | F16 | 596M | 5.39 GiB | 0.19 GiB |
| embeddinggemma-300m-qat-q8_0-unquantized | embeddings | Q8_0 | 303M | 1.22 GiB | 4.36 GiB |
| KaLM-embedding-multilingual-mini-instruct-v2.5 | embeddings | Q8_0 | 494M | 1.65 GiB | 3.93 GiB |
| nomic-embed-text-v1.5 | embeddings | F32 | 137M | 2.40 GiB | 3.18 GiB |
| jina-embeddings-v5-text-nano | embeddings | F16 | 212M | 2.30 GiB | 3.28 GiB |
| all-MiniLM-L6-v2 | embeddings | F32 | 23M | 1.13 GiB | 4.45 GiB |
| mxbai-embed-xsmall-v1 | embeddings | F32 | 24M | 1.13 GiB | 4.45 GiB |
| nomic-embed-text-v2-moeMoE | embeddings | F32 | 475M | 2.58 GiB | 3.00 GiB |
| bge-m3 | embeddings | Q8_0 | 567M | 4.37 GiB | 1.21 GiB |
| snowflake-arctic-embed-l-v2.0 | embeddings | BF16 | 568M | 4.86 GiB | 0.72 GiB |
| jina-reranker-v1-tiny-en | embeddings | F16 | 33M | 1.01 GiB | 4.57 GiB |
| Qwen3.5-9B-DFlash | embeddings | BF16 | 1.3B | 3.47 GiB | 2.11 GiB |
| Qwen3-Embedding-0.6B | embeddings | F16 | 596M | 5.39 GiB | 0.19 GiB |
| gte-small | embeddings | Q8_0 | 33M | 1.36 GiB | 4.22 GiB |
| nomic-embed-text-v1 | embeddings | F32 | 137M | 2.40 GiB | 3.18 GiB |
| LFM2.5-Embedding-350M | embeddings | F16 | 354M | 1.82 GiB | 3.76 GiB |
| nomic-embed-code | embeddings | IQ3_XS | 7.1B | 5.50 GiB | 0.08 GiB |
| LCO-Embedding-Omni-3B-2605 | embeddings | Q8_0 | 4.7B | 5.30 GiB | 0.28 GiB |
| bge-reranker-v2-m3 | embeddings | Q8_0 | 568M | 4.37 GiB | 1.21 GiB |
| LFM2.5-ColBERT-350M | embeddings | F16 | 353M | 1.82 GiB | 3.76 GiB |
| gte-large | embeddings | Q8_0 | 335M | 4.11 GiB | 1.47 GiB |
This page models a generic 6GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.