Best models for an 8GB card
The most capable models whose largest fitting quantization lands inside 8GB at a usable context — the most common consumer configuration there is.
From the file· live filter over real dataFrom the file· 40 models
How this is ranked
Ranked by parameter count among models with a quantization under 6.5 GiB, leaving room for the cache and runtime overhead. Bigger parameters at a lower quantization generally beats a smaller model at a higher one, down to about 4 bits.
Best models for an 8GB card
Best local models for codingBest local reasoning modelsBest models for long contextBest permissively licensed modelsBest models under 4GBLargest models with a runnable quantizationBest local speech recognition modelsBest local text-to-speech modelsBest local vision-language modelsBest local image generation models