Best models for an 8GB card

The most capable models whose largest fitting quantization lands inside 8GB at a usable context — the most common consumer configuration there is.

From the file· live filter over real dataFrom the file· 40 models

How this is ranked

Ranked by parameter count among models with a quantization under 6.5 GiB, leaving room for the cache and runtime overhead. Bigger parameters at a lower quantization generally beats a smaller model at a higher one, down to about 4 bits.

Best models for an 8GB card

#ModelParamsSmallest quantParameters
1Darwin-36B-OpusMoE34.7B4.75 GiB34.7B
2Aurora-Code-1MoE34.7B5.98 GiB34.7B
3OmniAtlas-Qwen3-30B-A3B31.7B5.98 GiB31.7B
4Qwen3-Omni-30B-A3B-Captioner31.7B5.98 GiB31.7B
5Skyfall-31B-v4.231.4B6.39 GiB31.4B
6Skyfall-31B-v4.2-heretic31.4B6.39 GiB31.4B
7Huihui-GLM-4.7-Flash-abliteratedMoE31.2B5.78 GiB31.2B
8GLM-4.7-Flash-DerestrictedMoE31.2B5.78 GiB31.2B
9Salience-1.5-FlashMoE31.1B5.98 GiB31.1B
10Huihui-Qwen3-VL-30B-A3B-Instruct-abliteratedMoE31.1B5.98 GiB31.1B
11Qwen3-VL-30B-A3B-Instruct-abliterated-v131.1B5.98 GiB31.1B
12Qwen3-VL-30B-A3B-Thinking-abliterated-v131.1B5.98 GiB31.1B
13TildeOpen-30B-Instruct-LV30.7B6.43 GiB30.7B
14Huihui-Qwen3-Omni-30B-A3B-Instruct-abliterated30.5B5.98 GiB30.5B
15Huihui-Qwen3-Coder-30B-A3B-Instruct-abliteratedMoE30.5B5.98 GiB30.5B
16Huihui-Qwen3-30B-A3B-Instruct-2507-abliteratedMoE30.5B5.98 GiB30.5B
17Qwen3-Coder-30B-A3B-Instruct-RTPurboMoE30.5B5.98 GiB30.5B
18Qwen3-30B-A3B-Gemini-Pro-High-Reasoning-2507-ABLITERATED-UNCENSOREDMoE30.5B5.98 GiB30.5B
19Qwen3-30B-A3B-Thinking-2507-Claude-4.5-Sonnet-High-Reasoning-DistillMoE30.5B5.98 GiB30.5B
20Qwen3-30B-A3B-YOYO-V5MoE30.5B5.98 GiB30.5B
21MiroThinker-v1.0-30BMoE30.5B5.98 GiB30.5B
22Huihui-Qwen3-30B-A3B-Thinking-2507-abliteratedMoE30.5B5.98 GiB30.5B
23Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoE30.0B6.07 GiB30.0B
24Huihui-granite-4.1-30b-abliterated28.9B5.73 GiB28.9B
25granite-4.1-30b-heretic28.9B5.73 GiB28.9B
26spoomplesmaxx-v2.1-30B28.9B5.73 GiB28.9B
27medgemma-27b-it28.8B5.83 GiB28.8B
28translategemma-27b-it28.8B5.83 GiB28.8B
29Gemma-3-27B-MeditronFO28.8B6.26 GiB28.8B
30Qwen3.5-28BMoE28.7B5.77 GiB28.7B
31Qwen3.6-28BMoE28.2B5.77 GiB28.2B
32Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled27.8B5.80 GiB27.8B
33Huihui-Qwen3.5-27B-abliterated27.8B5.80 GiB27.8B
34Qwen3.5-27B-Derestricted27.8B5.80 GiB27.8B
35Qwen3.5-27B-Engineer-Deckard-Gemini27.7B5.80 GiB27.7B
36gemma-3-27b-it27.4B6.06 GiB27.4B
37gemma-3-27b-it-qat-q4_0-unquantized27.4B6.06 GiB27.4B
38AtomicGPT-gemma3-27b27.4B5.83 GiB27.4B
39Nidum-Gemma-3-27B-it-Uncensored27.4B5.83 GiB27.4B
40gemma-3-27b-it-abliterated-refined-vision27.4B5.83 GiB27.4B
Spec sheetFrom the fileSpec sheetwhat these mean