Best local reasoning models

Models that reason before answering. They spend far more output tokens per question, which makes generation speed and context budget matter more than usual.

From the file· live filter over real dataFrom the file· 40 models

How this is ranked

Filtered to models their publishers identify as reasoning or thinking variants, ranked by downloads. Note that a reasoning model's effective speed is lower than its tokens-per-second suggests, because most of those tokens are thinking rather than answering.

Best local reasoning models

#ModelParamsSmallest quantSmallest quant
1Qwen3-30B-A3B-Thinking-2507MoE30.5B7.05 GiB7.05 GiB
2ThinkingCap-Qwen3.6-27B27.4B9.30 GiB9.30 GiB
3Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking39.5B13.76 GiB13.76 GiB
4MiniCPM5-1B-Claude-Opus-Fable5-Thinking1.1B0.43 GiB0.43 GiB
5MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking1.1B0.64 GiB0.64 GiB
6DeepSeek-R1-0528-Qwen3-8B8.2B2.11 GiB2.11 GiB
7Qwen3.6-27B-Heretic2-Uncensored-Finetune-Thinking27.4B9.77 GiB9.77 GiB
8Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-DistilledMoE36.0B12.34 GiB12.34 GiB
9Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1633.0B17.23 GiB17.23 GiB
10DeepSeek-R1-Distill-Qwen-7B7.6B2.59 GiB2.59 GiB
11DeepSeek-R1-Distill-Llama-70B70.6B14.79 GiB14.79 GiB
12DeepSeek-R1-Distill-Qwen-14B14.8B4.38 GiB4.38 GiB
13Ministral-3-14B-Reasoning-251213.9B3.21 GiB3.21 GiB
14DeepSeek-R1-Distill-Qwen-1.5B1.8B0.65 GiB0.65 GiB
15gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking31.3B6.66 GiB6.66 GiB
16QwQ-32B32.8B7.16 GiB7.16 GiB
17Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoE36.0B6.97 GiB6.97 GiB
18Llama3.3-8B-Instruct-Thinking-Heretic-Uncensored-Claude-4.5-Opus-High-Reasoning8.0B1.88 GiB1.88 GiB
19Qwen3-4B-Thinking-25074.0B1.01 GiB1.01 GiB
20Ministral-3-3B-Reasoning-25124.3B0.90 GiB0.90 GiB
21VibeThinker-3B3.1B1.19 GiB1.19 GiB
22DeepSeek-R1-Distill-Qwen-32B32.8B8.41 GiB8.41 GiB
23Huihui-ThinkingCap-Qwen3.6-27B-abliterated27.4B10.12 GiB10.12 GiB
24DeepSeek-R1-Distill-Llama-8B8.0B2.02 GiB2.02 GiB
25DeepSeek-R1-Distill-Llama-70B-abliterated70.6B14.29 GiB14.29 GiB
26Qwen3.5-9B-GLM5.1-Distill-v19.7B4.31 GiB4.31 GiB
27Wan2.2-Distill-Models14.3B4.95 GiB4.95 GiB
28MN-12B-Mag-Mell-R112.2B3.85 GiB3.85 GiB
29Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-Thinking9.4B2.55 GiB2.55 GiB
30Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-Hightop12.1B4.60 GiB4.60 GiB
31DeepSeek-R1MoE685B124.38 GiB124.38 GiB
32Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING9.4B2.28 GiB2.28 GiB
33Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED9.4B2.28 GiB2.28 GiB
34Qwen3-53B-A3B-2507-THINKING-TOTAL-RECALL-v2-MASTER-CODERMoE53.0B10.22 GiB10.22 GiB
35DeepSeek-R1-Distill-Llama-70B-heretic70.6B14.29 GiB14.29 GiB
36Phi-4-mini-reasoning3.8B1.06 GiB1.06 GiB
37Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled27.8B5.80 GiB5.80 GiB
38Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-heretic-v126.9B5.80 GiB5.80 GiB
39Qwen3-Next-80B-A3B-ThinkingMoE81.3B15.44 GiB15.44 GiB
40Phi-4-reasoning-plus14.7B3.33 GiB3.33 GiB
Spec sheetFrom the fileFrom the filewhat these mean