Best local reasoning models
Models that reason before answering. They spend far more output tokens per question, which makes generation speed and context budget matter more than usual.
From the file· live filter over real dataFrom the file· 40 models
How this is ranked
Filtered to models their publishers identify as reasoning or thinking variants, ranked by downloads. Note that a reasoning model's effective speed is lower than its tokens-per-second suggests, because most of those tokens are thinking rather than answering.
Best local reasoning models
Best local models for codingBest models for an 8GB cardBest models for long contextBest permissively licensed modelsBest models under 4GBLargest models with a runnable quantizationBest local speech recognition modelsBest local text-to-speech modelsBest local vision-language modelsBest local image generation models