Your smallest local model might be your best one - I measured 4 of mine
Article summary
Quick briefing — cleaned from the original RSS feed
I had seven models sitting in Ollama and no idea which one to use for what. So I stopped guessing and measured. 152 generations on a 16 GB laptop. Greedy decoding, deterministic grading - exact number, exact string, JSON field, regex. No LLM judge, so there is no second model's bias to audit. The result No model won every category. task type deepseek-r1:1.5b (1.1 GB) llama3.2:3b (2.0 GB) gemma:2b codellama 7b arithmetic (12) 10/12 2/12 2/12 3/12 extraction (9) 4/9 9/9 7/9 8/9 classification (8)…
1Key Takeaways
- I had seven models sitting in Ollama and no idea which one to use for what.
- Greedy decoding, deterministic grading - exact number, exact string, JSON field, regex.
- No LLM judge, so there is no second model's bias to audit.
- The result No model won every category.
2AIWedia Score
8.6/10
High relevance — worth your attention today
Based on source trust, recency, category impact, and story depth.
3Why it matters
Coding AI shifts how fast software ships and how much human review each change needs. DEV — ML reports that i had seven models sitting in Ollama and no idea which one to use for what.
Explore related
Browse toolsCoding AI news
Explore curated coding ai tools on AIWedia — compare, rank, and launch from our directory.
Full story on DEV — ML
Read full articleHeadlines aggregated via RSS for discovery on AIWedia. Original content © DEV — ML. We link to the source and do not republish full articles.