The Gauntlet — Ollama Model Leaderboard
Overview Code Generation Tool Calling RAG / Q&A Reasoning / Math | About Design Doc Scaling Curves

Four categories. One model per size that beats them all.

Every model tested across code generation, tool-calling, RAG, and reasoning — same seventeen models, four VRAM tiers, six-tier escalating gauntlet per category. Here's who came out on top overall, and what their shape actually looks like.

A model's total score can hide a lot — a narrow specialist and a genuine all-rounder can land on similar numbers. These four are the highest combined scorers in their VRAM tier, summed across all four categories out of 120 possible points. The shape of each polygon tells you whether that total came from consistency or from a few standout categories carrying a weaker one.

Sub-4GB120/120
gemma4:e2b
30303030 Code GenTool CallingRAG / Q&AReasoning
4-8GB105/120
qwen3:8b
20302530 Code GenTool CallingRAG / Q&AReasoning
8-12GB115/120
gemma4:12b-it-qat
30302530 Code GenTool CallingRAG / Q&AReasoning
12GB+114/120
gemma4:26b
29302530 Code GenTool CallingRAG / Q&AReasoning

Worth noticing: the 4-8GB champion, qwen3:8b, has the most lopsided shape of the four — a real dip in code generation, made up for by perfect scores in tool-calling and reasoning. Every other tier's champion is close to a full square. That asymmetry is the whole reason this project exists: a single "best" number would have hidden it completely.