Ollama Model Benchmark — Test Method Validation & DOE Design
1. Purpose
Surface well-performing but underrated Ollama models — the "Gemma 4 surprises" — for people hosting local LLMs on consumer hardware, via rigorous testing rather than anecdote or hype. The deliverable is a set of task-specific leaderboards (best-for-code, best-for-chat, best-for-summarization, etc.), each broken out by VRAM tier, so a reader finds the right model for their specific job and their actual hardware, not just "the best model overall."
Design pool structure (revised August 2026): models are grouped into four VRAM tiers based on measured runtime VRAM (ollama ps / nvidia-smi), never download size — confirmed repeatedly necessary (§6, §7 below) since the two can differ by 4–5×. Boundaries are inclusive on the lower bound, exclusive on the upper:
| Tier | Range | Represents |
|---|---|---|
| Sub-4GB | < 4GB | Budget/older consumer GPUs, most laptops |
| 4–8GB | [4, 8) GB | Mainstream consumer GPUs |
| 8–12GB | [8, 12) GB | Enthusiast-tier consumer GPUs (the original single-tier design target) |
| 12GB+ | ≥ 12GB | High-end/workstation-class consumer GPUs |
Each tier gets its own leaderboard and its own Reigning Champion (§3A.6) — a model is never compared across tiers, since "punches above its weight" only means something relative to what else fits the same hardware. This structure was adopted after four already-tested models turned out to span all four tiers (§7), which argued strongly against narrowing back down to one band and discarding that spread.
2. Test Method Validation (TMV)
2.1 Response variables (what we measure)
| Variable | Definition | Unit |
|---|---|---|
| Quality score | Rubric-scored or reference-scored correctness on the task | 0–100 (per rubric) |
| Tokens/sec (generation) | Output tokens generated ÷ generation time | tok/s |
| Time-to-first-token (TTFT) | Latency from request to first output token | ms |
| VRAM footprint | Peak VRAM used during inference (measured, not file size — see §6) | MB |
| RAM footprint | Peak system RAM used (for CPU-offload cases) | MB |
| Context utilization ceiling | Max context length before quality degrades or OOM | tokens |
| Completion rate | % of runs that produce a usable final answer within the num_predict cap | % |
| Failure-mode breakdown | % of failed runs by type — see taxonomy below | % per type |
| System-prompt leakage | Whether internal instructions/persona text appear in user-facing output | bool / rate |
Completion rate, failure-mode breakdown, and leakage were added after two practice runs (§6) surfaced that "hit the token cap" was masking at least three qualitatively different failure causes, and that leaked internal instructions is itself a quality signal a reader would want to know about.
Failure-mode taxonomy (observed across 6 models in practice runs, §6):
- Silent/enumeration loop — thinking trace degenerates into near-verbatim repetition (e.g., endlessly appending test values), cap is hit, response_text is empty. Most severe: zero usable output.
- Can't-commit loop — model drafts a correct, complete answer early, then re-drafts near-identical versions of it repeatedly instead of finalizing. Cap is hit; a usable answer may be buried in the trace but isn't cleanly returned.
- Legitimate overrun — coherent, correct, non-repetitive reasoning that simply needs more budget than allotted. Often recoverable by raising the cap; not a model defect.
- Truncated-but-usable — cap hit mid-answer, but the substantive content (e.g., code) completed before the cutoff; only trailing explanation is lost.
This taxonomy should be assigned per failed run for benchmark reporting, not just recorded as an "empty response" or "truncated" binary.
2.2 Repeatability (same conditions, same operator/machine)
- N = 5 runs minimum per condition cell before computing a mean
- Report mean, standard deviation, and coefficient of variation (CV%) per cell
- Flag any cell with CV% > 15% as unstable — investigate before trusting the mean
- Fixed seed where the model/runtime supports it; if not supported, note as a source of variance rather than ignoring it
2.3 Reproducibility (across conditions that should be neutral)
- Re-run a fixed "control" prompt set on a fixed "control" model/quant combo periodically across the whole test campaign to check for machine drift (thermal throttling, background load, driver updates)
- Log hardware state per run: GPU temp, clock speed, other processes running
2.4 Accuracy (are we measuring the right thing)
- Quality rubrics must be defined and locked before running tests (avoid post-hoc rationalization of scores)
- Where possible use reference-answer scoring (exact match, ROUGE/BLEU for structured tasks) over subjective rubric scoring, and reserve rubric scoring for open-ended tasks
- Blind scoring where feasible (score outputs without knowing which model/config produced them)
2.5 Test environment control
- Single physical machine (your workstation: 9950X / 96GB RAM / RTX 3090 Ti) for all runs in a given campaign — no cross-machine comparison without a documented hardware factor
- Ollama version pinned and logged per campaign
- No other GPU-consuming processes during runs
2.6 Failure handling and safety limits
Discovered necessary during the first practice run, when an uncapped generation ran 11 minutes on a single prompt (see §6):
num_predicthard cap on every generation call (practice value: 4096 tokens) — without this, a single runaway prompt can make a full screening pass across many models impractical- Wall-clock timeout per prompt (practice value: 120s), independent of the token cap, as a second safety net
- A run that hits either limit without producing a usable final answer is marked
completed: falsewith afailure_reason(see taxonomy in §2.1) — this is a distinct outcome from a low quality score, not folded into it - Capture the model's
thinkingoutput separately from its finalresponsewherever the API exposes it — for reasoning-capable models, the thinking trace is often the only way to diagnose why a run failed - Not all models expose thinking separately. Confirmed in §6: qwen3:8b and deepseek-r1:14b populate a distinct
thinkingfield via Ollama's API; phi4-reasoning:14b does not — its reasoning narration is embedded directly inresponse_text. The harness must not assume the field exists; check for it and fall back to treating the full response as mixed reasoning+answer when absent.
3. DOE Structure
3.1 Candidate factors
| Factor | Levels (example) | Notes |
|---|---|---|
| Model family | Llama, Qwen, Mistral, Gemma, Phi, Deepseek, etc. | Categorical — pool bounded to models with an 8–12GB-fitting quant |
| Model size | Whatever sizes fit 8–12GB (e.g., 3B–14B depending on quant) | Categorical; size alone doesn't determine fit — size × quant does |
| Quantization | Q4_K_M (primary/default), Q5_K_M, Q8_0 | Q4_K_M is the realistic default for this VRAM tier — treated as the baseline test condition, others as a depth factor in Stage 2 |
| Context length | 2K, 8K, 16K | Bounded by what's usable at 8–12GB without spilling to CPU offload |
| Task category | Code gen, chat/instruction-following, summarization, reasoning, RAG/retrieval | Primary axis — this is what the leaderboards are organized by |
| Prompt style | Zero-shot, few-shot | Secondary factor, tested within task category |
Task category is promoted to the primary axis because the deliverable is task-specific leaderboards, not one aggregate ranking. Everything else (quant, context, prompt style) is a factor tested within each task's leaderboard.
Resolved: failure triggers are model-idiosyncratic, not prompt-idiosyncratic. Six models were run against the identical 6-prompt set in practice (§6). Each model that failed, failed on a different prompt (qwen3:8b on the code-gen edge-case prompt; phi4-reasoning:14b on the SQL and both chat prompts; deepseek-r1:14b on the same code-gen prompt as qwen3:8b, but for a legitimate-overrun reason rather than a loop). There is no single "hard prompt" shared across models — trigger conditions are a property of each model's own architecture/training, not the prompt wording alone. This means the full task/prompt matrix in Stage 1 doesn't need bespoke "known-hard" prompts per model; a fixed, representative prompt set per task category is sufficient to surface each model's own failure tendencies.
3.2 Recommended staged approach
Stage 1 — Screening design (build the leaderboards) - Fix quantization at Q4_K_M (the realistic default for 8–12GB) and prompt style at zero-shot - Run every model in the 8–12GB pool across all task categories - Output: a first-pass leaderboard per task category, ranked by quality score (with speed/VRAM/completion rate shown alongside, not folded into the score) - This stage is explicitly designed to catch "Gemma 4"-style surprises — any model that ranks well above its size/family reputation gets flagged as a candidate for Stage 2 - Also the stage where failure-mode taxonomy (§2.1) gets assigned per model — a model with a narrow, severe failure mode (e.g., qwen3:8b's silent loop) is a different recommendation than one with a broad but recoverable one (e.g., deepseek-r1:14b's legitimate overrun)
Stage 2 — Focused design (depth on leaderboard candidates)
- For the top 3–5 models per task category (including any surprise standouts), vary quantization (Q4/Q5/Q8) and context length
- Answers the "is it worth the extra precision/VRAM" question per model, per task
- This is where trade-off findings live — e.g., "Model X at Q4 tops the code leaderboard, but drops out of the top 5 once context exceeds 8K"
- If a model showed a failure mode in Stage 1, test whether raising num_predict resolves it (relevant for "legitimate overrun" cases) or whether it persists regardless of budget (relevant for loop cases)
Stage 3 — Confirmation runs - Re-run the leaderboard-topping configs with higher N (e.g., N=10) to confirm findings aren't noise - Only Stage 3-confirmed results get published on the leaderboard as ranked; Stage 1/2 results are labeled exploratory/provisional - Any flagged failure mode from Stage 1 should also get its own confirmation pass (as done manually for qwen3:8b in §6 — 5/5 seeds failed) before being published as a characteristic of that model rather than a one-off
3.3 Effect analysis
- Main effects: does model size alone move quality/speed independent of other factors?
- Interaction effects: does quantization impact differ by model family (e.g., some architectures degrade more at Q4 than others)?
- Present as effect-size tables and interaction plots, not just raw score tables — this is what makes the site useful vs. a leaderboard
3A. Category deep-dive pilot: the "Gauntlet" methodology (code_gen first)
Rather than building shallow, single-difficulty prompt sets across every task category at once, the plan is to fully develop one category's test methodology first — code generation — then replicate whatever works for the other categories. The model is TFLtruck's Ike Gauntlet towing test: a fixed, standardized, escalating stress test, run identically across every competitor, rather than a one-off prompt with a single quality score.
3A.1 Design principles (borrowed directly from Ike Gauntlet)
- Fixed course. Identical harness settings for every model at every tier: temperature 0, fixed seed, a bare prompt with no hand-holding system prompt (removing "driver skill" from the test — what's measured is the model's own capability, not clever prompting).
- Escalating load, not a flat prompt set. A ladder of tiers, each a genuinely harder class of task — not just a longer or wordier version of the same task.
- Two independent measurement axes, not one blended score — see §3A.3.
- Versioned, not static. When every model in the active pool clears a tier cleanly, that tier has stopped differentiating and a harder tier gets added at the top — directly mirroring Ike Gauntlet 2.0 raising the trailer load from 8,000 to 10,800 lbs specifically because the original load "was not slowing these trucks down enough."
- A clear pass/fail headline (how far up the ladder a model climbed) sits on top of the detailed per-tier data, for at-a-glance leaderboard use.
3A.2 Tier ladder (v1, code_gen — fully defined and piloted)
- Trivial self-contained function — the floor/on-ramp, establishes every model can start
- Function with explicit edge-case handling — a known differentiator (this is the exact tier that broke qwen3:8b in practice runs, §6)
- Multi-function program with an interface contract between the pieces
- Bug-fix / edit-existing-code — a different skill than green-field generation (reading before writing)
- Schema-constrained or cross-referenced task — write code against a given spec/schema
- Strict-format output under complexity — output must conform to a JSON/function-call shape while solving something non-trivial, echoing the tool-calling capability-cliff research (§ open research, Aug 2026)
3A.2a Concrete tier definitions — code-gauntlet-v1, tiers 1–3
Stage A (§3A.3) requires a real, checkable pass/fail bar, not a vibe — each tier ships with an executable test suite.
Tier 1 — floor.
Prompt: "Write a Python function is_palindrome(s) that returns True if the string is a palindrome, ignoring case, spaces, and punctuation. Return False otherwise."
Tests: is_palindrome("racecar") == True; is_palindrome("A man a plan a canal Panama") == True; is_palindrome("No lemon, no melon") == True; is_palindrome("hello") == False; is_palindrome("") == True.
Tier 2 — edge-case handling (known differentiator).
Prompt: "Write a Python function that takes a list of integers and returns the second largest unique value. Handle edge cases." (unchanged from the practice-run prompt that broke qwen3:8b in §6 — kept identical so this tier's existing 5/5-confirmed failure data carries forward as-is)
Tests: second_largest_unique([5,3,9,9,7]) == 7; second_largest_unique([5,5,5]) is None; second_largest_unique([1]) is None; second_largest_unique([]) is None; second_largest_unique([2,2,3,3,4,4]) == 3; second_largest_unique([-1,-2,-3]) == -2.
Ruling (Aug 2026, after qwen3.5:9b pilot): the prompt never specifies how to handle insufficient-unique-values cases (return None vs. raise an exception). qwen3.5:9b chose to raise a well-documented ValueError instead of returning None, which fails these tests. Decision: keep the strict None-return contract as-is — per §2.4, rubrics are locked before testing and not loosened post-hoc to accommodate a specific model's result, and the Gauntlet's core principle (§3A.1) is one fixed bar for every model. The failure is genuine, reportable signal (a caller expecting a normal return value would hit an unhandled crash), not a test-design flaw. Going forward, any new tier prompt must state the edge-case contract explicitly (e.g. "return None if fewer than two unique values exist") to avoid this ambiguity recurring — tier 2 itself is grandfathered as historical data.
Tier 3 — multi-component interface contract.
Prompt: "Implement a simple library management system in Python with two components: (1) a Book class with attributes title, author, isbn, and available (bool, defaults to True); (2) a Library class with methods add_book(book), checkout_book(isbn) (marks unavailable and returns True, or returns False if missing/already checked out), return_book(isbn) (marks available and returns True, or False if missing), and find_by_author(author) (returns a list of matching Book objects). Write complete, working code for both classes."
Tests: add 3 books (2 sharing an author) → checkout one → verify unavailable → checkout same isbn again → expect False → return it → verify available → find_by_author returns the correct 2 books → checkout a nonexistent isbn → expect False.
Tier 4 — bug-fix (reading before writing). Prompt gives a broken function plus its intended docstring spec, and asks for a corrected version:
def apply_discount(price, discount_percent):
"""
Applies a percentage discount to a price.
discount_percent is 0-100 (e.g., 20 means 20% off).
Returns the discounted price, rounded to 2 decimal places.
Raises ValueError if discount_percent is negative or greater than 100.
Raises ValueError if price is negative.
"""
if discount_percent < 0:
return price
discount = price * discount_percent
new_price = price - discount
return new_price
"The function above has one or more bugs relative to its docstring. Fix it so it correctly implements the documented behavior. Return the complete corrected function."
Bugs seeded: missing /100 in the discount calculation; no validation for discount_percent < 0 or > 100; no validation for negative price; missing rounding.
Tests: apply_discount(100, 20) == 80.0; apply_discount(50, 0) == 50.0; apply_discount(99.99, 10) == 89.99; apply_discount(100, -5) raises ValueError; apply_discount(100, 150) raises ValueError; apply_discount(-10, 20) raises ValueError.
Tier 5 — schema-constrained cross-referencing.
Prompt gives two small datasets and an exact output contract:
"Given orders (list of dicts with keys order_id, customer_id, amount) and customers (list of dicts with keys customer_id, name), write a function top_customers(orders, customers, n) that returns the top n customers by total order amount. Each result must be a dict with exactly the keys name, total_spent, order_count, sorted descending by total_spent."
Test data: 5 orders across 3 customers (Alice: $250/2 orders, Bob: $170/2 orders, Carol: $75/1 order — no ties, to keep expected output deterministic).
Tests: top_customers(orders, customers, 2) returns Alice then Bob with exact key names and correct aggregates; n=10 returns all 3 in order; empty orders returns []. A model that uses different key names (e.g. customer_name instead of name) fails Stage A even if the logic is otherwise correct — that's the point of this tier.
Tier 6 — strict format under complexity (tool-call-shaped).
Prompt: "You are responding as a function call. Given this inventory list [4 items with name/quantity/price] and a low-stock threshold of 10, respond with ONLY a JSON object — no explanation, no markdown code fences — matching exactly this schema: {"low_stock_items": [...item names with quantity < threshold...], "total_value": <float, sum of quantity×price, rounded to 2 decimals>, "most_expensive_item": <name of highest unit-price item>}."
Grading is two-tier: strict pass — the raw response parses as JSON with zero cleanup (tests literal format discipline); lenient pass — parses only after stripping a markdown code fence. Both are recorded; only strict counts for Stage A, since format discipline under instruction is exactly what this tier measures (this is the tier most directly informed by the tool-calling capability-cliff research — small chat-tuned models often can't resist adding prose even when explicitly told not to).
Tests (on parsed content): low_stock_items == ["Widget", "Gizmo"]; total_value == 289.85; most_expensive_item == "Gizmo".
3A.2b Pilot confirmation result — gemma4:e2b, code-gauntlet-v1, full ladder (August 2026)
gemma4:e2b run against all six tiers, N=5 seeds each (temperature 0, num_predict cap 4096, timeout 120s), executed and scored against the real test suites in §3A.2a — not read for plausibility, actually run and asserted against:
| Tier | Stage A result | Avg tokens | Avg time |
|---|---|---|---|
| 1 — floor | 5/5 | 1,194 | 7.70s |
| 2 — edge-case (known differentiator) | 5/5 | 1,814 | 9.56s |
| 3 — multi-class interface contract | 5/5 | 1,824 | 9.72s |
| 4 — bug-fix | 5/5 | 1,204 | 6.32s |
| 5 — schema cross-reference | 5/5 | 2,066 | 10.79s |
| 6 — strict JSON format | 5/5 (all strict, none needed fence-stripping) | 827 | 4.45s |
30/30 — a complete, clean sweep of the entire ladder, including tier 6, which was specifically designed to be the tier most likely to trip up a chat-tuned small model (zero tolerance for explanatory prose or markdown fences). gemma4:e2b never once added commentary, never leaked a fence, and got every value correct.
This tells us two things, not one: gemma4:e2b is a genuinely strong, well-rounded coder at every sub-skill tested (generation, edge-cases, multi-component design, bug-fixing, schema discipline, format discipline) — and the current six-tier ladder has not yet found gemma4's ceiling. Per §3A.1, tier retirement is a pool-wide decision, not a single-model one — before concluding the ladder needs to go higher, the other confirmed-clean models from §6 (qwen3.5:9b, deepseek-r1:14b, mistral-small3.2:24b) should run the same six tiers. If they also clear cleanly, that's the trigger to add tier 7+; if any of them falls short of gemma4's 30/30, that's a real differentiation the leaderboard should report before the ladder needs to grow at all.
This is the first result to clear the confirmation bar set in §3A.6 — gemma4:e2b is accordingly named Reigning Champion, code-gauntlet-v1 (full ladder, 30/30), pending any model that also clears all six tiers with better quality/time, or a future model that survives a tier gemma4 doesn't when the ladder extends.
3A.2c Harness bug found and fixed during the qwen3.5:9b pilot (August 2026)
Running qwen3.5:9b through the same ladder surfaced a real bug in the harness's code-extraction logic, not a model failure — worth recording since it changes how every prior and future result should be trusted.
Bug 1 — "pick the longest fenced block" silently drops real definitions. qwen3.5:9b's response put the actual function in one code fence and a longer worked-example/demo block in a second fence; extract_code's original max(blocks, key=len) grabbed the demo block and discarded the real definition, producing a 5/5 NameError failure that had nothing to do with the model's actual code.
Fix attempt 1 — concatenate all fenced blocks initially seemed to resolve it, but surfaced a second issue: a model's own demo code can contain an intentional, uncaught exception (e.g. a comment-documented "raises ValueError" example call with no try/except around it). Concatenating and exec-ing the whole thing top-to-bottom means the demo code's crash kills the run before our own test suite gets to call the real function — again failing for the wrong reason.
Fix attempt 2 — AST-filtered extraction (adopted). extract_code now parses the concatenated fenced blocks with ast.parse, keeps only declarative nodes (FunctionDef, AsyncFunctionDef, ClassDef, Import, ImportFrom, Assign, AnnAssign), and discards everything else — bare print() calls, demo try/except blocks, any top-level executable statement. This isolates the real definition regardless of what demo/example code a model wraps around it, and is now the harness's standard extraction method going forward. Falls back to the raw concatenated text if AST parsing fails (e.g. genuinely malformed code — a real Stage A failure that should surface as such).
Practical implication: any prior result graded with the old "longest block" extraction should be treated as unverified until re-run with the AST-filtered version. gemma4:e2b's 30/30 (§3A.2b) was run under the original single-block extraction; since every gemma4 response only ever contained one clean code fence, re-extraction would produce an identical result, but this should be spot-checked, not assumed, before the record is cited elsewhere.
3A.2d Pilot confirmation result — qwen3.5:9b, code-gauntlet-v1, full ladder (August 2026)
| Tier | Stage A result | Avg tokens | Avg time |
|---|---|---|---|
| 1 — floor | 5/5 | 781 | ~7.5s |
| 2 — edge-case (known differentiator) | 0/5 (confirmed, correct diagnosis — see ruling above) | 730 | ~6.4s |
| 3 — multi-class interface contract | 5/5 | 1,246 | 10.2s |
| 4 — bug-fix | 5/5 | 702 | 5.9s |
| 5 — schema cross-reference | 5/5 | 846 | 7.0s |
| 6 — strict JSON format | 5/5 (all strict) | 860 | 7.1s |
25/30. qwen3.5:9b is the first model to actually fail a tier in the full-ladder pilot — and, notably, the only tier it fails is the exact same one that broke qwen3:8b in §6, though for a completely different, non-catastrophic reason (a defensible but non-conforming edge-case contract, not a repetition loop). Everything else — including tier 6's strict-format discipline, which was the tier expected to be hardest for chat-tuned models — was a clean 5/5.
Notably, qwen3.5:9b's token counts were perfectly deterministic within each tier (identical eval_count across all 5 seeds) — a different reproducibility profile than gemma4:e2b, whose token counts varied slightly seed-to-seed. Worth watching whether this holds up as a general qwen3.5 characteristic across other tasks.
This is real, useful differentiation for the leaderboard: gemma4:e2b remains Reigning Champion at 30/30; qwen3.5:9b is a strong, fast, mostly-clean second at 25/30, with a documented, specific, non-fatal edge-case-contract quirk rather than a severe failure mode.
3A.2e Harness gap found during the deepseek-r1:14b pilot: missing thinking-trace capture
The Code Gauntlet harness (built fresh from the practice-run harness) omitted thinking_text capture — a regression from §2.6's explicit requirement that reasoning-model thinking traces be captured separately, since they're often the only way to diagnose why a run failed. This went unnoticed through gemma4:e2b and qwen3.5:9b (neither is a thinking model in a way that mattered here) and surfaced only once deepseek-r1:14b produced a tier-5 failure with a completely empty response_text and no way to see what happened. Fixed: run_prompt now captures and returns thinking, stored per-run alongside response_text.
3A.2f Pilot confirmation result — deepseek-r1:14b, code-gauntlet-v1, full ladder (August 2026)
| Tier | Stage A result | Avg tokens | Avg time |
|---|---|---|---|
| 1 — floor | 5/5 | ~992 | ~13.6s |
| 2 — edge-case (known differentiator) | 5/5 | ~1,977 | ~26.8s (one seed hit the 4096 cap exactly but still passed — real definition completed before the cutoff, recovered correctly by the AST-filter fix in §3A.2c) |
| 3 — multi-class interface contract | 5/5 | ~1,592 | ~19.6s |
| 4 — bug-fix | 5/5 | ~2,590 | ~32.2s |
| 5 — schema cross-reference | 0/5 | 4,096 (every seed hit the cap exactly) | 50.9s |
| 6 — strict JSON format | 5/5 (all strict) | 512 | 6.3s |
25/30 — same score as qwen3.5:9b, but a categorically different failure. Diagnosis of the tier-5 failure (§3A.2e's fix in hand): the thinking_text for a representative run was 17,162 characters of coherent, methodical reasoning — walking through the test cases, even proactively considering tie-breaking scenarios the test doesn't require — that simply never reached the point of writing the actual function definition before the 4096-token cap cut it off. This is a clean legitimate overrun (§2.1): no loop, no incoherence, just a model that reasons more thoroughly than our fixed budget accommodates for this specific tier's complexity.
This is a genuinely important finding for the leaderboard, distinct from qwen3.5:9b's result despite the identical 25/30 score: a raw pass count alone would treat these as equivalent, but they aren't. qwen3.5:9b's failure is a fast, cheap, deliberate design choice a caller could work around by catching the exception. deepseek-r1:14b's failure is a capacity problem — the model needed more thinking budget than our fixed cap allowed, and there's no way to work around that within a single call short of raising num_predict. This is exactly why §2.1's failure-mode taxonomy exists instead of a single pass/fail number: the two 25/30s tell a reader completely different things about what to expect.
Also worth flagging: deepseek-r1:14b was markedly slower throughout (12–51s per run vs. gemma4/qwen3.5's 4–11s range) — consistent with its known verbose-reasoning profile from §6. A reader choosing based on raw wall-clock latency, not just pass/fail, would weight this heavily.
3A.2g Pilot confirmation result — mistral-small3.2:24b, code-gauntlet-v1, full ladder (August 2026)
Run to complete the four-model pool picture. Out-of-pool for the actual leaderboard (15GB measured VRAM, above the 8–12GB target tier — §1, §6) but included for orientation, same as earlier practice runs.
| Tier | Stage A result | Avg tokens | Avg time |
|---|---|---|---|
| 1 — floor | 5/5 | 295 | 5.3s |
| 2 — edge-case (known differentiator) | 5/5 | 411 | 7.3s |
| 3 — multi-class interface contract | 5/5 | 840 | 14.8s |
| 4 — bug-fix | 5/5 | 257 | 4.6s |
| 5 — schema cross-reference | 5/5 | 669 | 11.7s |
| 6 — strict JSON format | 0/5 | 46 | ~1.0s |
25/30 — third model in a row to land on exactly 25/30, and third categorically different failure. Tier 6 failed for two independent reasons in the same response, both confirmed from the raw output:
- Format discipline: the response was wrapped in a
```jsonfence despite an explicit instruction not to use one (strict_format_pass: False— needed fence-stripping to parse at all). This is exactly the failure tier 6 was designed to probe. - Silent arithmetic error:
low_stock_itemsandmost_expensive_item(filtering/comparison logic) were both correct, buttotal_value(a four-term sum) came out to 267.25 instead of 289.85 — every seed, identically. Because the model followed "no explanation" faithfully, there's no visible reasoning to show which term it miscalculated; the error is silent and undiagnosable beyond "wrong."
This is a genuinely distinct failure profile from the other three models' tier-6 behavior (all clean) and from mistral's own performance everywhere else (perfectly clean, fast, terse — 4.6–14.8s across tiers 1–5). A model can have excellent format instincts or excellent arithmetic and still fail this tier on the other axis; mistral-small3.2:24b failed on both simultaneously here.
3A.2h Pool summary — four models, six tiers (August 2026)
| gemma4:e2b | qwen3.5:9b | deepseek-r1:14b | mistral-small3.2:24b | |
|---|---|---|---|---|
| Tier 1 — floor | 5/5 | 5/5 | 5/5 | 5/5 |
| Tier 2 — edge-case | 5/5 | 0/5 (contract choice) | 5/5 | 5/5 |
| Tier 3 — interface | 5/5 | 5/5 | 5/5 | 5/5 |
| Tier 4 — bug-fix | 5/5 | 5/5 | 5/5 | 5/5 |
| Tier 5 — schema | 5/5 | 5/5 | 0/5 (legitimate overrun) | 5/5 |
| Tier 6 — strict format | 5/5 | 5/5 | 5/5 | 0/5 (fence + arithmetic) |
| Total | 30/30 | 25/30 | 25/30 | 25/30 (out-of-pool) |
Per §3A.1's pool-wide retirement rule (a tier retires only once every tested model clears it): tiers 1, 3, and 4 are now unanimous across all four models and are retirement candidates — they haven't differentiated anyone yet. Tiers 2, 5, and 6 have each caught exactly one model, for three completely different reasons (a documented design choice, a genuine capacity limit, and a compound format-plus-arithmetic failure) — these are doing real work and should stay as-is rather than escalate.
This is a strong signal the ladder is reasonably well-calibrated as a whole, even though no single model has found every tier's edge: three tiers are appropriately hard, three are appropriately easy, and — notably — three different models have each shown exactly one weakness, with gemma4:e2b the only model to clear all six. That reinforces its Reigning Champion status (§3A.6) rather than calling it into question.
3A.2i Cap-fairness experiment — extended-cap retest of deepseek-r1:14b, tier 5 (August 2026)
Motivated directly by §3A.2f: is a fixed 4096-token num_predict cap unfair to models with a more deliberate reasoning style, or does it correctly reflect a real limitation? Rather than debate this abstractly, ran a controlled retest: same model, same tier, same prompt, same N=5 seeds, only num_predict doubled to 8192 (timeout extended to 180s accordingly).
Result: 0/5, unchanged — but for a completely different, better-diagnosed reason. Every seed now completed comfortably under the new ceiling (4,988–5,181 of 8,192 tokens — no cap-hits at all, confirming budget was genuinely the constraint at 4096). But the model still fails Stage A, consistently, because of a real, unrelated implementation bug: top_customers([], customers, n) returns entries for every customer with total_spent: 0 instead of the required empty list — a logic error in how it handles the empty-orders case, nothing to do with token budget.
What this settles: the standard 4096 cap is not treating deepseek-r1:14b unfairly. Raising the cap didn't change the tier-5 verdict at all — it just swapped an ambiguous "ran out of room" result for a precise bug diagnosis. If the extended-cap retest had instead produced a clean pass, that would have been real evidence the standard cap was too harsh for this model; it didn't, which is itself the answer.
Adopted protocol, formalized in open decision #11: any Stage A failure diagnosed as legitimate_overrun or truncated_usable (§2.1) — the two failure modes where more budget is plausibly the fix — becomes eligible for a one-time 2× extended-cap retest, N=5, reported alongside the standard-cap result rather than replacing it. silent_loop and cant_commit_loop failures are not eligible for this retest, since more budget doesn't resolve a loop (a loop just loops longer — this is literally how the original uncapped run consumed 82,000 tokens in §6 before any cap existed at all). This keeps the standard-cap result as the locked, official ranking data while still answering "was this fixable with patience" wherever it's a live question.
3A.2j Silent-error diagnosis experiment — shown-work retest of mistral-small3.2:24b, tier 6 (August 2026)
Same disciplined approach applied to open decision #12: mistral-small3.2:24b's tier-6 arithmetic error (§3A.2g — total_value = 267.25 instead of 289.85, identical across all 5 seeds) had no visible reasoning to diagnose against, because the prompt explicitly demanded "no explanation, only JSON." Rather than guess whether this was a genuine computational limit or an artifact of the compressed format, ran a controlled retest: identical task, identical inventory data, only the "no explanation" constraint relaxed — the model was asked to show its calculation step by step before giving the final JSON on its own line.
Result: 5/5 correct, every intermediate step right. A representative response:
1. Widget: 5 * 2.50 = 12.50
2. Gadget: 15 * 9.99 = 149.85
3. Gizmo: 2 * 45.00 = 90.00
4. Doohickey: 30 * 1.25 = 37.50
12.50 + 149.85 + 90.00 + 37.50 = 289.85
Every per-item multiplication and the final sum are correct. Conclusion: this was a format-induced mental-arithmetic slip, not a genuine computational capability gap. The model can clearly do this arithmetic correctly — it just made a silent error when compressing the whole calculation into a single internal pass with no externalized intermediate steps, under a strict "answer only" constraint.
A real, reportable cost trade-off came with it: the shown-work version used ~246 tokens and 4.4–9.5s per run, versus the silent version's ~46 tokens and ~1.0s — correctness came at roughly 5× the token cost and latency. This is directly useful, actionable guidance for the leaderboard: for models prone to this failure mode, a production tool-calling pipeline doing arithmetic under strict-format constraints may want to let the model reason first (even to a discarded scratchpad) before emitting final structured output, at a real but bounded cost.
Adopted protocol: any Stage A failure where non-arithmetic sub-parts of a structured output are correct but a computed value is wrong, under a "no explanation" format constraint, is eligible for a one-time shown-work diagnostic retest — not counted toward the model's ranking (the standard-format result stays locked, per §2.4), purely used to distinguish "can't compute this" from "computed it wrong when forced to hide its work."
3A.2k Sub-4GB expansion pilot results — four new-family models (August 2026)
Following the sub-4GB tier expansion (§7.2a), all four new models were run through the full six-tier Gauntlet.
| Model | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 | Tier 6 | Total |
|---|---|---|---|---|---|---|---|
| granite4:3b | 5/5 | 0/5 | 5/5 | 5/5 | 5/5 | 0/5 | 20/30 |
| llama3.2:3b | 5/5 | 5/5 | 0/5 | 5/5 | 0/5 | 0/5 | 15/30 |
| alibayram/smollm3 | 5/5 | 1/5 | 4/5 | 5/5 | 1/5 | 0/5 | 16/30 |
| phi4-mini | 5/5 | 0/5 | 1/5 | 5/5 | 0/5 | 0/5 | 11/30 |
granite4:3b is the first real contender to gemma4:e2b found in the sub-4GB tier — it only fails the two hardest-differentiating tiers (2 and 6), and does so cleanly. Tier 2 adds a fourth distinct edge-case contract to the pool's growing catalog: return-None (gemma4 family, mistral-small3.2:24b, deepseek-r1:14b, llama3.2:3b), raise-ValueError (qwen3.5:9b), return-a-wrong-type-string (phi4-mini: "Not enough unique values"), and now granite4:3b's own raise message ("There must be at least two unique numbers") — a real, recurring finding that "handle edge cases" is genuinely underspecified and every model family answers it differently.
llama3.2:3b (15/30) shows genuine, cleanly-diagnosed logic bugs rather than reliability problems: tier 3 fails on an inverted conditional in return_book (if not book.available: return False — backwards, blocks returning checked-out books specifically), and tier 5 mixes a real type bug (customer['name'] indexing a tuple from .items() with a string key) with an unresolved syntax error in other seeds.
alibayram/smollm3 (16/30) runs consistently "hot" — high token usage even on tiers other models clear in a few hundred tokens (2,000–3,600 range vs. ~100–900 for llama3.2:3b/granite4:3b), and repeatedly hits the exact 4096 cap with an unterminated string literal syntax error, distinct from every other cap-related failure signature seen so far (not a repetition loop, not a legitimate-overrun-with-clean-truncation — a literal mid-string cutoff). Tier 6 fails for a third reason: the response doesn't parse as JSON at all, even after fence-stripping.
phi4-mini (11/30) is the weakest model tested in this pilot, and notably weaker than its own larger sibling phi4-reasoning:14b overall — but via a different failure profile. No can't-commit-loop or system-prompt leakage (phi4-reasoning's signature issues); instead, high seed-to-seed variance on tier 5 (three different exception types — 'id', unhashable type: 'dict' — across otherwise-identical runs) and a 4-of-5 syntax-error rate on tier 3.
3A.2l Harness bug: community-model filenames with / (August 2026)
alibayram/smollm3's run completed and printed a full, valid transcript, but the results JSON failed to save — MODEL.replace(':','_') left the / from the community-model namespace intact, and Windows interpreted it as a path separator into a nonexistent alibayram directory. The detailed per-run data (response text, extracted code, thinking traces) for this run was lost; only the console-transcript-derived pass/fail/token/time figures in §3A.2k's table are recoverable. Fixed: filename sanitization now also replaces / (MODEL.replace(':','_').replace('/','-')). Any future community-namespaced model (org/model format) will save correctly; this pilot's smollm3 data should be treated as directionally reliable (scores and error-message text were visible live) but not available for deeper post-hoc diagnosis the way every other model's data is.
3A.2m Open question: a possible genuine ambiguity in tier 6's prompt
Three independent models, three different specific partial misses on low_stock_items, all otherwise getting the comparison logic (most_expensive_item) right:
| Model | low_stock_items produced |
Missing |
|---|---|---|
| llama3.2:3b | {'Widget'} |
Gizmo |
| phi4-mini | {'Widget'} |
Gizmo |
| granite4:3b | {'Gizmo'} |
Widget |
Two models miss Gizmo (quantity 2), one misses Widget (quantity 5) — both are legitimately "low stock" under the threshold of 10, so there's no shared arithmetic error explaining all three. This could be: (a) a genuine ambiguity in how the inventory list is formatted/presented in the prompt causing small models to drop an item during parsing, (b) small-model attention/list-processing limitations independent of the prompt, unrelated across models by coincidence, or (c) something specific to very short, low-token responses (all three of these are among the terser tier-6 responses in the pool) not giving the model room to double-check its own filtering. Not yet resolved — flagged as a candidate for the same kind of controlled diagnostic retest used for open decisions #11 and #12, if it recurs with further sub-4GB testing.
3A.2n Harness fix: syntax errors in demo code unfairly penalizing correct functions (August 2026)
Running qwen3:8b through the full ladder — the first 4–8GB tier expansion beyond the lone qwen3.5:9b — surfaced a real, previously-undiscovered harness limitation. Tier 1 (the "floor" tier every other model in the entire pool has cleared cleanly, including every sub-4GB model) scored a shocking 1/5.
Diagnosis: the model's is_palindrome function was completely correct in all 5 seeds. The failure was a literal typo in the model's own throwaway demo code — print(is_pal, "racecar")) (meant to call is_palindrome(, wrote is_pal, instead, leaving a mismatched parenthesis). The AST-filtered extraction fix from §3A.2c handles a model's demo code crashing at runtime (skips it, keeps the real definition) — but a syntax error anywhere in the response makes ast.parse() fail on the whole blob before any filtering can happen at all, so the harness fell back to the raw, still-broken text and penalized a fully correct function for an unrelated cosmetic typo three paragraphs away from the actual answer.
Fix: extract_code now trims trailing lines and retries ast.parse() until it succeeds, rather than giving up entirely on the first SyntaxError. Since demo/example code almost always comes after the real definition in a model's response, this reliably isolates a syntax error to the throwaway tail without discarding a working function above it. Verified against the actual failing response before deploying, then confirmed via a targeted re-run: tier 1 went from 1/5 to a clean 5/5 with no other change.
Practical implication: as with §3A.2c, any prior result could theoretically be affected by this specific failure mode (a syntax error confined to trailing demo/example code). Re-auditing all prior Stage A failures against the new extraction wasn't done wholesale given time constraints — but this is now a permanent fix benefiting every future run, and it's the kind of check worth running before any result gets cited as final.
3A.2o Harness fix: Windows console encoding crashing valid model output (August 2026)
Testing gemma4:e4b (the sub-4GB → 12GB+ progression pilot, §3A.2p) surfaced a fourth harness bug. Tier 3 — the multi-class interface tier, cleared cleanly by every other gemma4 variant — scored 0/5 with the error 'charmap' codec can't encode character '\u2705': character maps to <undefined>.
Diagnosis: the model's generated code included decorative emoji in its own print() statements (✅/❌ for success/failure messages — a stylistic choice, not a bug in the logic). Since exec()'d code shares the harness process's stdout, and this Windows environment's console defaults to the legacy cp1252 encoding, any Unicode character outside that codepage crashes the whole test with UnicodeEncodeError — a harness environment limitation, not a real model failure. The actual Book/Library implementation was correct.
Fix: reconfigure sys.stdout/sys.stderr to UTF-8 with errors="replace" at harness startup. Verified against the actual failing case, then confirmed via a targeted re-run: tier 3 went from 0/5 to a clean 5/5 with no other change. No other model in the pool has shown this exact error signature in any actual harness run to date (only in unrelated ad-hoc debugging scripts), so this appears isolated to gemma4:e4b's stylistic tendency toward emoji in output — but the fix is now permanent for any future model that does the same.
3A.2p The gemma4 family across all four VRAM tiers — what does more VRAM actually buy you? (August 2026)
Motivated directly by the question of what bigger models "buy" a self-hoster: the gemma4 family already had candidates spanning every tier in the pool, needing only two more runs (gemma4:e4b, gemma4:26b) to complete a fully controlled, same-architecture progression from sub-4GB through 12GB+.
| Model | VRAM | Tier | Score |
|---|---|---|---|
| gemma4:e2b | 1.8GB | Sub-4GB | 30/30 |
| gemma4:e4b | 3.4GB | Sub-4GB | 25/30 (fails tier 6 — silent computational error, matches mistral-small3.2:24b's pattern from §3A.2j) |
| gemma4:12b-it-qat | 8.0GB | 8-12GB | 30/30 (Champion) |
| gemma4:12b | 8.4GB | 8-12GB | 30/30 |
| gemma4:26b | 17GB | 12GB+ | 29/30 (new Champion — beats mistral-small3.2:24b's 25/30; fails one tier-2 seed to a genuine NameError) |
The finding is not monotonic, and that's the headline result. More VRAM within this family does not reliably buy more capability: gemma4:e2b, the smallest model in the entire family at 1.8GB, is tied for the best score anywhere in the lineup. gemma4:e4b, at nearly double the VRAM, is the weakest gemma4 variant tested. The 8–12GB tier is where the family demonstrably peaks — two clean 30/30 sweeps — and the 17GB flagship doesn't exceed that ceiling, only approaches it. For a reader deciding whether to spend more hardware budget within this family specifically, the honest answer is: the 8–12GB tier is the real sweet spot, not "biggest you can afford."
This reframes the pool's earlier "gemma4 is undefeated" observation (§7.4): true in the sense that gemma4 holds the Champion title in three of four tiers now, but the family's own internal scaling curve shows real, non-trivial variance — a useful nuance for the leaderboard to report rather than a simple "gemma4 always wins, bigger is better" headline.
Practical implication for the pool structure: mistral-small3.2:24b was previously labeled "orientation only" and treated as out-of-scope for the leaderboard proper, on the reasoning that 12GB+ testing wasn't a real priority yet. Now that gemma4:26b is a legitimate, in-scope 12GB+ competitor, that framing no longer holds — mistral-small3.2:24b is reclassified as a real, dethroned pool member (25/30, runner-up) rather than orientation-only data.
3A.2q The hunt for a 30/30 in every VRAM tier (August 2026)
Motivated by wanting at least one clean-sweep reference point per tier, not just a champion by relative score. Three tiers found one almost immediately:
- Sub-4GB:
gemma4:e2b— 30/30 (already established) - 8–12GB:
gemma4:12b-it-qatandgemma4:12b— both 30/30 - 12GB+:
qwen3:32b— a clean 30/30 on the first attempt (22GB measured VRAM), clearing every tier including tier 5 (the empty-orders bug that's broken four other models) and tier 2 (the catastrophic-loop tier that broke its own smaller sibling, qwen3:8b). This also dethronesgemma4:26b(29/30) as the new 12GB+ Champion after less than a day.
4–8GB resisted every attempt. Four different models, two families, no clean sweep:
| Model | VRAM | Score | Failed tiers |
|---|---|---|---|
| qwen3.5:9b | 5.7GB | 25/30 | tier 2 only (raises ValueError) |
| qwen2.5-coder:7b | 5.1GB | 25/30 | tier 6 only (silent computational error) |
| qwen3:8b | 6.3GB | 20/30 | tiers 2, 5 |
| mistral:7b | 5.6GB | 15/30 | tiers 2, 5, 6 |
Two failed download attempts along the way are informative in their own right: codellama:7b measured at 8.2GB (not 4-8GB as its 3.8GB download size suggested — an unusually large, ~2.2x expansion, likely reflecting its older Aug-2023 architecture relative to the more efficiently-quantized modern models already in the pool) and was excluded from this tier as a result; deepseek-coder-v2:lite was ruled out before downloading once external documentation placed its actual runtime VRAM in the 10-12GB range despite the "lite" name.
This is now a real, reportable finding rather than bad luck: across every model tested in this specific size range, at least one tier consistently trips something up. mistral:7b contributed two new findings of its own: a fifth distinct edge-case contract for tier 2 (returns float('-inf') as a sentinel rather than None or raising), and a genuine new bug on tier 5 ('bool' object is not subscriptable — a real logic error, not a harness artifact). Its tier 6 failure (total_value = 167.5, expected 289.85) makes it the fourth model to show the silent-computational-error pattern under the "no explanation" format constraint, joining mistral-small3.2:24b, gemma4:e4b, and qwen2.5-coder:7b — strengthening the case that this specific tier is measuring something real and recurring about small/mid-size models under strict output constraints, not coincidence.
Open question for future work: is 4–8GB genuinely harder, or is this an artifact of the specific tasks chosen? Worth revisiting once a fifth candidate is tested, but not chased further in this session given diminishing returns from repeated downloads.
3A.2r Pilot confirmation result — qwen3:8b, code-gauntlet-v1, full ladder (August 2026)
| Tier | Stage A result | Notes |
|---|---|---|
| 1 — floor | 5/5 (corrected) | Originally 1/5 due to the harness bug in §3A.2n |
| 2 — edge-case (known differentiator) | 0/5 | Confirmed — matches the original round-1 repetition-loop finding for this model (§6) |
| 3 — multi-class interface | 5/5 | Clean |
| 4 — bug-fix | 5/5 | Clean |
| 5 — schema cross-reference | 0/5 | Mixed: one seed showed the empty-orders bug (see below); remaining seeds hit genuine cap-related syntax truncation |
| 6 — strict JSON format | 5/5 | Clean, all strict |
20/30 (corrected from a harness-penalized 16/30). This gives the 4–8GB tier its first real competitive comparison: qwen3.5:9b (25/30, Champion) vs. qwen3:8b (20/30) — same family, newer generation clearly stronger, consistent with the pattern already seen at 8–12GB where qwen3.5-generation models outperformed qwen3-generation ones on the same tasks.
A fourth model now shows the tier-5 empty-orders bug. deepseek-r1:14b, qwen3:14b, and now qwen3:8b — three different model sizes within reasoning-capable lineages — have each independently produced the identical mistake: top_customers([], customers, n) returning zero-value entries for every customer instead of []. This is no longer a curiosity; it's a real, recurring pattern across a meaningful slice of the pool, strengthening the case (open decision, see §7 pool notes) that this may reflect something genuinely non-obvious about the empty-input case for this class of model, not coincidence.
3A.3 Evaluation logic — a two-stage gate, not a blended score
Confirmed design decision: evaluation is sequential.
- Stage A — Completion gate (binary). Did the model produce a valid, complete solution at this tier? A model that fails this gate does not receive a quality score for the tier — it receives a diagnosed failure instead (§3A.4). This mirrors the general failure-mode taxonomy in §2.1: failure is a distinct outcome, never folded into a low score.
- Stage B — Quality and time, only among models that passed Stage A. Correctness/quality (is the code actually right, is it good) and efficiency (tokens/time to reach the correct answer) are only compared among models that cleared the gate — this avoids scoring failed runs on a scale where they don't belong.
- Where a model stops climbing the ladder (last tier passed) is itself a headline data point, analogous to a truck's demonstrated towing class.
3A.4 Failure is data, not noise
Explicit design principle: a tier that causes a model to fail is a successful test outcome, provided the failure mode is diagnosed (silent loop / can't-commit loop / legitimate overrun / truncated-but-usable — §2.1). An undiagnosed failure ("empty response, unknown cause") is a failure of the test methodology, not of the model. This is why thinking-trace capture, num_predict/timeout logging, and leakage detection (§2.1, §2.6) must all be active during every Gauntlet run — failures are expected, and the whole point is being able to explain them.
3A.5 Path forward
This tier ladder and evaluation logic is the pilot for code_gen specifically. Once run against the model pool and refined, the same methodology — fixed course, escalating tiers, dual-axis measurement, versioning, failure-as-data — should be adapted to the other task categories rather than reusing a flat prompt-per-category approach for them.
3A.6 Reigning Champion
Borrowed from Ike Gauntlet's Gold Hitch Award: each VRAM tier (§7) carries its own Reigning Champion — the model with the best confirmed Gauntlet record (furthest tier passed, then quality/time among ties) within that tier, for a given category. Displayed prominently on that tier's leaderboard; dethroned when a new model in the same tier beats its record. Champions are never compared across tiers — "punches above its weight" only means something relative to what else fits the same hardware (§1).
Important distinction: the Champion is a comparison anchor for the leaderboard narrative, not the calibration target for the ladder itself. Tier difficulty and tier-retirement (§3A.1) stay governed by the whole active pool's pass rate across all VRAM tiers combined, never by one model's specific strengths or blind spots — this matters precisely because §3.1/§6 already established that failure triggers are model-idiosyncratic, so a ladder tuned to one champion's weak points could miss another architecture's entirely.
Before any model is crowned Champion, its winning run gets the same confirmation rigor a failure requires (§2.2, and as done manually for qwen3:8b's 5/5 failure confirmation in §6) — an N-repeat pass, not a single clean run. A lucky win is not a stronger claim than an unlucky failure; both need the same bar.
The Champion is also the natural candidate for §2.3's reproducibility "control" run, since it will be re-run most often across the campaign — checking the Champion periodically doubles as a drift check for the whole test environment.
Current champions by tier, code-gauntlet-v1 (§7.3):
| Tier | Champion | Score | Status |
|---|---|---|---|
| Sub-4GB | gemma4:e2b | 30/30 | Confirmed, clean sweep — closest challenger is gemma4:e4b at 25/30 (§3A.2p); 6 models tested (gemma4:e2b 30, gemma4:e4b 25, granite4:3b 20, alibayram/smollm3 16, llama3.2:3b 15, phi4-mini 11) |
| 4–8GB | qwen3.5:9b | 25/30 | Confirmed, no clean sweep found — tied with qwen2.5-coder:7b (25/30); 4-8GB resisted every attempt to find a 30/30 across 4 models, 2 families (§3A.2q) |
| 8–12GB | gemma4:12b-it-qat | 30/30 | Confirmed — clean sweep, beat gemma4:12b (also 30/30) on a speed tie-break (§7.3) |
| 12GB+ | qwen3:32b | 30/30 | Reigning Champion, clean sweep — dethroned gemma4:26b (29/30), which had itself just dethroned mistral-small3.2:24b (25/30) (§3A.2q) |
All four tiers now have real within-tier competition, and three of four now have at least one confirmed 30/30 clean sweep. gemma4 holds two of the four Champion titles (Sub-4GB, 8–12GB) — its own internal scaling curve is non-monotonic (§3A.2p): the family's strongest showing is the 8–12GB tier, not the largest variant tested. 4–8GB remains the pool's one genuinely unresolved tier — no model tested there has cleared all six tiers (§3A.2q).
3B. Second category deep-dive: tool_calling (August 2026)
Chosen as the second task category over summarization (no clean binary pass/fail without an LLM-judge, a real methodology departure) and long-context retrieval (tests context-window handling more than a distinct capability). Tool-calling is the single most practically relevant capability for self-hosters that code_gen doesn't touch at all — home automation, personal assistants, and agentic workflows all depend on it — and two candidates in the existing pool (alibayram/smollm3, granite4:3b) were originally chosen partly for their tool-calling reputations without ever being tested on it directly.
3B.1 Mechanism — a genuine departure from code_gen
Tool-calling is not text generation with a different prompt; it uses Ollama's /api/chat endpoint with a structured tools array (JSON Schema per function), and the model's response carries a structured message.tool_calls array (function.name + function.arguments, arguments already parsed as JSON, not a string to parse) rather than free text. Verified directly against the live API before building anything: a matching request correctly returns tool_calls; an irrelevant request correctly omits tool_calls entirely and returns a normal content answer instead. This is mechanically simpler to grade than code_gen in one respect — no exec(), no extraction-from-prose step, just structural validation against the JSON Ollama already parsed — but introduces a new failure surface entirely absent from code_gen: whether the model's own Ollama-packaged chat template correctly supports the tool-calling protocol at all (§3B.3).
3B.2 Tier ladder — tool-gauntlet-v1
Mirrors code_gauntlet's escalating-difficulty philosophy (§3A.1) with a genuinely tool-calling-specific progression:
| Tier | Name | Tests | Mechanism |
|---|---|---|---|
| 1 | floor | Single unambiguous tool call, one tool offered | Correct function name, correct argument |
| 2 | restraint | A request where no tool should be called (tool available but irrelevant) | Model must NOT invoke it, and answer directly instead |
| 3 | selection | Three plausible tools offered (get_weather/get_time/convert_currency) |
Correct tool chosen among genuine near-duplicates |
| 4 | extraction | Messy natural-language request requiring correct type extraction (a spelled-out date → ISO format, a spelled-out number → integer) | Argument types correct, not just presence |
| 5 | chaining | A request genuinely requiring two sequential tool calls, the second depending on a synthetic result injected after the first | Correct multi-turn chaining, not just single-call correctness |
| 6 | strict schema | A tool with an enum-constrained field, a nested array, and a date, all in one call | Full schema-valid call under real constraint pressure |
Five of six tiers are single-turn (type: "single"); tier 5 is genuinely multi-turn (type: "chained") — the harness sends the first message, extracts the model's get_user_location call, appends both the assistant's tool-call message and a synthetic role: "tool" result message, then re-sends to see if the model correctly uses that injected result in a second call. This is the harness's most mechanically complex piece and was the first thing verified working during the pilot.
3B.3 Pilot results — four models (August 2026)
| Model | VRAM | Score | Notes |
|---|---|---|---|
granite4:3b |
2.9GB | 30/30 | Clean sweep, and remarkably efficient — 20-71 tokens per tier (e.g. tier 1: 20 tokens vs. qwen3.5:9b's 85; tier 6: 71 vs. 177), finishing the entire 30-run ladder in 10.6s total vs. qwen3.5:9b's 42s. Genuinely validates the "reliable function-calling" reputation this model was chosen for (§7.2a) |
qwen3.5:9b |
5.7GB | 30/30 | Clean sweep on the very first run of a brand-new harness — validates both the single-turn and chained mechanics work correctly, though notably more verbose than granite4:3b at every tier |
phi4-mini |
3.7GB | 5/30 | Only tier 2 (restraint, which requires not calling a tool) passes. Every tier requiring an actual call fails — but not from bad reasoning (see below) |
alibayram/smollm3 |
2.7GB | 5/30 | Same exact pattern as phi4-mini (only tier 2 passes) — but a genuinely different underlying cause (see below) |
Two models score 5/30 with the identical pass/fail pattern, but the diagnoses are completely different — and this matters for accurate reporting.
phi4-mini's raw failing response shows it attempting the correct call — "[get_weather{"city": "Seattle"}]" — but as plain text in message.content, never in the structured tool_calls field the API expects. This points to a chat-template/protocol-support gap in how this specific model is packaged for Ollama: the model appears to know what to call and with what argument, but its Ollama build doesn't speak the tool-calling protocol correctly.
alibayram/smollm3's failure is categorically different. Its thinking trace reasons through the request as though no tools were ever offered at all — "I need to check if there's a real-time weather service... since I can't access live data, I'll have to rely on general knowledge" — with no evidence it ever attended to the tools array in the request. This directly contradicts the reasoning that put this model in the pool in the first place (§7.2a: "chosen for its... purpose-built tool-calling focus"). Given this is a third-party community republish rather than the official Ollama library build (§7.2a's original provenance caveat), the most likely explanation is that this specific community build's chat template doesn't wire the tools field through to the model at all — a packaging gap, not necessarily evidence against the underlying model's actual trained capability. Worth flagging as an open question rather than a settled verdict on the model itself.
For a self-hoster's practical purposes, both failures look identical — tool-calling functionally does not work via Ollama's native API either way — but the causes are different enough that they shouldn't be reported as the same finding. This mirrors the discipline already established for code_gen (§3A.4, "failure is data, not noise"): a diagnosed 5/30 is a different, more useful claim than an unexplained one.
granite4:3b vs. alibayram/smollm3 is also a striking within-tier contrast: both are sub-4GB, both were originally selected specifically for tool-calling strength, and one delivers a flawless, efficient clean sweep while the other fails to engage with tool-calling at all. Family/packaging — not size — is doing all the work here.
Open question carried forward: does phi4-mini's template gap recur across the broader phi4 family, and would the official (non-alibayram) SmolLM3 packaging — if one existed on Ollama — behave differently? Not answerable without further testing, but worth keeping in mind before drawing conclusions about either underlying model's real capability.
3B.4 Sub-4GB tier complete (August 2026)
All six sub-4GB code_gen candidates now tested on tool-gauntlet-v1:
| Model | Score | Notes |
|---|---|---|
gemma4:e2b |
30/30 | Clean sweep, verbose (~350 tokens/tier avg) |
gemma4:e4b |
30/30 | Clean sweep — notably, this model was code_gen's weakest gemma4 variant (25/30, a silent computational error on tier 6), yet flawless here. Model strength genuinely varies by task category, not just by size or family |
granite4:3b |
30/30 | Clean sweep, remarkably efficient (~35 tokens/tier avg) |
llama3.2:3b |
10/30 | Clean tool selection (tiers 1, 3 pass) but fails restraint (tier 2 — calls get_weather even when asked the capital of Japan) and has genuine type-fidelity bugs: passengers extracted as the string "3" rather than the int 3 (tier 4), attendees returned as a JSON-encoded string rather than an actual array (tier 6). Three distinct, real problems in one model, not a single root cause |
phi4-mini |
5/30 | Protocol/template gap — attempts correct calls as plain text (§3B.3) |
alibayram/smollm3 |
5/30 | Doesn't attend to the tools array at all (§3B.3) |
Three of six sub-4GB models achieve a clean sweep on tool-calling — a strikingly different picture than code_gen's sub-4GB tier, where only gemma4:e2b reached 30/30. Tool-calling appears to cluster capable small models together more tightly than code generation does, while its failures are sharper and more varied in root cause (a behavioral over-eagerness plus type bugs in llama3.2:3b; two different protocol-level gaps in phi4-mini and alibayram/smollm3) rather than the more continuous, partial-credit failure spectrum code_gen tends to produce.
3B.5 4-8GB tier complete (August 2026)
| Model | Score | Notes |
|---|---|---|
qwen3.5:9b |
30/30 | Clean sweep — validated the harness itself in the original pilot (§3B.3) |
qwen3:8b |
30/30 | Clean sweep — genuinely striking given this was one of code_gen's weakest performers (catastrophic repetition-loop failures, 20/30 even after harness fixes). Strong confirmation that task category reveals a distinct capability profile, not a single underlying "model quality" axis |
mistral:7b |
19/30 | A third distinct failure texture: perfect on tiers 1, 4, 6 (structure/extraction/schema) but fails tiers 2, 3, 5 — every tier requiring a judgment call about whether or how many tools to invoke. Tier 2: calls get_weather when asked the capital of Japan (matches llama3.2:3b's over-eagerness). Tier 3: one seed calls all three available tools simultaneously rather than selecting one — an even more extreme version of the same pattern. Tier 5: never attempts the first call in the chain at all. The eagerness/judgment problem is concentrated specifically around "should I call anything, and how many" — not formatting or extraction, which are flawless |
qwen2.5-coder:7b |
5/30 | A third distinct protocol/template gap, and the closest-to-correct of the three found so far: the raw response is {"name": "get_weather", "arguments": {"city": "Seattle"}} — the exact correct JSON schema for a tool call, output as plain text rather than routed through the structured tool_calls field. The model clearly has the right training data for the format; its Ollama template just doesn't wire it through. Plausibly explained by this being a code-completion-specialized model never built with chat-style agentic tool use as a target use case |
4-8GB now has three distinct, precisely diagnosed protocol/template gaps across the whole pool (phi4-mini's bracket notation, alibayram/smollm3's total non-engagement, qwen2.5-coder:7b's correct-JSON-wrong-channel) — reinforcing that this is a real, recurring category of failure worth calling out on the leaderboard as its own diagnosis type, distinct from a genuine reasoning or capability gap.
3B.6 8-12GB tier complete (August 2026)
| Model | Score | Notes |
|---|---|---|
gemma4:12b-it-qat |
30/30 | Clean sweep |
gemma4:12b |
30/30 | Clean sweep — the gemma4 family is now 5-for-5 across every tier tested (e2b, e4b, 12b-it-qat, 12b all 30/30; only 26b/31b untested) |
qwen3:14b |
30/30 | Clean sweep, mirroring qwen3:8b's flawless result — a second confirmation that the qwen3 family's tool-calling strength holds across sizes, despite genuine weaknesses in code_gen (qwen3:14b scored only 21/30 there, tiers 2 and 5) |
deepseek-r1:14b |
5/30 | A fifth distinct failure signature. Doesn't attempt a tool call in any form — no bracket notation, no stray JSON — just a minimal, generic non-answer: "I suggest getting online to get real-time information." (34 tokens, no reasoning shown). Closer to alibayram/smollm3's total non-engagement than phi4-mini/qwen2.5-coder:7b's attempted-but-wrong-channel pattern, but even more terse — no evidence of considering the tool at all |
8-12GB now has three clean sweeps out of four models — the gemma4 family in particular is undefeated at tool-calling across every size tested so far, a genuinely different picture than code_gen where gemma4:e4b was the family's one real weak point. Combined with 4-8GB's qwen3:8b/qwen3:14b split from code_gen weakness, the emerging pattern across the whole pool is that tool-calling competence and code_gen competence are close to orthogonal — knowing a model is weak at one gives almost no signal about the other.
3B.7 12GB+ tier complete — full pool now tested (August 2026)
| Model | Score | Notes |
|---|---|---|
gemma4:26b |
30/30 | Clean sweep — gemma4's sixth consecutive 30/30 across every size tested (e2b, e4b, 12b-it-qat, 12b, 26b). Only 31b remains untested |
qwen3:32b |
30/30 | Clean sweep, mirroring its own flawless code_gen result — the one model in the pool that's genuinely excellent at both categories |
mistral-small3.2:24b |
25/30 | A third distinct chaining-failure variant: correctly makes the first call (get_user_location) but never follows through with the second (get_weather) using the result — different from mistral:7b, which never attempted the first call at all |
Every VRAM tier now has at least one confirmed tool-calling clean sweep, and most have several. All 17 code_gen-pool models have now been run through both categories.
3B.8 Full pool summary — tool-gauntlet-v1, all 17 models (August 2026)
| Tier | Clean sweeps (30/30) | Real failures |
|---|---|---|
| Sub-4GB | gemma4:e2b, gemma4:e4b, granite4:3b (3-way) | llama3.2:3b 10/30, phi4-mini 5/30, alibayram/smollm3 5/30 |
| 4-8GB | qwen3.5:9b, qwen3:8b (2-way) | mistral:7b 19/30, qwen2.5-coder:7b 5/30 |
| 8-12GB | gemma4:12b-it-qat, gemma4:12b, qwen3:14b (3-way) | deepseek-r1:14b 5/30 |
| 12GB+ | gemma4:26b, qwen3:32b (2-way) | mistral-small3.2:24b 25/30 |
The single biggest structural difference from code_gen: clean sweeps are the norm here, not the exception. Ten of seventeen models (59%) hit 30/30, and every tier has a multi-way tie at the top — a genuinely different ceiling than code_gen, where 30/30 was rare enough to be individually notable each time it happened. Tool-calling appears to have a real capability cliff rather than a smooth difficulty gradient: a model either has working, well-templated tool support and clears everything, or it has some specific gap (protocol, restraint, chaining) and fails hard on exactly that gap while remaining clean elsewhere. Per-tier Champions among these ties would need the same speed/token tie-break already established for code_gen (§7.3's gemma4:12b-it-qat vs. gemma4:12b precedent) — a natural task for the leaderboard site itself to compute from the raw JSON rather than a hand calculation here, though the qualitative pattern is already visible: granite4:3b and the gemma4 family are consistently the most token-efficient of the clean-sweep models at every tier they compete in.
gemma4 is the standout family for this category — six clean sweeps across every size from 1.8GB to 17GB, a much stronger and more consistent record than its own code_gen performance (where it was excellent but had one real gap at e4b). qwen3:32b is the only model confirmed excellent at both categories in this pool — everything else that's flawless at tool-calling has at least one real code_gen weakness (qwen3:8b, qwen3:14b), and vice versa.
Six distinct, precisely diagnosed failure modes now cataloged across the models that don't work cleanly: over-eager calling (llama3.2:3b, mistral:7b), bracket-notation text (phi4-mini), correct-JSON-wrong-channel (qwen2.5-coder:7b), total non-engagement (alibayram/smollm3), minimal non-answer (deepseek-r1:14b), and incomplete chaining (mistral-small3.2:24b, mistral:7b). This diagnostic granularity — knowing not just that a model fails but precisely how — is exactly the kind of finding the whole TMV/DOE methodology (§2) was built to produce, and it transfers cleanly to a second task category without any change to the underlying philosophy.
3C. Third category deep-dive: RAG / context-grounded Q&A (August 2026)
Chosen as the third task category over summarization and creative writing (both lack a clean binary pass/fail without an LLM-judge) and reasoning/math (a close second, but less tied to a concrete self-hosting use case). RAG — "point a model at my documents and ask questions" — is one of the most common local-LLM setups after chat itself, and tests a capability neither prior category touches: faithfulness to a given, bounded context, including the honesty to say "not stated" rather than fabricate an answer.
3C.1 Mechanism
Uses Ollama's /api/generate (like code_gen, not tool_calling's /api/chat) with a passage-plus-question prompt structure. Grading is deterministic string/JSON matching against known facts, mirroring code_gen's reference-answer approach rather than requiring an LLM judge.
3C.2 Tier ladder — rag-gauntlet-v1
| Tier | Name | Tests |
|---|---|---|
| 1 | floor | A directly-stated fact, straightforward extraction |
| 2 | restraint (refusal) | A plausible-sounding question whose answer is absent from the passage entirely — model must decline, not fabricate |
| 3 | distractor resistance | Passage contains two similar facts (a real answer and a decoy); question requires picking the correct one |
| 4 | multi-passage synthesis | Two short passages; the answer requires combining a fact from each (a start date from one, a duration from the other) |
| 5 | precise extraction | Several exact fields (name, number, date) pulled into strict JSON matching a schema |
| 6 | long-context needle | A single specific fact buried in the middle of a much longer passage (~600 words), testing retrieval without getting lost in volume |
Tier 2 is the RAG-specific analog to tool_calling's "restraint" tier (§3B.2) and arguably the single most practically important measurement in this category: does a model make things up when it doesn't know, or say so honestly?
3C.3 Two harness bugs caught and fixed before any data was trusted (August 2026)
Bug 1 — token cap too low for this model's reasoning style. The initial pilot (qwen3.5:9b, MAX_PREDICT_TOKENS=512) scored 0/5 on tier 2 with an empty response field. Before accepting this as a model failure, the raw API response was inspected directly: done_reason: "length", and the thinking field showed the model correctly working toward "the context does not state a current CEO... I cannot answer with a name" — reasoning its way to the exact right conclusion, then getting cut off by the cap before ever writing the final answer. This is the same legitimate_overrun failure mode already catalogued in code_gen (§2.1, §3A.2f). Fixed by raising the cap to 4096, matching code_gen's established precedent.
Bug 2 — a false-positive in the tier-2 test function itself. Even after fixing the cap, tier 2 initially still failed: the harness's own fabrication-detector regex (\bceo (is|was)\b) matched the substring "CEO is" inside a genuinely correct refusal ("...there is no information stating who Meridian Robotics' current CEO is"), which happened to echo the question's wording. The model had actually answered correctly; the test function was wrong. Fixed by requiring a real capitalized name token to follow "CEO is/was" before flagging fabrication, rather than matching on the bare phrase.
Both bugs were caught before any score was recorded as final — directly applying §8.3's lesson ("assume the bug is yours before the model's") on the very first pilot of this new category.
3C.4 A genuine prompt-ambiguity ruling — tier 4 (August 2026)
After both harness bugs were fixed, tier 4 (multi-passage synthesis) still scored 0/5 for qwen3.5:9b, but for a real, diagnosable reason this time: the model's thinking trace explicitly reasoned that computing "June 1 + 145 days" would require "external calendar knowledge" and therefore violate the "use ONLY the context" instruction, and declined to perform the arithmetic at all. This is structurally identical to code_gen's tier-2 edge-case-contract ambiguity (§3A.2a) — the instruction didn't make clear that deriving a computed value from facts that are stated is the whole point of a synthesis tier, not a violation of it.
Per §2.4 (rubrics locked before testing, but genuine ambiguities get fixed rather than grandfathered when caught during the pilot itself, before broader data collection), the prompt was revised to explicitly permit computing from stated facts: "You may combine or compute from facts that ARE stated... that is not outside knowledge." This is now the locked wording going forward.
3C.5 Pilot results — two models (August 2026)
| Model | Score | Notes |
|---|---|---|
gemma4:e2b |
30/30 | Clean sweep, fast throughout (1-3s per response, no hesitation) — extends its strong track record to a third consecutive category |
qwen3.5:9b |
20/30 | Two genuine, distinct, real failures remaining even after both harness fixes and the prompt-ambiguity fix (see below) |
qwen3.5:9b's tier 3 failure is a can't-commit loop — the exact failure mode first documented in phi4-reasoning:14b during the original practice runs (§6.2), now independently reappearing in a different model and a different task category. The thinking trace shows the model reaching the correct answer ("Elena Vasquez," correctly resisting the Marcus Webb decoy) fairly early, then spending the remainder of the entire 4096-token budget second-guessing itself — debating whether the passage's language technically supports naming her, whether inference counts as "using only the context," reconsidering, re-deciding — without ever committing the answer to the final response. A real, cleanly diagnosed failure, not a harness artifact.
qwen3.5:9b's tier 4 failure, after the prompt fix, changed character a third time. With the ambiguity resolved, the model no longer refuses to do the arithmetic — it attempts it directly, but does the day-counting by hand, step by step, and gets visibly lost in its own off-by-one reasoning ("Wait, if I add 1 day to June 1 -> June 2. Add 29 days -> June 30... no, let's look at it as..."), burning the full 4096-token budget on manual calendar arithmetic without ever reaching a final answer. This is a genuine capability limitation — manual date arithmetic under token pressure — distinct from both the harness bug and the prompt-ambiguity that preceded it in diagnosis.
gemma4:e2b handles both of these tiers cleanly and fast (314 and 348-515 tokens respectively, 1.8-2.9s, no hesitation visible in either), confirming these are real, model-specific weaknesses in qwen3.5:9b rather than flaws in the tier design itself — exactly the validation role phi4-mini played for tool_calling's tier design (§3B.3).
This also meaningfully extends the emerging gemma4 story from §3A.2p/§3B.8: three categories tested now, and gemma4:e2b has been at or near a clean sweep in every one of them — a genuinely strong, broad-capability small model, not merely a code_gen or tool_calling specialist.
3C.6 Sub-4GB tier complete — four more harness bugs caught (August 2026)
All six sub-4GB code_gen/tool_calling candidates now tested on rag-gauntlet-v1:
| Model | Score | Notes |
|---|---|---|
gemma4:e2b |
30/30 | Clean sweep |
gemma4:e4b |
30/30 | Clean sweep, remarkably terse on tier 3 (4 tokens) |
alibayram/smollm3 |
30/30 | Clean sweep — see harness Bug 5 below; scored an apparent 25/30 before the fix |
granite4:3b |
25/30 | Fails tier 4 only — confidently wrong (computes "September 15" instead of "October 24") |
llama3.2:3b |
25/30 | Fails tier 4 only — confidently wrong ("December 15") |
phi4-mini |
25/30 | Fails tier 4 only — confidently wrong ("August 1") |
Two more real harness bugs caught and fixed this batch, on top of the two from §3C.3 (four total for this category so far):
Bug 3 — missing refusal-phrase variants. granite4:3b's tier-2 response ("The current CEO of Meridian Robotics cannot be determined from the provided context") is an unambiguously correct refusal, but the REFUSAL_PHRASES list didn't include "cannot be determined" as a recognized hedge phrase, so it scored as an ambiguous failure. Fixed by adding the missing phrasing; re-confirmed 5/5 immediately after.
Bug 4 (5th overall for this category) — literal <think> tags polluting the response field. alibayram/smollm3 embeds its reasoning directly in response_text using literal <think>...</think> tags rather than Ollama's separate thinking API field — the same non-exposure pattern already documented for phi4-reasoning:14b in §2.6, but manifesting as literal markup this time rather than silent mixing. Tier 5's JSON parser choked on the leading <think> text and scored 0/5, even though the actual answer immediately following </think> was a perfect, exact match to every expected field. Fixed by stripping <think>...</think> blocks before attempting to parse — flipped the model's true score from an apparent 25/30 to 30/30.
The real, non-harness finding: tier 4 (date synthesis) is a sharp, consistent differentiator across three different model families. granite4:3b, llama3.2:3b, and phi4-mini all fail it identically in texture — none hesitate or show any self-doubt (unlike qwen3.5:9b's can't-commit loop on the same tier, §3C.5); all three commit immediately and confidently to a specific, wrong date. This "confidently wrong" failure signature is distinct enough from the can't-commit loop to be worth naming as its own diagnosis category going forward. Only the two gemma4 variants and the corrected alibayram/smollm3 handle this tier cleanly in the sub-4GB pool.
3C.7 4-8GB, 8-12GB, and 12GB+ tiers complete — the full 17-model pool (August 2026)
| Model | Tier | Score | Tier-4 texture |
|---|---|---|---|
qwen3:8b |
4-8GB | 25/30 | Off-by-one — "October 23" |
mistral:7b |
4-8GB | 25/30 | Wide miss — "September 15" |
qwen2.5-coder:7b |
4-8GB | 25/30 | Wide miss — "September 15" |
gemma4:12b-it-qat |
8-12GB | 25/30 | Off-by-one — "October 23" (4/5 seeds; 1 seed capped/empty) |
gemma4:12b |
8-12GB | 25/30 | Full cap exhaustion, all 5 seeds — empty response, no exposed thinking trace to diagnose from |
deepseek-r1:14b |
8-12GB | 30/30 | Clean sweep |
qwen3:14b |
8-12GB | 26/30 | Mostly cap exhaustion (4 empty, 1 clean success) |
mistral-small3.2:24b |
12GB+ | 30/30 | Clean sweep |
gemma4:26b |
12GB+ | 25/30 | Off-by-one — "October 23" |
qwen3:32b |
12GB+ | 26/30 | Off-by-one — "October 23" (after fixing harness Bug 6, below) |
Bug 6 (7th overall for this category, including the two from §3C.3) — another refusal-phrase gap, this time with an inserted adverb. qwen3:32b's tier-2 response ("The context provided does not explicitly state who Meridian Robotics' current CEO is") is a correct refusal, but the literal-phrase-list matcher missed it because "does not state" is not a contiguous substring of "does not explicitly state." Fixed by replacing the growing literal phrase list with a single regex tolerant of 0-2 inserted words between "not" and the hedge verb (not\s+(?:\w+\s+){0,2}(state|mention|specif|...)), which should catch this whole class of adverb-infixed refusal going forward rather than requiring another one-off patch each time a new phrasing appears.
The tier-4 arithmetic error has a precise, diagnosed mechanism for the "October 23" cluster. qwen3:32b's raw output shows the actual calculation: "remaining days: 145 - (30 + 31 + 31 + 30) = 23 days in October... Adding 23 days to October 1, 2023, results in October 23, 2023." The bug is now fully legible: these models treat June as a full 30-day block elapsed from the start date, rather than the 29 days actually remaining after June 1st itself — a systematic off-by-one from not correctly accounting for the start date being day one of the count rather than day zero. This is not scattered noise; it's the same specific counting error, independently reproduced by four separate model families (qwen3:8b, gemma4:12b-it-qat, gemma4:26b, qwen3:32b).
The "September 15" cluster remains undiagnosed at the same level of precision — these three models (granite4:3b, mistral:7b, qwen2.5-coder:7b) give short, terse, non-reasoning-exposed answers with no visible working, so the exact computational shortcut producing this specific date can't be reconstructed from the available data the way the "October 23" mechanism could. Worth a dedicated shown-work diagnostic retest (mirroring §3A.2j's protocol) if this pattern is worth chasing further.
3C.8 Full pool summary — rag-gauntlet-v1, all 17 models (August 2026)
| Tier | Clean sweeps (30/30) | Real failures |
|---|---|---|
| Sub-4GB | gemma4:e2b, gemma4:e4b, alibayram/smollm3 (3-way) | granite4:3b/llama3.2:3b/phi4-mini all 25/30 |
| 4-8GB | none | qwen3.5:9b 20/30, qwen3:8b/mistral:7b/qwen2.5-coder:7b all 25/30 |
| 8-12GB | deepseek-r1:14b | gemma4:12b-it-qat/gemma4:12b 25/30, qwen3:14b 26/30 |
| 12GB+ | mistral-small3.2:24b | gemma4:26b 25/30, qwen3:32b 26/30 |
Five of seventeen models (29%) hit 30/30 — a lower clean-sweep rate than tool_calling (59%) but higher than code_gen's 4-8GB desert (§3A.2q). Tier 4 is overwhelmingly the reason: it's the only tier that ever fails across the entire pool (tiers 1, 2, 3, 5, and 6 are unanimous-clean or nearly so across all 17 models), making it the single most informative tier in this category by a wide margin — exactly the kind of sharp discriminator §3A.1's escalating-difficulty design is meant to surface.
The two reproducible wrong-answer clusters are the standout finding of this whole category. Four models converge on "October 23" via an identical, now fully diagnosed off-by-one mechanism; three converge on "September 15" via an as-yet-undiagnosed but clearly shared shortcut. Independent convergence on the exact same specific wrong answer, across unrelated architectures, is strong evidence of a systematic computational bias rather than independent random error — directly comparable to code_gen's cross-model "empty orders" bug (§3A.2i, §7.3a), but now observed in a completely different task category, suggesting this kind of convergent-error pattern may be a general property of how LLMs approximate certain calculations, not a quirk specific to one prompt or one prior category.
gemma4's dominance across categories finally cracks here. After six consecutive clean sweeps in tool_calling and near-perfect results in code_gen, four of five gemma4 variants tested in RAG fail tier 4 specifically (only gemma4:e2b and gemma4:e4b, the two smallest, are clean). This is a genuinely interesting reversal: gemma4's larger variants are more prone to this specific arithmetic slip than its smaller ones — the opposite of what raw capability would predict, and a concrete counterexample to any assumption that "bigger gemma4 is strictly better."
deepseek-r1:14b and mistral-small3.2:24b are the pool's most well-rounded performers in this category — clean sweeps despite deepseek-r1:14b being mediocre-to-weak in both prior categories (25/30 code_gen, 5/30 tool_calling) and mistral-small3.2:24b being solid-but-not-exceptional elsewhere (25/30 code_gen, 25/30 tool_calling). Combined with qwen3:32b's cross-category strength and qwen3:8b/qwen3:14b's tool_calling excellence despite code_gen weakness, the emerging three-category picture reinforces §8.1's central lesson even more strongly: a model's standing in one category predicts almost nothing about another, and the specific model that's "best" depends entirely on which capability actually matters for the reader's use case.
3D. Fourth category deep-dive: reasoning/math (August 2026)
Chosen as the fourth task category as a natural follow-up to RAG's tier-4 arithmetic findings — genuinely distinct from all three tested categories (no code, no tool schema, no external context to lean on), and cleanly checkable with exact numeric or logical answers, avoiding the LLM-judge requirement that ruled out summarization and creative writing twice already.
3D.1 Tier ladder — reasoning-gauntlet-v1
Six word problems and puzzles with independently verified, exact answers, locked before testing per §2.4:
| Tier | Name | Problem | Answer |
|---|---|---|---|
| 1 | floor | Single-step subtraction word problem | 28 |
| 2 | multi-step | Chained multiplication/percentage calculation | 1242 |
| 3 | logical deduction | A constraint-satisfaction puzzle with a unique solution | Alice=red, Ben=blue, Cara=green |
| 4 | adversarial (CRT trap) | The classic bat-and-ball cognitive-reflection-test question | $0.05 (not the intuitive $0.10) |
| 5 | combinatorics | A committee-selection counting problem | 40 |
| 6 | distractor filtering | A word problem stuffed with irrelevant details and one easy-to-miss exclusion instruction | 53 |
Tier 4 is a well-known problem from cognitive psychology, famous for reliably fooling both humans and language models via a strong but wrong intuitive pull toward $0.10 — a real, established adversarial test rather than an invented one. Tier 6 was deliberately designed with a specific trap: two categories of "guests who won't attend" (3 confirmed no, 2 unsure) plus two genuinely irrelevant details (cake-vs-pie, favorite color), testing whether a model correctly applies a specific stated exclusion instruction ("assume the unsure guests do NOT attend") rather than only catching the more obvious "3 said no."
3D.2 Two harness bugs caught and fixed on the first two pilots (August 2026)
Pilot 1 (qwen3.5:9b, 25/30) failed only tier 2 — a genuine legitimate_overrun: full 4096-token cap exhaustion, completely empty response, consistent with this model's already-documented tendency toward extensive internal deliberation on complex multi-step tasks (§3A.2f, §3C.5). Not a harness bug — the hardest arithmetic problem in the ladder triggering a known behavioral pattern.
Pilot 2 (phi4-mini, initially 19/30, corrected to 20/30) surfaced a real regex bug on tier 3. One seed's response — fully correct, solving all three colors accurately — was marked failed because the proximity window in the logic-puzzle checker (30 characters between a name and its color) was too narrow for this model's verbose phrasing: "Cara's preferred colour is then left to be Green by process of elimination" has roughly 40 characters between "Cara" and "Green." Fixed by widening the window to 80 characters; re-confirmed 5/5 on the affected tier immediately after.
phi4-mini's two remaining genuine failures are real and well-diagnosed: tier 2 shows two different specific wrong numbers across seeds (1238, 1473) — real arithmetic errors under multi-step pressure — and tier 6 shows a precise, traceable mistake: the model computed balloons for 21 guests (24 − 3, forgetting to also exclude the 2 "unsure" guests) rather than the correct 19, landing on 42 + 15 = 57 instead of 38 + 15 = 53.
3D.3 Sub-4GB tier complete — a third convergent wrong-answer cluster (August 2026)
| Model | Score | Notes |
|---|---|---|
gemma4:e2b |
30/30 | Clean sweep |
gemma4:e4b |
30/30 | Clean sweep |
llama3.2:3b |
30/30 | Clean sweep — genuinely striking given this model was one of the weakest performers in both code_gen (15/30) and tool_calling (10/30) |
granite4:3b |
20/30 | Fails tiers 3 (a genuine logic error: gives Alice=blue, contradicting the stated constraint) and 6 |
phi4-mini |
20/30 | Fails tiers 2, 6 (§3D.2) |
alibayram/smollm3 |
26/30 | Fails tier 6 only (4/5 seeds) |
Tier 6 produces a third reproducible convergent-error cluster for this project, mirroring RAG's "October 23"/"September 15" pattern (§3C.7) exactly. granite4:3b, phi4-mini, and alibayram/smollm3 all land on the identical wrong total — 57 balloons — via the identical mechanism: correctly subtracting the 3 confirmed "no" guests but failing to also exclude the 2 "unsure" ones, despite the explicit instruction to do so. Three unrelated model families, one shared misreading of the same sentence. Combined with RAG's two clusters and code_gen's "empty orders" bug (§3A.2i, §7.3a), this is now the fourth independent instance across two different task categories of unrelated models converging on the exact same specific wrong answer — strong, mounting evidence that certain prompt structures reliably trigger a shared failure mode across architectures, not coincidental noise.
3D.4 4-8GB tier complete — the "57" cluster grows to five models (August 2026)
| Model | Score | Notes |
|---|---|---|
qwen3:8b |
30/30 | Clean sweep — including tier 6, the trap that fooled three sub-4GB models. Genuinely notable given this model's weak code_gen showing (20/30) contrasted with strong tool_calling and now reasoning/math results |
qwen3.5:9b |
25/30 | Fails tier 2 only, a genuine legitimate_overrun (§3D.2) |
mistral:7b |
25/30 | Fails tier 6 only — joins the "57" cluster |
qwen2.5-coder:7b |
25/30 | Fails tier 6 only — joins the "57" cluster, and notably arrived there via explicit Python-style code (total_balloons = total_balloons_needed + extra_balloons, printing 57) rather than prose arithmetic, showing the same guest-count error survives even when the model externalizes its calculation as code |
The "57" convergence cluster is now five models strong — granite4:3b, phi4-mini, alibayram/smollm3, mistral:7b, qwen2.5-coder:7b — spanning both VRAM tiers tested so far and multiple unrelated architectures. This is the single most reproducible finding across all four task categories to date: the same specific misreading of "assume the unsure guests do NOT attend" recurring identically regardless of model family, size, or even whether the model reasons in prose or code.
3D.5 8-12GB tier complete — the first perfect tier in the project's history (August 2026)
| Model | Score | Notes |
|---|---|---|
gemma4:12b-it-qat |
30/30 | Clean sweep |
gemma4:12b |
30/30 | Clean sweep — gemma4's fourth consecutive clean sweep in this category |
deepseek-r1:14b |
30/30 | Clean sweep — a second consecutive category (after RAG) where this model excels despite weak code_gen (25/30) and very weak tool_calling (5/30) results, reinforcing that its strength is specifically reasoning-heavy tasks |
qwen3:14b |
30/30 | Clean sweep, but at real cost — substantially more verbose than the other three (up to 3,635 tokens and 44s on tier 6, vs. the ~470 tokens/6s the others needed), a genuine efficiency gap even among four models that all reach the correct answer |
This is the first VRAM tier, in any of the four task categories tested across this entire project, where every single model achieves a perfect clean sweep. Every prior tier in every prior category had at least one model with a real, diagnosed weakness. Whether this reflects something specific about 8-12GB models' general capability level, this particular tier's difficulty calibration, or simple chance with only four models tested, is worth flagging as an open question rather than over-interpreting — but it's a striking, reportable milestone regardless.
3D.6 A sixth harness bug — comma thousands-separators breaking number matching (August 2026)
mistral-small3.2:24b's tier-2 pilot scored an apparent 0/5, but the raw response showed the model getting the answer exactly right: "Non-defective widgets = 1,350 - 108 = 1,242... Final Answer: 1,242." The contains_number test function required the digit sequence "1242" to appear with no interruption, and the comma thousands-separator in "1,242" broke that match entirely — a real harness bug, not a model failure. Fixed by stripping comma separators between digits before matching; re-confirmed 5/5 immediately after, correcting the model's true score from an apparent 25/30 to a clean 30/30.
3D.7 12GB+ tier complete — full 17-model pool, all four task categories now tested (August 2026)
| Model | Score | Notes |
|---|---|---|
mistral-small3.2:24b |
30/30 | Clean sweep (corrected, §3D.6) |
gemma4:26b |
30/30 | Clean sweep — gemma4's fifth consecutive clean sweep in this category, spanning sub-4GB, 8-12GB, and 12GB+. No other model family has come close to this level of cross-tier consistency in any category tested across the whole project |
qwen3:32b |
27/30 | Fails tier 6 on 3 of 5 seeds — but not the "57" convergence cluster. All three failing seeds hit the full 4096-token cap with a completely empty response (confirmed via raw token counts: 4096 tokens, 102-104s, empty response_text), a genuine legitimate_overrun consistent with this model's known tendency toward extensive deliberation on complex tasks (§3A.2f, §3D.2). Two of five seeds succeeded cleanly at far lower token counts (507-1235), showing the failure is a budget/consistency issue on this specific model, not a comprehension error |
3D.8 Full pool summary — reasoning-gauntlet-v1, all 17 models (August 2026)
| Tier | Clean sweeps (30/30) | Real failures |
|---|---|---|
| Sub-4GB | gemma4:e2b, gemma4:e4b, llama3.2:3b (3-way) | granite4:3b/phi4-mini 20/30, alibayram/smollm3 26/30 |
| 4-8GB | qwen3:8b | qwen3.5:9b/mistral:7b/qwen2.5-coder:7b all 25/30 |
| 8-12GB | gemma4:12b-it-qat, gemma4:12b, deepseek-r1:14b, qwen3:14b (4-way, the first fully-clean tier in the project) | none |
| 12GB+ | mistral-small3.2:24b, gemma4:26b | qwen3:32b 27/30 |
Ten of seventeen models (59%) hit 30/30 — matching tool_calling's clean-sweep rate and far exceeding code_gen's. gemma4 is the standout family of the entire project: five clean sweeps across five different sizes tested (e2b, e4b, 12b-it-qat, 12b, 26b) with zero exceptions — no other model family has managed a perfect record across every size tested in any of the four categories built this project.
The "57 balloons" convergence cluster settled at five models (granite4:3b, phi4-mini, alibayram/smollm3, mistral:7b, qwen2.5-coder:7b) — every one of them a sub-4GB or 4-8GB model. No 8-12GB or 12GB+ model fell into this specific trap, suggesting the error may correlate with smaller model size specifically, unlike RAG's convergence clusters which spanned models up to 22GB (qwen3:32b was in the "October 23" cluster there). This is a genuinely interesting cross-category contrast worth flagging: RAG's convergent error was size-independent; reasoning/math's was not.
deepseek-r1:14b's pattern across categories is now unambiguous: weak at code_gen (25/30) and very weak at tool_calling (5/30), but a clean sweep in both RAG and reasoning/math — a model whose strength is specifically and consistently reasoning-heavy, context-independent tasks, not generation or protocol-following.
All four task categories (code_gen, tool_calling, RAG, reasoning/math) are now complete across the entire 17-model, 4-VRAM-tier pool — 408 individual tier-runs, over 2,000 individual model queries, and 20 distinct harness bugs found and fixed along the way (5 in code_gen, 0 in tool_calling, 7 in RAG, and 8 across the two reasoning/math sessions including the regex-window and comma-separator fixes documented above).
4. Data schema (for harness output → website ingestion)
{
"run_id": "uuid",
"timestamp": "ISO8601",
"model": "string",
"model_family": "string",
"model_size": "string",
"quantization": "string",
"context_length": "int",
"task_category": "string",
"prompt_style": "string",
"prompt_id": "string",
"gauntlet_tier": "int | null -- rung on the escalation ladder, for categories using the Gauntlet methodology (see §7)",
"gauntlet_version": "string | null -- e.g. 'code-gauntlet-v1' -- which ladder revision produced this result",
"response_text": "string",
"thinking_text": "string",
"thinking_field_available": "bool",
"completed": "bool",
"failure_reason": "string",
"failure_mode": "string | null -- one of: silent_loop, cant_commit_loop, legitimate_overrun, truncated_usable, null if completed",
"system_prompt_leakage": "bool",
"quality_score": "float | null",
"tokens_per_sec": "float",
"ttft_ms": "float",
"eval_count": "int",
"num_predict_cap": "int",
"timeout_sec": "int",
"vram_mb": "float",
"ram_mb": "float",
"ollama_version": "string",
"hardware_notes": "string"
}
One row per run; aggregation (mean/std/CV, completion rate, failure-mode breakdown) computed at ingestion time for the site.
response_text, thinking_text, thinking_field_available, completed, failure_reason, failure_mode, and system_prompt_leakage were added after two practice runs (§6) — the original schema had no way to distinguish "the model answered badly" from "the model never answered," no visibility into why, and no way to flag that a model's failures aren't all the same underlying problem.
5. Open decisions before harness build
- ~~Which task categories~~ — partially resolved: research (Aug 2026) grounded a candidate list against real usage/benchmarks — code_gen, reasoning/math, chat/instruction-following, RAG/document Q&A, summarization, structured extraction, tool-use/function-calling (sharp 7–9B capability cliff found), and creative writing/roleplay (surprisingly the #1 real usage category for open-weight models per OpenRouter traffic data, though barely covered by academic benchmarks). Final launch list still to be confirmed; code_gen is the deep-dive pilot (§3A) and will inform how the others get built out.
- ~~Full list of models to include in the 8–12GB screening pool~~ — resolved and restructured, §7. The single 8-12GB band was widened to four VRAM tiers (sub-4GB / 4–8GB / 8–12GB / 12GB+) after testing revealed already-tested models spanning all four; each tier gets its own leaderboard and Champion (§3A.6). 10 candidate models measured; only the 8-12GB tier still needs broad within-tier testing (3 of 4 candidates untested).
- Quality scoring method: rubric only, reference-based, or a mix — and whether scoring is manual, self-scored via another LLM judge, or both. For code_gen specifically, Stage B correctness can likely be reference/test-based (does the code pass real tests) rather than subjective rubric scoring — an advantage of starting the pilot here.
- How to handle non-deterministic sampling — fixed seed + temperature 0 for repeatability runs, plus separate temp>0 runs to characterize real-world variance
- ~~VRAM measurement method~~ — resolved: measured runtime VRAM (
ollama ps/nvidia-smi), not published file size (§6) - ~~Whether failure triggers are prompt-specific or model-specific~~ — resolved: model-idiosyncratic, not shared across models (§3.1, §6)
- How to detect
system_prompt_leakageprogrammatically (practice runs found it by eye in phi4-reasoning:14b's output) — needs a heuristic (e.g., matching against known persona-instruction phrasing) or manual flagging at Stage 1 review - Whether to auto-detect failure_mode (e.g., via repetition-similarity scoring on the thinking/response text) or assign it manually during Stage 1 review — auto-detection would be needed to scale past a handful of models
- ~~Exact wording/test-suite for each rung of the code_gen tier ladder~~ — resolved: all six tiers (§3A.2a) have concrete prompts and executable test suites, piloted and confirmed 30/30 against gemma4:e2b (§3A.2b)
- ~~Where the top of the ladder should sit, and the numeric retirement rule~~ — resolved with data: all four pool models now tested (§3A.2h). Tiers 1, 3, 4 are unanimous 4/4-clean and are retirement candidates. Tiers 2, 5, 6 each caught exactly one model for three distinct reasons and should stay active — the ladder is well-calibrated as a whole even though only gemma4:e2b has cleared every tier. Retiring 1/3/4 and what replaces them (escalate those three specifically, or leave the ladder at 6 tiers with 3 active discriminators) is the next concrete decision.
- ~~Whether the fixed 4096-token
num_predictcap is fair across models with very different reasoning styles~~ — resolved via a controlled extended-cap retest, §3A.2i. Doubling the cap to 8192 for deepseek-r1:14b's tier-5 failure eliminated the budget exhaustion entirely (well under the new ceiling) but the model still failed Stage A 5/5, now for a genuine, unrelated logic bug (incorrect empty-input handling) rather than an ambiguous cap-hit. Conclusion: the fixed 4096 cap is not unfair — raising it doesn't change the ultimate verdict, it just replaces an ambiguous failure with a precisely diagnosed one. This validates the standard cap as the correct "fixed course" setting. Adopted protocol going forward: any Stage A failure diagnosed aslegitimate_overrunortruncated_usable(§2.1) is eligible for a one-time extended-cap (2×) retest, reported alongside the standard-cap result, never replacing it — failures diagnosed assilent_looporcant_commit_loopare not eligible, since more budget doesn't fix a loop. - ~~How to detect a "silent computational error"~~ — resolved via a controlled shown-work retest, §3A.2j. Same task, same data, only the "no explanation" constraint relaxed (model asked to show its calculation before the final JSON). Result: 5/5 correct, every intermediate step right, vs. 0/5 correct (consistently wrong the same way) under the silent format. Conclusion: this was a format-induced mental-arithmetic slip, not a genuine computational capability gap. Real cost trade-off found: shown-work used ~246 tokens / 4–9s vs. silent's ~46 tokens / ~1s — correctness came at a real latency/token cost. Adopted protocol: any Stage A failure where the model's output was correct on non-arithmetic sub-parts but wrong on a computed value, under a "no explanation" format constraint, is eligible for a one-time shown-work diagnostic retest (not counted for ranking, purely diagnostic) — mirrors the extended-cap protocol from #11 in spirit: isolate one variable, don't overwrite the locked standard-format result.
- Whether tier 6's
low_stock_itemsfilter has a genuine prompt ambiguity — three sub-4GB models (§3A.2m) each dropped a different single item from the expected two-item set, with no shared arithmetic explanation. Candidate for a controlled diagnostic retest (shown-work variant, per #12's protocol) if the pattern recurs with further testing.
6. Practice run findings (August 2026)
Two rounds of manual practice runs (6 prompts across 3 draft task categories: code_gen, reasoning, chat_instruction) were run before committing to a full harness build, specifically to let real model behavior reveal gaps in the design rather than guessing upfront.
6.1 Round 1 — qwen3:8b vs gemma4:e2b
- A "simple" prompt reliably broke a capable model. The prompt "write a function that returns the second largest unique value, handle edge cases" caused qwen3:8b to spiral into a degenerate token-repetition loop inside its thinking trace (observed: the model started enumerating an edge-case test list and never stopped appending the same value). Confirmed 5/5 failures across 5 different seeds at temperature 0 — this is a reproducible characteristic of this model on this prompt, not a fluke.
- gemma4:e2b handled the identical prompt cleanly — correct code, tested edge cases, complexity analysis, in 16 seconds.
- VRAM tiering by download size would have been wrong. gemma4:e2b's 7.2GB download used only 1.8GB of measured VRAM at runtime.
- Without capturing the model's
thinkingoutput, this failure would have looked like an unexplained empty response — no way to distinguish "ran out of budget on a genuinely hard problem" from "stuck in a repetition loop." This directly shaped §2.6 and the schema in §4. - The failure mode has a live production connection. The same blind-cap pattern (a fixed
num_predict, no thinking-model awareness) was found in TrueNorth'sllm_config.py, which had no persistent record of empty-response events. Fixed by adding failure logging with areasoning_cap_exhaustionheuristic flag, so future occurrences in production are now visible instead of silently vanishing.
6.2 Round 2 — four more models for orientation
The same 6-prompt set was run against qwen3.5:9b, phi4-reasoning:14b, deepseek-r1:14b, and mistral-small3.2:24b to see whether the Round 1 failure was a qwen3-specific quirk or a general reasoning-model risk.
| Model | Measured VRAM | Failures (of 6) | Failure-mode profile |
|---|---|---|---|
| qwen3:8b | 6.3 GB | 1 (5/5 on confirmation) | Silent/enumeration loop — zero output, thinking exposed |
| qwen3.5:9b | — | 0 | Clean |
| gemma4:e2b | 1.8 GB | 0 | Clean, fastest, smallest footprint |
| phi4-reasoning:14b | 12 GB | 3 | Can't-commit loop — drafts correct answer, re-rehearses 40+ times; thinking field NOT exposed (embedded in response); system-prompt leakage observed |
| deepseek-r1:14b | — | 1 | Legitimate overrun — coherent, correct reasoning, ran out of budget; thinking properly exposed |
| mistral-small3.2:24b | 15 GB | 0 | Clean, terse, non-reasoning; outside 8–12GB pool (orientation-only, not eligible for the leaderboard) |
Key findings from this round:
- Failure is not one phenomenon. Three qualitatively different failure modes appeared across four failing runs on two models — silent loop, can't-commit loop, and legitimate overrun — each with a different implication for whether the model is actually a bad recommendation or just needs a bigger token budget. This produced the failure-mode taxonomy in §2.1.
- Thinking-field exposure is inconsistent across architectures, even among models explicitly branded for reasoning. qwen3:8b and deepseek-r1:14b separate it; phi4-reasoning:14b does not. The harness cannot assume the field exists (§2.6).
- Failure triggers are model-idiosyncratic, not prompt-idiosyncratic — no single prompt in the 6-prompt set was universally hard; each failing model failed on a different prompt. This resolves the open question from Round 1 (§3.1).
- System-prompt leakage is a distinct, reportable quality dimension. phi4-reasoning:14b's visible output included what appear to be its own internal persona/instruction text ("You are Phi, a language model developed by Microsoft... you must give a disclaimer..."). This is independent of the loop problem and worth surfacing to readers on its own.
- VRAM range across nominally-similar model sizes remains large — 1.8GB to 15GB across six models in the 9–24B nominal size range — reconfirming that the 8–12GB pool must be built from measured VRAM, not marketing size class.
7. Model pool definition and VRAM tiers (August 2026)
Before this pass, the "8–12GB pool" from §1 had never actually been enumerated — every model tested so far (§3A.2b–g) was a convenience sample of whatever happened to be downloaded, not a defined pool. This section resolves open decision #2.
7.1 Measurement method
Same approach validated in §6: load each candidate with a trivial prompt (num_predict: 10), read ollama ps immediately after for the actual VRAM figure, then unload (keep_alive: 0) before measuring the next model to avoid overlap. Download size is recorded for reference only — never used for tier assignment, since it's repeatedly proven unreliable (up to ~5× mismatch, e.g. gemma4:e4b: 9.6GB download → 3.4GB actual).
7.2 Full candidate measurement (all locally available models)
| Model | Download size | Measured VRAM | Tier |
|---|---|---|---|
| gemma4:e2b | 7.2 GB | 1.8 GB | Sub-4GB |
| alibayram/smollm3 | 1.9 GB | 2.7 GB | Sub-4GB |
| granite4:3b | 2.1 GB | 2.9 GB | Sub-4GB |
| gemma4:e4b | 9.6 GB | 3.4 GB | Sub-4GB |
| llama3.2:3b | 2.0 GB | 3.1 GB | Sub-4GB |
| phi4-mini | 2.5 GB | 3.7 GB | Sub-4GB (right at the ceiling) |
| qwen3.5:9b | 6.6 GB | 5.7 GB | 4–8GB |
| qwen3:8b | 5.2 GB | 6.3 GB | 4–8GB |
| gemma4:12b-it-qat | 7.2 GB | 8.0 GB | 8–12GB |
| gemma4:12b | 7.6 GB | 8.4 GB | 8–12GB |
| deepseek-r1:14b | 9.0 GB | 10 GB | 8–12GB |
| qwen3:14b | 9.3 GB | 10 GB | 8–12GB |
| phi4-reasoning:14b | 11 GB | 12 GB | 12GB+ |
| mistral-small3.2:24b | 15 GB | 15 GB | 12GB+ |
| gemma4:26b | 17 GB | 17 GB | 12GB+ |
| qwen3:32b | 20 GB | 22 GB | 12GB+ |
Not included: qwen2.5:32b (≥19GB, clearly workstation-class rather than consumer-GPU territory, not measured); qwen3.5:cloud (cloud-routed, not a local model); qllama/bge-reranker-v2-m3 and nomic-embed-text (embedding/reranking models, not generation models — out of scope for the code_gen leaderboard and likely every other planned category).
7.2a Sub-4GB expansion (August 2026) — new families
Deliberately added four models from families with zero prior representation in the pool, to maximize the odds of finding genuinely differentiated findings (per §3.1/§6: failure triggers are model-idiosyncratic, so family diversity matters more than testing more variants of families already covered):
llama3.2:3b— Llama was completely absent from the pool despite being one of the most canonical open-weight familiesphi4-mini— a direct contrast case tophi4-reasoning:14b(§6's worst behavioral profile: can't-commit loop, no thinking-field separation, system-prompt leakage) — does the smaller, non-reasoning Phi variant avoid those issues?alibayram/smollm3— SmolLM3 is not in Ollama's official library, only as a community republish (87.5K downloads, legitimate and established, but third-party — worth noting as a provenance caveat if this model performs well and gets recommended). Chosen for its fully-open training pipeline and purpose-built tool-calling focus, directly relevant to the earlier 7–9B tool-calling capability-cliff research.granite4:3b— IBM's Granite family, Apache 2.0, a completely different (enterprise-oriented) lineage than anything else in the pool, specifically noted for reliable function-calling.
All four measured within the sub-4GB tier, though phi4-mini (3.7GB) sits right at the ceiling.
7.3 Per-tier Gauntlet status
Sub-4GB, 4–8GB, 8–12GB, and 12GB+ all now have real within-tier competition — three of the four have a confirmed 30/30 clean sweep; 4–8GB remains the pool's one unresolved tier (§3A.2q):
| Tier | Gauntlet-tested | Score | Status |
|---|---|---|---|
| Sub-4GB | gemma4:e2b | 30/30 | Reigning Champion — clean sweep; closest challenger gemma4:e4b at 25/30 |
| Sub-4GB | gemma4:e4b | 25/30 | Fails tier 6 only (silent computational error) |
| Sub-4GB | granite4:3b | 20/30 | Fails tiers 2, 6 |
| Sub-4GB | alibayram/smollm3 | 16/30 | Fails tiers 2, 3, 5, 6 |
| Sub-4GB | llama3.2:3b | 15/30 | Fails tiers 3, 5, 6 |
| Sub-4GB | phi4-mini | 11/30 | Fails tiers 2, 3, 5, 6 |
| 4–8GB | qwen3.5:9b | 25/30 | Reigning Champion (tied) — fails tier 2 only; no model in this tier has found a clean sweep (§3A.2q) |
| 4–8GB | qwen2.5-coder:7b | 25/30 | Tied with qwen3.5:9b — fails tier 6 only |
| 4–8GB | qwen3:8b | 20/30 | Fails tiers 2, 5 |
| 4–8GB | mistral:7b | 15/30 | Fails tiers 2, 5, 6 |
| 8–12GB | gemma4:12b-it-qat | 30/30 | Reigning Champion — clean sweep, won a tie-break vs. gemma4:12b on speed |
| 8–12GB | gemma4:12b | 30/30 | Runner-up — also clean, but slower (509.5s vs 487.0s total, 41,939 vs 41,570 tokens across all 30 runs) |
| 8–12GB | deepseek-r1:14b | 25/30 | Fails tier 5 (§3A.2f, §3A.2i) |
| 8–12GB | qwen3:14b | 21/30 | Fails tier 2 (1/5, borderline cap-exhaustion) and tier 5 (0/5) — see §7.3a |
| 12GB+ | qwen3:32b | 30/30 | Reigning Champion — clean sweep on the first attempt (§3A.2q) |
| 12GB+ | gemma4:26b | 29/30 | Runner-up — dethroned after less than a day; fails one tier-2 seed to a genuine NameError |
| 12GB+ | mistral-small3.2:24b | 25/30 | Fails tier 6 only |
7.3a qwen3:14b — a genuinely borderline result, and a cross-model pattern
qwen3:14b's tier 2 result (1/5) is distinct from every prior tier-2 failure: not a catastrophic loop like qwen3:8b's 5/5 failure on the identical prompt, but a model working right at the edge of the token budget — the one passing seed used 3,969 of 4,096 tokens (97%), and the four failing seeds all hit the exact cap with no function defined. Same family, same prompt, a smaller (8b) variant that failed completely and a larger (14b) variant that mostly-but-not-reliably succeeds — a real, reportable within-family scaling data point.
Tier 5 (0/5) split cleanly into two failure types: one cap-exhaustion (top_customers not defined) and four instances of a specific logic bug — top_customers([], customers, n) returning zero-value entries for every customer instead of an empty list. This is the exact same bug, reproduced character-for-character in the error message, as deepseek-r1:14b's tier-5 result under its extended-cap retest (§3A.2i). Two different model families independently converging on the identical edge-case mistake on this specific prompt is worth flagging as a possible signal that the empty-orders case is a genuinely easy-to-miss edge case in this task's phrasing, not a model-specific quirk — worth keeping in mind if tier 5 ever gets revised.
7.4 Current state and next steps
Three of four tiers now have a confirmed 30/30 clean sweep and real competitive depth: Sub-4GB (gemma4:e2b, 6 models tested), 8–12GB (gemma4:12b-it-qat, the best-populated tier with 4 models and two clean sweeps), and 12GB+ (qwen3:32b, dethroning two prior champions within the same session). 4–8GB is the pool's most-tested tier by model count (4 models, 2 families) but its one genuinely unresolved result: no model there has cleared all six tiers (§3A.2q) — worth revisiting with a fifth candidate before concluding this is a real property of the size range rather than a sampling artifact.
8. Lessons learned (for future readers — human or AI)
This project's raw output is a leaderboard. Its more durable output is what running it actually taught us about how to evaluate a model at all. If you're building on this work, or running a similar evaluation of your own, the following is worth internalizing before you trust a single number this document produced.
8.1 Capability does not generalize across task categories
Going in, it was reasonable to assume a model's overall "quality" would show up consistently — that a model weak at one thing would be weak at most things. That assumption did not survive contact with the data. qwen3:8b and qwen3:14b both had genuine, embarrassing code_gen failures (catastrophic repetition loops on a simple edge-case prompt — §6) and were flawless at every tier of tool-calling (§3B.6). gemma4:e4b had the single worst gemma4-family result anywhere in code_gen (§3A.2p) and swept tool-calling perfectly (§3B.4). Only one model in the entire 17-model pool, qwen3:32b, turned out to be genuinely excellent at both categories (§3B.8).
The implication for anyone using this data: a model's score in one category is close to useless as a predictor for another. If your use case involves both code generation and tool-calling, you need both numbers — there is no shortcut through a single "how good is this model" scalar. This is probably the single most important finding of the whole project, more important than any individual model's ranking.
8.2 A failure score without a diagnosis is close to meaningless
Two models scored identically early in the tool_calling work — phi4-mini and alibayram/smollm3, both 5/30 (§3B.3) — for completely different reasons. One attempted the correct call in the wrong format (a fixable, template-level problem). The other showed no evidence of attending to the available tools at all (a much deeper gap, and possibly specific to a community repackaging rather than the underlying model). A published number of "5/30" collapses that distinction entirely, and the distinction is exactly the information someone would need to decide whether to debug their Ollama config or switch models.
This pattern repeated constantly: mistral:7b's tier-2 failure (over-eager tool calling) and mistral-small3.2:24b's tier-5 failure (correctly starts a multi-step chain but doesn't finish it) are both "failures," but they tell a builder to fix completely different things. The two-stage evaluation gate (§3A.3) — pass/fail first, diagnosis always, quality/speed only among passes — exists specifically to prevent a bare number from standing in for an explanation. Never trust a score you haven't opened up and read the actual failing output for.
8.3 Your test harness is being tested too — assume the bug is yours before the model's
Five separate measurement-infrastructure bugs were found and fixed over the course of this project, and every single one of them initially looked like a genuine model failure:
- Naive code extraction picked the wrong fenced block, discarding a model's correct answer (§3A.2c)
- A missing thinking-trace capture silently lost diagnostic data for reasoning models (§3A.2e)
- Windows console encoding (cp1252) crashed on emoji in a model's own
print()output, penalizing correct code for an unrelated cosmetic choice (§3A.2o) - A syntax error confined to throwaway demo code — a typo three lines below a perfectly correct function — failed the entire response, because the harness gave up extraction on the first parse error instead of trimming the broken tail (§3A.2n)
- Community-model filenames containing
/silently broke output-file writes on Windows, losing an entire model's detailed run data with no error raised (§3A.2l)
None of these were exotic. Every one was found by actually reading a "failing" response before accepting the failure, and every one changed the recorded result once fixed — sometimes from a near-zero score to a perfect one (§3A.2n: qwen3:8b's tier 1 went from 1/5 to 5/5 purely from the harness fix, no change to the model). If a result looks surprising, especially if it looks surprisingly bad, the harness is at least as likely a suspect as the model — check it before you publish the number.
8.4 Escalating difficulty + a locked evaluation rubric produces genuinely differentiated, trustworthy data
The Ike Gauntlet-inspired design (§3A.1) — a fixed course, escalating difficulty, versioned when a tier stops differentiating models — did what it was supposed to do: tiers 1/3/4 in code_gen went unanimous-clean early and became retirement candidates, while tiers 2/5/6 reliably separated real capability differences (§3A.2h). Locking rubric decisions once made — e.g., committing to "raises vs. returns None" as a real, countable distinction even when a model's alternate choice is defensible (§3A.2c) — matters more than it might seem: it's what stops the evaluator from quietly moving the goalposts to be kind to a model that "almost" got it right. A test that can be argued into a pass after the fact isn't measuring anything.
8.5 Not every question gets a clean answer — publish the honest unresolved ones too
Despite real, sustained effort — four different models across two families — no model in the 4-8GB VRAM tier ever achieved a clean sweep on code_gen (§3A.2q). That's a real, reportable finding in its own right, not a gap to quietly omit or keep chasing indefinitely. A dataset that only reports clean, resolved findings is systematically biased toward looking more complete than the underlying reality — future readers should expect and tolerate "we don't know yet" as a legitimate outcome, not a failure of the evaluation itself.
8.6 Practical takeaways for anyone replicating this approach
- Design the harness to fail loudly, not silently. Several of the bugs above (the filename bug in particular) lost data with no visible error — worth actively testing your own measurement pipeline's failure modes before trusting its successes.
- Diagnose before you score. Build the habit of reading at least one raw failing response per new failure signature before recording a number — the fastest way to catch both harness bugs and to distinguish genuinely different failure modes that would otherwise look identical.
- Test more than one model before trusting a tier design. A tier that looks perfectly calibrated against one strong model (a 30/30 clean sweep) tells you almost nothing about whether it actually differentiates — the real signal comes from testing a genuinely weaker model next and confirming the tier can fail (§3B.3's
phi4-minipilot, deliberately chosen as a likely-weak second test case). - Keep VRAM/hardware framing separate from architecture framing. Grouping by measured runtime VRAM rather than marketing parameter count or download size (§7.1) repeatedly surfaced real, non-obvious results —
codellama:7b's unexpectedly large runtime footprint relative to its download size (§3A.2q) being one of the more concrete examples. - A "reigning champion" needs the same confirmation rigor as a failure. A lucky clean run isn't stronger evidence than an unlucky failed one; both need the same N-repeat bar before either gets trusted (§3A.6).
9. How this compares to the field (August 2026)
This project was never built in a vacuum, but it also wasn't built by copying an existing methodology — it's worth a deliberate look at what else exists, what's genuinely different here, and where a v2 could close real gaps rather than duplicate work that's already well-served elsewhere.
9.1 What's standard in the field right now
Frontier-model leaderboards — MMLU-Pro, GPQA Diamond, SWE-bench Verified, ARC-AGI-2, Humanity's Last Exam, and the various aggregator sites that track them — are the numbers that dominate launch blog posts and comparison articles. As of mid-2026, MMLU, HumanEval, and GSM8K are widely described as saturated (top models cluster at 88%+), with GPQA Diamond, SWE-bench Verified, and HLE now doing the real differentiating work at the frontier. Almost all of these are run against API-hosted, full-precision models — they tell a reader almost nothing about how a quantized model behaves on their own GPU.
Human-preference arenas (LMArena, formerly LMSYS Chatbot Arena) take a completely different approach: anonymized pairwise comparisons voted on by real users, aggregated into an Elo rating. This is a genuinely complementary signal to deterministic testing — it captures "which response do people actually prefer," which a fixed-answer gauntlet structurally cannot.
Domain specialists exist and are well-built. The Berkeley Function Calling Leaderboard (BFCL) is the closest real analog to this project's tool_calling category — it uses an AST-based structural evaluation across thousands of function-call examples, spans multiple programming languages, and (as of v3) includes multi-turn, multi-step state tracking with categories for missing-function and long-context scenarios. On the RAG side, Vectara's hallucination leaderboard (built on their HHEM detection model, recently extended with an LLM-as-judge framework called FaithJudge) and the broader RAGAS-style metric set (faithfulness, context precision, context recall, answer relevancy) are the closest analogs to this project's RAG category — though nearly all of them lean on an LLM judge rather than deterministic exact-match grading.
Local/self-hosted-aware benchmark sites exist too, and there are more of them than might be expected — VRAM-tier buying guides, Ollama-specific throughput comparisons on specific consumer GPUs (RTX 4080/4090/5090 write-ups are common), and aggregator sites that map existing benchmark scores onto hardware tiers. But the overwhelming majority of this content measures tokens-per-second, not correctness — or reports someone else's cloud-hosted benchmark number next to a VRAM figure, rather than running a fresh quality test on the actual quantized model a reader would download.
Failure-mode taxonomies exist in the academic literature, but a paper surfaced during this research states the state of the field plainly: failure-mode analysis "is currently mostly qualitative... rarely with mechanistic evidence about why the failure happened," and even AgentBench's structured per-task taxonomy — one of the few benchmarks that ships one at all — stays at the level of observable symptoms ("Context Limit Exceeded," "Invalid Format") rather than diagnosing the actual underlying mechanism.
9.2 What's genuinely different about this project
-
Real local, VRAM-measured, quantized models tested for capability, not throughput. This is the clearest gap being filled — the local-LLM benchmark space is dense with speed comparisons and thin on anyone actually checking whether the quantized model gets the right answer, tier by tier, on hardware a self-hoster would actually own.
-
The identical model pool run across four independent capability categories, with orthogonality tracked as a first-class finding. Nobody encountered in this research runs the same 17 models through unrelated categories specifically to demonstrate that a model's rank in one predicts almost nothing about another (§8.1). Category leaderboards elsewhere are typically published independently, without the cross-referencing that makes this project's central finding visible at all.
-
Mechanism-level failure diagnosis, not symptom-level. Given the literature's own assessment that this is rare even in dedicated failure-taxonomy research, tracing a specific arithmetic slip to "the model treats June as a full 30-day elapsed block instead of the 29 days actually remaining after June 1st" (§3C.7) is a genuinely deeper level of diagnosis than field-standard practice.
-
Convergent cross-model error discovery. Nothing found during this research specifically hunts for unrelated model architectures landing on the identical wrong specific answer. The "57 balloons" and "October 23" clusters (§3C.7, §3D.4) are closer to an empirical finding in error analysis than a typical benchmark result — and the project found four independent instances of this pattern across two different task categories, not one.
-
Full harness-bug transparency as part of the public record. Publishing every measurement bug found and fixed, with before/after scores (twenty of them across four categories — §8.3), is essentially unheard of in benchmark reporting. The convention elsewhere is to publish a final number and keep the measurement pipeline's debugging history private, if it's recorded at all.
-
TMV/DOE quality-engineering rigor transplanted into LLM evaluation. Locked rubrics decided before testing (§2.4), N-repeat confirmation before trusting a champion (§3A.6), and a two-stage pass/fail-then-quality gate (§3A.3) come from manufacturing and process-engineering discipline, not the bootstrap-confidence-interval and judge-agreement-kappa toolkit that dominates the academic RAG/agent-evaluation papers surfaced in this research.
9.3 Real gaps worth closing in a v2
Roughly ordered by how much each would add relative to the effort of building it:
-
Quantization as a tested factor. The original DOE design (§3.1) flagged Q4/Q5/Q8 as a Stage 2 variable and it was never built — every result in this project reflects Ollama's default quantization only. Local-LLM-focused sites treat this as a first-class axis; at minimum, re-testing each category's champions across quantization levels would answer a question this project currently can't.
-
Real inference-speed benchmarking as a first-class metric, not just a tie-break.
total_time_secis currently only used to break ties between equal scores. For a self-hoster, "fast enough" often matters as much as "best" once a capability floor is cleared — a dedicated throughput pass (tokens/sec under realistic load, not just total wall-clock per gauntlet run) would make the leaderboard more directly actionable. -
True long-context testing. RAG's tier 6 (§3C.2) tops out around 600 words. Established long-context evaluation goes to 32K–128K+ tokens — a genuinely harsher and more realistic test of how local models actually degrade as context fills, which is a common real-world failure point this project hasn't touched.
-
Genuine multi-turn agentic testing. Tool-calling's tier 5 (§3B.2) is a two-turn toy version of what BFCL v3's multi-turn, multi-step state tracking or tools like AgentBench do across much longer, statefully-tracked interactions. A v2 tool-calling category could extend the chaining tier into a real multi-step agent trajectory.
-
Cross-validation against published benchmark numbers for the same models. Does this project's code_gen ranking correlate with published HumanEval or LiveCodeBench scores for the exact same models? Where it diverges would itself be a genuine finding — quantization effects rarely show up in academic numbers that report full-precision weights, and this project is uniquely positioned to surface that gap since it already has the quantized-model data on hand.
-
A carefully-designed LLM-judge tier, held to the same rigor as everything else. A locked rubric, calibration checks against known-good and known-bad examples, and ideally cross-model judge agreement, would finally make summarization and creative-writing testing possible — twice declined in this project (§3B, §3D) for lacking a clean deterministic answer, but by external accounts the single most-used real-world local-LLM category. Getting this right is more work than anything else on this list, but it would close the project's largest remaining blind spot.
10. What to publish, what to hold back (August 2026)
This project's whole premise is transparency — TMV/DOE rigor only means something if someone else can check the work. But full transparency has a specific, documented cost that cuts directly against the project's own goals, and it's worth stating the tradeoff plainly rather than defaulting silently to one side.
10.1 The contamination risk is real, not hypothetical
§9.1 already notes that MMLU, HumanEval, and GSM8K are now considered saturated and unreliable at the frontier — models have plausibly seen the exact questions (or close paraphrases) during training, so a high score increasingly reflects memorization rather than the capability the benchmark was designed to measure. This is exactly the failure mode a public page risks recreating at a smaller scale: if this project's exact prompts and exact expected answers sit indefinitely on an indexed public page, a future model's training run can absorb them, and every tier that depends on that specific prompt stops measuring what it was built to measure. Tier 4 of reasoning-gauntlet-v1 (§3D.1) is the sharpest example — it's the classic, well-known bat-and-ball cognitive-reflection question, chosen specifically because it's a famous, well-documented trap; publishing the literal prompt text makes it trivially easy for a future model to have memorized "$0.05, not $0.10" without ever having done the arithmetic. BFCL's response to this exact problem (§9.1) is instructive: it maintains a rotating "Live" test set of fresh, user-contributed examples specifically so a fixed prompt pool can't be memorized into invisibility.
10.2 What gets published, and what doesn't
Published in full: the methodology itself — tier ladder structure and rationale, the two-stage evaluation gate, the escalating-difficulty design principles, every finding and cross-category pattern, the harness-bug log, the VRAM-tiering approach, and the model-by-model score data. None of this is the part at risk; copying a methodology isn't the kind of "cheating" this section is worried about, and a methodology nobody can inspect isn't worth much anyway.
Held back: the literal prompt text and literal expected answers for tiers where memorization would quietly break the tier — reasoning/math's word problems and puzzle text (§3D.1) most of all, since exact numbers are the easiest thing in the world to memorize, and RAG's specific passages (§3C.2) close behind, since a memorized passage defeats the entire point of testing context-faithfulness. Code_gen and tool_calling's prompts are lower-risk (the tasks require genuine synthesis even if the prompt itself were memorized, since the interesting failures are in execution, not recall) but the same caution applies by default.
The practical version of this, going forward: describe each tier's design publicly ("a constraint-satisfaction puzzle with a unique solution," "the classic bat-and-ball cognitive-reflection problem," "a word problem with two categories of excluded entities and two genuinely irrelevant details") without the literal prompt text, and keep a private, rotating pool of structurally-identical items — same difficulty, same trap, different specific numbers and names — so a tier that starts showing contamination symptoms (every model suddenly clean-sweeping a tier that used to differentiate) can be quietly refreshed without redesigning it from scratch. This is a genuine v2 task, not yet built — the current tier content for all four categories remains the original, single fixed version described in §3A–§3D.
11. Publishing the site (August 2026)
The leaderboard existed as a live FastAPI service on a home workstation for most of this project. This section documents the step from there to a real, publicly reachable site at llmgauntlet.com — an addition to the project that's as much infrastructure as anything already logged in §6-§7, and worth the same record.
11.1 Deployment architecture — freeze to static, not expose the workstation
The obvious approach — pointing a domain directly at the home workstation's FastAPI service — was rejected for two reasons. First, it would mean exposing a real Windows machine to the public internet indefinitely, either directly or via a tunnel, a real and ongoing security surface. Second, it would tie the site's uptime to the workstation staying powered on and connected, forever, for a project meant to be a durable public record.
Every page this site serves is actually static: the leaderboard data comes from JSON files regenerated by the aggregate scripts (§4), not computed live per visitor. This made the site a clean fit for full pre-rendering: a crawl script (freeze_site.py) walks every route the live FastAPI app serves — the four category leaderboards, all model-detail pages across all four categories, the /about and /design-doc pages — and saves each page's rendered HTML to a matching static file path. The resulting folder (75 pages, under 1MB total) deploys to Cloudflare Pages, which provides free hosting, a global CDN, and automatic HTTPS, with the custom domain attached directly since it was registered through Cloudflare's own registrar.
A real link-rewriting bug, caught and fixed before publishing. Model IDs like gemma4:e2b contain colons, which work fine in the live FastAPI route (/model/gemma4:e2b) but are unsafe as folder names on most filesystems and static hosts. The freeze script rewrites every colon to a dash consistently in both the saved file paths and every internal href. The first version of this rewriting logic tried all four category-specific URL prefixes on every page rather than only the one relevant to that page's own category — since the generic /model/ prefix is itself a substring of every category-specific prefix (/reasoning/model/, /rag/model/, etc.), this caused a double-substitution specifically for phi4-mini, the one model ID among all 17 with neither a colon nor a slash needing replacement, producing a broken double-trailing-slash link (/reasoning/model/phi4-mini//) on every category page. Fixed by having each page's crawl pass only its own category's prefix rather than trying all four speculatively — the same "assume the harness is at fault before the model or, here, the page" discipline from §8.3, now applied to the publishing layer instead of the testing layer.
11.2 A real privacy-policy violation, found and fixed on the live site
§10 establishes that RAG and reasoning/math hold back their literal prompts and expected answers from publication, to avoid training-data contamination. The first version of the live detail pages for these two categories did not actually honor that policy — they rendered the full response_text, thinking_text, and the per-tier detail diagnosis string for every tier, and since models routinely restate the full problem (exact numbers, exact passages) when answering, this fully exposed the content the site's own stated policy promised not to publish. The gap was caught by directly viewing a live reasoning detail page and noticing the literal word problem and its exact answer sitting in the rendered response text.
Fixed at the data layer, not just the template: the model-detail loading functions for these two categories no longer return detail, response_text, or thinking_text at all, so the redaction can't be undone by a future template change without someone deliberately re-adding those fields. The detail pages for these categories now show only tier name, pass/fail count, token count, and elapsed time, plus a visible note explaining why and linking to §10's fuller reasoning. Code_gen and tool_calling detail pages were deliberately left untouched, since §10.2 already marks those categories fully open.
This is the same kind of finding as every harness bug already logged in §8.3 — a piece of tooling not actually doing what it was supposed to, caught by looking at the real output rather than trusting the code's intent — just discovered on the publishing side of the project rather than the testing side.
11.3 The overall-champion radar chart — a new way to see §8.1's central finding
§8.1 states the project's most important finding in prose: a model's rank in one category predicts almost nothing about its rank in another. The site's original landing page (the code_gen leaderboard) didn't actually show this — a visitor had to read §8.1's argument and take it on faith, or manually cross-reference four separate leaderboard pages themselves.
The site's new landing page computes, for each VRAM tier, the single model with the highest score summed across all four categories (out of 120 total), and plots that model's four category scores as a radar chart — one polygon per tier, four axes per chart. The shape of the polygon does the work prose can't: a model that's consistently strong across every category produces a near-perfect square: a model whose total score is high but uneven produces a visibly lopsided shape.
The four overall champions, verified directly against the aggregate JSON files rather than computed by hand:
| VRAM tier | Overall champion | Total (of 120) | Shape |
|---|---|---|---|
| Sub-4GB | gemma4:e2b |
120/120 | Perfect square — the only model in the entire pool with a maximum score in every category |
| 4-8GB | qwen3:8b |
105/120 | Visibly lopsided — a real dip in code_gen (20/30) offset by two perfect categories (tool_calling, reasoning, both 30/30) |
| 8-12GB | gemma4:12b-it-qat |
115/120 | Near-square, one real dip (RAG, 25/30) |
| 12GB+ | gemma4:26b |
114/120 | Near-square, edges out qwen3:32b (113/120) by a single point |
The 4-8GB result is the most interesting of the four and the one the radar chart makes legible at a glance: qwen3:8b is the overall champion for its tier despite never being the single best model in any one category there (qwen3.5:9b and qwen2.5-coder:7b both individually score higher in code_gen) — its total comes specifically from consistency across the categories where it does compete well, which is exactly the kind of pattern a single leaderboard column can't show and §8.1's argument otherwise has to state in words rather than demonstrate.
11.4 The site's other public pages
Two pages exist alongside the leaderboards and this document, each serving a distinct purpose:
/about is written for a first-time visitor — human or AI — rather than for someone already deep in the methodology. It opens with the "57 balloons" convergent-error finding (§3D.3) as a concrete hook rather than an abstract pitch, states §8.1's orthogonality finding directly, and summarizes §10's publish/hold-back reasoning with a visual "transparency strip" showing which categories are fully open versus design-public-text-held-back.
/design-doc is this document itself, rendered from the markdown source directly into a page matching the site's visual design — the full methodology, available to anyone who wants to verify a specific claim rather than take §about's summary on faith. Both /about mentions of "the design doc" link here; before this page existed, those references pointed nowhere.
12. Family-scaling investigation: what does a larger model actually earn for its size? (August 2026)
Motivated by a direct question: does a model's rank on the existing four categories (§3A-§3D) actually tell a reader what a larger variant in the same family buys them, or does it just say "bigger scored the same as smaller"? gemma4 already has five sizes tested across all four categories (§11.3), and the aggregate picture is close to flat — §8.1's "capability does not generalize across categories" finding, examined here at the size axis within one family instead of across categories. The blunt version of the question: if the smallest model in the whole pool is already this good, what is the extra parameter budget of the largest variant actually reinforcing?
12.1 Grounding the hypotheses in the vendor's own documentation, not guesswork
Rather than design atomic-skill ladders speculatively, Gemma 4's own model card and technical report were consulted directly for size-differentiated claims — testable, vendor-stated design intent rather than assumption:
- Context window is explicitly size-gated. "Small models feature a 128K context window, while the medium models support 256K." A hard, documented, testable number, not a fuzzy capability claim.
- Agentic coding is called out as a flagship-tier investment. The largest dense variant is specifically credited with "substantial gains... along every dimension relevant to agentic coding" relative to the prior generation.
- Thinking mode exists at every size, but nothing in the documentation claims thinking quality is equal across sizes — an open, untested claim.
The existing tool_calling category's chaining tier (§3B.2) only tests 2 sequential calls, and every gemma4 size cleared it perfectly (§3B.8) — almost certainly a ceiling effect from a task too shallow to expose claim #2, not genuine evidence the sizes are equal there. Two new categories were built specifically to test claims #1 and #2 at real difficulty. Claim #3 (thinking-mode reasoning quality) remains open — see §12.4.
12.2 Context-length category — methodology and result
Design: a single fact ("needle") planted at the exact midpoint of a generated passage, escalating context length as the tier axis (2K/8K/16K/32K/64K/128K tokens) rather than task difficulty. Midpoint placement was a deliberate choice, not arbitrary — the documented "lost in the middle" effect means models reliably do well at passage start (early-token attention bias) and end (recency), while the middle is where genuine long-context handling shows through unmasked by either advantage. The passage itself is proceduralized — many short, varied, deterministic log entries (not repeated text a model could pattern-match around), generated fresh per seed, mirroring the "Needle in a Haystack" methodology used in established long-context evaluation.
Two harness bugs found and fixed before any data was trusted:
- Silent context truncation. Without an explicit
num_ctx, Ollama silently truncates a long prompt at its default context window with no error — confirmed directly: a genuine 9,533-token passage was silently cut to 4,099 tokens withnum_ctxunset. Fixed by explicitly settingnum_ctxon every request, sized to each tier's target token count with headroom. - Undersized response budget. The initial
num_predict: 30(reasonable for a short factual answer in isolation) caused every single request to return a completely empty response,done_reason: "length"— the model was spending its entire budget on internal reasoning before ever reaching a visible answer, and got cut off first. Raisingnum_predictto 1000 resolved this immediately; a representative response used 196 tokens total for a 13-character visible answer.
A third, more serious finding, caught while pushing deliberately past the documented ceiling: asking Ollama for num_ctx above a model's real compiled context length does not fail safely. Requesting 234,000 tokens against gemma4:e2b (real ceiling: 131,072, confirmed via ollama show) did not error or cleanly clamp — it silently truncated to an unrelated ~65,539-token prompt (suspiciously close to 2^16, not any number tied to the model's actual architecture) and returned an empty, broken response. The harness now queries each model's real ceiling via ollama show before every run, caps num_ctx at or below that real number regardless of the tier's nominal target, and cleanly skips (rather than silently corrupts) any tier that genuinely exceeds a given model's context window — itself real, reportable data once the full 17-model pool is tested, since not every model will have a 128K-class window.
Result: gemma4:e2b and gemma4:26b both scored a perfect 30/30, every tier, all the way to ~110,000 tokens of real context (the tier-6 target of 128,000 was not fully reached due to word-to-token calibration imprecision at that scale — a known, minor limitation, not a truncation bug). The one real, measured difference between the two sizes was speed, not accuracy — 26b ran roughly 2-3x slower than e2b at every tier, with identical retrieval correctness throughout.
12.3 Agentic-depth category — methodology and result
Design: a single realistic multi-step scenario (a customer-service lookup chain: customer name → customer ID → order ID → tracking number → carrier → delivery estimate → carrier contact), escalating chain depth as the tier axis (1 through 6 sequential, genuinely dependent tool calls) rather than task complexity. Each step's correct argument value is the actual output of the previous step, not just plausible-looking sequential calls — grading checks both the correct tool name and the correct argument value at every step, so a model that calls the right tool with a stale or hallucinated argument fails exactly as much as one that picks the wrong tool. Uses /api/chat (genuine multi-turn conversation state, injecting a synthetic tool-result message after each correct call), not /api/generate.
Same undersized-budget bug reappeared, more severely, on the harder version of the task. The initial num_predict: 500 was sufficient for the first version of the prompt, but after the prompt was revised to require two pieces of information instead of one (see below), the model needed genuinely more upfront planning — confirmed directly via the thinking field, which showed a still-in-progress 6-step strategic plan cut off mid-thought at exactly 500 tokens. Raised to 2000; resolved immediately.
A genuine test-design flaw was caught and fixed — not a model capability crack. The original single-goal task ("find the delivery date") let the model reasonably skip the carrier-contact step once it had enough information to answer the question asked — confirmed by the fact that both gemma4:e2b's deepest tiers failed at the identical step, every single time, with the identical substitution (check_delivery_estimate instead of get_carrier_contact), 100% consistent across all seeds. That signature is a decision, not confusion: the model correctly recognized the contact-lookup step was unnecessary busywork for the question actually asked. Per §2.4's standing rule (genuine ambiguities get fixed when caught during the pilot, not grandfathered), the prompt was revised to require both pieces of information explicitly, with an explicit required order (delivery date first, then carrier contact) to remove a second, order-related ambiguity.
Result, on the corrected task: gemma4:e2b and gemma4:26b both scored a perfect 30/30, every tier, 1 through 6 dependent steps. As with the context category, the only measured difference was speed — 26b ran consistently slower (roughly 1.5-2x) at every tier, with identical correctness throughout.
12.4 Where this leaves the original question
Two of the three vendor-documented, size-differentiated claims have now been tested at genuine difficulty, not just at the ceiling of the existing four categories' shallower tiers, and both came back flat: gemma4's extra parameter budget at the 26B size does not buy more long-context retrieval accuracy or deeper agentic tool-chaining reliability than the 1.8GB e2b variant already has. This is a real, honest negative result — not a failure of the test design (both categories are validated, harness-bug-free, and genuinely hard enough to have shown a gap if one existed), but a genuine finding about this specific family: at least on these two axes, size buys latency, not capability.
The third claim — thinking-mode reasoning quality, as distinct from thinking-mode availability (present at every size per the model card) — remains untested, and is a fundamentally harder problem than the first two: unlike context-length or chain-depth, there is no single obvious escalation axis, and a naive "harder word problems" ladder risks re-measuring the same thing §3D already covers rather than isolating reasoning quality specifically. Not yet resolved as of this writing — see the pending design discussion.
12.5 The third claim: thinking-mode reasoning quality, tested via analogy self-critique
§12.4 left one vendor-documented claim untested — thinking-mode reasoning quality, as distinct from mere availability — because unlike context length or chain depth, there is no obvious escalation axis, and a naive "harder word problems" ladder risks re-measuring what §3D already covers.
The design settled on is grounded in a general property: certain outputs are hard to produce but easy to recognize once produced — a genuinely good analogy, a well-designed interface, a concise summary, a useful insight buried in a pile of information. Explaining a complex concept to a specific audience via analogy sits squarely in this category, and adding one instruction sharpens it considerably: after giving the analogy, the model must identify specifically where its own analogy stops being accurate. A model producing a plausible-sounding metaphor without genuine cross-domain understanding will give a generic hedge or a technical-sounding but wrong limitation; a model that actually understands both domains will name the real, specific structural gap.
Six tiers, escalating domain distance between concept and audience, rather than escalating difficulty within one domain:
| Tier | Concept | Audience |
|---|---|---|
| 1 | Car engine cooling (thermostat-controlled) | Home HVAC technician |
| 2 | Database index | Librarian |
| 3 | Database index (same concept, different audience) | Warehouse manager |
| 4 | TCP reliable delivery (ACK/retransmission) | Restaurant kitchen manager |
| 5 | Gradient descent optimization | Professional gardener |
| 6 | Distributed consensus / split-brain avoidance | Ballet choreographer |
Each tier is locked before testing with three separately-graded criteria — never collapsed into one holistic quality score, keeping the same discipline as every other category in this project:
- Structural fidelity — a fixed checklist (3 items per tier) of the specific ideas a correct analogy must represent, graded by a locked-rubric LLM judge answering one yes/no question per item, never a holistic rating.
- Audience jargon — a fixed blocklist of source-domain terms that should not appear unexplained for that specific audience, checked deterministically.
- Self-critique accuracy — a pre-authored list (3 items per tier) of the analogy's real, known limitations, written before any model is tested. The model's stated limitation is graded by whether it genuinely matches one of these real gaps, not by whether it merely sounds humble.
Calibration before trusting the judge on real data, per §9.3's own standing bar for LLM-judge tiers: three hand-written variants per tier — GOOD (strong analogy, correct specific critique), MEDIOCRE (weak analogy, generic critique), and critically, GOOD_WRONG_CRITIQUE (the same strong analogy as GOOD, paired with a plausible-sounding but factually wrong critique). The decisive test: does the judge score GOOD_WRONG_CRITIQUE's analogy portion identically to GOOD's (confirming it isn't penalizing structure it shouldn't), while correctly failing it on critique accuracy specifically (confirming it isn't just rewarding confident-sounding humility)? On the hardened pipeline, this distinction held cleanly and unanimously on both piloted tiers before any real model was tested.
12.6 Harness bugs found and fixed
The same undersized-response-budget bug recurred four separate times before the pipeline was trustworthy, each instance independently diagnosed via the same signature already established in §12.2§12.3: done_reason: "length", eval_count exactly matching the cap, and a completely empty visible response. First in the judge's structural-fidelity check (500 — insufficient), then again after the harder two-goal critique-matching question (raised to 2000 — still insufficient on a denser critique), finally resolved at 4096, matching the ceiling already standard elsewhere in this project. The lesson generalizes past this one category: a locked-rubric LLM judge doing real reasoning needs the same generous budget as any other reasoning task — there is no such thing as a "simple yes/no question" that's safe to under-provision.
A real single-call non-determinism problem, caught by the calibration set itself. The first calibration run showed GOOD and GOOD_WRONG_CRITIQUE — sharing byte-identical analogy text — scoring differently on structural fidelity (2/3 vs 3/3) despite temperature: 0, which should be deterministic. Fixed by asking every judge question three times with different seeds and deciding by majority vote, mirroring the N-repeat discipline (§2.2) already applied to every other category. Re-running calibration under the hardened pipeline produced zero ambiguous votes and fully consistent scoring.
A self-inflicted bug: changing the response-parsing marker without updating the calibration text to match. The prompts sent to real models were revised to request a specific marker (LIMITATION:) so the analogy and critique portions could be reliably separated, but the calibration examples — authored earlier with a different marker phrase ("Where this breaks down:") — were not updated to match. This silently caused every critique-match check to skip entirely (votes=[]) rather than fail loudly, since the code's fallback for "no critique paragraph found" is indistinguishable from "marker just isn't there." Caught by noticing the calibration results had regressed to look nothing like the previous run, not by any error message.
A genuine tier-design flaw, caught on the very first real-model pilot, not the model's fault. Tier 1's jargon blocklist originally included "heat exchanger" as a term a model shouldn't use unexplained for a home HVAC technician audience. But a home HVAC technician's own equipment (furnaces) centers on heat exchangers — it's professional vocabulary they already own, not obscure jargon. The model used it correctly, every seed, specifically as a translation aid ("the radiator acts like your home's air handler or heat exchanger"), and was penalized for genuinely good, audience-appropriate analogy writing. Fixed by removing it from tier 1's blocklist, with the reasoning documented in the data file itself so the decision isn't silently lost.
12.7 Result: the first real crack found in the whole family-scaling investigation
An initial six-tier pass at N=1 against both models showed near-total agreement — identical failure shapes on tiers 2 through 6 (same jargon slips, same structural gaps, same critique misses) — except tier 1, where gemma4:26b passed and gemma4:e2b failed. Per §3A.6's standing rule that a result needs the same confirmation rigor whether it's a win or a loss, tier 1 was re-run at N=5 for both models before trusting it.
The confirmation held, cleanly and completely:
gemma4:e2b (5 seeds) |
gemma4:26b (5 seeds) |
|
|---|---|---|
| Structural fidelity | 3/3, every seed | 3/3, every seed |
| Audience jargon | clean, every seed | clean, every seed |
| Self-critique accuracy | fails, every seed | passes, every seed |
Both models build the analogy itself identically well — perfect, unanimous agreement across all 10 combined seeds on structural fidelity and jargon. They diverge with equal, total consistency on exactly one thing: whether the stated limitation actually names the real structural gap, or is a confident, technical-sounding generality that never touches it. gemma4:e2b's critiques for this tier consistently describe the two domains as "operating on different principles" in increasingly technical-sounding ways without ever identifying the specific setpoint-vs-passive-valve distinction that is the tier's actual pre-authored limitation; gemma4:26b's critiques name it directly, every time.
This is the first place in the entire family-scaling investigation (§12.1§12.3) where the larger model's extra parameter budget demonstrably buys something. Context-length retrieval (§12.2) and agentic chain depth (§12.3) both came back completely flat between the smallest and largest gemma4 variants — genuine, honest negative results. This result lands exactly where the model card's third, previously-untested claim pointed: not raw output-building capability, but the quality of reasoning about the output's own limitations.
Update: all six tiers now confirmed at N=5, not just tier 1. Testing continued tier by tier, each confirmed independently before moving to the next:
| Tier | gemma4:e2b (5 seeds) |
gemma4:26b (5 seeds) |
Where size mattered |
|---|---|---|---|
| 1 | fails (critique unmatched) | passes (critique matched) | Self-critique accuracy — clean win |
| 2 | 2/3 structural (misses write-cost tradeoff) | 3/3 structural | Structural completeness — clean win |
| 3 | 1/3 structural | 1/3 structural | No difference — shared weakness |
| 4 | jargon leak (1 term) | jargon leak (2 terms) | No difference — size slightly worse |
| 5 | jargon leak (2 terms), critique matched | jargon leak (1 term), critique matched | Jargon severity — modest win |
| 6 | jargon leak (4 terms), critique unmatched | jargon leak (1 term), critique unmatched | Jargon severity — substantial win |
The complete, honest finding is more nuanced than a single clean story. Size buys something real in this category, but it isn't one consistent advantage — it shows up as better self-critique reasoning on tier 1, more complete structural coverage on tier 2, and increasingly disciplined jargon restraint as domain distance grows (tiers 5—6), while tier 3 is a genuine shared limitation extra parameters don't fix, and tier 4 shows size can even be a mild liability (more jargon leaked, not less). This is a richer and more defensible result than "bigger model wins everywhere" would have been, and every cell in the table above is independently confirmed at N=5, not inferred from a single sample.
One real open question this raises: everything tested so far in §12 uses only gemma4, so this result cannot yet distinguish a genuine, general "size buys better self-critique and jargon discipline" pattern from something idiosyncratic to how this specific family was trained — exactly the kind of family-specific-vs-general distinction §3.1 and §6 already flagged as a real risk early in this project. Testing the same six-tier ladder against a second family with real size spread (qwen3:8b vs qwen3:32b, already deeply characterized across every other category in this project) is the natural next step before treating this finding as more than a single-family result.
12.8 Cross-family confirmation: does the finding generalize, or is it gemma4-specific?
§12.7 flagged an open question: everything tested to that point used only gemma4, so the finding could not yet distinguish a genuine, general "size buys better self-critique and jargon discipline" pattern from something idiosyncratic to how this one family was trained. The same six-tier ladder was run against a second family with real size spread already characterized elsewhere in this project — qwen3:8b and qwen3:32b — confirmed at N=5 on every tier, using the identical judge model (gemma4:26b) throughout.
The complete, confirmed comparison:
| Tier | gemma4 (e2b — 26b) |
qwen3 (8b — 32b) |
|---|---|---|
| 1 | Win — self-critique accuracy | No difference — both perfect (5/5) |
| 2 | Win — structural completeness | Size worse — structural (inverted) |
| 3 | No difference — shared weakness | Win — self-critique accuracy |
| 4 | Size worse — more jargon leaked | Size worse — more jargon leaked (identical terms: "ACK", "sequence number") |
| 5 | Modest win — less jargon leaked | Size worse — more jargon leaked (inverted) |
| 6 | Substantial win — 4 terms §1 term | No difference — identical jargon (3 terms each) |
The answer is genuinely more interesting than either "yes, it generalizes" or "no, it's gemma4-specific." The two families diverge almost completely in which direction size points: gemma4 shows a real advantage from scale on four of six tiers, with only one tier showing size as a liability. qwen3 shows the opposite shape — a real advantage from scale on only one tier, with size acting as a liability on three. This directly answers §12.7's open question: whether extra parameters buy better self-critique reasoning depends on which family you're in, not on model size as a universal property. A finding built entirely on gemma4 (as §12.5§12.7 originally were) would have told a real but incomplete story — "size helps here" — that a second family shows is not a general law at all.
One genuinely striking convergence worth flagging on its own: tier 4 produced not just the same direction of effect in both families (larger leaks more jargon), but the identical specific terms leaking in the identical smaller-leaks-one, larger-leaks-two pattern — "ACK" alone for both smaller models, "ACK" plus "sequence number" for both larger ones. Two unrelated architectures, trained by different labs, converging on the exact same jargon slip at the exact same tier is the same category of finding as the "57 balloons" and "October 23" convergent-error clusters documented earlier in this project (§3D.3§3D.4, §3C.7) — independent evidence that certain prompt structures reliably trigger a shared failure mode across architectures, this time in a completely different task category (analogy generation) than where the pattern was first observed (arithmetic word problems).
What this means for the project going forward: any single-family finding in this investigation should be read as a finding about that family specifically until confirmed on a second one, exactly the same discipline §3.1 and §6 established for prompt-vs-model idiosyncrasy early in the project, now extended to family-vs-general claims about model scaling.
12.9 A third qwen3 size point: non-monotonic scaling confirmed within a family, not just across families
§12.8's qwen3 comparison used only the two endpoints (8b, 32b). qwen3:14b was added as a genuinely cheap third point — already in the pool, no new downloads, same judge model, same six tiers, N=5 throughout — specifically to test whether the divergence found between 8b and 32b is a clean threshold effect (something changes once past a certain size and stays changed) or something messier.
It's messier, and clearly so:
| Tier | qwen3:8b |
qwen3:14b |
qwen3:32b |
|---|---|---|---|
| 1 | 5/5 | 5/5 | 5/5 |
| 2 | 3/3 structural | 2/3 (matches 32b) |
2/3 |
| 3 | 2/3 structural, critique fails | 1/3 structural (worse than both siblings), critique passes | 2/3 structural, critique passes |
| 4 | jargon: 1 term | jargon: 2 terms (matches 32b) |
jargon: 2 terms |
| 5 | jargon: 1 term, critique passes | jargon: 1 term (matches 8b), critique passes |
jargon: 2 terms, critique passes |
| 6 | jargon: 3 terms, critique fails | jargon: 3 terms, critique passes (better than both siblings) | jargon: 3 terms, critique fails |
14b does not sit consistently "between" its siblings. On tiers 2 and 4 it tracks with the larger model; on tier 5 it tracks with the smaller one; on tier 3 it falls below both; on tier 6 it beats both. This directly confirms, within a single family this time rather than across two, the same non-monotonic scaling pattern already documented for gemma4's own family in the original code_gen category (§3A.2p, where gemma4:e4b was the worst variant in its family despite sitting between two stronger-performing sizes). It is not a one-off: two unrelated families, two different task categories, both show a middle-sized variant that is neither a clean interpolation nor a monotonic step between its neighbors.
Why this matters for how the whole family-scaling investigation (§12) should be read: a two-point comparison (smallest vs. largest, as §12.1§12.8 mostly used) can only ever show a straight line between two values. It cannot distinguish a genuine, smooth scaling trend from a family whose middle sizes wobble unpredictably — which is exactly what this third point reveals. Any future single-family or cross-family comparison in this project that uses only two size points should be read as a lower bound on the real complexity of the scaling curve, not the whole shape of it.
12.10 A third family: mistral confirms the convergence and adds a genuinely new axis
§12.8 and §12.9 left one open item: whether the gemma4-vs-qwen3 divergence is a genuine two-family split or whether a third family would land somewhere new entirely. mistral:7b and mistral-small3.2:24b (both already in the pool, no new downloads) were run through the same six tiers, N=5, using the same argument-driven harness now capable of running an entire model unattended end-to-end (--model, --tiers, --seeds flags replaced the hand-edited constants used for the first two families).
A genuinely new characteristic showed up before any scoring comparison was even possible: mistral:7b produces near-identical response lengths across different seeds within a tier — byte-for-byte identical in three of six tiers, and a clean bimodal split (two distinct response lengths, no in-between) in the other two. Neither gemma4 nor qwen3 showed anything like this; every prior model varied meaningfully seed to seed. On tier 1 specifically, this bimodality was consequential, not cosmetic: 7b landed in a structurally-correct response (3/3, critique matched) on seeds 1-2, and a genuinely different, physically-backwards response (the thermostat closing when the engine overheats, rather than opening — a real reasoning error, confirmed by reading the actual text, not a grading artifact) on seeds 3-5. Seed 3's response also echoed part of the prompt's own formatting instruction literally into its output ("LIMITATION: (CAPITALS)"), a distinct instruction-following slip from the reasoning error.
The comparison:
| Tier | mistral:7b |
mistral-small3.2:24b |
Where they differ |
|---|---|---|---|
| 1 | 2/5 pass, bimodal (structural 0§3, critique False§True) | 5/5 pass, fully consistent | Clean win — and the instability itself disappears |
| 2 | 5/5 pass | 5/5 pass | Identical, both perfect |
| 3 | 0/5, structural 1/3, critique fails | 0/5, structural 1/3, critique passes | Real partial improvement, not enough to flip pass/fail |
| 4 | jargon: 2 terms ("ACK", "sequence number") | jargon: 3 terms (adds "checksum") | Size worse — larger leaks more |
| 5 | jargon clean, critique fails | jargon: 2 terms, critique passes | A genuine trade-off, not a clean win either way |
| 6 | structural 2/3, jargon: "Raft" | structural 3/3, jargon: "quorum" | Better structure, same jargon severity |
Tier 1 is the headline result for this family, and it's a different shape than anything gemma4 or qwen3 showed: not a modest score improvement, but an entire failure mode (bimodal instability) disappearing at the larger size. mistral-small3.2:24b was completely deterministic — byte-identical response length across all 5 seeds, on every one of the six tiers, no exceptions. This suggests within this family specifically, added capacity may sharpen the model's effective decoding confidence enough to eliminate seed-sensitivity outright, not just improve average scores.
Tier 4's jargon leak independently confirms the cross-family convergence a third time. mistral-small3.2:24b leaks "ACK" and "sequence number" — the exact same two terms gemma4:26b and both larger qwen3 sizes leaked on this identical tier — plus one additional term ("checksum") beyond what any other model has leaked here. Three unrelated architectures, three different labs, the same tier, the same core jargon slip, and the larger model leaking more every single time it was tested. This has moved well past coincidence and belongs in the same category as this project's other documented convergent-failure clusters (§3C.7, §3D.3§3D.4).
Where this leaves the three-family comparison overall: each family shows a genuinely different scaling shape on this axis. gemma4 mostly rewards scale (§12.7: 4 of 6 tiers). qwen3 mostly does not, and its middle size doesn't even interpolate cleanly between its own endpoints (§12.8§12.9: 1 win, non-monotonic). mistral rewards scale specifically by eliminating an instability failure mode, not by raising an already-stable score (§12.10, this section). Three families, three distinct stories — reinforcing §12.8's original conclusion that whether size helps depends on family-specific training, not on any universal scaling law, while also showing that "helps" itself can mean structurally different things (higher scores vs. more consistent scores) depending on the family.
13. A model-selection decision framework
13.1 Why this framework, and what it's answering
Every section before this one has measured whether a model CAN do something. This section asks a different question, motivated by a real, practical gap: teams choosing among locally-hosted models today have almost no principled way to decide, beyond intuition, published leaderboards that don't reflect their actual task, or defaulting to "biggest model that fits my VRAM." §13 turns roughly fifteen model comparisons across five categories and three families into an actual decision procedure — grounded entirely in evidence already gathered in this project, not new claims.
13.2 The central finding this framework is built on
Model capability does not vary along one axis ("bigger is better"). It varies along at least three genuinely independent axes, and a real decision procedure has to ask about each one separately:
- Task determinism — does this task have a single, checkable correct answer (code compiles, a tool call has the right arguments, a RAG answer matches the source), or does it require balancing competing, legitimate considerations (which fact matters to THIS reader, where does THIS analogy's mapping actually break down)?
- Family — which lab trained the model, and does that specific family show a known strength or weakness on this task shape? Size alone does not predict this — two families can respond to identical scaling in opposite directions, as documented directly in §12.8.
- Reliability — does this model produce consistent output across runs of the same task, or does it show real seed-to-seed instability that a single spot-check would miss entirely, as documented in §12.10?
13.3 The decision table
| Task shape | Governing axis | What the evidence shows | Recommendation |
|---|---|---|---|
| Single correct answer, tight spec (code generation, tool-calling, RAG lookup, arithmetic/logic) | Determinism | gemma4:e2b (1.8GB, the smallest model in the entire 17-model pool) scored a perfect 120/120 across all four original categories — better than every larger model tested, including gemma4:26b |
Pick the smallest model that clears your accuracy bar. Size buys latency cost, not correctness, on this task shape. |
| Long-context retrieval, deep multi-step tool chains | Determinism (still) | Completely flat across size for gemma4 — 30/30 vs 30/30 on both axes, tested to 128K tokens and 6-step dependent chains (§12.2§12.3) | Same as above: pick on latency/VRAM footprint alone. Do not pay for size expecting better long-context handling within this range. |
| Purpose-aware summarization, audience-specific analogy, anything requiring recognizing what to leave out | Family and size, jointly, inconsistently | gemma4 rewards scale on 4 of 6 analogy tiers; qwen3 only on 1 of 6, and non-monotonically §12.9's qwen3:14b does not sit between its own 8b/32b endpoints); mistral shows its own distinct pattern (§12.10) |
Do not assume size helps here — check the specific family's known pattern before committing. This is the one place "bigger" is even worth testing for your use case. |
| Anything where a single wrong output has real cost, not just "usually right" | Reliability | mistral:7b shows genuine bimodal instability on some tiers — right half the seeds, factually backwards the other half, at temperature 0. mistral-small3.2:24b eliminates this same instability, not just improves the average (§12.10) |
Before trusting any model in this category on a single run, check for known instability. If found, either move to the larger sibling or run N§5 and take a majority result. |
13.4 The team/router idea — a concrete, evidence-grounded starting point
The summarization category (§12's newest addition) already accidentally tests this: fact preservation (checkable, near the deterministic end of axis 1) and correct omission (a judgment call, near the ambiguous end) are graded as two separate skills within one task. That is a real, already-built pilot for "different sub-skills route to different models," not a hypothetical. A first, concrete router design:
- A fast, cheap classifier step (
gemma4:e2bis a strong, well-evidenced choice here, given its consistent top-tier accuracy on well-specified classification-shaped tasks) that looks at an incoming task and buckets it against §13.3's table. - Deterministic-leaning sub-tasks route to the smallest model that has cleared the bar for that shape.
- Judgment-leaning sub-tasks route to whichever family has evidence of strength on that specific shape — falling back to running the same sub-task on two or more models and reconciling if no family has been tested on it yet.
- A reliability check on any single-model output before it is trusted for anything consequential — informed directly by §13.3's fourth row.
This is a design sketch, not yet a built system. The natural next steps are: (a) formalize reliability as a first-class tested axis rather than an incidental observation (run N=10§20 on representative tiers and measure variance directly), and (b) build a minimal working classifier against §13.3's table as a genuine proof of concept, rather than a purely theoretical router.
14. The summarization category: three families, real purpose-aware differentiation
14.1 Design, grounded in real source documents
Motivated directly by summarization being one of local models' most common real-world uses, and explicitly deferred earlier in this project (§9.3) as needing a trustworthy LLM judge that did not yet exist. That infrastructure now exists, proven across the analogy category's three families (§12.5§12.10).
Three real source documents were used, not synthetic prompts: a real internal employee benefits guide, the Greek National Cybersecurity Authority's public handbook, and Who Moved My Cheese?. LMOS's knowledge_distiller and document_cognition_ingestor capabilities were reviewed directly before designing this category; the extraction approach was confirmed sound (no need to import it), and one genuinely reusable discipline was borrowed and folded into every tier's prompt: an explicit "do not speculate beyond the document excerpt" instruction, taken from document_cognition_ingestor's own distillation prompt. knowledge_distiller's five-purpose taxonomy (approval/alignment/training/proficiency/pitch) was reviewed and deliberately NOT reused — on inspection, those purposes are built for persuasive business narrative ("guide the audience to a conclusion they feel they chose"), not neutral informational fidelity, and importing them would have quietly shifted what this category measures.
Two tiers per document, each pair deliberately purpose-divergent — the two purposes require OPPOSITE inclusion/omission choices on the identical source text, so a model doing genuine purpose-aware summarization must track which facts matter for which reader, not just compress proportionally:
| Tier | Source | Purpose | Word limit |
|---|---|---|---|
| 1 | Benefits Guide | new hire choosing a medical plan | 60 |
| 2 | Benefits Guide | HR compliance note on eligibility/deadlines | 50 |
| 3 | Cybersecurity handbook | executive board approving budget | 50 |
| 4 | Cybersecurity handbook | security engineer building the process | 50 |
| 5 | Who Moved My Cheese? | book club plot summary | 50 |
| 6 | Who Moved My Cheese? | practical takeaway for a new job | 40 |
Three locked, separately-graded criteria, mirroring the analogy category's discipline exactly: fact preservation (a pre-authored checklist of load-bearing facts THIS purpose requires), correct omission (a pre-authored checklist of true, present facts a good summary for this purpose should leave out — the real test of purpose-awareness), and length compliance (deterministic word count).
14.2 Calibration caught a real design flaw before any model was tested
The GOOD / MEDIOCRE / GOOD_WRONG_OMISSION calibration set (mirroring §12.5's pattern) caught a genuine compound-checklist-item flaw on the very first run: a should_omit item describing a full six-step technical methodology, tested against a summary containing 5 of the 6 steps, unanimously graded as "correctly omitted" by the judge — done_reason: stop, not a budget cutoff, a genuine over-literal all-or-nothing reading. Fixed by rephrasing the checklist item to ask about the general category ("e.g. separately walking through...") rather than requiring an exhaustive match, the same lesson §12.6 already established about compound criteria creating loopholes. Re-verified clean after the fix.
14.3 Harness bugs: the same undersized-budget lesson, twice more
The recurring undersized-num_predict bug (§12.6, §12.10) hit this category too, on the actual summary-generation call this time, not just the judge. gemma4:26b hit done_reason: length empty responses on three of six tiers at the category's starting budget (2500); raised to 4096 resolved two of the three, but tier 2 needed a second escalation to 8192 before it cleared cleanly. Confirmed harness-stable afterward across two further full runs (qwen3:8b, mistral:7b, mistral-small3.2:24b) with zero empty responses at the 8192 setting.
14.4 Results: real, family- and size-differentiated content failures
gemma4 (e2b vs 26b): 26b resolves both of e2b's genuine gaps — tier 1's missing HSA-contribution fact, and tier 5's striking style-mismatch (never naming Sniff, Scurry, Hem, or Haw, defaulting to the abstracted lesson style even when a literal plot summary was explicitly requested). But 26b introduces its own narrow miss on tier 3: four of five seeds converge on near-identical phrasing that frames senior management as the direct implementer ("management must establish...") rather than the financial/organizational backer the source and checklist specifically emphasized — a genuine reproducible pattern, not noise, given the near-verbatim consistency across seeds.
qwen3:8b: clean on tiers 1, 2, 6; fails 3, 4, 5. A distinctive behavioral quirk found nowhere else: qwen3:8b appends a literal "(50 words)" self-annotation to its own responses, and on tier 3 this annotation text is itself what pushes the real word count over the stated limit — the model's own accounting habit undermines its own length compliance.
mistral:7b: severe, consistent length-limit violations on three of six tiers — tier 3 at 90 words against a 50-word limit (nearly double), tier 4 at 72, tier 5 at 66 — even though the underlying content is often strong (facts 3/3 on both tiers 3 and 5). Also reproduces the exact bimodal seed-instability signature already documented for this model in the analogy category (§12.10): tier 1 splits cleanly between a 65-word response (one seed) and a 53-word response (the other four).
mistral-small3.2:24b: the larger size dramatically improves length discipline — only two tiers run over their limit, and both by a handful of words (52 vs 50, 54 vs 50), nothing like 7b's 30§80% overages. But tier 6 shows a real, substantial content failure: 0 of 2 correct omissions — the summary includes both the literal character names and literal setting details the "practical takeaway for a new job" purpose specifically asked it to leave out. This is the precise mirror image of gemma4:e2b's original tier-5 failure: e2b over-abstracted when a literal plot was needed; mistral-small3.2:24b over-literalizes when the abstract lesson was needed. Two unrelated families, opposite failure directions, same underlying skill — correctly matching abstraction level to purpose — on the same source material.
14.5 What this adds to the decision framework
This category is the second, independent confirmation of §13.3's "purpose-aware judgment" row: every model tested here shows genuine, family-specific content failures that plain fact-checking would miss, and at least one case (mistral-small3.2:24b's tier 6) shows that even a larger model within a family that generally rewards scale is not immune to the underlying failure mode — it simply fails in the opposite direction. Reliability also shows up here independently of the analogy category: mistral:7b's bimodal instability reproduces on an entirely different task shape (summarization, not analogy generation), which is itself evidence this is a genuine, task-independent trait of this specific model rather than an artifact of one test design.