Each point is a composite score (0–3) averaged across 5 confirmed seeds: structural fidelity (how completely the analogy represents the concept), jargon discipline (fewer unexplained source-domain terms scores higher), and self-critique accuracy (does the stated limitation match a real, pre-authored gap). A rising line means size helps on that tier. A flat line means it doesn't. A line that dips in the middle means the scaling isn't even monotonic — confirmed real for qwen3, not sampling noise.
Worth reading tier by tier: gemma4 climbs on tiers 1, 2, 5, and 6 and barely moves on tier 3 -- a family that mostly rewards scale. qwen3 is the opposite story almost everywhere it differs from flat, and its middle size (14b) does not sit between its own endpoints on several tiers -- it tracks the larger model on some, the smaller one on others, and beats both of them on tier 6. That wobble is the real finding: a two-point comparison would have missed it entirely.