Family-Scaling Investigation
Overview Code Generation Tool Calling RAG / Q&A Reasoning / Math | About Design Doc Scaling Curves

Does a bigger model actually earn its size?

Analogy quality and self-critique accuracy, plotted against measured VRAM within two model families — gemma4 (2 sizes) and qwen3 (3 sizes). Six tiers, one per concept-audience pairing, escalating domain distance.

Each point is a composite score (0–3) averaged across 5 confirmed seeds: structural fidelity (how completely the analogy represents the concept), jargon discipline (fewer unexplained source-domain terms scores higher), and self-critique accuracy (does the stated limitation match a real, pre-authored gap). A rising line means size helps on that tier. A flat line means it doesn't. A line that dips in the middle means the scaling isn't even monotonic — confirmed real for qwen3, not sampling noise.

gemma4 (e2b → 26b) qwen3 (8b → 14b → 32b)
Tier 1: Car cooling to HVAC tech
0 1 2 3 e2b 26b 8b 14b 32b measured VRAM (log scale)
Tier 2: DB index to librarian
0 1 2 3 e2b 26b 8b 14b 32b measured VRAM (log scale)
Tier 3: DB index to warehouse manager
0 1 2 3 e2b 26b 8b 14b 32b measured VRAM (log scale)
Tier 4: TCP to kitchen manager
0 1 2 3 e2b 26b 8b 14b 32b measured VRAM (log scale)
Tier 5: Gradient descent to gardener
0 1 2 3 e2b 26b 8b 14b 32b measured VRAM (log scale)
Tier 6: Consensus to choreographer
0 1 2 3 e2b 26b 8b 14b 32b measured VRAM (log scale)

Worth reading tier by tier: gemma4 climbs on tiers 1, 2, 5, and 6 and barely moves on tier 3 -- a family that mostly rewards scale. qwen3 is the opposite story almost everywhere it differs from flat, and its middle size (14b) does not sit between its own endpoints on several tiers -- it tracks the larger model on some, the smaller one on others, and beats both of them on tier 6. That wobble is the real finding: a two-point comparison would have missed it entirely.