Scaling Almost Anything?

Scaling laws let us trade compute for capability, but parameter count is only a coarse measure of capacity. Through a least-squares toy problem and three LLM examples — LoRA rank, Transformer depth, and MoE experts — we show that larger can even hurt, and argue that what matters is how much capacity training actually exploits.

read_more

Neural Scaling Laws, Selfish Learning, and the Limits of Scale

Why do neural scaling laws work — and where might they break down? Drawing on biological and urban scaling, Braess's paradox, and recent LLM observations, we conjecture that selfish learning — neurons crowding into easy feature directions at the expense of global coordination — may become the next bottleneck for scale.