all_posts

2 posts · more coming soon.

filter:

Scaling Almost Anything?

Parameter count is only a coarse measure of capacity. Through a least-squares toy problem and three LLM examples — LoRA rank, Transformer depth, and MoE experts — we show that larger can even hurt, and argue that what matters is how much capacity training actually exploits.

read_more