Have you ever wondered why we place such heavy emphasis on the "large" in "Large Language Models (LLMs)"? The reason seems simple: larger models perform better. But where is the ceiling? Does a bigger architecture always lead to better results?

In this post, we revisit the idea of scaling from a broader perspective, drawing examples from biological and urban systems, where the benefits of scale rarely continue indefinitely. We argue that neural scaling laws for LLMs may also face a similar ceiling. In particular, we conjecture that one possible counter-force is selfish learning: neurons, layers, or sub-networks may fail to coordinate efficiently and instead crowd into greedily reducing their own local objectives. Several recent observations suggest that this is worth taking seriously.

1. The Neural Scaling Law

Figure 1: Test loss as a function of compute, dataset size, and parameters — Kaplan et al. scaling law
Figure 1. The Kaplan scaling law of LLMs [1].

Suppose you want to train a large model. Natural questions arise: how much data do you need, and how much compute should you secure? Much of our understanding comes from large experimental studies across models of different sizes, summarized in the so-called neural scaling laws. Although the field moves quickly enough that some details already feel dated, two representative ones are listed below.

  • Loss vs compute, number of tokens, and model sizes [1]. One influential observation is that test loss (cross-entropy) follows a predictable power law when any one of these three factors is scaled while the others are sufficiently large. As illustrated in Figure 1, the resulting trend is close to a straight line on a log-log plot. In other words, as we invest more resources into training, performance tends to improve in a predictable way.
  • Compute-optimal scaling laws — the Chinchilla scaling law [2]. Another influential result argues that compute-optimal training requires model and data size to scale evenly. Roughly speaking, about 20 training tokens are needed for each parameter. Following this principle, the authors trained Chinchilla, a 70B model. Although it is four times smaller than its predecessor Gopher (280B), it was trained on four times more data. At roughly the same total compute budget, Chinchilla outperformed Gopher, illustrating the computational optimality of this rule.

These empirical laws have provided invaluable guidance for building modern LLMs. But they also rely on an implicit conjecture: given enough data and compute, increasing the number of parameters will continue to improve performance. While this assumption is natural and often useful, the broader history of scaling laws suggests that the simple logic of "larger is better" is rarely universal.

2. Scaling Laws in a Broader Context

Figure 2: Kleiber's law and harbor seal logistic growth curve
Figure 2. Left: Kleiber's law; Right: A natural population of seals shows an S-shaped curve.

Scaling laws appear in various domains. Here are two examples inspired in part by the book Scale written by Geoffrey West.

  • Kleiber's law in biology. To build intuition, let us begin with a simplified version. If an animal's heat dissipation scales with surface area while its mass scales with volume, then its basal metabolic rate should scale roughly as $m^{2/3}$, where $m$ is body mass. The intuition is that maintaining body temperature requires dissipating heat, and surface area grows more slowly than volume as size increases. Kleiber's law refines this, suggesting an exponent closer to $3/4$ for biological reasons; see left of Figure 2. Either way, the broad message is similar: larger animals are often more energy-efficient in an aggregate sense.
  • Urban scaling laws. The intuition here is that increasing city size reduces the per-capita cost of infrastructure. A city twice the size may need only about 85% more infrastructure, while GDP and patent production can rise to roughly 2.15 times their original levels.

This raises a natural question. If larger animals are more energy-efficient, why do we not see animals at the size of Godzilla? If larger cities benefit from economies of scale, why do they not become frictionless or effectively free as they expand? The answer is that positive scaling laws are eventually met by opposing constraints.

  • Negative scaling laws for Godzilla. The weight a bone can support is proportional to its cross-sectional area. While an animal's body weight increases cubically with its size, the cross-sectional area only increases quadratically. Consequently, the legs of such a massive creature would eventually be unable to support its weight as it scales up.
  • Negative scaling laws for cities. As density grows, so do congestion, disease transmission, housing pressure, and competition for limited resources. Growth creates its own taxes.

In other words, scaling laws are rarely universal across all scales. A broader perspective comes from ecology — for example, the population of harbor seals (see right of Figure 2) or the growth of yeast populations in a dough. Initially, the population explodes exponentially due to abundant resources and a lack of natural predators. However, growth eventually hits a carrying capacity due to food scarcity, competition, and diseases. This transition transforms the growth plot into a characteristic "S" shape, formally known in ecology as the logistic curve [3].

3. What Could Be the Negative Factor for Neural Scaling Laws?

Figure 3: The traffic network in Braess's paradox
Figure 3. The traffic network in Braess's paradox.

Does "negative" scaling also happen for LLMs, just as in other domains? If we had an "ideal GPU" with infinite memory and zero energy consumption, does that imply we could scale a neural network infinitely? While no one knows the definitive answer, we can approach the question from a different angle: more is not necessarily better and in some cases, it can be provably worse.

  • Braess's paradox, named after the German mathematician Dietrich Braess, serves as an example in urban transportation studies. It states that adding a road to a specific transportation network can increase overall congestion. This has been empirically identified in several major metropolises including Seoul and New York City [4].

    Consider a network shown in Figure 3, where 4000 drivers travel from a start point (S) to an end point (E) via one of two midpoints, A or B. The travel time from S to B and from A to E is 45 minutes. The travel time from S to A and from B to E is proportional to the volume of traffic, defined as $t = T/100$, where $T$ is the number of drivers on the road.

    Given the symmetry of this network, the 4000 drivers split equally: 2000 choose route S–A–E and 2000 choose S–B–E. This results in a travel time of 65 minutes ($2000/100 + 45$) per driver — the socially optimal outcome.

    Now, suppose a new, near-instant (0-minute) shortcut is added between A and B. A single driver switching to S–A–B–E initially sees travel time drop to 40 minutes. However, as more drivers flock to this shortcut, the network shifts toward a new equilibrium. Eventually, every driver is forced onto S–A–B–E, and travel time climbs to 80 minutes ($4000/100 + 4000/100$). Paradoxically, by adding a road, the system becomes less efficient.

  • In game theory, the Price of Anarchy (PoA) [5] measures the gap between decentralized equilibrium and global optimality. If an additional option incentivizes more "selfish" behavior among individuals, the system suffers a performance penalty — precisely the logic behind Braess's Paradox.

This perspective suggests an intriguing analogy for learning systems. Perhaps adding more parameters, neurons, or layers can also make learning more selfish — with different parts of the model competing for the same easy gains rather than coordinating globally.

3.1 Known Selfishness in Learning Dynamics

Returning to machine learning, selfishness is not just a speculative idea: related phenomena have already been supported by rigorous theoretical analysis. Here, we use selfish learning to describe the tendency of neurons, layers, or sub-networks to crowd into simple and immediately rewarding tasks — namely, those that reduce loss quickly — instead of tackling harder but globally more important ones.

  • Least squares. In the classical data-fitting problem $\min_x \|Ax-b\|^2$, over-parameterization creates a striking hierarchy: directions in the row space of $A$ are learned quickly, while directions in the null space are learned very slowly under gradient descent. The gradient lives entirely in the row space of $A$, so optimization behaves selfishly — eagerly following directions that offer immediate progress while ignoring null-space directions.
  • Two-layer neural networks [6] and index models [7]. For these tasks, adding more neurons can actually slow convergence exponentially. The learning problem splits naturally into easy and hard features, characterized by different signal-to-noise ratios. In an over-parameterized regime, neurons fail to diversify: many pile into the easy directions — the low-hanging fruit — while harder directions are largely left unexplored.

This should perhaps not be too surprising. Optimization in modern machine learning is largely driven by first-order methods, which are inherently local and greedy: they follow the gradient toward the most immediate decrease in loss. From that perspective, selfishness may be less an exception than a natural outcome of gradient-based learning.

3.2 Do Large Models Also Exhibit Selfishness?

This is the central argument of this post.

We conjecture that this selfishness could eventually become the negative scaling factor for neural scaling laws. Suppose a network distinguishes data into "easy-to-learn" and "hard-to-learn" components. Because neurons of the same layer are symmetric in their roles and initialization, the vast majority will lean toward learning the high-signal, easy-to-learn factors. Much like the drivers in Braess's Paradox crowding a single shortcut, these neurons congest the simple features while leaving the challenging, nuanced dimensions of the data unmapped.

The next question is whether this phenomenon, though provable in simpler settings, also appears at scale. The answer is still far from settled. Yet there are already suggestive signs that modern large models may fail to use added capacity effectively.

  • Rank collapse in LoRA fine-tuning [8]. In LoRA-based fine-tuning, increasing the adapter rank does not always translate into better downstream performance. As observed when fine-tuning Llama 2 in [8], even with a rank of $r=32$, the learned update can still collapse into a single dominant direction, effectively wasting most of the available spectrum and driving the stable rank close to 1. Notably, this is not because a rank-1 solution is already sufficient — the authors show that improving spectral usage improves performance. A plausible interpretation: the dominant subspace already captures the easiest semantic structure, leaving remaining rank directions with little incentive to do meaningful additional work.
  • Under-utilized depth in large pretrained models [9]. Recent work shows that skipping an early layer in Llama 3.1 70B can affect outputs much more strongly than skipping certain deeper layers. Many deeper layers behave approximately like identity mappings, suggesting that part of the model's nominal depth contributes far less transformation than parameter count alone would imply. One possible explanation: early layers absorb much of the readily available structure, leaving later layers with less meaningful work to do.

In both cases, what is captured early — the rank-1 subspace in LoRA and the early layers in Llama pretraining — appears to crowd out the emergence of additional useful structure later in training. This is precisely the kind of selfishness we have in mind. We therefore conjecture that as models continue to scale, selfish learning may become a bottleneck for training.

4. Implications

If this conjecture is right, then the next frontier of scaling may not be simply adding more raw capacity, but learning how to allocate that capacity non-selfishly. There are already suggestive examples in this direction. For instance, the selfishness observed in two-layer neural networks can be provably reduced by imposing suitable structures [10], and recent work has also explored how to make better use of depth in large models [11, 12]. In a broader sense, many of these methods can be interpreted as introducing structure that helps neurons, layers, and sub-networks collaborate more effectively. We thus conjecture the next gains in scaling may come not only from adding capacity, but from teaching models to use capacity more cooperatively.

References

  1. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. "Scaling laws for neural language models." arXiv preprint arXiv:2001.08361 (2020).
  2. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., et al. "Training compute-optimal large language models." arXiv preprint arXiv:2203.15556 (2022).
  3. https://bioprinciples.biosci.gatech.edu/population-ecology-1/
  4. https://en.wikipedia.org/wiki/Braess's_paradox
  5. Koutsoupias, E., and Papadimitriou, C. "Worst-case equilibria." Computer Science Review 3, no. 2 (2009): 65–69.
  6. Xiong, N., Ding, L., and Du, S. S. "How over-parameterization slows down gradient descent in matrix sensing: The curses of symmetry and initialization." arXiv preprint arXiv:2310.01769 (2023).
  7. Xu, W., and Du, S. "Over-parameterization exponentially slows down gradient descent for learning a single neuron." In COLT 2023, pp. 1155–1198. PMLR, 2023.
  8. Lion, K., Zhang, L., Li, B., and He, N. "PoLAR: Polar-decomposed low-rank adapter representation." arXiv preprint arXiv:2506.03133 (2025).
  9. Csordás, R., Manning, C. D., and Potts, C. "Do language models use their depth efficiently?" arXiv preprint arXiv:2505.13898 (2025).
  10. Wei, Y., Zhang, L., Li, B., and He, N. "On the Benefits of Weight Normalization for Overparameterized Matrix Sensing." arXiv preprint arXiv:2510.01175 (2025).
  11. Zhang, Y., Li, B., He, N., and Giannakis, G. B. "ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling." arXiv preprint arXiv:2602.09009 (2026).
  12. Xie, Z., Wei, Y., Cao, H., Zhao, C., Deng, C., Li, J., Dai, D., et al. "mHC: Manifold-constrained hyper-connections." arXiv preprint arXiv:2512.24880 (2025).