Power-law and log-periodic degree tails for a family of probability generating function equations arising in evolving networks

arXiv:2607.12564 2026 Architecture 1 ideas extracted · analyzed Aug 30, 2026

What the math gives to ML

The paper gives an explicit, implementable branching-process construction whose stopped population size has a heavy tail with a non-constant multiplicatively periodic correction, rather than an ordinary pure power law. This provides a principled distribution for dynamically varying computational width or sparse expert fanout, with a tunable finite mean but occasional very large computations. The transferable asset is the exact geometric stopping rule and the resulting discrete-scale-invariant tail, which can replace ad hoc fixed-capacity or Gaussian width noise in sparse neural architectures. The safest first use is a capped stochastic MoE or tree-like MLP whose per-token fanout is sampled from this process and trained with an unbiased compute/load penalty.

Ideas from this paper

Unverified 2026

Branching-Fanout Sparse MoE

Replace a fixed top-k expert count with a stochastic, token-specific fanout generated by a supercritical binary branching process stopped at a geometric time. The resulting number of active experts has finite mean but a power-law tail with log-periodic modulation, allowing most tokens to use little compute while reserving larger computation for difficult or ambiguous tokens.

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Power-law and log-periodic degree tails for a family of probability generating function equations arising in evolving networks arXiv:2607.12564