A Unified Discrete and Continuous Theory of Core-Halo Complexity Maximizers
arXiv:2607.17907
2026
Regularization
1 ideas extracted · analyzed Aug 30, 2026
What the math gives to ML
The paper derives a variational mechanism in which generalized statistical-complexity maximizers have exactly two probability levels: a concentrated core and a diffuse halo. The optimization reduces to the core multiplicity, and the global maximum is predicted to use the smallest admissible core, namely one dominant state in the discrete case. This mechanism can be transferred to attention rows, mixture-of-experts routing distributions, and categorical policies as a structured alternative to ordinary entropy regularization. The main falsifiable signature is a two-level probability histogram, with the unconstrained regularizer favoring a one-element core.
Ideas from this paper
Unverified
2026
Add a statistical-complexity maximization term to attention rows or MoE routing distributions so that each probability vector is encouraged to contain a small dominant core and a nearly uniform low-probability halo. Unlike ordinary entropy regularization, this explicitly favors an intermediate concentration regime and predicts a two-level structure: one or a few large probabilities and all remaining probabilities close to one another. The regularizer should use a small coefficient because its…
Useful6/10
Difficulty4/10
Novelty7/10