A Structural Characterization of Entropy Functionals
arXiv:2608.13917
2026
Regularization
1 ideas extracted · analyzed Sep 1, 2026
What the math gives to ML
The paper supplies a structural rule for selecting entropy and divergence families rather than choosing Shannon, Rényi, or Tsallis by convention: admissibility is controlled by convexity of the transformed mean generator, specifically convexity of t \mapsto g(1/t) for increasing g and concavity for decreasing g. This criterion is equivalent to strict convexity of an associated Csiszár generator, which gives data processing under arbitrary Markov kernels and interpretable equality conditions. A practical transfer is a learnable family of divergence losses whose generator is constrained to remain convex, allowing task-adaptive regularization while retaining contraction under stochastic augmentations, token merging, pooling, or teacher-student channels.
Ideas from this paper
Unverified
2026
Replace a fixed KL or Jensen-Shannon penalty with a learnable Csiszár f-divergence whose generator is parameterized so that convexity is guaranteed. Apply it between teacher and student distributions, augmentation views, or intermediate representations; the loss cannot increase after a stochastic channel such as augmentation, pooling, token merging, or quantization, making the regularizer structurally compatible with information-discarding network operations.
Useful6/10
Difficulty4/10
Novelty5/10