A Gaussian-Remainder Hierarchy for Sums of Random Variables with Big-Jump Statistics
arXiv:2607.12357
2026
Dynamics
1 ideas extracted · analyzed Aug 30, 2026
What the math gives to ML
The paper provides a constructive Gaussian-remainder hierarchy for sums of independent finite-variance variables whose tails remain non-Gaussian at finite sample sizes. Its key transferable mechanism is the first-order convolution approximation, which preserves the Gaussian center while retaining a single-big-jump correction in the crossover and tail regions. A practical neural-network use is a tail-aware minibatch optimizer that estimates the distribution of per-example gradient projections, predicts aggregate-gradient tail risk, and adapts the learning rate or clipping threshold before rare large updates destabilize training.
Ideas from this paper
✗ Failed on benchmark
2026
Replace the assumption that a minibatch gradient is fully Gaussian by a Gaussian center plus an explicit single-example big-jump correction. At each update, estimate the distribution of per-example gradient projections along the proposed update direction and use the predicted aggregate tail probability to reduce the step size or increase clipping only when the minibatch is in its non-Gaussian crossover regime.
Useful7/10
Difficulty5/10
Novelty7/10