On The Linear Convergence of Bregman Proximal Gradient Methods with Applications to Kullback--Leibler regression
arXiv:2607.05539
2026
Optimization
1 ideas extracted · analyzed Aug 30, 2026
What the math gives to ML
The paper’s transferable contribution is a convergence framework for Bregman proximal gradient methods based on Restricted Relative Strong Convexity (RRSC), rather than ordinary Euclidean strong convexity. This suggests replacing fixed Euclidean updates with a domain-aware divergence for positive parameters, probability vectors, router probabilities, and normalized attention weights. The most practical first test is a smoothed Burg mirror optimizer whose smoothing prevents singular behavior near zero while retaining multiplicative-update geometry for KL-like or simplex-constrained objectives.
Ideas from this paper
✗ Failed on benchmark
2026
Use a smoothed Burg entropy as the mirror map in a proximal-gradient optimizer for positive or simplex-valued neural parameters. The optimizer performs a Bregman-proximal step instead of an additive Euclidean update, while the smoothing parameter avoids the singularity of ordinary Burg entropy at zero.
Useful7/10
Difficulty5/10
Novelty5/10