On The Linear Convergence of Bregman Proximal Gradient Methods with Applications to Kullback--Leibler regression

arXiv:2607.05539 2026 Optimization 1 ideas extracted · analyzed Aug 30, 2026

What the math gives to ML

The paper’s transferable contribution is a convergence framework for Bregman proximal gradient methods based on Restricted Relative Strong Convexity (RRSC), rather than ordinary Euclidean strong convexity. This suggests replacing fixed Euclidean updates with a domain-aware divergence for positive parameters, probability vectors, router probabilities, and normalized attention weights. The most practical first test is a smoothed Burg mirror optimizer whose smoothing prevents singular behavior near zero while retaining multiplicative-update geometry for KL-like or simplex-constrained objectives.

Ideas from this paper

Failed on benchmark 2026

Smoothed Burg Proximal Optimizer

Use a smoothed Burg entropy as the mirror map in a proximal-gradient optimizer for positive or simplex-valued neural parameters. The optimizer performs a Bregman-proximal step instead of an additive Euclidean update, while the smoothing parameter avoids the singularity of ordinary Burg entropy at zero.

Useful7/10
Difficulty5/10
Novelty5/10
Paper: On The Linear Convergence of Bregman Proximal Gradient Methods with Applications to Kullback--Leibler regression arXiv:2607.05539