Stability-calibrated Sinkhorn attention
Source paper: Uniform Statistical Convergence of Empirical Sinkhorn Potentials with Exponential and Polynomial Dependence on the Regularization Parameter arXiv:2608.29152 ⓘ · analyzed Sep 1, 2026
AI-generated research hypothesis, automatically tested. Not peer-reviewed.
Idea description
Replace independently normalized attention or routing weights with an entropic doubly stochastic transport plan, while choosing its regularization ε using the paper's explicit statistical-stability bound. Increase ε when residual inversion or minibatch fluctuations are amplified, and decrease it only when the estimated bound permits sharper assignments.
Formulas
Mathematical statement
The paper's polynomial-stability principle states that if the summed entropy integral is bounded by A₀ ε^(-p) sqrt(Hε), the function-class envelope is bounded by E₀ ε^(-a), and the population Sinkhorn residual has inverse stability constant K₀ ε^(-r), then with probability at least 1−δ, the quotient sup-norm error is at most C K₀ / sqrt(n) times [A₀ ε^(1−r−p) sqrt(Hε) + E₀ ε^(1−r−a) sqrt(log(4/δ))]. Here n is the number of tokens or routing items, ε is the entropic temperature, Hε is the entropy-complexity term for normalized kernel sections, A₀, E₀, and C are scale constants, and a, p, and r describe envelope, complexity, and inverse-residual stability dependence on ε. The quotient metric d∞([u],[v]) = inf over b in R of ||u−v−b||∞ = one half of osc(u−v) removes the arbitrary additive gauge of Sinkhorn potentials. Use the bound as a temperature controller, with K₀ ε^(-r) estimated by the local inverse sensitivity of the Sinkhorn residual.
Implementation notes
Integrate this into the attention-logit or MoE-router path before the final weighted aggregation. For a minibatch of n tokens, compute scores Sij = qiᵀkj, define Kij = exp(Sij/ε), set row marginals ri = 1/n and column marginals cj = 1/m, and run 5–20 log-domain Sinkhorn iterations to obtain potentials α and β and transport weights P. Center every potential by subtracting its median to fix the quotient gauge. Maintain an exponential moving average of the observed Sinkhorn residual R = (P1−r, Pᵀ1−c). Estimate the local inverse-stability factor K-hat by finite differences: perturb the centered potentials or logits by small random vectors ξ, recompute the residual, and use K-hat = max ||ξ||∞ / ||ΔR||∞ over several probes, with damping and clipping. Estimate H-hatε from empirical variance or a simple covering proxy of normalized kernel rows; initially use a constant calibrated on a held-out batch. Fit exponents a, p, and r by regressing the envelope, complexity proxy, and K-hat against log ε over a temperature grid. Compute B-hat(ε) from the displayed bound and update ε toward a target noise budget τ. The first experiment should use CIFAR-10 or a small language model, comparing softmax attention, fixed-temperature Sinkhorn attention, and calibrated Sinkhorn attention at equal FLOPs. Measure validation loss, attention-plan variance, entropy, residual size, and loss spikes under small minibatches. Success means lower plan variance and equal-or-better validation loss without more Sinkhorn iterations.
Verification
This idea has not been verified yet.
Verification happens in two stages: Stage 1 — a mechanism check on a toy system confirms the claimed mathematical phenomenon reproduces; Stage 2 — a benchmark implements the idea on a real (small) neural network task and compares it against a tuned baseline over 8 paired seeds with a permutation test.
Artifacts
Artifacts unavailable.