Loss-Parameterized Fisher Width Along Learning Trajectories

arXiv:2608.21561 2026 Dynamics 1 ideas extracted · analyzed Sep 1, 2026

What the math gives to ML

The paper supplies a dynamical use of Fisher geometry rather than merely treating the Fisher matrix as a curvature preconditioner: the Gaussian width of a probe set can be tracked along optimization trajectories and compared at matched training loss. Its key transferable asset is the separation between loss level and branch-dependent Fisher geometry, illustrated by the substantially displaced Adam trajectory versus GD/SGD. A practical adaptation is to use matched-loss Fisher width as a trajectory-control signal, regularizing or selecting optimizer updates that remain near a reference branch while preserving the usual loss decrease.

Ideas from this paper

Unverified 2026

Matched-Loss Fisher Branch Control

Use Fisher width as a branch coordinate in addition to training loss. During a short reference run with SGD, fit the expected Fisher-width curve as a function of loss, then add a soft penalty to Adam or another optimizer when its width at the same loss deviates from that reference branch. This directly tests whether optimizer-induced geometric displacement is responsible for differences in training dynamics or generalization.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Loss-Parameterized Fisher Width Along Learning Trajectories arXiv:2608.21561