Loss-Parameterized Fisher Width Along Learning Trajectories
arXiv:2608.21561
2026
Dynamics
1 ideas extracted · analyzed Sep 1, 2026
What the math gives to ML
The paper supplies a dynamical use of Fisher geometry rather than merely treating the Fisher matrix as a curvature preconditioner: the Gaussian width of a probe set can be tracked along optimization trajectories and compared at matched training loss. Its key transferable asset is the separation between loss level and branch-dependent Fisher geometry, illustrated by the substantially displaced Adam trajectory versus GD/SGD. A practical adaptation is to use matched-loss Fisher width as a trajectory-control signal, regularizing or selecting optimizer updates that remain near a reference branch while preserving the usual loss decrease.
Ideas from this paper
Unverified
2026
Use Fisher width as a branch coordinate in addition to training loss. During a short reference run with SGD, fit the expected Fisher-width curve as a function of loss, then add a soft penalty to Adam or another optimizer when its width at the same loss deviates from that reference branch. This directly tests whether optimizer-induced geometric displacement is responsible for differences in training dynamics or generalization.
Useful6/10
Difficulty5/10
Novelty6/10