Every idea extracted from recent arXiv mathematics papers — verified and unverified. Click an idea to open its full card; badges show the empirical verdict.
Replace scalar entropy penalties on attention maps with a matrix-valued heat-flow regularizer over a circular or periodic token coordinate. Each position stores a positive semidefinite matrix describing coupled heads, experts, or channels; heat smoothing is constrained by the sharp modified log-Sobolev and Bogoliubov–Kubo–Mori contraction rather than an arbitrary smoothing coefficient. This should suppress high-frequency routing noise while preserving positive matrix structure and reducing…
Replace ordinary Jacobian penalties in coordinate MLPs or deformation networks with a learned local rotation frame and a polyconvex energy of the relative stretch. Penalize \(U\), its cofactor, and its determinant through a convex function, while separately smoothing the rotation field through \(R^T\operatorname{Curl}R\). The intended benefit is resistance to fold formation and better conditioning than directly penalizing \(\|J-I\|^2\), especially for large deformations.
Add a multiscale texture regularizer to spatial feature maps by measuring Gaussian Difference-of-Gaussians responses at geometrically spaced scales. Weighting each scale according to a Besov smoothness exponent penalizes non-persistent high-frequency structure without forcing features to be globally smooth, so the network can retain edges and textures that survive across adjacent scales.
Replace disjoint-pair estimates of embedding covariance moments with a complete U-statistic over every distinct pair in a minibatch. For embeddings z, the degree-two kernel h(z_i,z_j)=(z_i^T z_j)^2 estimates the spectral moment tr(M^2), where M=E[zz^T]; complete symmetrization reduces the degenerate component of estimator variance from O(1/B) to O(1/B^2).
Replace unconstrained adversarial example generation with an invertible transport map that is the gradient of a convex potential. For each class, the map pushes a kernel-smoothed empirical distribution toward a least-favorable distribution inside a prescribed KL/Sinkhorn ambiguity radius, producing hard but globally coherent training examples rather than pointwise perturbations.
Represent an intermediate feature as a low-rank PSD matrix and compress it using nonnegative measurements \(\langle A_i,X\rangle\), while penalizing the empirical ratio between maximum and minimum measurement distortion over low-rank feature pairs. This directly discourages collapsed directions and excessively amplified directions in a covariance or Gram-feature bottleneck.
Add a loss term requiring a neural optimizer or recurrent module to decrease a nonnegative Lyapunov-like energy over M update steps, rather than forcing monotonic one-step decrease. The term includes an empirically estimated mismatch allowance, so stochastic or delayed updates are tolerated while persistent instability remains penalized.
Regularize a neural dynamical map so that its log-volume expansion is cohomologous to a constant rather than forcing the Jacobian determinant to be constant at every state. Learn a scalar potential that explains transient expansion and penalize only the non-telescoping component, which should reduce long-horizon gradient explosion or collapse while retaining useful average expansion.