ML: Regularization

Machine-learning ideas tagged Regularization in the ML taxonomy of the Math2NN corpus.

Unverified 2026

Nonlinear Fourier amplitude budget regularizer

Use the logarithmic transmission amplitude of an SU(1,1) scan as a differentiable spectral penalty. The paper's constant-one nonlinear Hausdorff–Young inequality provides a principled upper budget for this amplitude in terms of the input L^p norm, replacing an arbitrary spectral-weight penalty with a scale-aware constraint.

Useful6/10
Difficulty4/10
Novelty8/10
Paper: The nonlinear Hausdorff-Young inequality arXiv:2608.15895
Unverified 2026

Geodesic Matrix Divergence Loss

Replace a conventional covariance or density-matrix discrepancy with the geodesic quantum f-divergence between an example's predicted positive-definite matrix and its target matrix. Use t as a controllable interpolation between the standard Petz divergence at t=0 and the maximal divergence at t=1, with f(x)=x log x or another operator-convex power generator. The loss is suited to covariance-predicting networks, SPD-valued embeddings, and matrix-valued classifiers where eigenvector alignment…

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Geodesic Quantum $f$-Divergences arXiv:2608.15833
Unverified 2026

Actuator-Aware Envelope Scheduler

Use adaptive performance specifications to prevent a neural controller or policy from demanding output changes that exceed bounded actuator amplitude or action-rate limits. The target error envelope tightens when the policy has control authority and relaxes when saturation or rate clipping persists, instead of allowing the controller to destabilize while chasing an infeasible target. This converts actuator clipping into an explicit slow state that can be used by reinforcement-learning policies…

Useful6/10
Difficulty4/10
Novelty7/10
Paper: Output Feedback Adaptive Performance Control arXiv:2608.15758
Unverified 2026

Schatten-budgeted low-rank update aggregation

Replace ordinary Frobenius-norm clipping when merging rank-one LoRA or adapter updates with a Schatten-budget computed from the positive operators |A_k|. For p>=2, the paper's sharp rank-one inequality bounds the norm of the merged update, including interactions between updates that are missed by independent per-update clipping.

Useful6/10
Difficulty4/10
Novelty8/10
Paper: A Counterexample to the Tang Zhang Schatten Norm Conjecture and Sharp Positive Results arXiv:2608.15558
Unverified 2026

Polarized Generalized-Dual Curvature Regularization

Use generalized dual numbers to compute second- or third-order derivatives of the training loss along several parameter-space directions, then use polarization to recover mixed directional derivatives without forming a Hessian or third-order tensor. Add a bounded mixed-curvature penalty or use the resulting directional curvature to rescale updates in directions that are simultaneously sharp.

Useful6/10
Difficulty6/10
Novelty6/10
Paper: Efficient Computation of Arbitrary-Order Directional Derivatives in Multiple Directions via Generalized Dual Numbers arXiv:2608.15345
Unverified 2026

Fractional Hilbert-Scale Neural Regularizer

Add a fractional Sobolev penalty to the spatial output of a neural field or reconstruction CNN, rather than relying only on pixelwise weight decay or total variation. The fractional order s continuously controls high-frequency suppression, allowing an experiment to test whether s less than 1 preserves edges better than the classical integer-order penalty while still reducing noise and unstable oscillations.

Useful6/10
Difficulty4/10
Novelty6/10
Paper: Nonlocal Tikhonov Regularization: Hilbert Scales, Explicit Rates, and the Classical Limit arXiv:2608.15315
Unverified 2026

Jensen-Gap Regularization for Temporal Cascades

Add a mean-preserving periodic-input consistency penalty to a stacked leaky recurrent or state-space network. The penalty suppresses output shifts caused purely by hidden-state fluctuations and nonlinear curvature, improving invariance to temporal modulation while preserving the average input signal.

Useful6/10
Difficulty4/10
Novelty7/10
Paper: Gain of Entrainment in Nonlinear Cascades arXiv:2608.15214
Unverified 2026

Leibenson Nonlinear Diffusion Layer

Replace a standard graph-convolution propagation step with a short time integration of the nonlinear graph flow \(\partial_t u=\Delta_p(u^q)\). The pointwise power \(q\) and gradient exponent \(p\) create state- and edge-gradient-dependent propagation: small signals can be suppressed or amplified by \(q\), while large graph discrepancies receive nonlinear diffusion controlled by \(p\). Use nonnegative feature states and conservative edge fluxes so the layer inherits positivity and total-mass…

Useful6/10
Difficulty4/10
Novelty6/10
Paper: Leibenson's equation on graphs arXiv:2608.15168
Unverified 2026

Fractal Spectral Anti-Collapse Regularizer

Add a multiscale penalty to transformer token-mixing activations when they are simultaneously concentrated on a spatial or token subset and on a separated, irregular frequency subset. The penalty uses the fractal uncertainty scaling law to discourage hidden states from collapsing onto narrow token patterns and narrow spectral bands, potentially improving robustness to token masking and frequency-corrupted inputs.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Fractal Uncertainty and Quantitative Uniqueness for the Fourier Bessel Transform arXiv:2608.15126
Unverified 2026

Path-Coupled Neural Stability Margin

Replace fixed-path robustness testing with a coupled continuation procedure that increases an adverse perturbation while simultaneously optimizing a bounded corrective response, such as feature-gating, normalization, or a small adapter. Define the model's margin as the cumulative perturbation at which its equilibrium, prediction, or input-output Jacobian becomes singular or exceeds a prescribed gain threshold; train the corrective response to enlarge this margin subject to an explicit cost.

Useful6/10
Difficulty7/10
Novelty7/10
Paper: Voltage Stability Assessment with Path-Coupled Load Growth and Corrective Generator Response arXiv:2608.15122
Unverified 2026

Differentially Positive Recurrent Core

Constrain the Jacobian of a recurrent or state-space transition to preserve a prescribed cone of admissible hidden-state perturbations. This imports differential positivity into neural dynamics and makes long-run hidden trajectories order-preserving rather than allowing arbitrary sign-changing perturbation growth.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Birkhoff center and recurrent behavior of differentially positive systems on a homogeneous space arXiv:2608.14980
Unverified 2026

Local tent-space wavelet sparsifier

Insert a fixed or learnable multiresolution transform before a CNN or vision-transformer block and penalize its coefficients with the paper's local tent-space square function. The penalty couples coefficients belonging to the same spatial dyadic region and can remove localized multiscale feature packets, potentially producing structured sparsity and better denoising than independent l_1 shrinkage.

Useful6/10
Difficulty4/10
Novelty6/10
Paper: Wavelet and tent space characterizations of $h^p(\mathbb{R}^n)$, $0 < p \le 1$, with exact $L^2$ convergence, and the maximal class of admissible functions for Goldberg-type splittings arXiv:2608.14960
Unverified 2026

Information-Projection Constraint Loss

Replace raw squared penalties on generated feature means with the paper's nested information-projection statistic. A model output distribution is projected once onto structural constraints and once onto structural-plus-test constraints; their KL divergence produces a sample-size-scaled loss and an approximate chi-square p-value. This should help when constraints have different variances or are strongly correlated, because the KL geometry automatically adapts to their covariance instead of…

Useful6/10
Difficulty6/10
Novelty6/10
Paper: The Boltzmann structure of sampling: Intrinsic $p$-value and its emergent closed-form expression arXiv:2608.14608
Unverified 2026

Full-Field Phase-Boundary PINN

For neural eigenmode solvers on periodic domains, train the full field directly and impose Bloch phase coupling only at opposite cell boundaries, rather than differentiating a periodic factor with respect to q through a quadratic volume operator. The boundary formulation preserves reciprocal-lattice equivalence exactly through z=exp(iqa), reducing spurious eigenmodes caused by inconsistent q-dependent discretization.

Useful6/10
Difficulty5/10
Novelty9/10
Paper: Full-field and Bloch-periodic-factor discretizations: Accuracy and phantom modes arXiv:2608.14348
Unverified 2026

Continuous Ergodic Projection for Recurrent States

Given a learned recurrent dynamics map, estimate a state-dependent invariant measure from each trajectory and use integration against that measure as a projection onto long-term invariant features. Penalize discontinuities of this projection between nearby states and assign zero mass to trajectories whose feature norms escape, producing a principled distinction between convergent attractors and divergent rollouts.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Continuous pointwise ergodicity for semigroup actions on locally compact spaces arXiv:2608.14175
Unverified 2026

Data-Processing-Safe Learnable Divergence

Replace a fixed KL or Jensen-Shannon penalty with a learnable Csiszár f-divergence whose generator is parameterized so that convexity is guaranteed. Apply it between teacher and student distributions, augmentation views, or intermediate representations; the loss cannot increase after a stochastic channel such as augmentation, pooling, token merging, or quantization, making the regularizer structurally compatible with information-discarding network operations.

Useful6/10
Difficulty4/10
Novelty5/10
Paper: A Structural Characterization of Entropy Functionals arXiv:2608.13917
Unverified 2026

Support-cutoff sparse polynomial block

Separate low-support monomials, which involve only a few distinct input coordinates, from high-support monomials in a high-degree symmetric interaction layer. Compute the low-support orbit features exactly and prune, sample, or factorize the high-support tail, using the paper's cutoff scale as the initial sparsity rule.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Orbit compression and asymptotic contractivity for symmetric Bohnenblust--Hille inequalities arXiv:2608.13753
Unverified 2026

Causal DAG Wasserstein Alignment

Replace an unconstrained Wasserstein representation-matching loss with a graph-causal transport loss whose coupling at node k is conditioned only on the representations of its parents. This forces domain alignment, distillation, or augmentation consistency to respect the information flow of the model's DAG, reducing spurious matches that exploit descendants or globally visible features.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Graph Causal Optimal Transport and Wasserstein Distances arXiv:2608.13716
Unverified 2026

Refined pullback-metric Jacobian cap

Regularize a neural encoder so its local pullback metric is bounded by the refined Schwarz-lemma constant instead of using a generic Frobenius Jacobian penalty. For an encoder into a negatively curved latent space, penalize only singular directions whose squared expansion exceeds the curvature- and dilatation-dependent threshold.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: A refined Schwarz lemma for $V$-harmonic maps arXiv:2608.13682
Unverified 2026

Vortex-Criticality Controller for Phase RNNs

Represent recurrent hidden states as compact phases and monitor spacetime vortices, defined by wrapped phase differences around elementary space-time plaquettes. Add a feedback controller that increases relaxation toward the homogeneous phase when vortex activity becomes supercritical, while allowing larger recurrent gain when the system is excessively quiescent. This creates a falsifiable operating regime: useful computation should occur near, but below, the defect-proliferation transition…

Useful6/10
Difficulty6/10
Novelty8/10
Paper: Far-from-equilibrium topological phase transition in one dimension arXiv:2608.13658
Unverified 2026

Learned Busemann Stable Leaves

Learn an endpoint-conditioned scalar potential whose level sets represent states with the same asymptotic behavior, analogous to the paper's stable magnetic orthospheres. Train the dynamics to contract differences within a level set while preserving differences between distinct endpoint classes, producing a latent representation organized by stable manifolds rather than Euclidean proximity.

Useful6/10
Difficulty6/10
Novelty8/10
Paper: Surfaces with nonpositive magnetic curvature arXiv:2608.13534
Unverified 2026

Incompressible transport noise for feature maps

Replace independent additive noise on spatial feature maps with stochastic advection by divergence-free vector fields. The perturbation preserves spatial volume and feature mass, while the associated Stratonovich-to-Itô correction provides a tunable diffusion that preferentially damps high-frequency spatial fluctuations.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Global classical solutions by transport noise for reaction-diffusion systems with entropy dissipation arXiv:2608.13332
Unverified 2026

Positive Boundary-Diffusion Neural Layer

Replace an unconstrained deep residual recurrence by a discretized diffusion system over feature or token positions, with trainable source terms and analytically constrained boundary feedback. The state remains nonnegative under nonnegative inputs, while negative boundary gains enforce exponential decay of perturbations and prevent exploding activations in very deep stacks.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Positive stabilization of a pure diffusion system arXiv:2608.13203
Unverified 2026

Quadratic-budget Toeplitz long-range layer

Replace a dense translation-invariant interaction matrix with a positive-definite Toeplitz kernel K_n(e^f) whose log-spectrum is parameterized by a small number of Fourier coefficients with 1/|k| decay. Use the paper's explicit quadratic term as a spectral-volume budget, allowing long-range structure while discouraging uncontrolled determinant growth and ill-conditioning. Subtracting this term from a log-determinant regularizer leaves a residual intended to capture higher-order deviations from…

Useful6/10
Difficulty6/10
Novelty7/10
Paper: On Toeplitz determinants with slow Fourier decay arXiv:2608.13182