ML: Mlp

Machine-learning ideas tagged Mlp in the ML taxonomy of the Math2NN corpus.

408 ideas found

Unverified 2026

Lower-Order-Invariant High-Order Representation Loss

Add an auxiliary loss that makes selected representation coordinates insensitive to all subsets of fewer than d variables while retaining a d-way parity statistic. The objective discourages the network from solving a task through pairwise shortcuts and explicitly rewards a controlled high-order interaction.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: How far are $d$-dimensional copulas with uniform $(d-1)$-marginals from (total) independence? arXiv:2608.18286
Unverified 2026

Veronese Projective Feature Layer

Replace or augment the first embedding layer for antipodally identified inputs with the normalized traceless quadratic map from the Veronese construction. Because q and -q produce exactly the same feature, the layer enforces projective invariance by construction rather than learning it from augmented examples. The resulting matrix-valued features can be flattened, projected, or processed by an equivariant linear layer.

Useful5/10
Difficulty2/10
Novelty6/10
Paper: Normal Curvature and the Projective Systole arXiv:2608.18002
Unverified 2026

Mixed-norm Brascamp–Lieb product block

Construct a four-branch neural interaction whose inputs are affine projections of a two-dimensional latent coordinate and whose output is the weighted product prescribed by the theorem. Normalize this product by the corresponding branch L1 masses, yielding a feature whose mixed norm is theoretically bounded up to the inequality constant. Use the normalized interaction as an architecture component or as a replacement for an unconstrained multiplicative fusion layer.

Useful5/10
Difficulty5/10
Novelty9/10
Paper: Mixed-norm Brascamp-Lieb inequalities arXiv:2608.17952
Unverified 2026

Accessibility regularizer for power-law experts

When a neural field learns power-law exponents, penalize exponent configurations whose Newton support violates the paper's finite-distance accessibility condition. This discourages combinations of exponents that create excessively strong joint singularities while preserving anisotropic scaling when it is supported by the data.

Useful5/10
Difficulty3/10
Novelty8/10
Paper: Newton Support Functions and Metric Completion of Singular Conformal Metrics at Corners arXiv:2608.17714
Unverified 2026

Sharp Schatten Certificate for Adapter Fusion

Replace the ordinary triangle-inequality budget for merging m linear residual branches or LoRA updates by the sharp quasi-reverse Minkowski certificate. During training, penalize or constrain the Schatten norm of the aggregate absolute update, which certifies the norm of the actually merged update with factor C_{p,m} rather than the loose factor m. This is especially attractive for p=2, where the certificate controls Frobenius energy and can be implemented with standard matrix operations.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Sharp Quasi-Reverse Minkowski Inequality for Schatten Norms arXiv:2608.17565
Unverified 2026

Data-consistent contractive residual adapters

Replace an unconstrained residual adapter around a neural linear layer by a contractive operator whose action interpolates observed feature perturbations and remains bounded in operator norm. The adapter is trained adversarially over this structured uncertainty set, producing perturbations tied to empirical feature data rather than arbitrary isotropic noise.

Useful5/10
Difficulty5/10
Novelty4/10
Paper: Operator-based data embedding for data-driven control of continuous-time systems from noisy data arXiv:2608.17518
Unverified 2026

Transversal encoded gates

Construct multiplicative neural gates directly on encoded tensors so that operands are multiplied coordinatewise without decoding between every operation. Polynomial evaluation makes this operation algebraically consistent with multiplication, allowing redundant gated MLPs or bilinear layers to retain fault tolerance while reducing the frequency of expensive correction steps.

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Fault-Tolerant Quantum Computation with Adversarial Errors arXiv:2608.16857
Unverified 2026

Moment-Cone Interaction Regularizer

Use the lifted convex hull as a training-time regularizer for pairs of nonnegative neural features, encouraging their empirical second- and third-order interaction statistics to lie in the paper's moment cone. This constrains correlations, squares, and cubic cross-moments jointly through PSD inequalities instead of merely penalizing large activations.

Useful5/10
Difficulty4/10
Novelty8/10
Paper: Nonnegative Quadratics over a Quadrant with a Bilinear Constraint arXiv:2608.16836
Unverified 2026

Weighted CKN feature-stability regularizer

Apply a weighted Caffarelli–Kohn–Nirenberg deficit to selected intermediate feature channels. The penalty discourages features that obtain large weighted responses only by becoming sharply localized or highly sensitive to small input perturbations. It can be evaluated with input-Jacobian estimates and added to the ordinary task loss.

Useful5/10
Difficulty5/10
Novelty5/10
Paper: Sharp $L^2$-Caffarelli--Kohn--Nirenberg and weighted Poincaré inequalities on half-spaces and orthants and their stability arXiv:2608.16803
Unverified 2026

Schatten-Stable Noncommutative Matrix Layer

Replace an ordinary elementwise interaction between two feature matrices by a noncommutative functional-calculus layer \(\varphi(A,B)\), where \(A\) and \(B\) are Hermitian channel operators that need not commute. Add a soft penalty on \([A,B]=AB-BA\), and use a Besov-smooth parameterization of \(\varphi\) so that perturbations are controlled in Schatten \(p\)-norm for \(p\leq2\). This creates a principled matrix interaction module that can remain stable when feature operators or graph…

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Commutator estimates for functions of noncommuting self-adjoint operators arXiv:2608.16731
Unverified 2026

Hyperuniform Collocation and Minibatch Sampling

Use spatially correlated training points whose low-frequency structure factor vanishes instead of iid points. For neural fields, PINNs, image-coordinate MLPs, or spatially indexed minibatches, this should suppress long-wavelength quadrature and gradient-estimation noise while preserving the represented target dynamics. The finite-order prediction is that a design with structure factor S(k)=O(|k|^{2q}) produces lower variance for smooth losses than iid sampling, especially as the domain or batch…

Useful5/10
Difficulty5/10
Novelty6/10
Paper: Hyperuniform Delone Realizations and Rigidity arXiv:2608.16547
Unverified 2026

Hermite compound-Poisson feature noise

Replace ordinary additive or multiplicative activation noise with a nonnegative count-valued perturbation generated by the Hermite operator kernel. For a nonnegative feature x, sample an integer N whose distribution is exactly the operator's weight sequence and feed N/n to the next layer; the parameter alpha controls an additional even-jump component and therefore changes the noise geometry independently of the ordinary Poisson component.

Useful5/10
Difficulty5/10
Novelty4/10
Paper: Complete asymptotic expansion for a Durrmeyer variant of operators based on Hermite polynomials arXiv:2608.16272
Unverified 2026

Quasi-analytic response head

For a neural model predicting a scalar response as a function of a continuous dynamical parameter, replace an unconstrained MLP output head by an analyticity-constrained spectral head. Train it on observations covering a positive-measure subset of the parameter interval and regularize the remaining coefficients so that the learned response satisfies a quasi-analytic derivative-growth bound; the intended benefit is reliable continuation from sparse parameter coverage rather than ordinary…

Useful5/10
Difficulty4/10
Novelty8/10
Paper: Rigidity of Mather's $β$-function on a KAM set for analytic billiards-like maps and unique quasi-analytic continuation arXiv:2608.15401
Unverified 2026

Correlated-Gaussian Orbit Fingerprint

Replace a polynomial layer's single-replica output statistics with a finite fingerprint computed from several correlated Gaussian replicas. Train the fingerprint to be invariant under orthogonal reparameterizations while remaining discriminative between genuinely different polynomial maps, preventing models from collapsing distinct tensor functions that have identical marginal output laws. This is a practical symmetry-aware regularizer or auxiliary embedding for tensorized MLPs and polynomial…

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Finite Gaussian Reconstruction of Polynomial Orbits: From Correlated Moments to Oscillatory Periods arXiv:2608.14475
Unverified 2026

Jordan spectral feature scaling

Group neural features into small Hermitian matrix elements and scale each group with the paper's tracial spectral Lp norm rather than independently normalizing scalar channels. This introduces a coupled spectral geometry while remaining implementable with ordinary eigendecompositions in the associative Hermitian-matrix special case.

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Spectral nonassociative $\mathrm{L}^p$-spaces for $\mathrm{JBW}^*$-algebras arXiv:2608.14231
Unverified 2026

Spectral Phase-Lifted Features

Augment a hidden representation with positively homogeneous interaction features built from approximate eigenmodes of a linear layer. Fractional products of mode magnitudes and phases provide nonlinear channels whose transformation laws are inherited from the spectrum of the underlying operator, potentially representing oscillatory or multiplicative dynamics more compactly than a generic MLP.

Useful5/10
Difficulty7/10
Novelty8/10
Paper: The spectrum of operator extensions to free Banach Lattices arXiv:2608.11437
Unverified 2026

Gaussian Moment-Window Activation Regularizer

Whiten intermediate feature vectors and constrain several gauge moments to remain in the dimension-dependent interval predicted by the paper's Gaussian/log-concave comparison. Apply the penalty only to moderate orders, where the paper gives a uniform bound independent of the particular log-concave distribution; this should suppress heavy activation tails without forcing all features to be exactly Gaussian.

Useful5/10
Difficulty3/10
Novelty6/10
Paper: Moment comparisons, Sudakov inequalities and entropy of centroid bodies arXiv:2608.10853
Unverified 2026

Authority-Limited Removal Gating

Construct a branching residual network whose active computational paths reproduce according to a fixed offspring/connectivity law, while a controller can only remove paths using an age- or depth-dependent hazard \(u(a)\). Use the resulting bound as a diagnostic and gating schedule: removal can suppress unstable activity and reduce compute, but it should not be expected to cross the reproduction-driven propagation barrier unless the network's expansion operator is also changed.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Removal-Only Actuation in Age-Structured Branching Populations: Fundamental Limits of Equilibrium Placement arXiv:2608.10641
Unverified 2026

Mean-Curvature Relaxation Layer

Insert a small number of differentiable graphical mean-curvature-flow steps between a neural network's raw vector-field prediction and its task loss. The relaxation performs geometry-aware smoothing rather than isotropic Gaussian smoothing, and it can enforce fixed boundary values after every step.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Well-posedness for the mean curvature flow on the half-space and on bounded domains arXiv:2608.08901
Unverified 2026

Adaptive Isotropic Latent Coordinates

Add the paper's joint shape-and-mass distortion objective to a neural coordinate map whose output is a three-dimensional latent representation. Penalize anisotropic local Jacobians through a log-distortion term and penalize nonuniform latent occupancy through a density-gradient term, while learning the radii of an ellipsoidal latent target domain. This should discourage folds and collapsed regions without forcing every dataset into a fixed spherical latent prior.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Adaptive Volumetric Parameterization of Simply Connected 3-Manifolds with Applications arXiv:2608.08672
Unverified 2026

Plucker Compound-Rank Regularizer

Construct a symmetric feature-interaction or Jacobian matrix A_theta whose desired rank is t, then regularize its t-th compound matrix toward rank one. This transfers the paper's identity that a rank-t matrix has a rank-one t-th compound, while the rank-one factor encodes Plucker coordinates of the kernel subspace.

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Brehm-Wintner-Conley Dimension, Plücker Coordinates, and Generalized Dziobek-Williams Equations for Central Configurations arXiv:2608.07771
Unverified 2026

Worst-Subset Conditioning Regularizer

Train an overcomplete linear or MLP layer so that square subsets of its output rows remain numerically invertible after neuron pruning or routing failures. Penalize sampled subsets with unusually small least singular values, using the paper's entropy exponent to quantify the severity expected from random redundancy.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Extreme least singular values of random row submatrices with bounded-density subgaussian entries arXiv:2608.07410
Unverified 2026

Adaptive Lambda-Quantile Prediction Head

Replace a fixed quantile output with a Lambda-quantile head that receives a predictive sample set and applies a learned value-dependent threshold \(\Lambda(x)\). Unlike ordinary quantile regression, the model can use a low threshold in one value range and a high threshold in another, which is useful when error costs or calibration requirements vary across the output domain. Start with a piecewise-constant or monotone spline parameterization, then test whether allowing controlled…

Useful5/10
Difficulty5/10
Novelty6/10
Paper: Lambda-quantiles under the microscope arXiv:2608.07122
Unverified 2026

Auxiliary-energy neural optimizer

Replace the direct nonlinear loss step by a scalar-auxiliary-variable discretization of a gradient flow. The optimizer maintains an auxiliary value representing the square root of the nonlinear energy, so the coupled update has a discrete modified-energy decrease even when the step size is not restricted by the local curvature of the loss.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: A Thermodynamically Consistent Cahn-Hilliard-Navier-Stokes Model for Tumor Growth arXiv:2608.06099