ML: Mlp

Machine-learning ideas tagged Mlp in the ML taxonomy of the Math2NN corpus.

408 ideas found

Unverified 2026

WKB-Stokes Sparse Mixer

Replace a dense channel-mixing matrix in a sequence layer with alternating diagonal propagation and sparse unipotent Stokes jumps. The diagonal part carries independently controlled exponential phases, while the unipotent factors implement cheap residual-like mode conversion without changing determinant or requiring a dense matrix multiply. Constrain the phase magnitudes and jump coefficients during training to obtain a reversible, norm-monitorable mixer.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: Holonomy Asymptotics along Quartic Differential Rays arXiv:2608.04729
Unverified 2026

Circular Morera Regularizer

Add a multiscale circular-integral penalty to a complex-valued neural field f_theta: R^2 -> C. The penalty directly tests the local contour condition that characterizes holomorphic functions, providing a derivative-free alternative to explicitly penalizing the Cauchy-Riemann residual.

Useful5/10
Difficulty3/10
Novelty7/10
Paper: An Infinitesimal Circular Morera Theorem arXiv:2608.04540
Unverified 2026

Bohnenblust–Hille coefficient regularization

Replace ordinary coefficient decay in a degree-d polynomial neural layer with the Bohnenblust–Hille coefficient quasi-norm, whose exponent p=2d/(d+1) is dimension-independent and strictly below 2 for d>1. Combine this penalty with a sampled torus supremum penalty so the layer is constrained both in its realized function amplitude and in the coefficient geometry predicted by the inequality.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Tightness of and counterexamples to several quantum estimates arXiv:2608.04411
Unverified 2026

Type-B fermionic equivariant layer

Represent each of n signed tokens with two or more anticommuting feature channels and build equivariant outputs from exterior products rather than unconstrained tensor products. Penalize or project out positive-degree signed-permutation invariants, approximating the coinvariant quotient so that the layer retains order-sensitive orientation information without learning redundant invariant directions.

Useful5/10
Difficulty6/10
Novelty6/10
Paper: Type $B$ fermionic coinvariant rings arXiv:2608.02881
Unverified 2026

Topology-Guided Capacity Allocation

Use the layer at which persistent connected components and holes disappear to allocate capacity nonuniformly across a network. If representations simplify much earlier than desired, widen the responsible layers or insert an additional block; if simplification is excessively delayed, avoid spending parameters there. This turns persistent-homology COM into an actionable architecture-search signal rather than a post-hoc visualization.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Topological Simplification in Predictive Coding Networks arXiv:2608.02816
Unverified 2026

Pascal-simplex anti-collapse router

Replace an unconstrained collection of coefficients over degree-nR compositions by a signed simplex-indexed coefficient tensor satisfying the paper's local cancellation equations. Anchor the balanced coefficient and use the resulting discrete unique-continuation principle to prevent the learned tensor from collapsing onto a tiny set of compositions, while still allowing structured sparsity below the full simplex size. Apply the tensor to a signed residual feature mixture or to expert logits…

Useful5/10
Difficulty6/10
Novelty9/10
Paper: Discrete Unique Continuation on Simplex arXiv:2608.02707
Unverified 2026

Pair-monotone implicit optimizer

Replace the explicit parameter update \(\theta_{k+1}=\theta_k-\eta\nabla L(\theta_k)\) with an approximate generalized proximal step defined by a simple map \(v\). The map is chosen so that the gradient operator and v satisfy an empirical pair-monotonicity condition, allowing larger stable outer steps and reducing oscillations in stiff or highly curved neural-network training.

Useful5/10
Difficulty7/10
Novelty6/10
Paper: Convergence and Stability Analysis of a Generalized Proximal Point Algorithm and Its Inexact Version arXiv:2608.01584
Unverified 2026

Entropy-Certified Interaction Supports

Replace a dense third-order channel-interaction tensor by a fixed sparse support selected through the paper's uniform-marginal infeasibility certificate. Supports with a large dual margin have an effective entropy base below the channel alphabet size, suggesting fewer independent interaction slices and cheaper contractions. Use the certificate either during architecture search or as a pruning score for an already-trained tensorized layer.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Slice and Partition Rank Criteria for Polynomial Zero-Avoidance arXiv:2607.29490
Unverified 2026

Pfaffian activation budget

Use a tanh MLP with an explicitly tracked Pfaffian-chain complexity and select its width and input sparsity using the paper's zero-count bound. The bound limits the number of regular decision-boundary crossings along one-dimensional data-space restrictions, so it provides a principled way to discourage excessively oscillatory fits beyond ordinary weight decay.

Useful5/10
Difficulty4/10
Novelty8/10
Paper: Khovanskii's Bezout-type Theorem for Pfaffian Functions: A Self-Contained Proof, and Applications arXiv:2607.29267
Unverified 2026

Boundary-Equalized Conformal Neural Coordinates

Train an MLP coordinate map so that its local scale distortion is smooth in the interior and approximately constant on the boundary of the parameter domain. This implements the Chebyshev-Darboux-Milnor principle as a regularizer for neural parameterizations, potentially reducing boundary stretching and improving interpolation quality on learned geometric domains.

Useful5/10
Difficulty5/10
Novelty6/10
Paper: On the construction of geographical maps: Lagrange, Chebyshev, Darboux and Milnor arXiv:2607.29263
Unverified 2026

Distributional spectral-preconditioned features

Replace or augment a singular scalar activation \(\sigma\) with a distributionally regularized activation \(g\) whose Fourier transform is multiplied by \((i\rho)^\alpha\). This suppresses the problematic low-frequency singular component and can produce better-conditioned random-feature or first-layer representations, while a residual raw-activation branch prevents loss of standard approximation behavior.

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Radon Measure Representations for Infinite-Width Neural Networks with Singular Activations arXiv:2607.29258
Unverified 2026

Flow-Constrained Hierarchical Policy

Parameterize a tree-structured policy through realization weights satisfying sequence-form flow conservation, instead of independently predicting probabilities at every node. Conditional action probabilities are recovered by dividing a child sequence weight by its parent weight, guaranteeing globally consistent probabilities and avoiding invalid or contradictory branch masses. This is suitable for hierarchical RL policies, adaptive computation trees, and neural routers with sequential gating…

Useful5/10
Difficulty4/10
Novelty6/10
Paper: Baseball, An Extensive-Form Game-Theoretic Duel arXiv:2607.29041
Unverified 2026

Incoherent Frame Averaging for Tensor Layers

Add randomized orthogonal frame mixing and an incoherence penalty to tensorized neural layers so that predictions and gradients are less controlled by a small coordinate block. The goal is to retain the bulk, approximately Gaussian behavior of tensor contractions while preventing rare coherent directions from dominating training.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Regularized Bulk Universality versus Bounded-Disorder Nonuniversality for Annealed Complexity of Spherical $p$-Spin Landscapes arXiv:2607.27613
Unverified 2026

Jacobi Moment Spectral Regularizer

Regularize the Gram spectrum of selected neural layers so that its low-order moments match the spectral moments generated by a truncated q-boson Jacobi operator. Unlike a simple Frobenius or spectral-norm penalty, this controls several parts of the singular-value distribution simultaneously and can discourage harmful spectral tails without forcing all singular values to be equal.

Useful5/10
Difficulty5/10
Novelty6/10
Paper: Sharp Bounds on Ground State Energy of the SYK Model arXiv:2607.27185
Unverified 2026

q-Boson Spectral Initialization

Initialize a neural layer with singular values taken from the finite spectral measure of the paper's q-boson Jacobi operator instead of using Xavier or ordinary orthogonal initialization. The resulting layer has a deliberately shaped singular-value distribution and an explicit finite-size spectral edge, allowing initialization to target stable signal propagation while retaining spectral diversity.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: Sharp Bounds on Ground State Energy of the SYK Model arXiv:2607.27185
Unverified 2026

Noise-crossing band-pass neuron

Replace selected ReLU or sigmoid units with a stochastic binary crossing activation that fires only when exactly one of two independent noise thresholds is crossed. The resulting expected activation is low for inputs far below or far above the noise distribution and maximal near its median, creating an analytically controlled band-pass and potentially reducing saturation-driven instability.

Useful5/10
Difficulty4/10
Novelty6/10
Paper: Reconstructing Backpropagation from Forward Fluctuations in Noise-modulated Neural Networks arXiv:2607.26483
Unverified 2026

Derangetropy Rank Warp

Insert a distribution-free rank warp before selected MLP or attention projections. For each scalar activation, replace its empirical rank u by the cumulative interval map induced by the Type-III derangetropy kernel, optionally followed by Gaussian or affine output calibration. The transform is invariant to strictly increasing reparameterizations of the feature and contracts the marginal toward central ranks, potentially reducing sensitivity to heavy tails and outliers.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Derangetropy Operators arXiv:2607.24705
Unverified 2026

Sharp Curl-Helicity Regularizer

Add a scale-invariant inequality penalty to a neural vector-potential model on a discretized round 3-sphere. The penalty enforces the theorem's sharp lower bound between the L^{3/2} norm of the predicted magnetic field B=curl A and its helicity H=<B,A>, discouraging pathological high-frequency or spatially concentrated fields that fit observations but have implausible geometry. A divergence-free gauge and Killing-form initialization make the constraint numerically well-conditioned.

Useful5/10
Difficulty5/10
Novelty9/10
Paper: The sharp curl-Sobolev inequality arXiv:2607.23827
Unverified 2026

Shared-support heterogeneous-strength adapters

Replace a collection of dense task-specific linear layers with a common sparse structural matrix and task-specific edge strengths. All tasks share the same learned connectivity pattern, but retain independent values on active connections, allowing parameter sharing without forcing identical interactions.

Useful5/10
Difficulty5/10
Novelty5/10
Paper: Estimation of multiple precision matrices under shared support with heterogeneous edge strengths arXiv:2607.23577
Unverified 2026

Magnitude-Euler Path Signature Regularizer

Represent the computation graph of an MLP as a directed acyclic Lawvere metric space and compute a truncated, length-resolved Euler signature of its active paths. Add a penalty that separates signatures between classes while suppressing signatures that are insensitive to labels, thereby encouraging globally distinct computation routes without changing layer widths or degree statistics.

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Magnitude homology and Euler characteristics of directed acyclic graphs arXiv:2607.23357
Unverified 2026

Orthogonal symmetric pair embedding

For every unordered pair of scalar features, construct invariant coordinates from the elementary symmetric quantities s=x+y and q=xy, then feed a truncated orthogonalized polynomial basis in (s,q) to the neural network. Estimate the basis by weighted Gram-Schmidt or Cholesky whitening under the paper's triangle weight, so polynomial channels have low redundancy and controlled scale instead of requiring an unconstrained MLP to learn both symmetry and decorrelation.

Useful5/10
Difficulty3/10
Novelty7/10
Paper: Symmetric Jacobi Polynomials on a Triangle and Their Spectral Algebra arXiv:2607.22751
Unverified 2026

Gap-Graded Neural Collocation

Replace uniform PINN or neural-operator collocation by a graded point distribution concentrated in narrow regions between nearly touching interfaces. Use the paper's distance-dependent mesh scale to determine point spacing, and switch to a gap-dependent minimum scale when the separation becomes too small for the global mesh.

Useful5/10
Difficulty4/10
Novelty6/10
Paper: High-order finite element method for perfect conductivity and linear elasticity with nearly touching inclusions arXiv:2607.22128
Unverified 2026

Convex-Ordered Count Head

Equip a neural-network count head with a mean parameter and a dispersion parameter from the Conway-Maxwell-Poisson family, then enforce a mean-preserving convex-order relationship between predictions. This provides a principled way to make the predictive count distribution more or less tail-dispersed while retaining the same predicted mean, potentially improving calibration on overdispersed or underdispersed count data.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Lorenz and convex ordering of parasite burden distributions with density-dependent deaths arXiv:2607.21931