Unverified
2026
Add a coordinate-free regularizer that prevents a batch of unit-normalized embeddings from concentrating almost entirely on one side of a hyperplane passing through their spherical centroid. Sample random directions tangent to the estimated centroid, measure the soft fraction of embeddings in each corresponding hemisphere, and penalize fractions below the spherical Grünbaum constant. This targets directional mode collapse while preserving rotational invariance.
Useful5/10
Difficulty3/10
Novelty7/10
Unverified
2026
Use the paper's degree-sensitive crown inequality to penalize or constrain router assignments that create medium- or high-degree tokens or experts. The resulting router favors a controlled population of low-degree, medium-degree, and high-degree nodes rather than allowing a few hubs to absorb most interactions, which can stabilize sparse attention or mixture-of-experts load balancing.
Useful5/10
Difficulty5/10
Novelty5/10
Unverified
2026
Estimate the spatial distribution of minibatch embeddings using normalized residuals, then use the resulting spatial depth as a bounded confidence weight on each example's loss. Examples whose embeddings are spatially central receive near-unit weight, while isolated or adversarial examples are automatically downweighted without estimating covariance matrices or choosing a dimension-dependent bandwidth.
Useful5/10
Difficulty4/10
Novelty7/10
Unverified
2026
Replace an isotropic Fourier-feature map with a fractional low-pass map whose order is selected from the estimated intrinsic Frostman dimension of the training samples. The layer represents a coefficient vector f in the ambient domain, applies the multiplier |k|^{-s}, and evaluates the smoothed function on the observed fractal-like data support. The theorem provides a geometry-dependent bound preventing high-frequency coefficient energy from producing arbitrarily large responses on concentrated…
Useful5/10
Difficulty5/10
Novelty6/10
Unverified
2026
Treat batches of samples, modalities, or MoE experts as components of a differentiable mixture and add the paper's topology-sensitive RPA free energy to the training objective. Learn a low-dimensional topology descriptor for each component, map it to an effective structure factor, and use the resulting free energy either to promote specialization or to penalize unwanted phase separation in representations.
Useful5/10
Difficulty5/10
Novelty8/10
Unverified
2026
Add a local curvature penalty to graph learning or GNN training that penalizes sampled node signals with negative discrete Bakry–Émery curvature. The regularizer targets graph bottlenecks and irregular diffusion geometry, and can be applied either to a learned adjacency matrix or to the task-relevant hidden representations propagated by a fixed graph.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Build a linear state-space or recurrent layer in a learned pseudo-unitary coordinate frame $\Theta(t)$, and penalize the covariant coefficient $P_{m,\Theta}$ instead of penalizing $\Theta'(t)$ or transition-matrix norms directly. The regularizer is sensitive to meaningful variation of the represented Hamiltonian but is invariant to redundant gauge representations, potentially reducing unstable latent modes without forcing every parameter matrix to be small.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
Augment a representation-learning objective with penalties enforcing the paper's four-point metric inequalities, and use an exponential snowflake kernel instead of unconstrained dot-product similarity. The experiment tests whether geometrically valid similarities improve retrieval or attention stability at equal model size and compute.
Useful5/10
Difficulty5/10
Novelty6/10
Unverified
2026
Represent the sequence of hidden states through a residual or state-space network as a polygonal curve and penalize turns according to their signed moment arm relative to the curve's input and output states. This targets bends that most strongly reduce endpoint separation, rather than applying an unweighted total-curvature penalty. The expected benefit is better long-range signal transport and less folding of hidden trajectories at comparable parameter count.
Useful5/10
Difficulty4/10
Novelty7/10
Unverified
2026
Use a two-gradient predictor-corrector average as the gradient supplied to Adam, retaining trajectory smoothing while avoiding the three or four gradient evaluations required by full RK3. Vary the mixing coefficient to test whether the reported regularization comes from gradient averaging itself rather than from high-order integration.
Useful5/10
Difficulty4/10
Novelty5/10
Unverified
2026
Regularize a scalar feature field on a 2D grid by interpreting each feature value as the uniformizing variable of a hyperbolic ring and penalizing violations of local orthogonal-ring angle closure. Unlike a raw Laplacian penalty, this constrains the representation through positive hyperbolic radii and geometrically meaningful edge compatibility.
Useful5/10
Difficulty5/10
Novelty8/10
Unverified
2026
Construct a recurrent or state-space block with two learned transition matrices A and B representing two commuting update directions. Besides penalizing noncommutation and deviation from isometry, penalize the negative spectrum of the paper's core operator H(A,B), encouraging a structured overlap of one-step and two-step ranges. Compare this against an orthogonal-RNN baseline and against commutation-only regularization on long-horizon sequence tasks.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
For a learned phase-space layer, estimate its symplectic Fourier bandwidth R and divide its output gain by the theorem's support-dependent factor R raised to an exponent determined by the Schatten index p. This creates a resolution-aware normalization: layers with larger phase-space bandwidth are automatically damped when p is not equal to 2, while the Hilbert-Schmidt case p = 2 remains unscaled.
Useful5/10
Difficulty5/10
Novelty6/10
Unverified
2026
Generate structured augmentations of categorical sequences using the paper's adjacent crystal rewrites, then enforce prediction consistency across the resulting orbit. Unlike arbitrary random swaps, the rewrite preserves paired subsequences and modifies only the unmatched portion, making it appropriate for exchangeable discrete codes or explicitly permutation-equivariant inputs.
Useful5/10
Difficulty4/10
Novelty6/10
Unverified
2026
Use eigenvector delocalization as a mask-quality criterion rather than selecting a random sparse graph blindly. Penalize masks whose normalized adjacency has concentrated leading eigenvectors or disconnected or weakly connected components, while preserving the power-law distance prior. This creates a sparse routing graph that is less likely to trap information in local regions.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
Add a loss that prevents an intermediate feature map from being simultaneously concentrated inside a porous spatial region and a porous frequency region. The regularizer is based on the fractal uncertainty inequality: if frequency support is restricted to a porous set Y, then the fraction of feature energy inside a porous spatial set X is at most C h^beta; violations of this bound are penalized.
Useful5/10
Difficulty4/10
Novelty7/10
Unverified
2026
Use the derivative of fractional feature energy with respect to its order as a regularizer for intermediate representations. This penalizes unstable scale behavior rather than simply suppressing all high frequencies, so it can preserve useful detail while discouraging uncontrolled changes across spatial scales.
Useful5/10
Difficulty4/10
Novelty8/10
Unverified
2026
Represent tokens, features, or attention states by normalized rank-one matrices and train the network to preserve their Schatten-p distance profiles over complex phase rotations. Because the paper proves that equality of all distances \(\|\lambda e-v\|_p\) identifies \({\rm Tr}(e^*v)\), this regularizer preserves matrix overlap geometry under a learned transformation.
Useful5/10
Difficulty6/10
Novelty8/10
Unverified
2026
Represent MoE experts as leaves of a balanced ternary tree and regularize the hierarchical boundary of each expert's assignment mask. At fixed routing mass x, the ternary martingale isoperimetric theorem supplies the explicit minimum one-variation T_3(x), so the router can be penalized according to an occupancy-dependent profile rather than a uniform parent-child disagreement cost. This should favor coherent, stable routing regions while preventing small expert supports from obtaining…
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Use a barycentric rational activation or filter whose interpolation nodes are periodically zoomed into the range of preactivations or eigenvalues actually encountered by the network. Protect the layer from catastrophic poles by monitoring the associated generalized eigenproblem and penalizing poles close to the active input interval. This targets rational networks whose expressivity comes from localized poles but whose training is destabilized by denominator zeros.
Useful5/10
Difficulty5/10
Novelty6/10
Unverified
2026
Regularize a point-cloud or graph neural network so that two augmented versions of the same sample induce filtered proximity graphs with approximately interleaved Reeb graphs. The network is encouraged to preserve multiscale connectivity in learned scalar features, not merely pointwise feature similarity or final predictions. Use an approximate interleaving loss for small graphs and the cheaper H0 persistence-distance surrogate for larger batches.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
Add a topological loss that preserves the winding number of a complex numerator field predicted by a neural network. The loss is invariant to positive rescaling of the field, so it penalizes vortex creation or destruction rather than harmless amplitude changes.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Regularize a neural network using exact finite-difference interaction terms at a chosen perturbation scale, while retaining the covering decomposition of a composition f∘g. Instead of penalizing only the total mixed difference, separately penalize selected covering terms containing large subsets or overlapping subsets, which targets higher-order and nonlocal interactions without computing Hessians.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a Microscopic Dynamical Entropy-inspired regularizer to a VAE or sequential world model. Instead of maximizing only the entropy of the latent marginal, maximize latent marginal entropy plus an estimate of the log-volume of unresolved variables compatible with each latent state, thereby preferring representations that summarize predictable macroscopic structure while assigning nuisance detail to the residual channel.
Useful5/10
Difficulty5/10
Novelty6/10