Unverified
2026
Parameterize a learned token metric as a nonnegative sum of sparse integral rank-one projections with unimodular support, rather than learning an unconstrained dense positive-semidefinite matrix. Graph-incidence covectors give an immediately implementable support family, while nonnegative coefficients guarantee positive semidefiniteness by construction.
Useful5/10
Difficulty5/10
Novelty5/10
Unverified
2026
Add a learnable curved augmentation trace to latent features and penalize excessive overlap between its translated tubular neighborhoods. The regularizer uses the paper's curvature-driven bound as a scale-dependent target: nearby translations may overlap at order delta, while translations at distance r should overlap only at order delta squared divided by r. This encourages feature perturbations to form a non-flat, coverage-efficient manifold rather than collapsing onto a line or a small set of…
Useful5/10
Difficulty5/10
Novelty8/10
Unverified
2026
Add a regularizer that rewards each neuron's expected absolute response to random sign perturbations, normalized by the neuron's l2 norm so ordinary weight scaling cannot trivially increase the objective. Use the paper's distance-sensitive Khintchine lower bound to penalize filters close to the two-coordinate extremal set, promoting distributed and perturbation-stable feature extraction.
Useful5/10
Difficulty3/10
Novelty8/10
Unverified
2026
Insert an overcomplete sparse feature bottleneck into an MLP or embedding stream: encode an activation h with z = ReLU(W^T h + b), then reconstruct or continue computation from Wz. Normalize dictionary columns and train them to remain nearly tight and low-coherence, while choosing a negative bias from an estimate of worst-case cross-feature interference. The hypothesis is that this gives cleaner, more stable feature supports than an ordinary L1 sparse autoencoder at the same latent width.
Useful5/10
Difficulty5/10
Novelty4/10
Unverified
2026
Construct a Fourier layer whose active frequencies lie on several nonparallel polygonal patches or thin annular sectors, and cap repeated difference vectors generated by pairs of patches. The bounded-multiplicity geometry limits how many input frequency pairs can contribute to the same output frequency, potentially reducing spectral aliasing and gradient variance in nonlinear Fourier mixing.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
Train a neural approximation to a scale-dependent effective action, energy functional, or field while penalizing the residual of a known continuous-symmetry Ward identity. Select the regulator, smoothing scale, or architecture hyperparameter at the minimum Ward residual, and require that the residual decreases when model capacity or derivative-expansion order increases.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace or augment a scalar periodic positional coordinate with a normalized bank of odd Fourier harmonics, keeping every position on the same-radius sphere. The resulting representation has an explicit translation-invariant similarity kernel, allowing the frequency count and spectral weighting to control how sharply attention distinguishes nearby versus distant phases.
Useful5/10
Difficulty3/10
Novelty3/10
Unverified
2026
Build a fixed multiscale router that maps 2D coordinates to 3D voxel coordinates using the paper's X-shaped self-similar refinement. Use the router to run a 3D feature field or volumetric token mixer over a 2D-organized tensor, while retaining a mathematically controlled locality bound instead of an arbitrary flattening permutation. The first target is a 3D neural field or small voxel classifier where the router replaces either a dense 3D feature table or a naive raster-order token layout.
Useful5/10
Difficulty6/10
Novelty6/10
Unverified
2026
Restrict a fine-tuning adapter or output head to the subspace invariant under a prescribed monodromy, analogous to the paper's unbroken flavor lattice. This removes update directions intentionally changed by the domain-loop transformation, producing a parameter-efficient adapter with an explicit algebraic constraint.
Useful5/10
Difficulty4/10
Novelty8/10
Unverified
2026
Insert a fixed reversible lattice shear into a residual network so successive blocks follow a structured monodromy orbit rather than using unrelated learned transformations. Apply the transformation to a small learned subspace of hidden channels while leaving the remaining channels unchanged. This creates deterministic phase-dependent feature mixing with no additional trainable parameters.
Useful5/10
Difficulty5/10
Novelty8/10
Unverified
2026
Add a geometric loss to a neural scalar field on hyperbolic latent coordinates, requiring the shifted Hessian \(\nabla^2v-vg\) to remain positive definite while matching the self-shrinker curvature equation. A boundary trace on a finite approximation of the ideal boundary conditions the solution, encouraging a canonical hyperbolically convex extension instead of arbitrary interpolation.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
Add a Fourier-domain residual loss whose per-frequency weight is determined by the geometric overlap of a convex bandwidth domain with its reflection about that frequency. Frequencies close to the boundary receive larger weight through \(\omega_\Omega^{-d}\), forcing the network to model fragile spectral components instead of optimizing only the high-energy interior. Use clipping or a bounded transform of the singular weight so that a few boundary bins cannot dominate training.
Useful5/10
Difficulty3/10
Novelty6/10
Unverified
2026
Use PLMS endpoint parameters to impose an explicit penalty or constraint on lower- and upper-tail dependence between learned representation coordinates. This targets rare-event co-activation directly, rather than relying on covariance or average correlation to control extreme latent behavior.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Generate pairs of latent variables with exactly uniform marginals but non-Gaussian, asymmetric dependence by applying a randomly chosen PLMS map to one uniform latent coordinate. The coupling can expose a model to controlled concordant, discordant, or piecewise-dependent examples without changing either marginal distribution.
Useful5/10
Difficulty4/10
Novelty8/10
Unverified
2026
Train a scalar neural field on a bounded convex domain with a restricted half-Laplacian residual and an explicit strict-concavity barrier. The paper's theorem motivates requiring the learned potential to have negative-definite Hessian throughout the domain, while the nonlocal residual gives the model a global Cauchy-process-style inductive bias rather than only local smoothness.
Useful5/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace or augment the first embedding layer for antipodally identified inputs with the normalized traceless quadratic map from the Veronese construction. Because q and -q produce exactly the same feature, the layer enforces projective invariance by construction rather than learning it from augmented examples. The resulting matrix-valued features can be flattened, projected, or processed by an equivariant linear layer.
Useful5/10
Difficulty2/10
Novelty6/10
Unverified
2026
When a neural field learns power-law exponents, penalize exponent configurations whose Newton support violates the paper's finite-distance accessibility condition. This discourages combinations of exponents that create excessively strong joint singularities while preserving anisotropic scaling when it is supported by the data.
Useful5/10
Difficulty3/10
Novelty8/10
Unverified
2026
Replace a time-homogeneous recurrent update by a sequence of parameterized maps f_t, and regularize late-time pairs of updates to approximately commute: applying block f_t followed by f_r should agree with applying f_r followed by f_t. This should make long-horizon predictions robust to local time-step reorderings and schedule perturbations, while proximal statistics provide a diagnostic for whether trajectories repeatedly approach one another rather than diverging permanently.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace rejection sampling or coordinate random walks for adversarial and augmentation perturbations in a convex feasible set with Hit-and-Run: choose a random direction through the current perturbation, compute the exact feasible chord, and sample uniformly on that chord. The paper's spectral-gap result predicts faster global exploration when the perturbation polytope is rounded or whitened, while preserving feasibility at every step.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Apply a weighted Caffarelli–Kohn–Nirenberg deficit to selected intermediate feature channels. The penalty discourages features that obtain large weighted responses only by becoming sharply localized or highly sensitive to small input perturbations. It can be evaluated with input-Jacobian estimates and added to the ordinary task loss.
Useful5/10
Difficulty5/10
Novelty5/10
Unverified
2026
Regularize probability-valued network outputs in the square-root representation rather than directly penalizing density curvature. This suppresses sharp oscillations while avoiding the severe scaling of derivative penalties involving \(\nabla\rho/\rho\) near vacuum regions.
Useful5/10
Difficulty3/10
Novelty6/10
Unverified
2026
Replace an ordinary elementwise interaction between two feature matrices by a noncommutative functional-calculus layer \(\varphi(A,B)\), where \(A\) and \(B\) are Hermitian channel operators that need not commute. Add a soft penalty on \([A,B]=AB-BA\), and use a Besov-smooth parameterization of \(\varphi\) so that perturbations are controlled in Schatten \(p\)-norm for \(p\leq2\). This creates a principled matrix interaction module that can remain stable when feature operators or graph…
Useful5/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace the symmetric Euclidean contrastive loss between embeddings with a two-point quadratic contrast whose displacement is generated by a local affine connection and measured using the metric at the source endpoint. Because the metric and transport need not be compatible, the loss can be asymmetric, allowing the model to represent directional relations between examples.
Useful5/10
Difficulty6/10
Novelty5/10
Unverified
2026
Replace ordinary isotropic residual noise in a normalized continuous-depth block with projected Brownian forcing on the unit sphere. Apply a shared random symmetric quadratic drift to all tokens, plus a small token-specific tangent perturbation; the shared term preserves structured antipodal dynamics while the independent term removes persistent symmetry and cluster degeneracy. This is intended as a controlled stochastic regularizer, not merely additive Gaussian noise.
Useful5/10
Difficulty5/10
Novelty7/10