✓✓ Beats tuned baseline
2026
Add a geometry-guided infill operator to a population optimizer used for black-box neural-network tuning. Fit a local Jacobian from recent parameter perturbations and validation-residual vectors, generate a damped Gauss-Newton candidate for exploitation, and sample exploratory candidates in the same Jacobian-derived metric. The host optimizer retains selection, population survival, covariance adaptation, and its total evaluation budget; only a configurable fraction of new candidates is replaced…
Useful6/10
Difficulty5/10
Novelty7/10
△ Mechanism confirmed, baseline not beaten
2026
Use the observed label-availability indicator as an auxiliary supervision signal when labels are preferentially missing for uncertain or difficult examples. Train the classifier with a joint likelihood containing both the class-label likelihood for labeled examples and a missingness likelihood whose probability depends on the classifier's posterior uncertainty.
Useful6/10
Difficulty4/10
Novelty5/10
✗ Mechanism failed
2026
Replace AdamW or SGD updates on simplex-valued routing probabilities with a logarithmic-barrier mirror step. The update remains strictly positive, avoids projection-induced zero coordinates, and can approach a boundary solution asymptotically while retaining the paper's theoretically motivated O(log k/k) convex convergence behavior.
Useful6/10
Difficulty5/10
Novelty6/10
△ Mechanism confirmed, baseline not beaten
2026
Replace unconstrained adversarial example generation with an invertible transport map that is the gradient of a convex potential. For each class, the map pushes a kernel-smoothed empirical distribution toward a least-favorable distribution inside a prescribed KL/Sinkhorn ambiguity radius, producing hard but globally coherent training examples rather than pointwise perturbations.
Useful6/10
Difficulty6/10
Novelty5/10
✓✓ Beats tuned baseline
2026
Replace one deterministic residual update with a short cyclic composition of learned vector fields evaluated for randomized, short run times. Because finite compositions of noncommuting flows generate directional-derivative and Lie-bracket terms, changing the cycle order gives the network an explicit, low-cost way to learn drift directions that are unavailable from the individual vector fields alone.
Useful6/10
Difficulty5/10
Novelty7/10
✗ Failed on benchmark
2026
Represent an intermediate feature as a low-rank PSD matrix and compress it using nonnegative measurements \(\langle A_i,X\rangle\), while penalizing the empirical ratio between maximum and minimum measurement distortion over low-rank feature pairs. This directly discourages collapsed directions and excessively amplified directions in a covariance or Gram-feature bottleneck.
Useful6/10
Difficulty5/10
Novelty6/10
△ Mechanism confirmed, baseline not beaten
2026
Compress the hidden state of a stable neural state-space layer using low-rank controllability and observability Gramians. States that are difficult to excite from the input or weakly visible at the output are removed, producing a smaller recurrent state with a principled input-output preservation criterion.
Useful6/10
Difficulty5/10
Novelty6/10
✗ Failed on benchmark
2026
Replace the usual hand-designed expert-load penalty with a heterogeneous survival penalty derived from a susceptibility distribution. Each expert receives an availability factor q_e=G(A_e), where A_e is its cumulative recent routing pressure and G_e is a learned or fixed mixture of exponentials; highly used experts are suppressed smoothly, while heterogeneous experts can have different resistance to pressure. The mixture produces adaptive curvature and long-tailed penalties that may reduce…
Useful6/10
Difficulty4/10
Novelty6/10
✗ Failed on benchmark
2026
Regularize a circular recurrent kernel by directly controlling the growth rate and phase velocity of its Fourier modes. This converts replay-speed selection into a low-dimensional spectral control problem and can suppress unstable or excessively slow modes without adding recurrent parameters.
Useful6/10
Difficulty6/10
Novelty8/10
✗ Failed on benchmark
2025
Replace a conventional scalar activation by a geometrically indexed family of affine pieces whose slope changes with the logarithmic magnitude of the input. The same two endpoint parameters are reused across all scales, giving a compact, explicitly scale-aware activation that can represent different responses for exponentially separated activation magnitudes.
Useful6/10
Difficulty4/10
Novelty7/10
Unverified
2026
Replace Fourier or sinusoidal one-dimensional coordinate features with a trainable radical layer f(x)=sum_i c_i sqrt(P_i(x)), where every P_i is a strictly positive quadratic. A nonzero scalar output formed by such a layer has at most 2n distinct real zeros, providing an explicit bound on sign changes and suppressing uncontrolled ringing. Use the radical features as an input embedding for a conventional MLP or neural implicit field.
Useful5/10
Difficulty4/10
Novelty8/10
Unverified
2026
Train a score-based diffusion model on a domain with partially reactive constraints, replacing hard rejection or large boundary penalties by a Robin boundary condition. Samples approaching the constraint boundary acquire a hazard proportional to accumulated boundary local time, giving a continuous interpolation between reflection and absorption.
Useful5/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace an unconstrained Cartesian-product router over heterogeneous branches with a router whose joint expert or state assignments obey a finite-group conservation rule. Branch i emits a distribution over labels in its own subgroup H_i of a common finite abelian group G; only tuples whose group sum is zero are retained. This gives an exact, differentiable structural prior for modular arithmetic, multi-relational graphs, multi-view fusion, or any setting where latent labels compose by a…
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Initialize directional prototypes, cosine-classifier weights, or angular attention directions with a spherical t-design rather than iid random vectors. Exact matching of spherical moments through degree t should provide uniform angular coverage and reduce initialization anisotropy, especially when the number of prototypes is small.
Useful5/10
Difficulty4/10
Novelty5/10
Unverified
2026
Replace ad hoc Gaussian or Laplace noise injection with a Laplace majorant calibrated to the observed finite-range sub-Gamma parameters of a neural perturbation. For convex perturbation losses, the calibration guarantees that the expected loss under the scaled Laplace noise upper-bounds the expected loss under every centered random perturbation satisfying the same Bernstein-type MGF constraint.
Useful5/10
Difficulty6/10
Novelty6/10
Unverified
2026
Replace independently random unit-normalized prototypes in a spherical classifier, vector-quantizer, or prototype contrastive head with prototypes selected to cover the sphere evenly. The method directly targets the paper's finding that weak-signal spherical K-means preserves initialization-induced Voronoi structure, reducing redundant prototypes and making early assignments less dependent on random seed.
Useful5/10
Difficulty4/10
Novelty5/10
Unverified
2026
Represent augmentation centers or training examples in a low-dimensional torus and accumulate the geometric overlap of their augmentation neighborhoods. Add a finite-horizon penalty that detects latent locations with insufficient accumulated coverage, while constraining center uniformity so that coverage optimization does not collapse all samples to one location.
Useful5/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace nearest-center or k-modes assignment on a pool of binary neural-network samples with responsibility thresholding followed by a coordinate-wise dominance screen. Only retain a candidate mode when its assigned samples are sufficiently explained by that mode and, at every bit position, the responsibility-weighted majority agrees with the proposed center; otherwise mark the mode unreliable or discard it.
Useful5/10
Difficulty4/10
Novelty7/10
Unverified
2026
Apply a regularizer that penalizes feature disagreement under a finite set of known transformations. The paper's spectral-gap inequality gives a quantitative reason that this local consistency penalty controls distance from the subspace invariant under the transformation group, while the task loss prevents undesirable collapse.
Useful5/10
Difficulty5/10
Novelty5/10
Unverified
2026
Add a mixed Fourier-L1 penalty to a particle or molecular neural network so that frequencies involving selected coordinate blocks are penalized by products of per-coordinate weights, rather than only by one isotropic norm. This should favor representations that capture pairwise or blockwise structure efficiently in high-dimensional configuration spaces, especially for wavefunctions, molecular energies, and other permutation-structured functions.
Useful5/10
Difficulty6/10
Novelty7/10
Unverified
2026
Represent an original neural-network block and a proposed rewritten block as constrained optimization formulations over inputs and trainable parameters, then certify that the rewrite preserves feasibility and the ordering of losses over a bounded domain. This gives a compiler or pruning pipeline a formal reject/accept gate instead of relying only on numerical regression tests.
Useful5/10
Difficulty7/10
Novelty8/10
Unverified
2026
Regularize a scalar network output so that its superlevel sets are approximately quasiconcave in input or latent space. Instead of penalizing the full Hessian, penalize positive curvature only in directions orthogonal to the output gradient, matching the paper's projected-Hessian and weighted 1-Laplacian structure.
Useful5/10
Difficulty6/10
Novelty6/10
Unverified
2026
Use the graph-coloring stability concept to route graph nodes to experts. Nodes rank experts by router logits, adjacent nodes are constrained to use different experts, and a blocking cycle is a directed cycle in which every node prefers the expert currently assigned to the next node. Eliminate profitable feasible cycles or penalize their existence so routing reaches a locally stable assignment instead of oscillating between equally plausible expert allocations.
Useful5/10
Difficulty6/10
Novelty9/10
Unverified
2026
Represent each example or minibatch by two positive semidefinite feature maps, such as teacher and student covariance operators, and penalize their noncommutative operator-valued f-divergence rather than only a scalar KL or Frobenius distance. The matrix-valued penalty preserves directional disagreement in feature space and is compatible with positive postprocessing, making it a candidate replacement for covariance matching in distillation and representation regularization.
Useful5/10
Difficulty5/10
Novelty6/10