Unverified
2026
Use coordinate hit-and-run rather than isotropic Gaussian random walks to generate latent negatives or augmentation trajectories inside a convex latent domain K. At each step, select one coordinate and resample the entire feasible chord along that coordinate; the paper's l0-isoperimetric theorem predicts that sets of non-negligible mass cannot be separated by severe coordinate-only bottlenecks when K is well-conditioned relative to an unconditional body Q.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Partition neural modules into two empirically identified reliability or noise classes and restrict their communication graph to a two-block stochastic block model. Allocate a fixed connectivity budget across within-class and cross-class edges using a water-filling update that favors block pairs producing the largest increase in validation utility. The resulting layer is sparse and modular, with a testable prediction that optimal connectivity concentrates on a few block pairs rather than…
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace fixed LoRA factors with a rank-adaptive moving subspace whose columns are augmented using derivative information from several Runge–Kutta stages. The optimizer integrates a matrix-valued gradient-flow approximation inside this enlarged left/right basis, allowing high-order motion of the adapter subspace while retaining a low-rank parameterization.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Add a Maximum Entropy on the Mean penalty to an inverse-model output or neural latent code using an empirical prior library of plausible vectors. The penalty selects the least-KL distribution over prior samples whose mean equals the network prediction, encouraging reconstructions to lie in statistically plausible regions without requiring a differentiable density estimator.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Use the paper's MBM-GP construction to predict input-dependent big-M constants for ReLU disjunctions during neural-network verification. Exact activation-bound optimization is performed only at a small subset of input points, while a Gaussian-process upper confidence bound supplies conservative bounds elsewhere, reducing verifier preprocessing and potentially tightening the MILP compared with one global worst-case constant.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Partition trainable parameter blocks into specialists that receive a fixed task or data-domain assignment and generalists that stochastically sample tasks at every update. Estimate local ruggedness from the correlation between losses at nearby parameter perturbations, then increase the generalist fraction when this correlation is low and increase specialization when the landscape is smooth. The mechanism mirrors the paper's permanent-specialist versus stochastic-generalist allocation while…
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Use the change in the policy-induced reachable set as a trust-region constraint, rather than limiting only parameter distance or KL divergence. A policy update is accepted when its predicted finite-horizon zonotope remains sufficiently close to the previous reachable tube and does not cross the safety boundary, yielding a dynamics-aware step-size ceiling.
Useful6/10
Difficulty7/10
Novelty8/10
Unverified
2026
Replace constant friction and optimizer noise with a velocity-dependent friction gamma(u) and noise amplitude tied by a fluctuation-dissipation relation. High-speed momentum states can be damped and randomized differently from low-speed states, creating controlled transient exploration while preserving a known equilibrium momentum distribution.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
For a neural stochastic state-space model or discrete diffusion sampler, monitor whether learned transition logits admit a global scalar potential on the active latent manifold. Penalize residual cycle affinities in the conditional sector, but leave reset cycles unpenalized so the model can retain useful dissipative mixing. The distinctive prediction is a linear decrease of integrability error with residual cycle current and a quadratic decrease of entropy production near autonomous…
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a local Jacobian spectral regularizer and an initialization sweep to steer a looped transformer away from uncontrolled near-unit dynamics. The goal is to prevent examples from entering a fold-critical regime with very long relaxation times, or alternatively to deliberately target a controlled critical regime when adaptive test-time compute is useful.
Useful6/10
Difficulty6/10
Novelty5/10
Unverified
2026
Treat a slowly varying block of neural-network parameters as a coarse-grained stochastic process and continuously estimate both its covariance spectrum and its linear response to small artificial perturbations. Use the fluctuation–response mismatch as a feedback signal to tune injected parameter noise or minibatch size; the thermal Einstein relation is imposed only when a calibrated equilibrium-like regime is desired, while antisymmetric response components are retained as admissible…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace a conventional optimizer step by a three-phase cyclic update in which successive parameter blocks or gradient components are exposed to two low-noise phases and one high-noise, chemically driven phase. Treat the loss decrease as mechanical work, phase-dependent gradient-noise scales as reservoir temperatures, and an auxiliary drive as chemical free energy. Adapt the drive toward a target positive cycle affinity rather than increasing the learning rate indefinitely, creating a measurable…
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Represent a higher-order neural computation as a bipartite incidence graph between node features and hyperedges, and assign each node-hyperedge incidence an anchor probability or learned anchor score. Add a regularizer that maximizes the predicted size of the surviving (k,n)-core under random node, hyperedge, or token dropout, thereby preventing structured pruning or routing from disconnecting essential higher-order computations. At inference, retain only incidences belonging to the predicted…
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Replace a purely smooth momentum update by a second-order parameter dynamics with short, explicitly scheduled impulses at the beginning of each training window. The impulse is chosen to produce the required parameter displacement while the smooth gradient force handles local relaxation; this directly transfers the paper's linear-versus-quadratic short-time work mechanism.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace the usual momentum state in an optimizer with a persistent Ornstein-Uhlenbeck-driven velocity subject to a dry-friction threshold. Correlated forcing can help traverse shallow noisy regions, while the friction term suppresses parameter motion when the effective force is small, potentially reducing update noise and improving late-stage stability.
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Model the integer token loads of a mixture-of-experts layer as a canonical occupancy system with a fixed total number of tokens. A distributed routing phase persists while the normalized load is below a critical value; beyond that point, the excess load is either allowed to condense into a designated overflow expert or penalized if expert collapse is undesirable. The key benefit is an explicit transition criterion and finite-batch fluctuation diagnostic for routing collapse.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a small dynamical state on the transformer module graph and use it to control adaptive computation, but reject controller parameters whose discrete-time update has latent roots outside the unit disk. The state can modulate halting thresholds, residual-block gains, and memory gates; the certificate applies to the controller integrator and prevents unstable oscillations or exploding internal control signals during long adaptive-depth rollouts.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Augment an optimizer with two slowly and periodically modulated controls, such as learning rate and momentum or learning rate and gradient-noise scale. The optimizer state then traces a loop in control space; nonzero curvature can create a net parameter displacement that depends on loop orientation, even when the controls return to their initial values. Use curvature estimates to select loops that produce useful descent while penalizing loops with excessive dissipation.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace an unconstrained recurrent hidden-state channel with a two-dimensional oscillator constrained to the supercritical Hopf normal form. A learned control parameter can place the channel below threshold for decaying dynamics or above threshold for sustained periodic dynamics, while the cubic term bounds the amplitude and prevents recurrent-state explosion.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Treat groups of neural-network states or experts as metastable sectors and estimate both sector imbalance and inter-sector connectivity from minibatch routing or trajectory transitions. At balanced sector usage, the effective two-sector spectral splitting becomes a direct estimate of connectivity: a large splitting indicates that the sectors are still strongly communicating, whereas a small splitting indicates genuine specialization or incipient collapse into disconnected modes.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a temperature-response constraint to stochastic neural predictors so that changes in inverse temperature cannot produce disproportionately large changes in expected loss or energy. This converts the nonequilibrium fluctuation-response inequality into a measurable robustness monitor and a regularizer for beta-conditioned stochastic representations.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace a static MoE load-balancing penalty with a two-stage capacity allocator. First compute each expert's technically feasible token capacity from latency, memory, and overflow constraints; then redistribute capacity using cumulative proportional fairness so experts that were repeatedly under-served receive more capacity later. Constrain the redistribution by an explicit efficiency budget, so fairness cannot silently cause an uncontrolled increase in routing loss or expert compute.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
For a neural dynamical predictor, train or maintain several independently initialized models and aggregate their multi-step states using the signed displacement along the locally unstable forecast direction. The key mechanism is cancellation of opposite unstable-manifold errors: ordinary averaging should reduce this component at rate N^{-1/2} when errors are independent and centered, while robust aggregation should be activated when validation residuals show heavy tails or persistent bias.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Augment SGD or Adam with a short-window estimate of optimizer trajectory entropy production obtained from forward and reverse minibatch or noise paths. Reduce the learning rate when estimated dissipation rises sharply, and increase it only when dissipation remains controlled, avoiding the rare-event sensitivity of exponential work estimators.
Useful6/10
Difficulty6/10
Novelty7/10