Architecture ideas

Research ideas extracted from mathematics papers, categorized as Architecture.

Unverified 2026

Positive Reflected Bellman Layer

Replace an unconstrained spatial aggregation in a neural PDE surrogate or controlled-dynamics model with a fixed-branch expectation layer. Each output is a maximum over controls of a nonnegative weighted average of next-state values, with reflected overshoots attenuated by Robin factors. Increasing any input value therefore cannot decrease the output, giving a hard monotonicity and positivity property instead of relying on a penalty.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: A Positivity-Preserving Expectation Scheme for Hamilton--Jacobi--Bellman Equations with Oblique Robin Boundary Conditions arXiv:2608.11936
Unverified 2026

Ambiguity-Flat Weyl Feature Layer

Replace a random cyclic filter bank or patch projection with the Weyl–Heisenberg orbit of one normalized learnable prototype. Regularize the prototype so that all nonzero shift and modulation correlations have a large and nearly equal magnitude, maximizing the smallest eigenvalue of the induced feature Gram matrix and preventing poorly observed feature directions.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Uniformly Stable Minimal Weyl--Heisenberg Measurements Approaching the SIC Benchmark arXiv:2608.11850
Unverified 2026

Spherical-design directional heads

Use a fixed spherical t-design as the direction codebook for a directional attention or feature-aggregation module instead of independently sampled random directions. Equal weights provide exact zero mean and isotropic second moments, while exactness for spherical polynomials up to degree t reduces directional aliasing and seed-dependent anisotropy.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Minkowski Polytopes of Spherical Designs: High-Order Isotropy and Quantitative Sphericity arXiv:2608.11570
Unverified 2026

Smith-Reduced Chain Encoder

Replace an ordinary hierarchical graph encoder with a finite chain-complex encoder whose learned boundary maps satisfy \(\partial_{k-1}\partial_k=0\). Compute Smith normal form on the integer incidence matrices and treat unit-labelled cell pairs as refinement overhead: cancel or gate those pairs before message passing, while preserving non-unit labels that encode genuinely nontrivial structure. The resulting representation should be insensitive to arbitrary cell subdivision while retaining…

Useful6/10
Difficulty6/10
Novelty6/10
Paper: A Chain- and Diagram-Level Semantics for Morphological Calculus Refinement, monodromy, and bivector orbit decompositions arXiv:2608.11325
Unverified 2026

Deep-spectrum community features

Replace the usual top-eigenvector positional encoding in a graph neural network with a density-selected spectral subspace. The selector explicitly searches below the leading eigenvectors, where community information may survive after latent geometric modes have consumed the largest eigenvalues. The selected coordinates can be concatenated to node features or used as a bias in graph attention.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Spectral graph clustering with inhomogeneous latent geometry arXiv:2608.11321
Unverified 2026

Finite-Difference Conditional Neural Operator

Replace a tensor-product network over a low-dimensional state and a large distribution embedding with a neural operator that consumes the distribution vector once and outputs values on a finite-difference grid in the low-dimensional state. Train it with the governing PDE residual, explicit boundary residuals, and optional signed shape constraints, allowing the network to preserve numerical structure that a generic MLP would learn only implicitly.

Useful6/10
Difficulty5/10
Novelty5/10
Paper: Mastering Stochastic OLG Models in Continuous Time arXiv:2608.11134
Unverified 2026

Equivariant Critical-Mode Branch Seeding

When a symmetry-frequency block becomes critical, initialize or perturb the network specifically along its critical representation rather than injecting isotropic noise into all hidden channels. This creates trainable branches for the symmetry patterns predicted by the bifurcation calculation and can expose useful periodic solutions that ordinary symmetry-preserving training fails to reach.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Local and Global Equivariant Bifurcation for Periodic Weyl and Riesz Fractional Equations arXiv:2608.11101
Unverified 2026

Reversal-Assisted MoE Routing

Replace one-shot top-k expert assignment with a capacity-constrained stochastic routing process in which tokens have a temporary routing direction and can reverse it at rate gamma. Tokens preferentially move through short vacancy clusters, while reversals break persistent directed congestion and should delay or eliminate expert-level jams. This creates a tunable routing phase diagram rather than relying only on an auxiliary load-balancing loss.

Useful6/10
Difficulty6/10
Novelty8/10
Paper: Jamming transition in an active exclusion process arXiv:2608.11041
Unverified 2026

Successive Orthogonal Innovation Blocks

Add a neural feature, adapter, or expert block only through the component of its outputs that is orthogonal to the span of all previously installed blocks. Quotient coefficient directions that produce nearly identical outputs with an SVD or pseudoinverse, so the new block contributes intrinsic representational dimensions instead of duplicating old features. The expected benefit is a smaller effective architecture and better-conditioned block expansion.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Successive Schur-Riesz Analysis for Approximation arXiv:2608.10757
Unverified 2026

Unit-circle-root Toeplitz mixer

Replace a freely learned finite impulse-response mixing kernel with a matrix polynomial whose roots are constrained to the unit circle. The resulting block-Toeplitz operator has an explicitly positive semidefinite spectral construction, while increasing the polynomial degree gives a systematic capacity knob for approximating matrix-valued frequency responses.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Pure matrix states on block Toeplitz matrices arXiv:2608.10701
Unverified 2026

Tree-Hall Collision-Free Router

Replace independent top-k routing by a tree-structured hypergraph assignment layer. Each candidate route is a singleton or pair of resources, and the router selects exactly q_e routes for every tree edge e while ensuring that no resource is consumed twice. This removes capacity collisions before expert computation instead of repairing them with token dropping or load-balancing penalties.

Useful6/10
Difficulty7/10
Novelty7/10
Paper: A Necessary and Sufficient Hall Condition for Hypergraphs arXiv:2608.10193
Unverified 2026

Measure-Lifted Entropy Amplifier

Replace a deterministic latent state with a probability measure over latent states, represented by particles or weighted prototypes. Apply the learned latent transition to every particle, so one base trajectory map induces a dynamics on distributions; use an entropy-preservation or entropy-growth regularizer to prevent collapse of the ensemble. The mechanism predicts that any positive base-state trajectory entropy can generate unbounded distinguishability in the ideal measure space through…

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Entropies of compact subsets and supported measures arXiv:2608.09702
Unverified 2026

Erdelyi-Kober Log-Scale Mixer

Replace generic cross-scale mixing with a fixed-shape or lightly parameterized Erdelyi-Kober fractional convolution over logarithmic scale. The fractional order controls how strongly nearby scales are emphasized, while the exponential tail parameter controls the receptive field over distant scales, providing an interpretable alternative to dense cross-scale attention.

Useful6/10
Difficulty4/10
Novelty8/10
Paper: Boundedness of Erdélyi--Kober Integrals and Mellin Fractional Integrals on Weighted Lebesgue Spaces arXiv:2608.09401
Unverified 2026

Scalene Nilpotent-Symmetry Network

Augment a sequence network with a learned staggered matrix-product-operator symmetry and penalize its commutator with the network map. Unlike ordinary equivariance, the auxiliary operator need not define a self-commuting transfer-matrix family: it can be discovered through cross-commutation with a second alternating operator, while nilpotency supplies a finite hierarchy of symmetry constraints. The model should preserve generalized symmetry sectors and exhibit lower commutator error on…

Useful6/10
Difficulty7/10
Novelty8/10
Paper: Scalene Yang--Baxter triples as a source of hidden symmetries beyond the ordinary Yang--Baxter equation arXiv:2608.09081
Unverified 2026

Single-Node Observable Leaky-RNN

Construct a sparse recurrent network with positive edge weights and Leaky-ReLU updates so that one selected hidden node, observed over a finite time window, contains enough information to reconstruct the full hidden state. Add an auxiliary decoder from the observed trajectory to the initial state or current state, and use graph rewiring or edge-growth until every hidden node has a directed path to the sensor within the observation horizon.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: On the Observability and Controllability of Leaky-ReLU Networks arXiv:2608.09059
Unverified 2026

Vector-Balanced MoE Routing

Replace count-only MoE load balancing with greedy balancing of aggregate token-feature vectors. A token is assigned to the expert for which adding its feature vector produces the smallest increase in that expert's squared aggregate norm, encouraging experts to receive complementary semantic mixtures rather than identical token counts.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Max-$k$-Cut via Node Features arXiv:2608.08499
Unverified 2026

Minimax-balanced progressive MoE splitting

Grow a mixture-of-experts layer by splitting one expert into two children while conserving its routing mass, and choose the split ratio to minimize the worst imbalance over all intermediate expert counts. Use the paper's sharp threshold as a hard design target: with n experts, some intermediate stage must have capacity ratio at least D_n = 2^{1-1/\lceil n/2\rceil}, so schedules substantially below this are impossible rather than merely difficult to discover. Initialize child router logits with…

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Optimal Finite Interval Discrepancy via Binary Refinement arXiv:2608.08431
Unverified 2026

Sobolev-Orthogonal MLP Features

Replace raw polynomial or Fourier-like features in a small MLP with basis functions orthonormal under a Sobolev inner product that jointly measures feature magnitude and input derivative magnitude. This explicitly controls feature smoothness while preserving decorrelation, potentially improving conditioning and reducing the need for large derivative-regularization coefficients.

Useful6/10
Difficulty4/10
Novelty7/10
Paper: A Riemann-Hilbert representation for Sobolev orthogonal polynomials arXiv:2608.08397
Unverified 2026

Granularity-Aware Feasible Routing

Replace a continuous allocation or routing decision with a lattice-valued decision whose unit size is explicitly normalized by total capacity. Round allocations downward rather than to the nearest lattice point, preserving per-example capacity feasibility, and train or evaluate against the resulting granularity ratio rather than treating discretization as an implementation detail.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Bid Lattices and the Value of Flexibility:A Granularity Ratio for Capacity Markets arXiv:2608.08371
Unverified 2026

Nielsen Quaternionic Hyperbolic Latent Layer

Represent each recurrent latent state as a pair of unit quaternions \((q_1,q_2)\in\mathrm{SU}(2)^2\), and evolve it with a composition of elementary Nielsen maps corresponding to a chosen hyperbolic matrix \(A\in\mathrm{SL}(2,\mathbb{Z})\). The layer exactly preserves the group manifold and Haar volume, preserves the commuting locus \(q_1q_2=q_2q_1\), and reproduces toral hyperbolic dynamics there, giving a structured long-horizon prior instead of an unconstrained matrix recurrence.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Quaternionic Extensions of Hyperbolic Toral Automorphisms arXiv:2608.08252
Unverified 2026

Hurwitz–Radon signed bilinear mixer

Replace a learned dense bilinear map with a structured family of signed orthogonal matrices. Given feature vectors y,z in R^n, produce r interaction features h_a = y^T H_a z / sqrt(n), where the H_a form a Hadamard/Clifford-like family; the resulting bilinear map has operator norm at most one when r is within the Hurwitz–Radon limit. Learn only channel projections and optional scalar gates around this fixed mixer.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Hilbertian Kahane--Salem--Zygmund Inequalities: Extremizers and Quantitative Gaps arXiv:2608.08246
Unverified 2026

Excursion-Adaptive Temporal Tokenization

Replace a uniformly sampled trajectory sequence by a binary temporal partition whose intervals are split only when the observed trajectory makes an excursion larger than a threshold. Encode one summary token per retained leaf, optionally including duration and endpoint displacement, so smooth trajectory regions receive fewer tokens while rapidly changing regions retain resolution.

Useful6/10
Difficulty4/10
Novelty6/10
Paper: Sharp Wasserstein Convergence Rates for Empirical Path Laws of Itô Processes arXiv:2608.07879
Unverified 2026

Parallel Phase Oscillator SSM

Replace real diagonal state-space channels with complex damped oscillators whose hidden states encode both amplitude and phase. Train with parallel causal convolution and deploy with the equivalent one-step recurrence, allowing the same layer to support efficient batched training and low-memory streaming inference.

Useful6/10
Difficulty5/10
Novelty4/10
Paper: Phase State Space Models: Parallel, Surrogate-Free Training of Spiking Networks arXiv:2608.07754
Unverified 2026

Dynamic Hyperedge Token Mixer

Replace dense token-to-token attention in selected layers with communication through a small number of multi-token hyperedges. Each hyperedge aggregates its incident token states and broadcasts the resulting message back to those tokens, allowing higher-order interactions while reducing the number of pairwise links. Reconstruct hyperedges periodically from cumulative token displacement so stable tokens retain useful groups while rapidly changing tokens are regrouped.

Useful6/10
Difficulty6/10
Novelty5/10
Paper: HPSO: Particle Swarm Optimization with Hypergraph-Based Topology arXiv:2608.07587