Architecture ideas

Research ideas extracted from mathematics papers, categorized as Architecture.

Unverified 2026

Fractional-to-Local RG Residual Block

Replace a single local message-passing or convolution operator by a spectrally controlled mixture of fractional and ordinary diffusion. The exponent σ is learned or scheduled, while a crossover gate forces the model to change parameterization near the renormalization-group threshold σ*=2, allowing long-range propagation when useful without retaining an unnecessarily nonlocal operator at short scales.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: The $6-ε$ Expansion for Long-Range Lee--Yang and Percolation Criticality arXiv:2608.15120
Unverified 2026

Laguerre Memory Convolution

Replace the length-L learned convolution kernel in a causal sequence layer with K Laguerre basis functions, where K is much smaller than L and the basis parameter controls the decay time scale. The layer retains a long receptive field but learns only K coefficients, while FFT or a fixed state-space realization evaluates the resulting convolution efficiently. This is especially appropriate for audio, sensor streams, and long-context regression where the desired impulse response is smooth or…

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Impulse Response Estimation via Laguerre-Fourier Expansion arXiv:2608.14769
Unverified 2026

Scalar Residual Thermodynamic Head

Attach a small temperature-pressure residual head to a pretrained structural encoder instead of relearning the full free-energy surface. Predict one scalar Gibbs free energy and obtain entropy, volume, and other thermodynamic responses by automatic differentiation, enforcing that all outputs derive from a common potential.

Useful6/10
Difficulty4/10
Novelty6/10
Paper: Universal Thermodynamic Interatomic Potentials for Crystalline Materials arXiv:2608.14502
Unverified 2026

Åberg Percolation Routing

Replace a dense neural interaction graph by a dynamically activated graph whose edge $(u,v)$ is retained only when its effective coupling exceeds the local spacing of response modes. The network remains sparse below the connectivity transition but becomes globally communicating once a giant component forms, providing a controllable alternative to arbitrary magnitude pruning.

Useful6/10
Difficulty6/10
Novelty8/10
Paper: Small-world structure of quantum computer hardware arXiv:2608.13855
Unverified 2026

Warm-Started Perron Positional Encoding

Add a distributed spectral positional encoding to a graph neural network, graph transformer, sparse-attention model, or MoE router by computing the dominant eigenvector of the current weighted adjacency matrix with a few warm-started power iterations. Unlike a Fiedler-vector feature, this encoding uses only local neighbor aggregation, is naturally nonnegative for nonnegative adjacency weights, and can be updated incrementally when the graph or edge weights change.

Useful6/10
Difficulty4/10
Novelty4/10
Paper: Adjacency-Based Spectral Proxy Control of Mobile Communication Agents arXiv:2608.13616
Unverified 2026

Differentiable Octahedron Message-Passing Layer

Replace a generic learned update on a triangular feature lattice by a max-plus octahedron recurrence, optionally softened with log-sum-exp. The layer propagates information between two time slices while preserving the paper's characteristic tropical local consistency, which may provide a parameter-efficient inductive bias for grid reasoning, image patches, or graph layouts.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Skew Hives, Skew Skeps, Skew Schur Log-Concavity arXiv:2608.13544
Unverified 2026

Synchronized spherical dimension bridge

Insert a hyperspherical adapter that splits an embedding into several unit-sphere blocks, changes the dimension of each block, and recombines them with a synchronized spherical join. Train the adapter to preserve pairwise angular distances, while using the paper's max-distortion composition principle to avoid uncontrolled accumulation of blockwise errors.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: The Gromov-Hausdorff Distance Between Consecutive Spheres arXiv:2608.13264
Unverified 2026

Quadratic-budget Toeplitz long-range layer

Replace a dense translation-invariant interaction matrix with a positive-definite Toeplitz kernel K_n(e^f) whose log-spectrum is parameterized by a small number of Fourier coefficients with 1/|k| decay. Use the paper's explicit quadratic term as a spectral-volume budget, allowing long-range structure while discouraging uncontrolled determinant growth and ill-conditioning. Subtracting this term from a log-determinant regularizer leaves a residual intended to capture higher-order deviations from…

Useful6/10
Difficulty6/10
Novelty7/10
Paper: On Toeplitz determinants with slow Fourier decay arXiv:2608.13182
Unverified 2026

Adaptive spectral-gap toroidal encoding

Replace ordinary absolute positional embeddings with coordinates on a learned flat torus and use dual-lattice Fourier characters as positional features. Control the covariance of the coordinate fundamental domain so that the paper's inequality guarantees a lower bound on the smallest nonzero positional frequency, preventing the learned periodic coordinate system from developing arbitrarily weak or nearly constant modes.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Spectral and Isoperimetric Bounds on Flat Tori arXiv:2608.13052
Unverified 2026

Resonance-Gated Triadic Fourier Layer

Replace unconstrained spectral mixing with a three-component triadic interaction whose strength is determined by the quadratic phase mismatch R(xi,xi_1). Near-resonant products receive high weight because their phases remain coherent, while strongly nonresonant products are attenuated. The resonance bandwidth can be fixed from the frequency grid or learned as a positive parameter.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: On solitary wave solutions with two-frequency parameters to the three-component system of quadratic nonlinear Schrödinger equations arXiv:2608.12983
Unverified 2026

Intermittent Multi-Mode Memory Gate

Add a bounded routing state to an RNN, state-space model, or mixture-of-experts layer, with several neutral fixed points representing persistent modes. The state moves between modes when far from a fixed point but escapes each mode only polynomially when close to it, creating controllable long memory without setting a linear eigenvalue arbitrarily close to one. A temperature parameter selects between an entropy-rich phase using many modes and a low-entropy phase concentrated near one preferred…

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Thermodynamic formalism for intermittent maps with multiple neutral fixed points and phase transitions arXiv:2608.12784
Unverified 2026

Cactus-Graph Phase Budgeting

Build neural computation graphs with explicitly phase-budgeted serial and parallel branches, treating serial compositions as SRG products and parallel residual branches as SRG sums. Allocate phase centers theta_i so that every loop or branch aggregate stays away from -1, enabling stability-aware architecture search and constructive control of branch gains.

Useful6/10
Difficulty7/10
Novelty8/10
Paper: The $θ$-Symmetric SRG with Applications to Stability of Cactus Dynamic Networks arXiv:2608.12591
Unverified 2026

Bloch-Husimi attention

Replace unconstrained attention score vectors by normalized SU(2) coherent-state responses of a positive operator on an (N+1)-dimensional spin space. Each query produces a smooth bounded response over a fixed spherical grid, while values are aggregated normally. The coherent-state kernel imposes geometric structure and exposes a controllable concentration parameter N.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Isospectral majorization and isoperimetric inequalities for coherent states on the Bloch sphere arXiv:2608.12248
Unverified 2026

Positive Reflected Bellman Layer

Replace an unconstrained spatial aggregation in a neural PDE surrogate or controlled-dynamics model with a fixed-branch expectation layer. Each output is a maximum over controls of a nonnegative weighted average of next-state values, with reflected overshoots attenuated by Robin factors. Increasing any input value therefore cannot decrease the output, giving a hard monotonicity and positivity property instead of relying on a penalty.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: A Positivity-Preserving Expectation Scheme for Hamilton--Jacobi--Bellman Equations with Oblique Robin Boundary Conditions arXiv:2608.11936
Unverified 2026

Ambiguity-Flat Weyl Feature Layer

Replace a random cyclic filter bank or patch projection with the Weyl–Heisenberg orbit of one normalized learnable prototype. Regularize the prototype so that all nonzero shift and modulation correlations have a large and nearly equal magnitude, maximizing the smallest eigenvalue of the induced feature Gram matrix and preventing poorly observed feature directions.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Uniformly Stable Minimal Weyl--Heisenberg Measurements Approaching the SIC Benchmark arXiv:2608.11850
Unverified 2026

Spherical-design directional heads

Use a fixed spherical t-design as the direction codebook for a directional attention or feature-aggregation module instead of independently sampled random directions. Equal weights provide exact zero mean and isotropic second moments, while exactness for spherical polynomials up to degree t reduces directional aliasing and seed-dependent anisotropy.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Minkowski Polytopes of Spherical Designs: High-Order Isotropy and Quantitative Sphericity arXiv:2608.11570
Unverified 2026

Smith-Reduced Chain Encoder

Replace an ordinary hierarchical graph encoder with a finite chain-complex encoder whose learned boundary maps satisfy \(\partial_{k-1}\partial_k=0\). Compute Smith normal form on the integer incidence matrices and treat unit-labelled cell pairs as refinement overhead: cancel or gate those pairs before message passing, while preserving non-unit labels that encode genuinely nontrivial structure. The resulting representation should be insensitive to arbitrary cell subdivision while retaining…

Useful6/10
Difficulty6/10
Novelty6/10
Paper: A Chain- and Diagram-Level Semantics for Morphological Calculus Refinement, monodromy, and bivector orbit decompositions arXiv:2608.11325
Unverified 2026

Deep-spectrum community features

Replace the usual top-eigenvector positional encoding in a graph neural network with a density-selected spectral subspace. The selector explicitly searches below the leading eigenvectors, where community information may survive after latent geometric modes have consumed the largest eigenvalues. The selected coordinates can be concatenated to node features or used as a bias in graph attention.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Spectral graph clustering with inhomogeneous latent geometry arXiv:2608.11321
Unverified 2026

Finite-Difference Conditional Neural Operator

Replace a tensor-product network over a low-dimensional state and a large distribution embedding with a neural operator that consumes the distribution vector once and outputs values on a finite-difference grid in the low-dimensional state. Train it with the governing PDE residual, explicit boundary residuals, and optional signed shape constraints, allowing the network to preserve numerical structure that a generic MLP would learn only implicitly.

Useful6/10
Difficulty5/10
Novelty5/10
Paper: Mastering Stochastic OLG Models in Continuous Time arXiv:2608.11134
Unverified 2026

Equivariant Critical-Mode Branch Seeding

When a symmetry-frequency block becomes critical, initialize or perturb the network specifically along its critical representation rather than injecting isotropic noise into all hidden channels. This creates trainable branches for the symmetry patterns predicted by the bifurcation calculation and can expose useful periodic solutions that ordinary symmetry-preserving training fails to reach.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Local and Global Equivariant Bifurcation for Periodic Weyl and Riesz Fractional Equations arXiv:2608.11101
Unverified 2026

Reversal-Assisted MoE Routing

Replace one-shot top-k expert assignment with a capacity-constrained stochastic routing process in which tokens have a temporary routing direction and can reverse it at rate gamma. Tokens preferentially move through short vacancy clusters, while reversals break persistent directed congestion and should delay or eliminate expert-level jams. This creates a tunable routing phase diagram rather than relying only on an auxiliary load-balancing loss.

Useful6/10
Difficulty6/10
Novelty8/10
Paper: Jamming transition in an active exclusion process arXiv:2608.11041
Unverified 2026

Successive Orthogonal Innovation Blocks

Add a neural feature, adapter, or expert block only through the component of its outputs that is orthogonal to the span of all previously installed blocks. Quotient coefficient directions that produce nearly identical outputs with an SVD or pseudoinverse, so the new block contributes intrinsic representational dimensions instead of duplicating old features. The expected benefit is a smaller effective architecture and better-conditioned block expansion.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Successive Schur-Riesz Analysis for Approximation arXiv:2608.10757
Unverified 2026

Unit-circle-root Toeplitz mixer

Replace a freely learned finite impulse-response mixing kernel with a matrix polynomial whose roots are constrained to the unit circle. The resulting block-Toeplitz operator has an explicitly positive semidefinite spectral construction, while increasing the polynomial degree gives a systematic capacity knob for approximating matrix-valued frequency responses.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Pure matrix states on block Toeplitz matrices arXiv:2608.10701
Unverified 2026

Tree-Hall Collision-Free Router

Replace independent top-k routing by a tree-structured hypergraph assignment layer. Each candidate route is a singleton or pair of resources, and the router selects exactly q_e routes for every tree edge e while ensuring that no resource is consumed twice. This removes capacity collisions before expert computation instead of repairing them with token dropping or load-balancing penalties.

Useful6/10
Difficulty7/10
Novelty7/10
Paper: A Necessary and Sufficient Hall Condition for Hypergraphs arXiv:2608.10193