Architecture ideas

Attention variants, state-space and recurrent cells, normalization and token-mixing schemes — each tested against the standard block it replaces.

Unverified 2026

Unitary-dilated stochastic router

Replace a softmax transition or mixture-of-experts router by probabilities obtained from squared amplitudes of an isometric latent transition. Each input state is mapped to an orthogonal latent subspace, and summing probability over the latent index produces the desired expert or next-state distribution. The latent amplitudes can retain information that would be destroyed by directly averaging expert outputs, while normalization is guaranteed by construction.

Useful5/10
Difficulty6/10
Novelty5/10
Paper: The Born Representation Theorem and the Unistochastic Theorem arXiv:2608.04354
Unverified 2026

Trace-Free Hodge Feature Mixer

Build a parameter-free spectral channel mixer whose channels are arranged as components of an l-form and whose multiplier is the trace-free Beurling--Ahlfors transform. At every nonzero spatial frequency it mixes the exact and coexact channel subspaces with opposite signs, preventing a uniform channel-direction bias and preserving a structured cancellation property. Insert it as a residual branch before a convolution, MLP, or attention block, with one learned scalar gate controlling its…

Useful5/10
Difficulty6/10
Novelty8/10
Paper: The trace-free Beurling--Ahlfors transform and the Bourgain--Brezis problem for Hodge systems arXiv:2608.04237
Unverified 2026

Repair-cost detector for incompatible similarity predictions

Use the paper's lower bound on nearest-correlation repair cost to detect when a neural network's pairwise similarity predictions contain too much globally incompatible off-diagonal energy. Instead of projecting every predicted matrix onto the correlation cone, train the network to reduce the repair-risk statistic or trigger expensive repair only when a cheap diagnostic predicts substantial distortion.

Useful5/10
Difficulty4/10
Novelty6/10
Paper: Correlation Matrices in High Dimensions: The Elliptope as a Sample-Correlation Ensemble arXiv:2608.04162
Unverified 2026

Entropy-Minimum POVM Router

Replace a scalar softmax classifier or MoE router with a positive-operator-valued measurement computed from learned class or expert density matrices. The resulting operators are positive semidefinite and sum exactly to the identity, so routing probabilities remain normalized for every input state while retaining matrix-valued uncertainty and correlations between latent directions.

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Unifying quantum measurement constructions via a relative-entropy minimum change principle arXiv:2608.04055
Unverified 2026

Type-B fermionic equivariant layer

Represent each of n signed tokens with two or more anticommuting feature channels and build equivariant outputs from exterior products rather than unconstrained tensor products. Penalize or project out positive-degree signed-permutation invariants, approximating the coinvariant quotient so that the layer retains order-sensitive orientation information without learning redundant invariant directions.

Useful5/10
Difficulty6/10
Novelty6/10
Paper: Type $B$ fermionic coinvariant rings arXiv:2608.02881
Unverified 2026

Topology-Guided Capacity Allocation

Use the layer at which persistent connected components and holes disappear to allocate capacity nonuniformly across a network. If representations simplify much earlier than desired, widen the responsible layers or insert an additional block; if simplification is excessively delayed, avoid spending parameters there. This turns persistent-homology COM into an actionable architecture-search signal rather than a post-hoc visualization.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Topological Simplification in Predictive Coding Networks arXiv:2608.02816
Unverified 2026

Cocycle-Twisted Attention

Attach each token or graph node a discrete grade a in a finite group A, and modify attention value composition with a normalized group 2-cocycle rather than independent pairwise gates. The cocycle provides a globally consistent projective interaction rule, so composing three messages gives the same result under either parenthesization. This may improve relational reasoning while reducing the number of freely learned interaction parameters.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Zesting and the relative complexity of Reshetikhin-Turaev invariants arXiv:2608.02795
Unverified 2026

Injective Boundary-Aware Disk Pooling

Replace fixed-radius image blur or pooling with disk averages whose radius is proportional to the distance from each pixel to the image boundary. Compute the transform at every spatial location and train a lightweight decoder to reconstruct the pre-transform feature map, using reconstruction error as an anti-collapse regularizer. This creates a scale-adaptive smoothing layer with an injectivity motivation in the continuum while providing larger context in the image interior.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Variable-Radius Disk Transforms and an Area-Integral Problem of Zalcman arXiv:2608.02546
Unverified 2026

Tempered Log-Memory State Mixer

Replace or augment an exponential state-space memory branch with a causal convolution whose lag-j weight is exp(-lambda j) ell(j)/j. The 1/j boundary provides broad logarithmic memory, while lambda supplies an explicit finite memory scale and prevents uncontrolled accumulation from an untempered long-memory kernel.

Useful5/10
Difficulty5/10
Novelty5/10
Paper: Limit Theorems for Tempered Linear Processes with Innovations in the Domain of Attraction of a Stable Law arXiv:2608.01674
Unverified 2026

Holonomy-preserving quad-mesh augmentation

Use the paper's explicit square insertion surgery to generate new quad-mesh examples with altered local valence patterns but unchanged genus, unchanged non-target vertices, and unchanged rotational-holonomy subgroup. Train a mesh GNN with consistency loss or label-preserving augmentation across the original and surgically modified meshes, forcing predictions to depend on global structure rather than accidental local tessellation.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Which Holonomy Signatures Are Realizable? A Complete Answer for Closed Surfaces arXiv:2608.01444
Unverified 2026

Noise-aware Walsh-Hadamard bottleneck

Insert a Walsh-Hadamard transform before a quantized categorical or activation bottleneck and assign coordinate-dependent quantization precision using the attenuation spectrum of a quaternary symmetric noise model. Coordinates corresponding to tensor-product frequencies with many nonzero indices are attenuated by higher powers of \(\delta\), so their quantization can be made coarser with little effect on the reconstructed post-noise representation. This creates a structured, fast transform…

Useful5/10
Difficulty4/10
Novelty6/10
Paper: Frequency Coding over Noisy Sampling arXiv:2608.00539
Unverified 2026

Contractive 9-Way Hierarchical Positional Encoding

Replace a large flat positional-embedding table with a recursively decoded nine-way address whose child transformations contract coordinates by exactly 1/3. Encode an input position using features attached to the address prefix at several depths, guaranteeing that increasing depth produces a geometrically localized representation and that an infinite valid address cannot ambiguously represent two distinct points. This is especially suitable for 2D vision tokens, maps, point clouds, or…

Useful5/10
Difficulty4/10
Novelty4/10
Paper: Hex9: A Quasi-Authalic, Quasi-Continuous Hexagonal DGGS on the Reference Ellipsoid arXiv:2608.00022
Unverified 2026

Entropy-Certified Interaction Supports

Replace a dense third-order channel-interaction tensor by a fixed sparse support selected through the paper's uniform-marginal infeasibility certificate. Supports with a large dual margin have an effective entropy base below the channel alphabet size, suggesting fewer independent interaction slices and cheaper contractions. Use the certificate either during architecture search or as a pruning score for an already-trained tensorized layer.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Slice and Partition Rank Criteria for Polynomial Zero-Avoidance arXiv:2607.29490
Unverified 2026

Collision-Aware Graph Edge Router

Use the model's non-monotonicity result to make graph connectivity a learned resource rather than assuming that every extra edge helps. An edge router assigns transmission scores but also charges a source-side collision cost for exposing an infected node to many susceptible neighbors. The resulting router can prune edges that increase competition and reduce useful reachability.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: The Zombie Infection Model arXiv:2607.29409
Unverified 2026

G2 structured latent code

Use the signed-base expansion as a compact discrete-continuous latent parameterization for a VAE or autoencoder. A short binary sequence produces exponentially refined coordinates, while a learned Markov prior captures correlations between successive latent bits. The decoder receives the resulting bounded real coordinates instead of an unconstrained Gaussian latent vector.

Useful5/10
Difficulty5/10
Novelty6/10
Paper: Fractal random variables defined by probability distributions of digits of their $G_2$-representation having two bases with different signs arXiv:2607.29327
Unverified 2026

Distributional spectral-preconditioned features

Replace or augment a singular scalar activation \(\sigma\) with a distributionally regularized activation \(g\) whose Fourier transform is multiplied by \((i\rho)^\alpha\). This suppresses the problematic low-frequency singular component and can produce better-conditioned random-feature or first-layer representations, while a residual raw-activation branch prevents loss of standard approximation behavior.

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Radon Measure Representations for Infinite-Width Neural Networks with Singular Activations arXiv:2607.29258
Unverified 2026

Flow-Constrained Hierarchical Policy

Parameterize a tree-structured policy through realization weights satisfying sequence-form flow conservation, instead of independently predicting probabilities at every node. Conditional action probabilities are recovered by dividing a child sequence weight by its parent weight, guaranteeing globally consistent probabilities and avoiding invalid or contradictory branch masses. This is suitable for hierarchical RL policies, adaptive computation trees, and neural routers with sequential gating…

Useful5/10
Difficulty4/10
Novelty6/10
Paper: Baseball, An Extensive-Form Game-Theoretic Duel arXiv:2607.29041
Unverified 2026

Determinantal Exclusion Router

Replace independent softmax expert choices with a collision-free Markov router whose particles occupy expert positions on a one-dimensional or circular index lattice. A particle can move only to an empty neighboring expert, and the move rate contains a product of sine ratios that globally repels nearby assignments; this should reduce expert collapse and produce more evenly spread routing without requiring a separate pairwise diversity loss.

Useful5/10
Difficulty7/10
Novelty8/10
Paper: Exact Results for the Symmetric Dyson Exclusion Process arXiv:2607.28807
Unverified 2026

Covariance-complement uncertainty head

Represent uncertainty of a graph-structured neural feature field through dual covariance rather than explicitly storing a dense primal covariance. Recover calibrated primal marginal variances from dual statistics using the paper's covariance-complement identity.

Useful5/10
Difficulty4/10
Novelty8/10
Paper: Accelerated Random-Sweep Gibbs Sampling for Gaussian Graphical Models via Dual Normal Factor Graphs arXiv:2607.28706
Unverified 2026

Twisted-Shift Feature Mixer

Build a neural feature-mixing layer from a truncated shift S and a diagonal phase operator T satisfying TS=qST, with |q|=1. The relation forces moving one position in the graded feature basis to multiply the phase operator by q, providing a compact inductive bias for periodic, phase-sensitive, or cyclic data.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: On the diversity of twisted commuting operators arXiv:2607.28372
Unverified 2026

Binomial-thinning (s,S) capacity controller

Replace continuously fluctuating conditional-computation decisions with a fixed-charge (s,S) controller for the number of active experts or channel groups. If the currently provisioned capacity falls below s, activate capacity up to S; otherwise retain the current capacity, preventing repeated small routing or kernel-launch decisions. Binomial thinning models the random subset of provisioned experts or channels that are actually available after token load, dropout, failures, or admission limits.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: The optimality of an (s, S) hiring policy on a workforce planning problem with fixed recruitment costs and binomial turnover arXiv:2607.28171
Unverified 2026

Coset-aware MoE routing repair

Add an integer-lattice feasibility layer after ordinary top-1 or top-2 MoE routing. The router first produces its usual expert assignments, then minimally changes a small number of low-confidence assignments so the batch count vector lies in a prescribed lattice or desired coset, eliminating persistent modular load imbalance that ordinary auxiliary losses may not detect.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: On the number of factorable induced subgraphs arXiv:2607.27870
Unverified 2026

q-Boson Spectral Initialization

Initialize a neural layer with singular values taken from the finite spectral measure of the paper's q-boson Jacobi operator instead of using Xavier or ordinary orthogonal initialization. The resulting layer has a deliberately shaped singular-value distribution and an explicit finite-size spectral edge, allowing initialization to target stable signal propagation while retaining spectral diversity.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: Sharp Bounds on Ground State Energy of the SYK Model arXiv:2607.27185
Unverified 2026

Dispersive Analytic Smoothing Block

Insert a short gKdV-inspired spectral flow between neural blocks to regularize rough feature maps without using an isotropic low-pass filter. The module applies a Fourier dispersive phase and derivative-coupled polynomial residual updates, with an optional finite factorial dilation penalty to encourage analytic-looking features.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Instantaneous analytic smoothing of rough data for the modified and cubic gKdV equations arXiv:2607.27115