SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization

arXiv:2608.12443 2026 Optimization 1 ideas extracted · analyzed Sep 1, 2026

What the math gives to ML

The transferable contribution is a group-relative policy-gradient baseline that uses structural diversity rather than treating sampled solutions uniformly. Given multiple solutions from one policy and one problem instance, it constructs parameter-free embeddings from existing encoder node states, then gives greater baseline weight to structurally dissimilar peers. This can reduce redundant samples and prevent the learning signal from collapsing onto only the best trajectory. The most promising transfer is a plug-in variance-reduction rule for sampled sequence policies, including combinatorial solvers and other grouped preference-optimization systems.

Ideas from this paper

Mechanism confirmed, baseline not beaten 2026

Diversity-Weighted Leave-One-Out Policy Baseline

Replace the usual best-sample or uniform group baseline in sampled-policy training with a leave-one-out baseline weighted toward structurally dissimilar solutions. Diverse peers contribute more independent information, while near-duplicate trajectories contribute less redundant signal.

Useful7/10
Difficulty4/10
Novelty6/10
Paper: SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization arXiv:2608.12443