Weighted Resolvent-Equivariant Attention / report_bench_2026-09-02T232715.md

Mechanism confirmed, baseline not beaten

Raw ⬇ ZIP

Стенд-проверка (stage-2) · промт оператора:

(универсальный)

Ответ агента:

{ "worked": false, "confidence": 10, "verdict": "The required registered benchmark could not be rerun in this environment because the bench package and README are unavailable, so this is not new valid bench evidence. The existing local artifact reports a correct sequence-track comparison with no task-metric win: baseline MSE 0.195805 versus idea MSE 0.195998, paired permutation p=0.7309. Its mechanism signal was directional but small (1.58x commutator reduction), therefore the idea is not demonstrated as a benchmark win.", "metrics": { "baseline": "Existing artifact only: sequence/transformer_tiny, tuned lr=0.001, 8 epochs; 8-seed MSE mean 0.1958046276.", "idea": "Existing artifact only: lambda=0.3, lr=0.001, 8 epochs; 8-seed MSE mean 0.1959976610; paired delta +0.0001930334; p=0.7309.", "environment_check": "The mandatory injected TRACKS command returned ('tabular','vision','sequence','dynamics'), but subsequent import bench failed and the specified bench README/path was absent." }, "bench_report": { "bench_version": 1, "track": "sequence", "model": "transformer_tiny", "metric_direction": "lower is better", "n_seeds": 8, "baseline": { "best_cfg": { "lr": 0.001, "epochs": 8, "lam": 0.0 }, "sweep": [ { "cfg": { "lr": 0.001, "epochs": 8, "lam": 0.0 }, "mean": 0.19084082171320915 }, { "cfg": { "lr": 0.003, "epochs": 8, "lam": 0.0 }, "mean": 0.2167150229215622 }, { "cfg": { "lr": 0.006, "epochs": 8, "lam": 0.0 }, "mean": 0.26577289402484894 } ], "full": { "mean": 0.19580462761223316, "std": 0.019276107610397688, "per_seed": [ 0.21272796392440796, 0.16291894018650055, 0.18857084214687347, 0.19914554059505463, 0.2006099969148636, 0.1974794715642929, 0.22919276356697083, 0.17579150199890137 ], "n": 8 } }, "idea": { "best_cfg": { "lr": 0.001, "epochs": 8, "lam": 0.3 }, "mean": 0.19599766097962856, "std": 0.018529453120078188, "per_seed": [ 0.21339234709739685, 0.1631648980973962, 0.19162587821483612, 0.19664713740348816, 0.1998937577009201, 0.20089472830295563, 0.22621849179267883, 0.1761440485715866 ], "n": 8 }, "comparison": { "delta_mean": 0.000193033367395401, "idea_wins": 3, "n_pairs": 8, "per_seed_diffs": [ 0.0006643831729888916, 0.0002459585666656494, 0.0030550360679626465, -0.0024984031915664673, -0.0007162392139434814, 0.0034152567386627197, -0.002974271774291992, 0.0003525465726852417 ], "p_value": 0.7309, "mde": 0.0019204674874472793, "mde_rel_pct": 0.9808080181079926, "verdict": "no measurable effect", "system_worked": false }, "mechanism_signature": { "prediction": "commutator regularization lowers trained attention commutator; multi-step compatibility follows from resolvent identity", "baseline_comm_rms": { "mean": 0.03365310910157859, "std": 0.005540028408003687, "per_seed": [ 0.028813593089580536, 0.04024209827184677, 0.037993812933564186, 0.02756293211132288 ], "n": 4 }, "idea_comm_rms": { "mean": 0.02128551318310201, "std": 0.0034537152021512133, "per_seed": [ 0.018550011795014143, 0.024316378869116306, 0.02507482608780265, 0.01720083598047495 ], "n": 4 }, "observed_reduction_factor": 1.5810334856429433, "confirmed": true } }, "how_to_run": "/home/maxwelhelp/main/bin/python3 bench_weighted_attention.py", "files": [ "bench_weighted_attention.py", "bench_report.json", "experiment.py", "results.json", "REPORT.md" ], "limitations": "No fresh run was possible: bench import failed outside the mandatory injected command, and /home/maxwelhelp/all/math2nn/bench/README.md plus the package directory were absent. The available artifact is therefore not independently verifiable in this session; additionally, its baseline tuning and mechanism signature used four-seed subsets and its idea sweep varied learning rate but not nearby lambda settings.", "system_verdict": "partial", "practical_verdict": "no_effect", "mechanism_ok": 1, "system_judged": true }