Coordinate Path-Integral Joint Gibbs Policy / report_bench_2026-08-31T125803.md
Mechanism confirmed, baseline not beaten
Стенд-проверка (stage-2) · промт оператора:
(универсальный)
Ответ агента:
{ "worked": false, "confidence": 10, "verdict": "The Coordinate Path-Integral Joint Gibbs intervention was evaluated on the registered constrained_two_agent_dynamics track. It significantly worsened standard test MSE: baseline mean 0.0062669871 versus idea mean 0.0247058901, paired delta +0.0184389030, permutation p=0.0081. The trained-model path-order signature measured near-zero disagreement, but the mechanism did not improve forecasting performance.", "metrics": { "baseline": "Best configuration lr=0.006, epochs=12, hidden=32, wd=0.0; full mean MSE 0.006266987067647278, std 0.0012038076406143412, per-seed [0.0054755136370658875, 0.008236748166382313, 0.007714319974184036, 0.006880192551761866, 0.005664798431098461, 0.006498826202005148, 0.004625890403985977, 0.005039607174694538].", "idea": "Best configuration lr=0.006, epochs=12, hidden=32, wd=0.0, lam=0.2, K=7; mean MSE 0.024705890100449324, std 0.008283505340451159, per-seed [0.01789645291864872, 0.03382120281457901, 0.017041007056832314, 0.02942308597266674, 0.03449169918894768, 0.03302520141005516, 0.012732679024338722, 0.019215792417526245]." }, "bench_report": { "bench_version": 1, "track": "constrained_two_agent_dynamics", "model": "custom_gru_cell", "metric_direction": "lower is better", "n_seeds": 8, "baseline": { "best_cfg": { "lr": 0.006, "epochs": 12, "hidden": 32, "wd": 0.0 }, "sweep": [ { "cfg": { "lr": 0.001, "epochs": 12, "hidden": 32, "wd": 0.0 }, "mean": 0.0322869922965765 }, { "cfg": { "lr": 0.003, "epochs": 12, "hidden": 32, "wd": 0.0 }, "mean": 0.012963124550879002 }, { "cfg": { "lr": 0.006, "epochs": 12, "hidden": 32, "wd": 0.0 }, "mean": 0.0070766935823485255 } ], "full": { "mean": 0.006266987067647278, "std": 0.0012038076406143412, "per_seed": [ 0.0054755136370658875, 0.008236748166382313, 0.007714319974184036, 0.006880192551761866, 0.005664798431098461, 0.006498826202005148, 0.004625890403985977, 0.005039607174694538 ], "n": 8 } }, "idea": { "mean": 0.024705890100449324, "std": 0.008283505340451159, "per_seed": [ 0.01789645291864872, 0.03382120281457901, 0.017041007056832314, 0.02942308597266674, 0.03449169918894768, 0.03302520141005516, 0.012732679024338722, 0.019215792417526245 ], "n": 8 }, "comparison": { "delta_mean": 0.018438903032802045, "idea_wins": 0, "n_pairs": 8, "per_seed_diffs": [ 0.012420939281582832, 0.025584454648196697, 0.009326687082648277, 0.022542893420904875, 0.028826900757849216, 0.026526375208050013, 0.008106788620352745, 0.014176185242831707 ], "p_value": 0.0081, "mde": 0.0069627855763603015, "mde_rel_pct": 111.10260004053649, "verdict": "idea worse (significant)", "system_worked": false }, "mechanism_signature": { "custom_track": { "name": "constrained_two_agent_dynamics", "file": "bench/custom_tracks/constrained_two_agent_dynamics.py", "domain": "dynamics" }, "mechanism_signature": { "predicted_rms": 2.836034980191471e-07, "observed_rms": 2.836034980191471e-07, "confirmed": true }, "idea_trials": [ { "cfg": { "lr": 0.006, "epochs": 12, "hidden": 32, "wd": 0.0, "lam": 0.2, "K": 7 }, "mean": 0.024705890100449324 } ] } }, "how_to_run": "/home/maxwelhelp/main/bin/python3 stage2_registered.py", "files": [ "stage2_registered.py", "bench_report.json" ], "limitations": "The registered track is a constrained two-agent dynamics forecasting task rather than a full multi-agent RL replay environment, so exploration return, collision rate, and runtime scaling were not tested. The path-integral intervention was implemented as a differentiable consistency loss over recurrent-model gradients; the measured signature confirms near-zero path disagreement but does not establish a performance benefit.", "system_verdict": "failed", "practical_verdict": "harms", "mechanism_ok": 0, "system_judged": true }