Agent autonomously designed 18 improvement methods across 4 phases, each building on prior results. Iterative hypothesis → experiment → analysis loop.
Improvement Phases (vs HES baseline 49.58%)
Phase 2
6 methods
Phase 3
4 methods
Phase 4
σ-sweep
Phase 5
gating
P2
D
stratified_hes
50.42%
+0.84
P3
K
esdiffw
51.67%
+2.09
P4
M
esdiffw_s40
51.25%
+1.67
P5
O
nhighgate
52.29%
+2.71
P5
R
shiftgate
51.25%
+1.67
▶ P2: 6 hypotheses → stratification best
▶ P3: entropy–score diff metrics → +2.09
▶ P4: σ-sweep, 0.3 optimal (0.2 collapses)
▶ P5: gating variants → nhighgate +2.71
★ Counting outlier tokens >> averaging
★ Fully autonomous loop: read results → hypothesize → code → GPU → verify — 18 methods, 33 jobs, 125h