Policy: cold-start grace, sticky finish thresholds, optional EMA score smoothing - #19
tydinjarman wants to merge 4 commits into
Conversation
A worker younger than cold_start_grace_seconds cannot be judged stuck/off-track/drift: hermes boots silently for 5-10s (and buffers NDJSON on Windows), so early assessments run on empty evidence and Jev scores needs_human high, escalating healthy runs. FOREMAN_COLD_START_GRACE_SECONDS env knob; default 0 preserves old behavior. +2 tests.
Jev scores are independent per assessment and oscillate across backends (observed hermes on mraize: ready_to_finish 0.37 mid-work then 0.24 idle on an unchanged repository). Once completion thresholds are met at any iteration, persist finish_thresholds_met in state.json; a later idle assessment with resolved verification FINISHes instead of burning the worker budget on an already-complete job.
Jev scores are independent per assessment and oscillate across backends (observed on mraize hermes: impl 0.75/0.45 alternating while work was stable), so raw-score thresholds whipsaw and the worker budget burns before FINISH. Smoothed estimates (persisted in state.json, alpha knob, default off) dampen the oscillation; idle assessments then FINISH via the sticky thresholds instead of churning workers. +4 tests using the recorded mraize sequence as data.
|
Thanks for sharing the production observations behind these changes. I reviewed this branch against current Since this PR was opened, Foreman’s metric-driven global policy was refactored into pluggable responsibilities with The underlying problems are worth addressing, but they should be reconsidered within the new boundaries:
Could you reopen this work as one or more focused PRs from current |
|
Superseded by three focused PRs, rebuilt from current
No FactoryAssessment changes, no helper scripts. Closing this one. |
Title: Policy: cold-start grace, sticky finish thresholds, optional EMA score smoothing
Three policy improvements from production runs with the Hermes worker backend
(3/3 consecutive FACTORY_FINISHED on our fleet reference box; prior runs
escalated at the worker cap before these changes):
1. Cold-start grace (
FOREMAN_COLD_START_GRACE_SECONDS, default 0 = old behavior)Workers younger than the window cannot be judged
worker_stuckorwork_off_track. Before this, a worker that had produced no streamed outputyet (backends that buffer stdout) was killed in its first seconds; every
supervisor assessment interval after spawn was a false-positive risk.
2. Sticky finish thresholds
finish_thresholds_metis persisted in state.json. Once the finishrequirements are met at any assessment, a later noisy/idle assessment FINISHes
instead of churning new workers. Motivation: per-assessment scores are
independent and oscillate (observed: impl 0.75 at it4 paired with tests 0.16,
then impl 0.45 with tests 0.43 at it5) — a single-instant three-way
conjunction is statistically fragile.
3. EMA score smoothing (
FOREMAN_SCORE_SMOOTHING_ALPHA, default 1.0 = off;FOREMAN_USE_SMOOTHED_SCORES)Thresholds the SMOOTHED estimate instead of the raw score. Default-off keeps
upstream behavior identical. With alpha=0.4 the observed oscillating sequence
converges into a ~0.1 band instead of whipsawing. Tests use a recorded
production sequence as fixture data (+4 tests).
Tests: tests/test_policy_grace.py (+2), tests/test_policy_ema.py (+4);
existing policy tests unchanged and passing.