Skip to content

Policy: cold-start grace, sticky finish thresholds, optional EMA score smoothing - #19

Closed
tydinjarman wants to merge 4 commits into
thruwire:mainfrom
tydinjarman:pr1-policy
Closed

tydinjarman wants to merge 4 commits into
thruwire:mainfrom
tydinjarman:pr1-policy

Conversation

@tydinjarman

Copy link
Copy Markdown
Contributor

Title: Policy: cold-start grace, sticky finish thresholds, optional EMA score smoothing

Three policy improvements from production runs with the Hermes worker backend
(3/3 consecutive FACTORY_FINISHED on our fleet reference box; prior runs
escalated at the worker cap before these changes):

1. Cold-start grace (FOREMAN_COLD_START_GRACE_SECONDS, default 0 = old behavior)

Workers younger than the window cannot be judged worker_stuck or
work_off_track. Before this, a worker that had produced no streamed output
yet (backends that buffer stdout) was killed in its first seconds; every
supervisor assessment interval after spawn was a false-positive risk.

2. Sticky finish thresholds

finish_thresholds_met is persisted in state.json. Once the finish
requirements are met at any assessment, a later noisy/idle assessment FINISHes
instead of churning new workers. Motivation: per-assessment scores are
independent and oscillate (observed: impl 0.75 at it4 paired with tests 0.16,
then impl 0.45 with tests 0.43 at it5) — a single-instant three-way
conjunction is statistically fragile.

3. EMA score smoothing (FOREMAN_SCORE_SMOOTHING_ALPHA, default 1.0 = off;

FOREMAN_USE_SMOOTHED_SCORES)
Thresholds the SMOOTHED estimate instead of the raw score. Default-off keeps
upstream behavior identical. With alpha=0.4 the observed oscillating sequence
converges into a ~0.1 band instead of whipsawing. Tests use a recorded
production sequence as fixture data (+4 tests).

Tests: tests/test_policy_grace.py (+2), tests/test_policy_ema.py (+4);
existing policy tests unchanged and passing.

Tydin Jarman added 4 commits September 22, 2026 08:49
A worker younger than cold_start_grace_seconds cannot be judged
stuck/off-track/drift: hermes boots silently for 5-10s (and buffers NDJSON
on Windows), so early assessments run on empty evidence and Jev scores
needs_human high, escalating healthy runs. FOREMAN_COLD_START_GRACE_SECONDS
env knob; default 0 preserves old behavior. +2 tests.
Jev scores are independent per assessment and oscillate across backends
(observed hermes on mraize: ready_to_finish 0.37 mid-work then 0.24 idle on
an unchanged repository). Once completion thresholds are met at any
iteration, persist finish_thresholds_met in state.json; a later idle
assessment with resolved verification FINISHes instead of burning the
worker budget on an already-complete job.
Jev scores are independent per assessment and oscillate across backends
(observed on mraize hermes: impl 0.75/0.45 alternating while work was
stable), so raw-score thresholds whipsaw and the worker budget burns
before FINISH. Smoothed estimates (persisted in state.json, alpha knob,
default off) dampen the oscillation; idle assessments then FINISH via the
sticky thresholds instead of churning workers. +4 tests using the recorded
mraize sequence as data.
@JARosen

JARosen commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Thanks for sharing the production observations behind these changes. I reviewed this branch against current main. Its isolated test suite passes, but it is not ready to merge in its current form.

Since this PR was opened, Foreman’s metric-driven global policy was refactored into pluggable responsibilities with ForemanResult and responsibility-owned directives. This branch still modifies the older FactoryAssessment policy path and now conflicts with both config.py and policy.py. It also contains several one-off patch/debug scripts that cause the full repository lint check to fail.

The underlying problems are worth addressing, but they should be reconsidered within the new boundaries:

  • cold-start handling likely belongs in worker-health behavior;
  • sticky completion needs careful semantics so one transient high score cannot permanently authorize finishing after the evidence changes;
  • score smoothing should not mutate assessment results globally or implicitly couple unrelated responsibilities.

Could you reopen this work as one or more focused PRs from current main, without the helper scripts, and implement each behavior through the responsibility architecture? Smaller separate PRs would make the safety semantics and production evidence much easier to review.

@tydinjarman

Copy link
Copy Markdown
Contributor Author

Superseded by three focused PRs, rebuilt from current main per your review:

No FactoryAssessment changes, no helper scripts. Closing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants