Skip to content

Completion: sticky finish with evidence fingerprint - #31

Closed
tydinjarman wants to merge 1 commit into
thruwire:mainfrom
tydinjarman:rework/completion-stickiness
Closed

tydinjarman wants to merge 1 commit into
thruwire:mainfrom
tydinjarman:rework/completion-stickiness

Conversation

@tydinjarman

Copy link
Copy Markdown
Contributor

Supersedes part of #19, per your review: the "sticky completion needs careful semantics" change, rebuilt from current main through the responsibility architecture, with no helper scripts.

What it does

  • When core.completion observes the finish thresholds met, it records an evidence fingerprint (_StickyCompletion: completion-relevant scores, worker count, verification state) owned by the CompletionResponsibility instance.
  • A later assessment whose scores dip below the thresholds — the noisy idle reassessments semantic scoring produces — may still FINISH, but only while the fingerprint still holds: no new workers have run since capture, and no recorded score has regressed beyond a 0.20 tolerance (_STICKY_EVIDENCE_TOLERANCE), with needs_verification similarly bounded from rising.
  • If the evidence moves on, the authorization is discarded (self._sticky = None) instead of carried forward — one transient high score can never permanently authorize finishing after the evidence changes.
  • The responsibility instance persists across evaluation cycles within a run (built once per FactoryPolicy); responsibility re-routing rebuilds it, which safely resets the sticky state.

Validation

When core.completion observes the finish thresholds met, it records an
evidence fingerprint (the completion-relevant scores, worker count, and
verification state). A later assessment whose scores dip below the
thresholds -- the noisy idle reassessments semantic scoring produces --
may still finish the run, but only while that fingerprint still holds:
no new workers have run and no recorded score has regressed beyond a
0.20 tolerance. If the evidence moves on, the authorization is discarded
instead of carried forward, so one transient high score can never
permanently authorize finishing after the evidence changes.

State is owned by the CompletionResponsibility instance (which persists
across evaluation cycles within a run); re-routing responsibilities
resets it, failing closed.

@JARosen JARosen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I do not see a reachable runtime path that consumes the sticky state. It is captured in the same responsibility call that proposes FINISH, and the runtime immediately applies that directive and exits. If a higher-priority directive wins, it either terminates the run or starts another worker, which invalidates the fingerprint.

Please add an end-to-end runtime test demonstrating a real sequence where completion is captured, the run remains active without new work invalidating it, and a later assessment uses the sticky authorization. If no such lifecycle exists, I think this feature should be closed rather than retaining currently unreachable state.

@tydinjarman

Copy link
Copy Markdown
Contributor Author

Thank you for pushing on this — I traced the full lifecycle through the runtime and agree: there is no reachable path that consumes the sticky state. Closing per your suggestion rather than retaining unreachable state.

The trace, for the record:

  1. Capture site (CompletionResponsibility.directives()): _capture_sticky() only runs when state.active_workers is empty, latest_intervention is not STOP_WORKER, and finish_ready and verification_resolved all hold — and the same call then proposes FINISH at priority 700.
  2. Selection is forced (FactoryPolicy.evaluate() always picks max (priority, confidence)); application is unconditional (_watch_loop calls _apply after every assessment, and _apply(FINISH) sets status to FINISHED, which exits the loop). No later assessment can ever run after a FINISH proposal.
  3. What can outrank FINISH @700 in the capture iteration (worker-dependent directives can't be proposed with no active workers):
    • core.human-escalation ESCALATE @1000 → run ESCALATED, terminal.
    • Iteration-limit ESCALATE @950 → terminal.
    • quality.documentation START_WORKER @750 → start_worker() appends to state.workers — the only non-terminal survivor.
    • core.verification START_VERIFIER @600 → loses to FINISH.
    • STEER/STOP/CONTINUE @900 can't coincide with a capture (_worker_warning returns None with no active workers, and CompletionResponsibility returns [] whenever workers are active); RETRY_WORKER @800's early-return branch sits before the capture code.
  4. Consumption is then impossible: the sticky branch requires _sticky_holds(), i.e. len(state.workers) == record.worker_count. Worker records are append-only (_run_worker's finally only removes from active_workers), so Documentation's START_WORKER — the sole non-terminal outcome — permanently invalidates the fingerprint. Every other outcome ends the run.
  5. The docstring's motivating scenario ("noisy idle reassessments whose scores dip below thresholds") can't occur either: the first idle assessment meeting thresholds applies FINISH immediately, so there is never a second idle assessment for stickiness to span.

The existing unit tests pass only because they call directives() directly, bypassing the runtime loop — exactly the gap you flagged. No end-to-end test can be written to pass without inventing a new lifecycle, which would be forcing a design just to keep the feature alive. The concept may return if a real pre-commit/approval-stage lifecycle ever exists.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants