Skip to content

Runtime: optional runtime-owned test execution - #29

Open
tydinjarman wants to merge 2 commits into
thruwire:mainfrom
tydinjarman:rework/runtime-test-verification
Open

tydinjarman wants to merge 2 commits into
thruwire:mainfrom
tydinjarman:rework/runtime-test-verification

Conversation

@tydinjarman

Copy link
Copy Markdown
Contributor

Supersedes part of #20, per your review: this is the focused "optional runtime-owned test execution" change, rebased on current main, with no Hermes or policy changes and no helper scripts.

What it does

  • Opt-in hook in src/foreman/runtime.py: when a coding worker completes, optionally run a verification command with a timeout. Disabled by default (verify_tests_on_complete: bool = False, FOREMAN_VERIFY_TESTS_ON_COMPLETE; test_command_timeout: float = 120.0, FOREMAN_TEST_COMMAND_TIMEOUT).
  • Timeout cleanup fully reaps the subprocess (your explicit demand): POSIX kills the whole process group (child starts with start_new_session=True) via SIGTERM→SIGKILL then always wait()s — no zombies, no orphaned children; Windows uses terminate/kill + wait.
  • Surfaces a structured VerificationOutcome (completed/timed_out/launch_failed, returncode, pid, counts, summary) and emits a structured TEST_RESULT event — evidence only, never a gate. Never raises for command failures.
  • Summary counts parsed with the same failure-first handling ("1 failed, 3 passed"), kept local to this change so it stands alone.

Validation

  • 11 new tests in tests/test_runtime_verification.py: success path, failure-first parse, timeout path asserting returncode is not None and os.kill(pid, 0) → ProcessLookupError (fully reaped on POSIX), launch failure, disabled-by-default no-op, structured event emission. Full suite 180 passed, ruff check clean.
  • This is 3 of 3 focused replacements for Evidence: untracked-file contents, pytest summary parsing, optional runtime test-verification hook #20. Docs section placed at a distinct anchor so the three merge cleanly in any order.

Add an opt-in hook that runs a verification command when a coding worker
completes, with a timeout. Disabled by default via typed config
(`verify_tests_on_complete`, `test_command_timeout`, default 120s).

On timeout the subprocess is fully reaped: POSIX kills the whole process
group (child starts with start_new_session=True) via SIGTERM then
SIGKILL and always wait()s -- no zombies, no orphaned children;
Windows uses terminate/kill + wait. Never raises for command failures;
surfaces a structured VerificationOutcome (completed/timed_out/
launch_failed, returncode, counts, summary) and emits a TEST_RESULT
event as evidence only, never as a gate.

@JARosen JARosen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The verification subprocess currently starts after WORKER_COMPLETED has already entered the supervisor queue. That creates a lifecycle race: Foreman can assess the completion event, select FINISH, and close while verification is still running. Because the worker has already been removed from active_workers, shutdown does not wait for that task; cancellation of run_verification_command() also does not terminate and reap its subprocess.

Please make verification part of the worker terminal lifecycle—finishing before the terminal event is emitted, or represented as an explicitly tracked active task—and guarantee subprocess cleanup on cancellation. An end-to-end test should demonstrate that the factory cannot finish before the TEST_RESULT is recorded.

…ycle

The race: WORKER_COMPLETED entered the supervisor queue before runtime
verification ran, so the supervisor could assess the completion, select
FINISH, and close while verification was still active - with the worker
already removed from active_workers, shutdown never waited, and cancelling
run_verification_command() left its subprocess unterminated.

Fix: _run_worker() now runs _verify_tests_after_worker() while the worker
is still tracked in active_workers, and the TEST_RESULT event is recorded
before the terminal WORKER_COMPLETED event (which now carries a nested
verification summary: status, returncode, passed/failed/errored/skipped,
summary). Since the worker task remains in active_workers throughout,
close()/_terminate_active() waits for in-flight verification instead of
racing past it. run_verification_command() catches asyncio.CancelledError,
terminates and reaps the subprocess (and its POSIX process group) via the
existing _terminate_process_tree(), then re-propagates; _run_worker()
captures verification-time cancellation so lifecycle bookkeeping and the
terminal event still complete before re-raising.

Tests: 182 passed (13/13 in test_runtime_verification.py), including two
new ones: cancellation of run_verification_command() reaps the child (PID
verified gone via /proc), and an end-to-end two-worker test proving
TEST_RESULT precedes WORKER_COMPLETED precedes FINISH, and that shutdown
waits for a still-in-flight verification on another worker (its TEST_RESULT
lands after FACTORY_FINISHED, uncancelled).

@tydinjarman tydinjarman left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed the lifecycle race (pushed as commit 80d4b2e).

You were right about the race: WORKER_COMPLETED entered the supervisor
queue before verification ran, so the supervisor could assess the
completion, select FINISH, and close while verification was still active —
with the worker already removed from active_workers, shutdown never
waited, and cancelling run_verification_command() left its subprocess
unterminated.

Verification is now part of the worker terminal lifecycle: _run_worker()
runs _verify_tests_after_worker() while the worker is still tracked in
active_workers, and the TEST_RESULT event is recorded before the
terminal WORKER_COMPLETED event (which now carries a nested verification
summary: status, returncode, passed/failed/errored/skipped, summary). Since
the worker task remains in active_workers throughout,
close()/_terminate_active() waits for in-flight verification instead of
racing past it. run_verification_command() catches
asyncio.CancelledError, terminates and reaps the subprocess (and its
POSIX process group) via the existing _terminate_process_tree(), then
re-propagates; _run_worker() captures verification-time cancellation so
lifecycle bookkeeping and the terminal event still complete before
re-raising. docs/runtime.md documents the terminal-lifecycle ordering and
cancellation cleanup.

Test evidence: full suite 182 passed (13/13 in
test_runtime_verification.py), including two new tests — cancellation of
run_verification_command() reaps the child (PID verified gone via
/proc), and an end-to-end two-worker test proving TEST_RESULT
precedes WORKER_COMPLETED precedes FINISH, and that shutdown waits for
a still-in-flight verification on another worker (its TEST_RESULT lands
after FACTORY_FINISHED, uncancelled). Ruff clean.

One note: test_integration.py::test_noisy_events_are_coalesced flakes
rarely (~2–5% of runs) in a way unrelated to this change — the diff is
provably inert on that path (verification-gated code never enabled by that
test). Left alone as out of scope.

@JARosen

JARosen commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Thank you for addressing the original lifecycle race and cancellation cleanup so thoroughly. The revised ordering is much safer. Before another implementation round, I think we should pause on two broader issues.

First, the coding backend remains in active_workers after its process has completed while the runtime-owned test subprocess runs. Periodic assessment can therefore select STEER_WORKER or STOP_WORKER against a completed backend, and neither action controls the verification subprocess. That needs an explicit verification-in-progress lifecycle or a policy/runtime representation that cannot target the completed coding process.

Second, now that #28 has merged, this should reuse the shared pytest-summary parser rather than maintain a second parser in runtime.py.

There is also a product-boundary question I should resolve before asking you to revise again: core runtime execution of a fixed python -m pytest command is a hard-coded, Python-specific workflow step, while Foreman otherwise treats independent verification as a policy outcome and aims to remain repository-agnostic. A pluggable evidence collector or responsibility-owned configurable probe may be the better home for this capability.

Please hold off on another rework for the moment. I would rather settle that boundary clearly than send you through incremental patches.

@JARosen

JARosen commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Following up on my request to pause: Foreman now supports centrally configured command evidence providers, selected by checks and executed before the completion assessment. That gives test execution a repository-agnostic home within the existing evidence architecture.

I think this largely supersedes the runtime-specific hook proposed here. Do you see any capability from this PR that remains uncovered? If so, a focused follow-up through the evidence-provider path would be welcome. Otherwise, I suggest closing this as superseded. Thank you for the work on lifecycle ordering and subprocess cleanup.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants