Skip to content

Feed structured test evidence into the assessment (test_results is hardcoded []) #17

Description

@anchapin

Context

From an external jevify audit I ran across this repo (plus its PR #14, which I've contributed to): foreman is already Jev-shaped by design — fast typed decider over slow generative workers — so this isn't a greenfield proposal. It's a specific gap found while reading observation.py and responsibilities/builtin.py.

The gap

ObservationBuilder.build() sets test_results=[] unconditionally, so the tests_sufficient check (gating FINISH) and VerificationResponsibility (gating START_VERIFIER) are judged purely from diff text and worker output tails. A worker that never ran tests can clear the bar on confident prose.

Proposal

  1. Deterministic extraction first (code, not Jev): parse the latest worker's stdout/stderr tails with a small registry — pytest summary lines, go test, cargo test, npm test — into {framework, passed, failed, errored, duration_s, raw_tail}. Dumb regex on the framework's summary format; it can't be gamed by prose because it requires the actual summary line.
  2. One Score question added to VerificationResponsibility.checks(): test_evidence_strength over rubric 0–3 (no evidence → parseable passing results covering changed files), sharing the existing single system_one call.
  3. directives() keeps its threshold shape but consumes the score: should_verify also fires when test_evidence_strength ≤ 1 even if needs_verification is borderline; ready_to_finish additionally requires evidence ≥ 2.

Failure mode: parser misses a framework → evidence 0 → spurious verification cycles. Mitigate with an explicit test_framework_unknown flag that falls back to today's Noul path.

Related

The same audit suggested per-changed-file risk Scores once per verification decision, appended to the verifier worker's mission ("review these files in risk order") — happy to split that into a separate issue if this one resonates.

Validation

Before building: an offline replay over 20–30 persisted .foreman/runs/*/events.jsonl, reconstructing the final assessment with and without parsed test evidence (~$0.06 of model calls). If variant (b) doesn't improve accuracy ≥10 points or halve ECE on tests_sufficient/ready_to_finish, park this.

Filed from the jevify investigation; happy to shape it to whatever fits the roadmap — including a PR if the direction looks right.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions