Context
From an external jevify audit I ran across this repo (plus its PR #14, which I've contributed to): foreman is already Jev-shaped by design — fast typed decider over slow generative workers — so this isn't a greenfield proposal. It's a specific gap found while reading observation.py and responsibilities/builtin.py.
The gap
ObservationBuilder.build() sets test_results=[] unconditionally, so the tests_sufficient check (gating FINISH) and VerificationResponsibility (gating START_VERIFIER) are judged purely from diff text and worker output tails. A worker that never ran tests can clear the bar on confident prose.
Proposal
- Deterministic extraction first (code, not Jev): parse the latest worker's stdout/stderr tails with a small registry — pytest summary lines,
go test, cargo test, npm test — into {framework, passed, failed, errored, duration_s, raw_tail}. Dumb regex on the framework's summary format; it can't be gamed by prose because it requires the actual summary line.
- One Score question added to
VerificationResponsibility.checks(): test_evidence_strength over rubric 0–3 (no evidence → parseable passing results covering changed files), sharing the existing single system_one call.
directives() keeps its threshold shape but consumes the score: should_verify also fires when test_evidence_strength ≤ 1 even if needs_verification is borderline; ready_to_finish additionally requires evidence ≥ 2.
Failure mode: parser misses a framework → evidence 0 → spurious verification cycles. Mitigate with an explicit test_framework_unknown flag that falls back to today's Noul path.
Related
The same audit suggested per-changed-file risk Scores once per verification decision, appended to the verifier worker's mission ("review these files in risk order") — happy to split that into a separate issue if this one resonates.
Validation
Before building: an offline replay over 20–30 persisted .foreman/runs/*/events.jsonl, reconstructing the final assessment with and without parsed test evidence (~$0.06 of model calls). If variant (b) doesn't improve accuracy ≥10 points or halve ECE on tests_sufficient/ready_to_finish, park this.
Filed from the jevify investigation; happy to shape it to whatever fits the roadmap — including a PR if the direction looks right.
Context
From an external jevify audit I ran across this repo (plus its PR #14, which I've contributed to): foreman is already Jev-shaped by design — fast typed decider over slow generative workers — so this isn't a greenfield proposal. It's a specific gap found while reading
observation.pyandresponsibilities/builtin.py.The gap
ObservationBuilder.build()setstest_results=[]unconditionally, so thetests_sufficientcheck (gating FINISH) andVerificationResponsibility(gating START_VERIFIER) are judged purely from diff text and worker output tails. A worker that never ran tests can clear the bar on confident prose.Proposal
go test,cargo test,npm test— into{framework, passed, failed, errored, duration_s, raw_tail}. Dumb regex on the framework's summary format; it can't be gamed by prose because it requires the actual summary line.VerificationResponsibility.checks():test_evidence_strengthover rubric 0–3 (no evidence → parseable passing results covering changed files), sharing the existing singlesystem_onecall.directives()keeps its threshold shape but consumes the score:should_verifyalso fires whentest_evidence_strength ≤ 1even ifneeds_verificationis borderline;ready_to_finishadditionally requires evidence ≥ 2.Failure mode: parser misses a framework → evidence 0 → spurious verification cycles. Mitigate with an explicit
test_framework_unknownflag that falls back to today's Noul path.Related
The same audit suggested per-changed-file risk Scores once per verification decision, appended to the verifier worker's mission ("review these files in risk order") — happy to split that into a separate issue if this one resonates.
Validation
Before building: an offline replay over 20–30 persisted
.foreman/runs/*/events.jsonl, reconstructing the final assessment with and without parsed test evidence (~$0.06 of model calls). If variant (b) doesn't improve accuracy ≥10 points or halve ECE ontests_sufficient/ready_to_finish, park this.Filed from the jevify investigation; happy to shape it to whatever fits the roadmap — including a PR if the direction looks right.