Skip to content

feat(review): capture fired signals + reversal mapping for the slop and quality gate scores - #8235

Closed
RealDiligent wants to merge 2 commits into
JSONbored:mainfrom
RealDiligent:fix/critical-issue-gate-score-capture-8223
Closed

feat(review): capture fired signals + reversal mapping for the slop and quality gate scores#8235
RealDiligent wants to merge 2 commits into
JSONbored:mainfrom
RealDiligent:fix/critical-issue-gate-score-capture-8223

Conversation

@RealDiligent

Copy link
Copy Markdown
Contributor

Summary

Scope

  • The PR title follows type(scope): short summary Conventional Commit format, for example fix(api): restore profile access checks.
  • This PR is focused and does not mix unrelated backend, UI, MCP, docs, dependency, and deploy changes.
  • This follows CONTRIBUTING.md and does not reintroduce GitHub Pages, VitePress, site/, or CNAME.
  • I linked a currently open issue this PR resolves (e.g. Closes #123) — a linked open issue is required for every contributor PR.

Validation

  • git diff --check
  • npm run typecheck
  • npm run actionlint
  • npm run test:coverage locally; codecov/patch requires ≥99% coverage of the lines AND branches you changed (aim for 100% on your diff so CI variance does not fail near the threshold). Global coverage is a non-blocking trend with a loose 90% backstop, not the gate.
  • npm run test:workers
  • npm run build:mcp
  • npm run test:mcp-pack
  • npm run ui:openapi:check
  • npm run ui:lint
  • npm run ui:typecheck
  • npm run ui:build
  • npm audit --audit-level=moderate
  • New or changed behavior has unit/integration tests for new branches, fallback paths, and sanitizer boundaries

If any required check was skipped, explain why:

  • Ran the writer's full suite (vitest run test/unit/configured-gate-blocker-signals.test.ts — 20 tests green, including the 6 new calibration: capture writers + corpus mappings for the slop and quality gate scores #8223 cases: both firing arms per knob with crossing and non-crossing outcomes plus the default slop threshold; every never-evaluated arm (off/advisory-slop modes, null score/threshold); the normalization bounds; the no-diff-content invariant; and the rejecting-SignalStore fail-open path) plus the full root npm run typecheck. actionlint/workers/mcp/ui checks are untouched surfaces; CI runs them all.

Safety

  • No secrets, wallet details, hotkeys, coldkeys, user PATs, private keys, raw trust scores, private rankings, or private maintainer evidence are exposed.
  • Public GitHub text stays sanitized, low-noise, and does not imply compensation guarantees or optimization tactics.
  • Auth, cookie, CORS, GitHub App, Cloudflare, or session changes include negative-path tests.
  • API/OpenAPI/MCP behavior is updated and tested where needed.
  • UI changes use live API data or real empty/error/loading states, not production mock/demo fallbacks.
  • Visible UI changes include a UI Evidence section below with JPG/JPEG or PNG screenshots arranged as organized, captioned, clickable thumbnails. SVG screenshots are not used as review evidence. Review-only screenshots or recordings are not committed to the repository.
  • Public docs/changelogs are updated where needed; changelogs are only edited for release-prep PRs.

UI Evidence

Not applicable — backend capture wiring only (no UI, docs, or extension surface touched; no registry entries, no knob movement, per the issue's Boundaries).

Notes

  • The corpus becomes visible in the trend automatically once events flow — the rule-row surfaces key off recorded event types, no additional wiring needed (the issue's second deliverable).

@RealDiligent
RealDiligent requested a review from JSONbored as a code owner July 23, 2026 13:36
@superagent-security

Copy link
Copy Markdown
Contributor

Superagent didn't find any vulnerabilities or security issues in this PR.

…nd quality gate scores (JSONbored#8223)

slopGateMinScore and qualityGateMinScore gate real verdicts but recorded no
signal.rule_fired events, so no corpus could ever form for them and the knob
registry could never govern their thresholds.

Mirror JSONbored#8101/JSONbored#8104's capture wiring exactly: recordGateScoreSignals(env,
policy, repo, pr) fires one rule per knob (slop_gate_score,
quality_gate_score) at the env-bearing gate path, with the SAME filter each
pure evaluation applies -- slop evaluates only in block mode with a non-null
risk (buildSlopGateBlocker's own arm, including the default 60 threshold);
quality evaluates whenever its mode is not off with both a score and a
threshold present (buildQualityGateWarning's arm), pass AND fail alike, since
a threshold backtest needs both outcomes. Metadata carries the score
normalized to [0,1] (both are 0-100 integers per normalizeScore, divided by
100 to be confidence-equivalent for buildConfidenceThresholdClassifier
replays -- documented beside the writer) plus the detection's own detail
string as rawSignal, never diff content, per JSONbored#8130's raw-context posture for
computed-score rules. Best-effort writes that can never affect the verdict.

Reversal labeling: both ids join CONFIGURED_GATE_BLOCKER_SIGNAL_CODES via the
new GATE_SCORE_SIGNAL_CODES constant, with the justification the issue
requires recorded beside it -- slop carries direct gate authority in block
mode; quality is advisory-only but registry-governable, and a reversal labels
the overall bot outcome its score contributed to, exactly the corpus label
the drift/loosening evaluators consume.

Tests: both firing arms per knob (crossing and non-crossing outcomes, the
default slop threshold), every never-evaluated arm (off/advisory-slop modes,
null score/threshold), normalization bounds, the no-diff-content invariant,
and the rejecting-SignalStore fail-open path.
@codecov

codecov Bot commented Jul 23, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.39%. Comparing base (0c8d3d6) to head (3d2ed91).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8235      +/-   ##
==========================================
- Coverage   92.10%   90.39%   -1.71%     
==========================================
  Files         776       99     -677     
  Lines       78328    26103   -52225     
  Branches    23668     5127   -18541     
==========================================
- Hits        72147    23597   -48550     
+ Misses       5062     2233    -2829     
+ Partials     1119      273     -846     
Flag Coverage Δ
shard-1 79.08% <66.66%> (+24.16%) ⬆️
shard-2 35.33% <57.14%> (-17.27%) ⬇️
shard-3 40.96% <100.00%> (-13.80%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
src/queue/processors.ts 95.73% <100.00%> (+<0.01%) ⬆️
src/rules/advisory.ts 98.09% <100.00%> (+0.12%) ⬆️

... and 677 files with indirect coverage changes

@loopover-orb loopover-orb Bot added the gittensor:feature Gittensor-scored feature linked to a feature issue — scores a 0.25x multiplier. label Jul 23, 2026
@loopover-orb

loopover-orb Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

Caution

🛑 LoopOver review result - reject/close recommended

Review updated: 2026-07-23 13:56:04 UTC

3 files · 1 AI reviewer · 1 blocker · CI green · clean

🛑 Suggested Action - Reject/Close

Review summary
The AI review returned blocking findings for this change but did not include a separate narrative summary. Review the blockers below before deciding this PR.

Blockers

  • src/rules/advisory.ts: `recordGateScoreSignals` reads `policy.slopGateMode`/`policy.slopGateMinScore` directly without first calling `applyMergeReadinessGate(policy)` (as `recordConfiguredGateBlockerSignals` does two functions above it), so when a repo sets `mergeReadinessGateMode: block` with `slopGateMode` left unset/advisory, the gate genuinely evaluates and can block on slop via the composite override (since `buildSlopGateBlocker` only ever reads `policy.slopGateMode` directly and has no knowledge of the composite itself), but this function's `slopMode === "block"` check sees the un-overridden mode and skips the write — breaking the PR's own stated invariant that this uses 'the same filter each pure evaluation applies' and silently dropping corpus evidence for exactly the case calibration: capture writers + corpus mappings for the slop and quality gate scores #8223 is meant to capture.
Nits — 5 non-blocking
  • The new test file's import reformatting (`import { recordGateScoreSignals,\n RAW_CONTEXT_MAX_DIFF_CHARS, ...`) is stylistically inconsistent with the rest of the file's multi-line import formatting.
  • Consider naming `8223` used only in comments with a shared reference (e.g. linking to the issue in one place) rather than repeating the raw number across several doc comments, per the external size/magic-number note.
  • No test exercises the `mergeReadinessGateMode: block` + unset `slopGateMode` interaction for `recordGateScoreSignals`, which is exactly the scenario the missing `applyMergeReadinessGate` call would have caught.
  • Fix the blocker by computing `const effective = applyMergeReadinessGate(policy);` at the top of `recordGateScoreSignals` and reading `effective.slopGateMode`/`effective.slopGateMinScore` (quality fields can stay off `policy` since readiness is excluded from the composite, per the existing doc comment on `applyMergeReadinessGate`).
  • Add a test mirroring `recordConfiguredGateBlockerSignals`'s composite-aware behavior: `mergeReadinessGateMode: "block"` with `slopGateMode` unset should still fire `slop_gate_score`.

Why this is blocked

  • src/rules/advisory.ts: `recordGateScoreSignals` reads `policy.slopGateMode`/`policy.slopGateMinScore` directly without first calling `applyMergeReadinessGate(policy)` (as `recordConfiguredGateBlockerSignals` does two functions above it), so when a repo sets `mergeReadinessGateMode: block` with `slopGateMode` left unset/advisory, the gate genuinely evaluates and can block on slop via the composite override (since `buildSlopGateBlocker` only ever reads `policy.slopGateMode` directly and has no knowledge of the composite itself), but this function's `slopMode === "block"` check sees the un-overridden mode and skips the write — breaking the PR's own stated invariant that this uses 'the same filter each pure evaluation applies' and silently dropping corpus evidence for exactly the case calibration: capture writers + corpus mappings for the slop and quality gate scores #8223 is meant to capture.
📋 Copy for AI agents — paste into your coding agent
Fix the following blocker(s) from this PR review:

1. src/rules/advisory.ts: \`recordGateScoreSignals\` reads \`policy.slopGateMode\`/\`policy.slopGateMinScore\` directly without first calling \`applyMergeReadinessGate\(policy\)\` \(as \`recordConfiguredGateBlockerSignals\` does two functions above it\), so when a repo sets \`mergeReadinessGateMode: block\` with \`slopGateMode\` left unset/advisory, the gate genuinely evaluates and can block on slop via the composite override \(since \`buildSlopGateBlocker\` only ever reads \`policy.slopGateMode\` directly and has no knowledge of the composite itself\), but this function's \`slopMode === "block"\` check sees the un-overridden mode and skips the write — breaking the PR's own stated invariant that this uses 'the same filter each pure evaluation applies' and silently dropping corpus evidence for exactly the case \#8223 is meant to capture.

Decision drivers

  • ❌ Code review — 1 blocker (1 reviewer)
  • ❌ Gate result — Blocking (Repo-configured hard blocker found.)
Context & advisory signals — never blocks the verdict
Signal Result Evidence
Linked issue ✅ Linked #8223
Related work ✅ No active overlap found No same-issue or scoped active PR overlap found.
Change scope ✅ 20/20 Low review scope from cached public metadata (1 linked issue).
Validation posture ✅ 25/25 PR body includes validation/test evidence.
Contributor workload ✅ 10/10 Author activity: 343 registered-repo PR(s), 136 merged, 36 issue(s).
Contributor context ✅ Confirmed Gittensor contributor RealDiligent; Gittensor profile; 343 PR(s), 36 issue(s).
Improvement ✅ Minor risk: clean · value: minor · LLM: moderate
Linked issue satisfaction

Addressed
The PR adds recordGateScoreSignals mirroring the required pattern (createSignalStore, same filters as the pure evaluations, normalized [0,1] confidence documented beside the writer, detail string as rawSignal not diff, best-effort catch), wires it beside recordConfiguredGateBlockerSignals at the env-bearing gate path, and extends the reversal list via a new GATE_SCORE_SIGNAL_CODES constant with ju

Review context
  • Author: RealDiligent
  • Role context: outside_contributor
  • Public audience mode: oss maintainer
  • Lane context: Repository is configured for direct PR review.
  • Public profile languages: Python, Ruby, JavaScript, Svelte, TypeScript, Markdown, MDX, Rust
  • Official Gittensor activity: 343 PR(s), 36 issue(s).
  • PR-specific overlap: none found.
Contributor next steps
  • Keep the PR focused and include validation evidence before maintainer review.
Signal definitions
  • Related work = same linked issue, overlapping active PRs, or title/path similarity.
  • Change scope = cached public metadata such as size labels, draft state, and review-burden hints.
  • Validation posture = whether the PR provides enough public validation/test evidence for maintainer review.
  • Contributor workload = public contributor activity and cleanup pressure, not a repo-wide quality failure.
  • Contributor context = public GitHub/Gittensor identity context; non-Gittensor status is not a blocker.
🧪 Chat with LoopOver

Ask LoopOver a question about this PR directly in a comment — grounded only in the same cached, public-safe facts shown above, never a new claim.

  • @loopover ask &lt;question&gt; answers contribution-quality Q&A with source citations and freshness.
  • @loopover chat &lt;question&gt; answers in natural prose from cached decision-pack facts via local inference (maintainer/collaborator; read-only).
  • A plain-language @loopover mention with a real question is routed to the closest matching read-only command automatically — no exact syntax required.

Full command reference: https://loopover.ai/docs/loopover-commands

🧪 Experimental — new and may change.

🟩 Safe / merged · 🟦 Advisory · 🟨 Held for review · 🟥 Blocked / closed


💰 Earn for open-source contributions like this. Gittensor lets GitHub contributors earn for the work they already do — register to start earning →.

Checked by LoopOver, a quiet PR intelligence layer for OSS maintainers.

  • Re-run LoopOver review

@loopover-orb

loopover-orb Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

LoopOver is closing this pull request on the maintainer's behalf (AI reviewers agree on a likely critical defect: src/rules/advisory.ts: `recordGateScoreSignals` reads `policy.slopGateMode`/`policy.slopGateMinScore` directly without first calling `applyMergeReadinessGate(policy)` (as `recordConfiguredGateBlockerSignals` does two functions above it), so when a repo sets `mergeReadinessGateMode: block` with `slopGateMode` left unset/advisory, the gate genuinely evaluates and can block on slop via the composite override (since `buildSlopGateBlocker` only ever reads `policy.slopGateMode` directly and has no knowledge of the composite itself), but this function's `slopMode === "block"` check sees the un-overridden mode and skips the write — breaking the PR's own stated invariant that this uses 'the same filter each pure evaluation applies' and silently dropping corpus evidence for exactly the case #8223 is meant to capture.). This is an automated maintenance action — to pursue this change, please open a new pull request with the issues resolved. Closed PRs may be analyzed later to improve review accuracy, but they are not automatically reopened or re-reviewed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gittensor:feature Gittensor-scored feature linked to a feature issue — scores a 0.25x multiplier.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

calibration: capture writers + corpus mappings for the slop and quality gate scores

1 participant