Skip to content

feat: add JEV requirement audit pilot - #12

Open
lemarier wants to merge 3 commits into
mainfrom
david/jev-audit-pilot
Open

lemarier wants to merge 3 commits into
mainfrom
david/jev-audit-pilot

Conversation

@lemarier

@lemarier lemarier commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Requirement references and assertion counts do not establish that a test covers its stated claim. Add a pilot that prepares bounded TypeSafe Choice requests and evaluates responses for coverage gaps, uncertainty, false reassurance, latency, and recorded cost. Stale or malformed responses remain explicit review items; all results are advisory. Valid usage remains measurable even when an answer is malformed.

Include an explicitly invoked live runner with a six-call maximum, no retries or redirects, socket timeouts, bounded responses, private output creation, and partial-result preservation. Credentials come from the environment or a literal assignment in a file outside the repository. Normal checks remain offline; this change adds no audit CI activation.

Include six original development cases and a separate 24-case corpus from four other reviewed PR groups. The expanded corpus has 18 review-derived cases and six synthetic controls, pinned source excerpts, licenses, and screening criteria set before the run. Labels remain provisional and correlated within PRs; group summaries and annotation hashes make that limitation visible.

Validation: just check passes 52 offline tests and git diff --cached --check passes. All 57 expanded source excerpts match pinned public originals. The unchanged question on jev-1.13.0 matched 22/24 expanded labels (17/18 review-derived, 5/6 synthetic): 11/11 gap choices and 9/10 supported choices, with no confident false reassurance. At threshold 0.8, 23/24 cases still require review. Median recorded latency was 369.9515 ms; 48,819 input tokens cost an estimated $0.002050398 at the published $0.042 per million input tokens with free output. These are screening results, not evidence of general accuracy or savings.

The live run exposed hundredth-rounded probabilities totaling 0.99. Preserve raw values and accept such distributions only when their rounding intervals can contain a unit sum; flag these rows. The batch stopped, the saved response was reevaluated offline after the compatibility fix, and only unattempted cases were then sent. No paid request was retried. Review-workload counts include both flagged gaps and uncertainty; normal tests remain offline and audit CI remains disabled.

Copilot AI lite review requested due to automatic review settings September 21, 2026 12:36

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Understand this PR’s impact

Explore downstream dependencies and potential security impact with Blast Radius.

View blast radius →

📝 Walkthrough

Walkthrough

Added documentation and fixture data for a JEV requirement-coverage audit pilot. Added an offline tool that prepares hashed requests and evaluates recordings. Added a separate bounded live runner with credential validation, fixed HTTPS requests, partial-recording persistence, and failure stopping. Added offline and mocked live tests for validation, metrics, limits, security boundaries, and CLI behavior.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~60 minutes

Merge Risk: 🔵 Low · up to 77d13

A future label or probability-count change could leave this test validating incomplete data instead of failing. Make the conversion strict before merging.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 8.70% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 46 functions across 4 files. (2 skipped: 2… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely identifies the main change: adding a JEV requirement-audit pilot.
Description check ✅ Passed The description explains the change, scope, safety constraints, validation commands, results, limitations, and unperformed CI activation. The required content is present, although the template heading…
Full details: Docstring Coverage

Explanation

Docstring coverage is 8.70% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 46 functions across 4 files. (2 skipped: 2 unsupported.)

  • Fix all pre-merge checks with AI

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-21T12:41:09.421633Z 0c2daff PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0c2daff6d1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tools/jev_audit.py Outdated
row.update(status="answered", choice=choice, confidence=confidence, correct=choice == case["expected"], action=action)
except Invalid as error:
row.update(status="invalid", reason=str(error))
usage_complete = False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve complete usage when only the answer is malformed

When every recorded response has valid usage but one answer fails validation (for example, an unexpected choice or malformed probability distribution), usage_from has already added that response's tokens before this handler unconditionally clears usage_complete. The report therefore exposes all tokens in known_usage but suppresses estimated_recorded_cost_usd, losing cost measurements for precisely the malformed-response experiments the evaluator supports. Only missing or invalid usage should make usage incomplete; answer validity should be tracked independently.

Useful? React with 👍 / 👎.

@lemarier lemarier changed the title feat: add offline JEV requirement audit pilot feat: add JEV requirement audit pilot Sep 21, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: QUIET

Plan: Advanced

Run ID: f7998427-75d5-45cf-afb0-33f7e24162cb

📥 Commits

Reviewing files that changed from the base of the PR and between f8674a4 and f97ccf8.

📒 Files selected for processing (8)
  • README.md
  • docs/jev-audit.md
  • tests/fixtures/jev-audit/NOTICE.md
  • tests/fixtures/jev-audit/cases.json
  • tests/test_jev_audit.py
  • tests/test_jev_audit_live.py
  • tools/jev_audit.py
  • tools/jev_audit_live.py

Included review availability: Your plan provides up to 10 included reviews per hour; 3 remain after this review.

Comment thread tools/jev_audit.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
tests/test_jev_audit.py-186-186 (1)

186-186: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Make this lockstep conversion strict.

If audit.LABELS and probabilities diverge, zip() uses the shorter iterable and dict() silently omits unmatched entries. Add strict=True so the test fails immediately.

Proposed fix
-                distribution = dict(zip(audit.LABELS, probabilities))
+                distribution = dict(zip(audit.LABELS, probabilities, strict=True))

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: QUIET

Plan: Advanced

Run ID: 9ac5b5a1-209b-4fac-bc0c-913159dcca28

📥 Commits

Reviewing files that changed from the base of the PR and between f97ccf8 and 77d13b1.

📒 Files selected for processing (6)
  • docs/jev-audit.md
  • tests/fixtures/jev-audit/NOTICE.md
  • tests/fixtures/jev-audit/expanded.json
  • tests/test_jev_audit.py
  • tests/test_jev_audit_live.py
  • tools/jev_audit.py

Included review availability: Your plan provides up to 10 included reviews per hour; 3 remain after this review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants