Skip to content

fix(eval,cli): honor --quiet and keep metric labels behind export redaction - #3454

Merged
kwakayama merged 3 commits into
mainfrom
fix/eval-quiet-and-label-redaction
Aug 7, 2026
Merged

kwakayama merged 3 commits into
mainfrom
fix/eval-quiet-and-label-redaction

Conversation

@kwakayama

@kwakayama kwakayama commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up to #3453. Three review findings landed after that PR was already in the merge queue, so they ship here.

1. --quiet regression (user-facing)

#3453 moved eval report output off cliLogger onto console.log so the ● glyph could mark metric assertions and nothing else. That also moved the output off the logger's level. --quiet/-q raises the canonical log level to WARN (cli/utils/index.ts:150), which used to suppress every cliLogger.info line in the report — so veryfront eval --quiet printed nothing. After #3453 it printed the entire report.

printLine and printBlankLine now check isQuiet() directly. Warnings still print, matching the WARN level quiet mode selects.

Confirmed by writing the test first: it failed on the merged code with Result: 1/1 passed (100%) where [] was expected, and passes now.

2. Metric labels bypassed the export redaction boundary

redactEvalReportForExport strips explanation and evidence from metric results unless includeMetricEvidence is set, so that detail does not reach third-party exporters by default. The new label restates the metric's configured parameter — the asserted tool name, the expected text, the regex — which is the same class of detail evidence carries, and it passed straight through.

Two gaps, both fixed:

Both now leave on the same terms as evidence. Terminal output is unaffected — this boundary only governs what reaches an exporter.

CodeRabbit's version of this finding was to drop the configured text and pattern from labels entirely. I did not take that: naming the parameter is the point of the feature, and it is what the terminal output was asked for. Scoping it to the export boundary keeps the CLI useful and respects the policy the repo already defines.

3. Tool names were not elided

elide() was applied to answer.contains text and answer.regex patterns but not to tool names, so a long tool name could wrap a metric line. Applied consistently now, with the elision test extended to cover all five inline parameters.

Testing

  • src/eval/, cli/commands/eval/, src/extensions/eval/, extensions/ext-eval-report-mlflow/ — 21 passed, 253 steps, 0 failed
  • deno lint, deno fmt --check, docs:api-reference:check — passing

Both fixes are covered by tests written before the fix and verified failing first.

Summary by CodeRabbit

  • New Features
    • Added quiet mode support to the evaluation CLI, suppressing informational reports and metric output when enabled.
    • Long tool names in evaluation metric labels are now shortened with an ellipsis.
  • Bug Fixes
    • Exported evaluation reports now remove metric labels by default, including summary metrics, unless metric evidence is enabled.
  • Documentation
    • Updated evaluation API reference links and clarified metric evidence redaction behavior.

…action

Review follow-ups on the metric-label change.

- Moving eval output off `cliLogger` onto `console.log` also moved it off
  the logger's level, so `veryfront eval --quiet` printed the full report
  where it previously printed nothing. `printLine` and `printBlankLine`
  now check `isQuiet()` directly, with a test asserting no report output.
- `EvalMetricResult.label` and `EvalMetricSummary.label` restate the
  metric's configured parameter, which is the same class of detail
  `evidence` carries. `redactEvalReportForExport` stripped evidence but
  passed labels straight through to third-party exporters, and never
  redacted summary metrics at all. Both now leave on the same terms as
  evidence. Terminal output is unaffected.
- Tool names are elided like the other inline parameters, so a long tool
  name cannot wrap a metric line.
@kwakayama
kwakayama requested a review from kojiwakayama as a code owner August 7, 2026 11:30
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@kwakayama, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 2 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5ed6f893-f86b-47c4-9451-ff9942048322

📥 Commits

Reviewing files that changed from the base of the PR and between 222f179 and ca4fb10.

📒 Files selected for processing (2)
  • docs/api-reference/veryfront/eval.md
  • src/eval/report.ts
📝 Walkthrough

Walkthrough

The eval CLI suppresses informational output in quiet mode. Agent metric labels elide long tool names. Exported reports remove metric labels when metric evidence is disabled. Tests and API source references are updated.

Changes

Eval output controls

Layer / File(s) Summary
Metric label length bounds
src/eval/metric-labels.ts, src/eval/metric-labels.test.ts
Agent tool names now use configured elision. Tests cover five metric label formats.
Report metric redaction
src/extensions/eval/eval-report-exporter.ts, src/extensions/eval/eval-report-exporter.test.ts, src/eval/types.ts, src/eval/report.ts, docs/api-reference/veryfront/eval.md, docs/api-reference/veryfront/extensions.md
Export redaction removes labels from record and summary metrics when metric evidence is disabled. Comments, type documentation, and source references are updated.
Quiet eval CLI output
cli/commands/eval/command.ts, cli/commands/eval/command.test.ts
The eval CLI skips report and metric lines in quiet mode. The integration test verifies output suppression and cleanup.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

Suggested reviewers: kojiwakayama

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 20.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the two primary changes: restoring quiet-mode behavior and redacting metric labels from exports.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/eval-quiet-and-label-redaction

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
docs/api-reference/veryfront/extensions.md (1)

654-655: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Document the label redaction contract.

includeMetricEvidence: false now removes label from record metrics and summary metrics. These rows only update source links. Update the public EvalReportExportRedaction documentation, or its generated source comment, to state that labels follow metric evidence visibility.

As per coding guidelines, public behavior changes must update relevant documentation and generated references.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/api-reference/veryfront/extensions.md` around lines 654 - 655, Update
the public EvalReportExportRedaction documentation or its generated source
comment to state that when includeMetricEvidence is false, label fields are
removed from both record and summary metrics while source links remain. Ensure
the relevant generated API reference reflects this contract.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@cli/commands/eval/command.test.ts`:
- Around line 1258-1287: Update the test setup around runEvalCommand to capture
the original values of VERYFRONT_API_TOKEN, VERYFRONT_PROJECT_SLUG,
VERYFRONT_EVAL_EXPORT, VERYFRONT_EVAL_EXPORTERS, and XDG_CONFIG_HOME before the
try block, then restore each value in finally before removing projectDir and
configHome; preserve whether each variable was originally unset rather than
restoring empty strings.

In `@src/extensions/eval/eval-report-exporter.ts`:
- Around line 176-180: Replace the em dash in the comment near the
redacted.label deletion with a colon or comma, preserving the comment’s meaning
and ensuring no em dash or en dash remains in the TypeScript file.

---

Nitpick comments:
In `@docs/api-reference/veryfront/extensions.md`:
- Around line 654-655: Update the public EvalReportExportRedaction documentation
or its generated source comment to state that when includeMetricEvidence is
false, label fields are removed from both record and summary metrics while
source links remain. Ensure the relevant generated API reference reflects this
contract.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: e7d8361a-71dc-4c11-b04d-9b6748e8d094

📥 Commits

Reviewing files that changed from the base of the PR and between 375412a and ab5cbff.

📒 Files selected for processing (7)
  • cli/commands/eval/command.test.ts
  • cli/commands/eval/command.ts
  • docs/api-reference/veryfront/extensions.md
  • src/eval/metric-labels.test.ts
  • src/eval/metric-labels.ts
  • src/extensions/eval/eval-report-exporter.test.ts
  • src/extensions/eval/eval-report-exporter.ts

Comment thread cli/commands/eval/command.test.ts
Comment thread src/extensions/eval/eval-report-exporter.ts Outdated
Follow-ups from review.

`includeMetricEvidence` now governs metric labels as well as evidence
payloads, on both record and summary metrics. Both public declarations of
`EvalReportExportRedaction` say so, and the generated API reference is
regenerated to match.

AGENTS.md forbids em dash and en dash characters in public comments. Two
comments added by this work used them; both now use a colon or a comma.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/eval/report.ts`:
- Around line 287-288: Update the comment near the summary-row grouping logic to
state that metrics share a summary row only when their complete key matches:
result.name, result.family, and result.severity. Preserve the existing behavior
of dropping the row when no single label describes all grouped metrics.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 2ec8ac95-a6cf-4c3d-bca7-013571792865

📥 Commits

Reviewing files that changed from the base of the PR and between ab5cbff and 222f179.

📒 Files selected for processing (5)
  • docs/api-reference/veryfront/eval.md
  • docs/api-reference/veryfront/extensions.md
  • src/eval/report.ts
  • src/eval/types.ts
  • src/extensions/eval/eval-report-exporter.ts
🚧 Files skipped from review as they are similar to previous changes (2)
  • docs/api-reference/veryfront/extensions.md
  • src/extensions/eval/eval-report-exporter.ts

Comment thread src/eval/report.ts Outdated
Metrics share a summary row when name, family, and severity all match,
not on name alone. The comment said "same name", which understated the
condition it was explaining.
@kwakayama
kwakayama added this pull request to the merge queue Aug 7, 2026
Merged via the queue into main with commit 3f3dbcd Aug 7, 2026
31 checks passed
@kwakayama
kwakayama deleted the fix/eval-quiet-and-label-redaction branch August 7, 2026 12:14
@kwakayama kwakayama mentioned this pull request Aug 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant