Skip to content

feat: no-tool correctness probes for abstention behavior (#107) - #134

Draft
askmy-stack wants to merge 1 commit into
mainfrom
cursor/no-tool-probes-c2a5
Draft

askmy-stack wants to merge 1 commit into
mainfrom
cursor/no-tool-probes-c2a5

Conversation

@askmy-stack

Copy link
Copy Markdown
Owner

Summary

Implements #107: probes for when agents must not call tools.

  • New ProbeKind.no_tool (e.g. email tools available, intent is “What is 2 + 2?”)
  • Model outcomes: unnecessary_tool vs wrong_tool (distinct)
  • Unnecessary write/destructive calls → high_severity_unnecessary + HIGH SEVERITY message
  • Metrics: no_tool_correctness_rate, unnecessary_tool_call_rate, unnecessary_high_severity_rate
  • Report section No-tool failures
  • Offline: structural validation only; model/fake-runner scores abstention
  • Docs + examples/probes/no_tool_math.json

Type of change

  • Enhancement / feature
  • Documentation
  • Tests / CI

Test plan

  • ruff / mypy / pytest
  • CI on this PR

Closes #107

Open in Web Open in Cursor 

Add ProbeKind.no_tool so agents are scored on when not to call tools.
Model outcomes distinguish unnecessary_tool vs wrong_tool; write/destructive
unnecessary calls are high severity. Metrics and report section cover
no-tool failures; offline remains structural-only.

Co-authored-by: Abhinaysai Kamineni  <askmy-stack@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P1] No-tool correctness probes (when agents must not call tools)

2 participants