Skip to content

feat: add optional LLM semantic judge for ambiguous tool pairs (#81) - #156

Draft
askmy-stack wants to merge 2 commits into
mainfrom
cursor/semantic-judge-c2a5
Draft

askmy-stack wants to merge 2 commits into
mainfrom
cursor/semantic-judge-c2a5

Conversation

@askmy-stack

Copy link
Copy Markdown
Owner

Summary

Implements #81 — optional LLM semantic judge for ambiguous tool comparisons.

Changes

  • New tool_semantics.judge module:
    • Opt-in judge_tool_pair / judge_ambiguous_pairs via existing ModelRunner
    • Findings always labeled MODEL-BASED
    • Cannot clear deterministic breaking/critical changes (blocked_by_deterministic)
    • Safe prompts: name / short description / risk / parameter names only
  • Docs: docs/semantic-judge.md (when to use vs deterministic / embeddings)
  • Fake-runner unit tests only (no live CI calls)

Test plan

  • pytest tests/test_judge.py
  • ruff / mypy
  • CI green

Closes #81

Open in Web Open in Cursor 

cursoragent and others added 2 commits September 16, 2026 15:00
Opt-in MODEL-BASED judgments via ModelRunner. Deterministic breaking
changes cannot be cleared by same-task claims. Fake-runner tests only.

Co-authored-by: Abhinaysai Kamineni  <askmy-stack@users.noreply.github.com>
Co-authored-by: Abhinaysai Kamineni  <askmy-stack@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optional LLM semantic judge for ambiguous tool comparisons

2 participants