Repository navigation
feat(tool-router): closed-set Jev tool routing benchmark - #43
Conversation
Pure TypeScript under lib/tool-router: catalog validation, a BM25-lite lexical baseline over name+snippet, a Jev router that asks one Choice question over every tool plus none (two-stage category then tool when the catalog exceeds the documented 255-option limit), a deterministic mock transport, and metrics/report helpers. Provider errors, timeouts and out-of-set answers become unavailable, never a default pick. Model is pinned to jev-1.13.0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Forty synthetic tools across eight categories, each with a one-line snippet and a realistic parameter schema, plus six near-duplicate distractor pairs. Thirty synthetic tasks with an expected tool, an acceptable set and a category; five expect none. Nothing here comes from a real repository or service. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
scripts/tool-router-bench.ts routes every fixture task with the lexical baseline and with Jev, prints per-category Markdown tables (top-1, top-3, none precision/recall, unavailable count, mean latency) and a context-bytes table, and writes the full result JSON. Mock mode uses the deterministic lexical stand-in and measures only the harness; live mode reads TYPESAFE_API_KEY from the environment and exits non-zero if any task was unavailable. Includes the checked-in mock results. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Explains what the experiment is and is not (routing evidence, not authorization; independent community work), the single Choice question shape and pinned jev-1.13.0, why top-3 comes from the choice distribution instead of score questions, the baseline, the metric definitions, how to run mock and live, and the mock tables. Cross-links from README and the existing graph-router guide. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
node:test coverage for the tool router: fixture shape (40 tools, 8 categories, 6 distractor pairs, 30 tasks, 5 none, snippet and schema size bounds), catalog validation, deterministic baseline hand cases, payload shape with pinned model and no schemas, none as an explicit choice, error/timeout/cancel/out-of-set answers becoming unavailable, two-stage routing over the option limit, mock determinism, metric arithmetic and context-byte accounting. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One live run over the 30 synthetic tasks: 25/25 top-1, 5/5 none, 0 unavailable, 280 ms mean, 43,251 input tokens. Records that the provider concentrated probability on one option for 23 of 30 tasks, so top-3 is effectively top-1 on this catalog, and that snippets plus the routed schema average 23.3% of loading every schema. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Resolve the critical ESM CLI failure and the outstanding routing and error-handling issues.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (3)
What changed in this PR
Adds a standalone closed-set Jev tool-routing benchmark with synthetic fixtures, lexical comparison, context metrics, and mock/live CLI support.
Changes:
- Implements single- and two-stage routing with explicit
noneandunavailableoutcomes. - Adds routing library, lexical baseline, metrics, tests, and synthetic fixtures.
- Adds benchmark CLI, live results, documentation, and README cross-reference.
Outstanding findings:
scripts/tool-router-bench.ts:30— critical, 1 vote: replace ESM-incompatible__dirnameusage.scripts/tool-router-bench.ts:64— moderate, 2 votes: mark token usage unknown when requests fail.lib/tool-router/route.ts:143,149— moderate, 1 vote each: reject invalid probabilities and confidence values.lib/tool-router/route.ts:210— moderate, 1 vote: remove abort listeners when transports never settle.lib/tool-router/route.ts:235— moderate, 2 votes: preserve the selected category on second-stage failures.
| File | Description |
|---|---|
tests/tool-router-bench.test.ts |
Benchmark behavior tests |
scripts/tool-router-bench.ts |
Mock/live benchmark CLI |
README.md |
Benchmark overview link |
lib/tool-router/types.ts |
Routing data types |
lib/tool-router/route.ts |
Closed-set routing logic |
lib/tool-router/mock.ts |
Deterministic mock transport |
lib/tool-router/metrics.ts |
Accuracy and context metrics |
lib/tool-router/index.ts |
Public exports |
lib/tool-router/catalog.ts |
Catalog and task validation |
lib/tool-router/baseline.ts |
Lexical ranking baseline |
fixtures/tool-router/tasks.json |
Synthetic routing tasks |
fixtures/tool-router/catalog.json |
Synthetic tool catalog |
docs/tool-router.md |
Cross-reference documentation |
docs/tool-router-results.live.json |
Recorded live benchmark results |
docs/tool-router-bench.md |
Benchmark guide, results, and caveats |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
@copilot Fix the code for all comments in this review thread. When a review comment includes a suggested change, apply the suggestion exactly. Do not make changes beyond what is described in the linked review thread. |
Co-authored-by: BunsDev <68980965+BunsDev@users.noreply.github.com>
Implemented all requested fixes from that review thread in commit |
| !NAME.test(raw.name) || | ||
| raw.name === NONE_OPTION || | ||
| typeof raw.category !== "string" || | ||
| !NAME.test(raw.category) || | ||
| raw.category === NONE_OPTION || |
|
@copilot Fix the code for all comments in this review thread. When a review comment includes a suggested change, apply the suggestion exactly. Do not make changes beyond what is described in the linked review thread. |
Co-authored-by: BunsDev <68980965+BunsDev@users.noreply.github.com>
Rejected |


What does this change?
Adds a tool-router benchmark. Given a task and a catalog of 40 tools, each described by a one-line snippet, can Jev pick the right tool? The benchmark compares Jev with a deterministic lexical ranker and measures how many context bytes an agent saves by loading only the chosen tool's full schema.
lib/tool-router/: closed-set routing library. It uses onechoicequestion pinned tojev-1.13.0, withnoneas a real option. It never falls back to a default: an invented id, error, timeout or cancellation returnsunavailable. A two-stage category-then-tool path covers catalogs larger than 255 options; it is tested with 300 tools and was not needed here.fixtures/tool-router/: a synthetic catalog of 40 tools and 30 tasks, 5 of which expectnone.scripts/tool-router-bench.ts: CLI for mock and live runs. It includes a BM25-lite lexical baseline over the samename + snippettext Jev sees.docs/tool-router-bench.md: guide, mock results and live results. This is unrelated to the existing/tool-routerworkspace, which routes one graph edge at a time; seedocs/tool-router.md.A routing pick is evidence about which tool a task is asking for. It does not authorize running anything, and nothing here executes a tool.
Live result (one run,
jev-1.13.0, 30 requests, 0 unavailable)nonerecall (5)Context: all 40 schemas are 19,000 bytes. Snippets plus Jev's top-1 schema are 4,433 bytes, 23.3% of that.
Source:
docs/tool-router-results.live.json. Caveats, also stated in the doc:How was it checked?
pnpm typecheck: exit 0pnpm exec tsx --test tests/tool-router-bench.test.ts: 11/11 pass, offline. A test asserts that the serialized request never contains full schemas (properties).op run; the key never enters any file. Every response reportedjev-1.13.0.Draft: no verdicts change. Proposal-review is untouched.
🤖 Generated with Claude Code