Skip to content

feat(tool-router): closed-set Jev tool routing benchmark - #43

Merged
BunsDev merged 9 commits into
mainfrom
feat/tool-router
Oct 3, 2026
Merged

BunsDev merged 9 commits into
mainfrom
feat/tool-router

Conversation

@BunsDev

@BunsDev BunsDev commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

What does this change?

Adds a tool-router benchmark. Given a task and a catalog of 40 tools, each described by a one-line snippet, can Jev pick the right tool? The benchmark compares Jev with a deterministic lexical ranker and measures how many context bytes an agent saves by loading only the chosen tool's full schema.

  • lib/tool-router/: closed-set routing library. It uses one choice question pinned to jev-1.13.0, with none as a real option. It never falls back to a default: an invented id, error, timeout or cancellation returns unavailable. A two-stage category-then-tool path covers catalogs larger than 255 options; it is tested with 300 tools and was not needed here.
  • fixtures/tool-router/: a synthetic catalog of 40 tools and 30 tasks, 5 of which expect none.
  • scripts/tool-router-bench.ts: CLI for mock and live runs. It includes a BM25-lite lexical baseline over the same name + snippet text Jev sees.
  • docs/tool-router-bench.md: guide, mock results and live results. This is unrelated to the existing /tool-router workspace, which routes one graph edge at a time; see docs/tool-router.md.

A routing pick is evidence about which tool a task is asking for. It does not authorize running anything, and nothing here executes a tool.

Live result (one run, jev-1.13.0, 30 requests, 0 unavailable)

Top-1 (25 tool tasks) Top-3 none recall (5) Mean ms
Lexical baseline 36% (9/25) 60% (15/25) 80% (4/5) —
Jev 100% (25/25) 100% (25/25) 100% (5/5) 280

Context: all 40 schemas are 19,000 bytes. Snippets plus Jev's top-1 schema are 4,433 bytes, 23.3% of that.

Source: docs/tool-router-results.live.json. Caveats, also stated in the doc:

  • This is one run of 30 synthetic tasks. It is not a capability claim or a calibration.
  • Jev put all probability mass on a single option for most tasks, so top-3 is effectively top-1. Read top-3 as "the expected tool got non-trivial mass".

How was it checked?

  • pnpm typecheck: exit 0
  • pnpm exec tsx --test tests/tool-router-bench.test.ts: 11/11 pass, offline. A test asserts that the serialized request never contains full schemas (properties).
  • Live run with 1Password op run; the key never enters any file. Every response reported jev-1.13.0.
  • The catalog, tasks and snippets are synthetic, written for this experiment. There are no API keys and no private data in the diff.

Draft: no verdicts change. Proposal-review is untouched.

🤖 Generated with Claude Code

BunsDev and others added 6 commits September 22, 2026 03:35
Pure TypeScript under lib/tool-router: catalog validation, a BM25-lite
lexical baseline over name+snippet, a Jev router that asks one Choice
question over every tool plus none (two-stage category then tool when
the catalog exceeds the documented 255-option limit), a deterministic
mock transport, and metrics/report helpers. Provider errors, timeouts
and out-of-set answers become unavailable, never a default pick. Model
is pinned to jev-1.13.0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Forty synthetic tools across eight categories, each with a one-line
snippet and a realistic parameter schema, plus six near-duplicate
distractor pairs. Thirty synthetic tasks with an expected tool, an
acceptable set and a category; five expect none. Nothing here comes
from a real repository or service.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
scripts/tool-router-bench.ts routes every fixture task with the lexical
baseline and with Jev, prints per-category Markdown tables (top-1,
top-3, none precision/recall, unavailable count, mean latency) and a
context-bytes table, and writes the full result JSON. Mock mode uses
the deterministic lexical stand-in and measures only the harness; live
mode reads TYPESAFE_API_KEY from the environment and exits non-zero if
any task was unavailable. Includes the checked-in mock results.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Explains what the experiment is and is not (routing evidence, not
authorization; independent community work), the single Choice question
shape and pinned jev-1.13.0, why top-3 comes from the choice
distribution instead of score questions, the baseline, the metric
definitions, how to run mock and live, and the mock tables. Cross-links
from README and the existing graph-router guide.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
node:test coverage for the tool router: fixture shape (40 tools, 8
categories, 6 distractor pairs, 30 tasks, 5 none, snippet and schema
size bounds), catalog validation, deterministic baseline hand cases,
payload shape with pinned model and no schemas, none as an explicit
choice, error/timeout/cancel/out-of-set answers becoming unavailable,
two-stage routing over the option limit, mock determinism, metric
arithmetic and context-byte accounting.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One live run over the 30 synthetic tasks: 25/25 top-1, 5/5 none, 0
unavailable, 280 ms mean, 43,251 input tokens. Records that the
provider concentrated probability on one option for 23 of 30 tasks, so
top-3 is effectively top-1 on this catalog, and that snippets plus the
routed schema average 23.3% of loading every schema.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
typesafe-ai-playground Ready Ready Preview Oct 3, 2026 6:10am UTC

Request Review

@BunsDev
BunsDev marked this pull request as ready for review September 23, 2026 11:45
Copilot AI lite review requested due to automatic review settings September 23, 2026 11:45

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Resolve the critical ESM CLI failure and the outstanding routing and error-handling issues.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 2 Medium severity

Open (3)
What changed in this PR

Adds a standalone closed-set Jev tool-routing benchmark with synthetic fixtures, lexical comparison, context metrics, and mock/live CLI support.

Changes:

  • Implements single- and two-stage routing with explicit none and unavailable outcomes.
  • Adds routing library, lexical baseline, metrics, tests, and synthetic fixtures.
  • Adds benchmark CLI, live results, documentation, and README cross-reference.

Outstanding findings:

  • scripts/tool-router-bench.ts:30 — critical, 1 vote: replace ESM-incompatible __dirname usage.
  • scripts/tool-router-bench.ts:64 — moderate, 2 votes: mark token usage unknown when requests fail.
  • lib/tool-router/route.ts:143,149 — moderate, 1 vote each: reject invalid probabilities and confidence values.
  • lib/tool-router/route.ts:210 — moderate, 1 vote: remove abort listeners when transports never settle.
  • lib/tool-router/route.ts:235 — moderate, 2 votes: preserve the selected category on second-stage failures.
File Description
tests/​tool-router-bench.test.ts Benchmark behavior tests
scripts/​tool-router-bench.ts Mock/live benchmark CLI
README.md Benchmark overview link
lib/​tool-router/​types.ts Routing data types
lib/​tool-router/​route.ts Closed-set routing logic
lib/​tool-router/​mock.ts Deterministic mock transport
lib/​tool-router/​metrics.ts Accuracy and context metrics
lib/​tool-router/​index.ts Public exports
lib/​tool-router/​catalog.ts Catalog and task validation
lib/​tool-router/​baseline.ts Lexical ranking baseline
fixtures/​tool-router/​tasks.json Synthetic routing tasks
fixtures/​tool-router/​catalog.json Synthetic tool catalog
docs/​tool-router.md Cross-reference documentation
docs/​tool-router-results.live.json Recorded live benchmark results
docs/​tool-router-bench.md Benchmark guide, results, and caveats

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/tool-router-bench.ts Outdated
Comment thread lib/tool-router/route.ts Outdated
Comment thread scripts/tool-router-bench.ts Outdated
@BunsDev

BunsDev commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator Author

@copilot Fix the code for all comments in this review thread.

When a review comment includes a suggested change, apply the suggestion exactly.

Do not make changes beyond what is described in the linked review thread.

Co-authored-by: BunsDev <68980965+BunsDev@users.noreply.github.com>

Copilot AI commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

@copilot Fix the code for all comments in this review thread.

When a review comment includes a suggested chan...

Implemented all requested fixes from that review thread in commit ce74f43.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved catalog key handling and routing correctness/resource issues remain.

Review effort: Lite
Findings: 1 High severity

Open (1)
Resolved since last review (3)

Comment on lines +27 to +31
!NAME.test(raw.name) ||
raw.name === NONE_OPTION ||
typeof raw.category !== "string" ||
!NAME.test(raw.category) ||
raw.category === NONE_OPTION ||
@BunsDev

BunsDev commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

@copilot Fix the code for all comments in this review thread.

When a review comment includes a suggested change, apply the suggestion exactly.

Do not make changes beyond what is described in the linked review thread.

Co-authored-by: BunsDev <68980965+BunsDev@users.noreply.github.com>

Copilot AI commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

@copilot Fix the code for all comments in this review thread.

When a review comment includes a suggested chan...

Rejected __proto__ in both catalog name and category validation and added regression coverage. Targeted tests and typecheck pass; commit: 3deb7c7.

@BunsDev
BunsDev merged commit d750fc6 into main Oct 3, 2026
5 checks passed

This branch was successfully deployed

1 active deployment
Preview — 5b7d3f0c Deployed Oct 3, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants