Skip to content

feat: parallel probes, completion cache, cost, scale bench (#100) - #164

Draft
askmy-stack wants to merge 1 commit into
mainfrom
cursor/parallel-cache-cost-c2a5
Draft

askmy-stack wants to merge 1 commit into
mainfrom
cursor/parallel-cache-cost-c2a5

Conversation

@askmy-stack

Copy link
Copy Markdown
Owner

Summary

Implements #100: parallel model-backed probe workers, completion caching, optional token/cost summaries, and an informational scale benchmark harness.

Changes

  • ProbeCompletionCache — key = model + prompt/tools + snapshot hash + temperature + seed
  • evaluate_probes_with_model(..., workers=, cache=) — deterministic result order
  • Cost/token aggregation from provider usage metadata
  • scripts/bench_scale.py + run_scale_benchmark for 10/100/500/1000/5000 (not CI-gated)
  • CLI: --workers, --cache-dir
  • Docs: docs/performance.md

Test plan

  • pytest tests/test_perf_cache.py (cache hit/miss, workers order, cost, small scale)
  • Full pytest + ruff + mypy
  • CI green on this PR
Open in Web Open in Cursor 

Add --workers/--cache-dir for model-backed probes, token/cost summaries,
and an informational scale harness (10–5000 tools) outside PR CI gates.

Co-authored-by: Abhinaysai Kamineni  <askmy-stack@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants