Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1 AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1 and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.

Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
  from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
  and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)

Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
  open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
  budget is set; hidden and non-breaking against backends without /usage

Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.

Tests: usage endpoint, severity thresholds, and two widget e2e cases.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.

Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.

Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1 force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0 Compare July 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to main July 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
Collaborator Author

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assert TokenBudgetUsage.from_totals(849, 1000).severity == "warning"
assert TokenBudgetUsage.from_totals(850, 1000).severity == "critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into main Jul 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancement User-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants