Skip to content

fix(telemetry): Keep gen_ai spans and maxTokens headroom on failure - #525

Draft
sentry-junior[bot] wants to merge 3 commits into
mainfrom
fix/telemetry-genai-resilience
Draft

sentry-junior[bot] wants to merge 3 commits into
mainfrom
fix/telemetry-genai-resilience

Conversation

@sentry-junior

@sentry-junior sentry-junior Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Fixes two related gaps from the WARDEN-Y / junior#1724 investigate path.

Why traces/conversations were missing

  • Junior's Warden workflow runs Analyze and Report as separate GHA steps (two processes → two root traces). That split is expected.
  • The Analyze process still lost its whole skill/gen_ai span tree on this failure while errors+logs survived. Custom pi/claude otel attrs were dumping unbounded prompt JSON onto spans; Sentry's own AI integrations truncate to ~20KB.
  • Request-time input + max_tokens > context failures also never recorded gen_ai.request.max_tokens before the provider rejected, so traces couldn't show the budget math.

What changed

  • Truncate gen_ai.input/output/system (+ tool arg/result) payloads to Sentry's GenAI message byte budget, with truncation metadata.
  • Cap Pi catalog maxTokens with 96k input headroom rooted in observed WARDEN-Y failing prompts (~51–80k) + undercount pad — not a flat 128k generation ceiling. Sonnet/GPT-shaped catalogs unchanged.
  • Record gen_ai.request.max_tokens / context_window on the invoke span before the provider call.
  • Tag root span with warden.action.mode so analyze vs report traces are distinguishable.
  • Flush Sentry after analyze-mode trigger work so short-lived analyze processes don't race process exit.

Not in this PR

  • Full AI Conversations transcript wiring for pi/openrouter (still Anthropic-integration-shaped product surface).
  • Re-enabling compaction (separate question; wouldn't fix first-call max_tokens overflow anyway).

Checks

  • vitest otel + pi-model-tokens + pi + runner: 52 passed

Requested by greg.pstrucha.

--

View Junior Session [Sentry]

sentry-junior Bot and others added 2 commits August 28, 2026 23:06
Truncate oversized GenAI span attributes to Sentry's message budget so
analyze-mode transactions stop dropping their skill/model tree. Cap Pi
catalog maxTokens with observed-input headroom, tag action mode on the
root span, record request max_tokens before provider calls, and flush
after analyze triggers.

Co-Authored-By: Greg Pstrucha <greg.pstrucha@sentry.io>
oxlint consistent-type-definitions requires interface for object shapes.
@vercel

vercel Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated (UTC)
warden-prod Skipped Skipped Aug 28, 2026 11:32pm

Request Review

Comment thread packages/warden/src/sdk/otel.ts
Multi-part assistant turns with large tool_call payloads could still exceed
the 20KB GenAI attribute budget after only truncating parts[0].text. Drop
trailing parts until the empty envelope fits, then hard-fallback to a
placeholder if needed.

This branch was successfully deployed

1 active and 1 inactive deployments
Preview – warden — fe42f201 Deployed Aug 28, 2026 by vercel[bot]
Preview – warden-prod — fe42f201 Deployed Aug 28, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants