Skip to content

test: make stream capture tests reach their bounds and correct docs - #35

Merged
aikins01 merged 3 commits into
telemetry-dev:mainfrom
lofoneh:fix/stream-capture-bounds
Oct 4, 2026
Merged

aikins01 merged 3 commits into
telemetry-dev:mainfrom
lofoneh:fix/stream-capture-bounds

Conversation

@lofoneh

@lofoneh lofoneh commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Closes #31

Tests

  • The five OpenAI chat recovery tests no longer set maxAttributeLength. Each one now prefills content to the real 48 KiB budget, so the 256-char value actually gets rejected and recovery happens right at the boundary. I checked each test by breaking the mechanism it's named for (separator refund, role field refund, recoverable id capture, id release, shell reserved once). The old tests passed all five breakages; the new ones fail each.
  • Tests 3–5 sit 1–2 bytes under 48 KiB. That slack is the tool-call path's own conservative reservation; it's the same with no rejection at all, so it isn't a leak.
  • The Anthropic fallback-tail fixture now rejects slices, so an iterations[::-1] scan fails.

Docs

  • packages/sdk/README.md: an older core keeps the same fixed retention bounds, and only the mask is assumed. I also fixed the matching LEGACY_CAPTURE_POLICY comment in openai and anthropic. maxAttributeLength stays because the policy type requires it.
  • TS openai README: transcription is 65,536 serialized UTF-16 units, not 64 KiB. The Python one really is UTF-8 bytes, so I left it.
  • Python openai README: Responses capture uses 48 KiB and the core default of 1,024 items. TS Responses really is 1,000, so its README was already right.

The five OpenAI chat recovery tests relied on maxAttributeLength, which no longer shrinks the fixed 48 KiB retention budget, so their reject and refund paths never ran. Prefill content up to the budget instead. The Anthropic fallback-tail fixture now rejects slices, so a full iterations[::-1] scan fails.
Retention uses the fixed bounds with any core; only the mask is assumed. TypeScript transcription counts serialized UTF-16 length, and Python Responses capture uses the core default of 1,024 items.

@aikins01 aikins01 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, the recovery tests now hit the real 48 KiB budget, and they fail when the separator, string, or field refund is removed, which the old versions didn't catch. The Anthropic fixture now catches a slice-based full scan too.

One sibling spot still needs the same correction: the root README (lines 119–120) gives one set of bounds for both SDKs. Python Responses uses 48 KiB and 1,024 items while TypeScript uses 1,000, and TypeScript transcription counts 65,536 UTF-16 units of serialized JSON while Python keeps 64 KiB of UTF-8. Could you split those there to match the package READMEs?

@lofoneh

lofoneh commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Good catch, fixed in bb4875d. The root README now splits Responses and transcription by language. Chat and Anthropic stay as one line since both SDKs use 48 KiB and 1,000 there.

@aikins01 aikins01 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for the quick fix, the README now matches both SDKs

@aikins01
aikins01 merged commit ed0a97d into telemetry-dev:main Oct 4, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

sdk: restore stream recovery tests and align documented capture bounds

2 participants