Repository navigation
Conversation
The Node OTLP exporter streams with Transfer-Encoding: chunked and sends no Content-Length. The Receiver read Content-Length only, so every Node export came back 400 with zero spans; two TH-8103 children had to put a de-chunking relay in front of it. Receiver now accepts both framings. It also records one entry per accepted export (path, lower-cased headers, flattened resource attributes) via requests(), so a contract test can assert X-Api-Key/X-Secret-Key and project_name through the shared harness instead of a private recorder. Verified: 6 harness tests pass, and the TanStack example's real Node exporter delivered 4 chunked exports (4 spans, collector path, both auth headers, project_name/project_type) with no relay. Refs: TH-8339, TH-8103
New package typescript/packages/traceai_claude_agent_sdk. wrapQuery() wraps the SDK's exported query() and turns its message stream into conversation, assistant_turn, tool_execution, mcp_tool and subagent spans. Each span keeps the Python claude_agent.span_kind string and sets gen_ai.span.kind and fi.span.kind. The conversation span falls back to CHAIN because TS FISpanKind has no CONVERSATION. The attribute names are copied from the Python _attributes.py, and a test fails when one is missing. Tokens and cost come from the result message. Cost is also written to gen_ai.cost.total, the key the collector promotes. The session id goes to session.id. The wrapper also records fork origin, provider (custom when ANTHROPIC_BASE_URL is not an anthropic.com host), abort and errors. Prompts, tool inputs and outputs, and assistant text are hidden unless hideInputs/hideOutputs or FI_HIDE_* opt in. This is stricter than the Python package. @anthropic-ai/claude-agent-sdk (Anthropic Commercial Terms) is a peer dependency, ^0.3.142. It is never imported at runtime or bundled. The contract test packs the tarball and checks it holds no SDK file or native binary. Tests: 40 jest tests, plus shared-harness contract tests. The contract tests send fixture journeys and the real SDK 0.3.289 / 0.3.142 CLI, pointed at a loopback Messages mock through options.env .ANTHROPIC_BASE_URL, through fi-core to the harness Receiver. No Anthropic call. pnpm-lock.yaml gets only this importer and the 15 packages/snapshots it needs. pnpm 10.30.1 also re-resolved other importers' ts-node peers, and those changes were left out. Refs: TH-8235
…y false resolveContentPolicy() used fi-core's rule (value.toLowerCase() === "true"), so FI_HIDE_INPUTS=1, =yes or =" true" turned prompt and tool capture ON. Only an explicit "false" (trimmed, any case) now opts in; every other value, including the empty string, keeps the hidden default. The traceConfig option still takes precedence. README documents the rule. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
The Query proxy only intercepted next/return/throw/asyncIterator. Query.close() (sdk.d.ts:3236-3243, the SDK's abort path) and Symbol.asyncDispose (`await using`) went straight to the original, so a query stopped mid-run never finished its spans and the trace was never exported. Both are now intercepted: the open spans are ended as cancelled (ERROR, claude_agent.cancelled=true), then the call is forwarded. After a completed stream the call changes nothing. Tests: jest close()/asyncDispose cases on the fake Query, and a real-SDK contract scenario (run_real_sdk.mjs SCENARIO=close) that stops two real CLI queries mid-tool. That scenario exported 0 spans before the fix. The real-SDK contract test now also asserts the x-secret-key header. The fake Query gains close/asyncDispose/backgroundTasks and a per-message delay. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
… result In a streaming-input session one query() yields a result per user turn. The next main-loop turn fell back to the query start time (spans.ts startTurn: nextTurnStartMs ?? this.startTimeMs), so the second turn overlapped the first. onResult now sets the main scope's nextTurnStartMs, the same way a user message does. Test: a two-result streaming-input journey (new fixture, numbers taken from the real CLI against the mock) checks turn 2 starts at or after turn 1's end; it failed by 28 ms before the fix. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
…sk_notification A foreground subagent moved to the background arrives as system/task_updated with patch.is_backgrounded (sdk.d.ts:6059; task_id only), or follows an app call to Query.backgroundTasks() (sdk.d.ts:3234). Both were ignored: the subagent span ended at the "running in the background" tool_result and the subagent's later turns were parented to an ended span. - task_started records task_id -> tool_use_id; task_updated with patch.is_backgrounded marks that subagent background. - The Query proxy intercepts backgroundTasks(toolUseId?): matching open foreground subagents are marked before the call is forwarded, and unmarked if it rejects or resolves false. The caller gets the original promise. - A task_notification with status failed/stopped now makes the subagent span ERROR in the foreground case too (it ended OK before when the tool_result itself was not an error). README documents background subagents. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
R1: total_cost_usd is a running total that a resumed, continued or forked session carries forward (sdk.d.ts:5679), and the collector promotes gen_ai.cost.total / gen_ai.usage.* on any span and Observe sums them per trace and per session.id. The wrapper copied the cumulative value, so each resume re-counted earlier spend (real CLI: 0.00018 then 0.00036 promoted for 0.00036 spent). - A process-local LRU (1000 sessions) maps session id -> last running totals. The promoted keys carry the cumulative value minus that baseline: zero for a new session, the session's own totals for resume/continue, the parent's for a fork. A drop within one query (/clear) counts both sides. - No known baseline (resume after restart, continue+forkSession, or a first result below the saved totals): no promoted key and no claude_agent.cost.total_usd; claude_agent.usage.baseline_unknown=true. - Running totals always go on unpromoted claude_agent.cumulative.* keys. - options.continue is now labelled session.is_resumed=true / is_new=false. - A result with nothing usable, or a zeroed crash result, writes no usage key and leaves the baseline alone. R2: tokens came from result.usage, which is main-loop only and per turn in streaming input (sdk.d.ts:5683). They now come from the latest result's modelUsage summed over models (input, output, cache read, cache creation), with the same delta handling. Streaming input stays one conversation span per query() (recorded in the README; Python's ClaudeSDKClient makes one per user turn). Tests: usage.test.ts (21 jest: streaming totals, multi-model modelUsage, no usage on turn/tool/subagent spans, absent usage, resume/continue/fork deltas, unknown baselines, /clear, crash result, LRU). Real-CLI contract scenarios resume, continue_fork, restart (two processes, shared HOME) and streaming assert the promoted keys sum to the final cumulative totals; all four failed before the fix. Fixture results now carry modelUsage. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
…op tsbuildinfo R7 review nits: - package.json repository/bugs/homepage pointed at github.com/futureagi/traceai (copied from strands). They now match traceai_anthropic: future-agi/traceAI, with this package's directory. - The module-level knownProviders Set kept every provider passed to an instrumentation alive forever. It now holds WeakRefs, pruned on read and add, and dedupes a provider set twice. shutdown() with no argument still flushes every live provider. - The tarball shipped dist/**/*.tsbuildinfo; files now excludes it. Tests: package.test.ts (metadata against the sibling; 20 dropped providers are released after a forced GC, failed with 20 held before; a live provider is tracked once). The pack contract test rejects *.tsbuildinfo (failed on dist/esm/tsconfig.esm.tsbuildinfo before). approved: Nikhil 2026-10-03 blanket Refs: TH-8235
…ask_id notifications, partial baselines N1: close()/asyncDispose/abort after a result with no open turn, tool or subagent ends the query as returned (conversation OK), the normal end of a streaming-input session. A close with work still in flight stays cancelled. N2: task_notification without tool_use_id (optional, sdk.d.ts:5997) is matched through the task_started task_id mapping. N3: an unknown usage baseline covers only the first result; later results in the same query promote their delta against it (flag stays true). N4: resume/fork with resumeSessionAt has no known baseline. README notes the /clear undercount edge. Tests first: 8 jest cases and the new real-CLI streaming_close contract test failed on 06a3a20 and pass now. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
…ion stays in one process The usage baseline is process-local. Resuming one session from several processes can promote spend made in another process again. Documented, with claude_agent.cumulative.* as the exact source for those deployments. Rejected: a cross-process store (shared state and a new dependency) and dropping promotion on every resume (loses resumed spend in the common single-process case). approved: Nikhil 2026-10-03 blanket Refs: TH-8235
… result ends OK on close The N1 shortcut keyed off resultSeen, which never resets, so an abort in a later streaming turn between model steps (turn span ended, next assistant message not yet out) was recorded as OK. Track idleAfterResult instead: a result sets it, any later assistant/user message clears it. close(), asyncDispose and AbortController aborts mid-turn are cancelled again. README names the residual gap (abort after a new prompt is sent, before the SDK yields anything for it). Comments in queryWrapper.ts refreshed. M2 (reset the in-query baseline on conversation_reset) not taken: whether the CLI's running totals restart there is unverified; the README keeps the documented /clear undercount. Test first: 3 new jest cases (close, asyncDispose, abort) failed on a8831d6. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
… the idle-after-result state Background subagent messages (parent_tool_use_id set) that arrive after the main loop's result no longer clear idleAfterResult, so close() after that subagent's task_notification ends the conversation OK; a subagent still open stays cancelled through nothingInFlight(). Test first: the new post-result background subagent case failed on a1dfe62. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
…ssion state drive the idle-after-result state R4-1: a background task_notification is queued as a main-thread command that runs its own turn (sdk.d.ts:5440, 4603), so close() after one is a cancellation again (supersedes 005a70f's expectation). R4-2: a main-loop stream_event (includePartialMessages) after a result also starts a turn. system/session_state_changed is authoritative when sent (sdk.d.ts:5858): 'idle' after a result ends OK, 'running' cancelled. README updated. Test first: 3 new jest cases failed on 005a70f. approved: Nikhil 2026-10-03 blanket Refs: TH-8235
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
New TypeScript package
@traceai/claude-agent-sdkfor TH-8235. It wrapsquery()from@anthropic-ai/claude-agent-sdkand keeps parity with the existing Python package. The Python package is unchanged.claude_agent.conversation(CHAIN),claude_agent.assistant_turn(LLM),tool.<name>(TOOL),claude_agent.subagent.<type>(AGENT under theTask/Agenttool span), MCP tools (mcp_tool/TOOL).claude_agent.span_kindis kept for parity.fi.span.kindandgen_ai.span.kindcarry the Future AGI kind; the collector readsfi.span.kindfirst.query()call. Tokens come from the result'smodelUsagesummed over every model, cost fromtotal_cost_usd. Both are session running totals, so resume/continue/fork promote a delta against an in-process per-session baseline. With no known baseline the first result promotes nothing and later results in the same query promote their delta (claude_agent.usage.baseline_unknown=true). Running totals are always onclaude_agent.cumulative.*. Exact only while a session stays in one process (README). Turn, tool and subagent spans carry no usage keys.session.idcomes from the SDK session id unless the app already set one through fi-core context.options.continueis labelled as a resumed session.FI_HIDE_INPUTS=false/FI_HIDE_OUTPUTS=false(trimmed, any case) ortraceConfigopts in;1,yes,true, empty all keep content hidden. Stricter than the Python package, which records the prompt and tool I/O unconditionally.close(),Symbol.asyncDisposeandAbortControllerend every open span as cancelled while work is in flight. Right after a result, before the main loop starts another turn and with nothing still running, they end the query normally (status OK); this is the usual way to end a streaming-input session.session_state_changeddecides it when the CLI sends one. Background subagents (run_in_background,task_started/task_updatedis_backgrounded,Query.backgroundTasks()) stay open untiltask_notification.peerDependency(^0.3.142) and a devDependency for tests only. A pack test asserts that the tarball holds no SDK source, native binary or*.tsbuildinfo.Deviations from the accepted architecture (with evidence)
Agent(sdk.d.ts:1562). BothTaskandAgentare treated as subagent tools.FISpanKindhas noCONVERSATION(fi-semantic-conventions SemanticConventions.ts:607), so the conversation span is CHAIN, as the spec required.tool_use/tool_result). The spec requires userhooksto pass through unchanged, so no hooks are injected.result.usageis main-loop only and per turn in streaming input (sdk.d.ts:5683); tokens now come frommodelUsage(sdk.d.ts:5687) with per-session deltas (sdk.d.ts:5679).query(); Python'sClaudeSDKClientpath makes one per user turn. Totals cover every turn.traceai_anthropic, because thetraceai_openai_agentsconfig emitsdist/esm/src/index.jswhile itspackage.jsonpoints atdist/esm/index.js.register({ batch: true })does not batch on sdk-trace-base 2.x. This package flushes explicitly.Tests (run on 4cf737f, clean tree)
pnpm test(jest--no-cache) in the package: 101 passed (5 suites: wrapper, usage, parity, collector, package);tsc --noEmitclean.PYTHONPATH=python/tests uv run --no-project --python 3.11 --with pytest --with opentelemetry-api --with opentelemetry-sdk --with opentelemetry-exporter-otlp-proto-http pytest typescript/packages/traceai_claude_agent_sdk/contract -q --noconftest -o addopts='': 12 passedReceiverfrom test(harness): shared OTLP harness for TH-8103 contract tests #203: collector path,X-Api-Key/X-Secret-Key,project_name/project_type=observe.options.env.ANTHROPIC_BASE_URL): one Read tool;close()+asyncDisposemid-tool; resume; continue + fork; resume after a process restart (shared HOME); streaming input with two user turns; streaming input ended byclose()after its result. The usage scenarios assert that promoted cost/tokens summed over all traces equal the final cumulative totals.Review
Independent fresh-context review of e14df89: CHANGES_REQUESTED, findings R1-R7. Fixed test-first on top of e14df89; every new test failed on the old code first.
options.continueunhandledcontinuelabelled resumedresult.usage(main loop, per turn)modelUsageacross models, same delta handlingclose()/Symbol.asyncDisposeskip the tracerFI_HIDE_*fails open on1/yes/truefalseopts intask_updated/backgroundTasks()backgrounding ignoredtask_notification; failed/stopped status is ERRORtraceai_anthropic's; WeakRef registry;filesexcludes tsbuildinfoRound 2 on 06a3a20, run as separate installed cards (Opus 5.5):
Fixed in cc1a581 (test first: 8 jest cases and a new real-CLI contract test failed with the fix stashed) and in a8831d6 (docs):
close()/asyncDisposeafter a successful result marked the conversation ERROR "cancelled"streaming_closescenariotask_notificationwithout the optionaltool_use_id(sdk.d.ts:5997) left a background subagent opentask_startedtask_id mappingresumeSessionAtused the latest saved totals;/clearedgeresume/fork withresumeSessionAthas no known baseline; the/clearundercount is documentedclaude_agent.cumulative.*is exact for every deployment. A cross-process store (shared state, new dependency) and dropping promotion on every resume (loses resumed spend in the common case) were both rejectedRound 3 on a8831d6 (installed pr-reviewer t_2411725e): CHANGES_REQUESTED.
resultSeen, which never resets. An abort in a later streaming turn, between model steps, was therefore recorded as OK. Fixed in a1dfe62: anidleAfterResultflag is set by a result and cleared by any later assistant or user message. Test first: 3 cases (close, asyncDispose, abort) failed on a8831d6. The README names the one remaining gap: an abort after a new prompt is sent, before the SDK yields anything for it.conversation_reset. Not taken: whether the CLI's running totals restart there is unverified. The documented/clearundercount stays.pr-verifier t_46d2f7cf on a8831d6: CHANGES_REQUESTED. It confirmed M1 and M2, and found N2-N4 and V1 resolved. Its fix note: clear the idle state only on main-loop messages. A background subagent's messages after the result must not turn a later close() into a cancellation. Applied in 005a70f; the new test failed on a1dfe62.
Round 4 (pr-verifier t_2915cba8 on a8831d6..005a70f): M1 fixed for close, asyncDispose and AbortController. M2 is acceptable to defer. Two P3 findings were fixed test first in 4cf737f (3 new cases failed on 005a70f):
task_notificationruns its own main-thread turn (sdk.d.ts:5440), soclose()after one is a cancellation again.system/session_state_changedis authoritative when the CLI sends it.stream_eventafter a result starts a turn.Round 5 (pr-verifier t_2cc44bc8 on 005a70f..4cf737f): VERIFIED with no regressions and no blocking findings. Its two P3 follow-ups are listed under Limits.
Review is complete at 4cf737f.
Video demo
A narrated terminal demo (5:46) was recorded at the reviewed head
4cf737f9d2d100118da08cda40c3242ea70ee51b. It is a real terminal recording, not a slideshow.What it runs:
@anthropic-ai/claude-agent-sdk0.3.289 against a local loopback mock of the Messages API. There is no live Anthropic call.Where to find it: the video, captions, transcript and media check are attached privately to Linear issue TH-8235 as
th-8235-demo.mp4,th-8235-demo.srt,transcript.mdandmedia-verification.md. The mp4's sha256 isebec6535…467d63d99.Chapters:
/tracer/v1/traceswith both auth headers; usage keys on the conversation onlytotal_cost_usdclose()andasyncDisposemid-tool end all 6 spans ERROR cancelled; a streamingclose()right after the result ends OKrecordUsageThe demo driver is
run_demo.sh+demo_spans.py. It is read-only and reuses the contract module's real-SDK runner.Limits
startup()is not wrapped.claude_agent.cumulative.*is exact in every case.task_notificationis treated as starting a main-loop turn. The SDK types don't say whether ambient orskip_transcripttasks do; the error is on the conservative side (cancelled).session_state_changedis only sent when the CLI opts in (CLAUDE_CODE_EMIT_SESSION_STATE_EVENTS), so the authoritative idle path is usually off.conversation_reset) is deferred: it would only ever undercount, and the CLI behaviour is unverified.future-agi/docsPR is not opened here.Stacked on #203 (shared harness, includes the chunked-body fix
3eaadc8). Not merged.Linear: TH-8235. approved: Nikhil 2026-10-03 blanket