Skip to content

feat(claude-agent-sdk): @traceai/claude-agent-sdk for TypeScript (TH-8235) - #209

Open
nik13 wants to merge 15 commits into
devfrom
feat/th-8235-claude-agent-sdk-ts
Open

nik13 wants to merge 15 commits into
devfrom
feat/th-8235-claude-agent-sdk-ts

Conversation

@nik13

@nik13 nik13 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

New TypeScript package @traceai/claude-agent-sdk for TH-8235. It wraps query() from @anthropic-ai/claude-agent-sdk and keeps parity with the existing Python package. The Python package is unchanged.

  • Spans: claude_agent.conversation (CHAIN), claude_agent.assistant_turn (LLM), tool.<name> (TOOL), claude_agent.subagent.<type> (AGENT under the Task/Agent tool span), MCP tools (mcp_tool/TOOL). claude_agent.span_kind is kept for parity. fi.span.kind and gen_ai.span.kind carry the Future AGI kind; the collector reads fi.span.kind first.
  • Tokens and cost go on the conversation span only, and hold only the spend that is new in this query() call. Tokens come from the result's modelUsage summed over every model, cost from total_cost_usd. Both are session running totals, so resume/continue/fork promote a delta against an in-process per-session baseline. With no known baseline the first result promotes nothing and later results in the same query promote their delta (claude_agent.usage.baseline_unknown=true). Running totals are always on claude_agent.cumulative.*. Exact only while a session stays in one process (README). Turn, tool and subagent spans carry no usage keys.
  • session.id comes from the SDK session id unless the app already set one through fi-core context. options.continue is labelled as a resumed session.
  • Content is off by default. Only FI_HIDE_INPUTS=false / FI_HIDE_OUTPUTS=false (trimmed, any case) or traceConfig opts in; 1, yes, true, empty all keep content hidden. Stricter than the Python package, which records the prompt and tool I/O unconditionally.
  • close(), Symbol.asyncDispose and AbortController end every open span as cancelled while work is in flight. Right after a result, before the main loop starts another turn and with nothing still running, they end the query normally (status OK); this is the usual way to end a streaming-input session. session_state_changed decides it when the CLI sends one. Background subagents (run_in_background, task_started/task_updated is_backgrounded, Query.backgroundTasks()) stay open until task_notification.
  • License boundary: the SDK is under Anthropic Commercial Terms, not OSI. It is a peerDependency (^0.3.142) and a devDependency for tests only. A pack test asserts that the tarball holds no SDK source, native binary or *.tsbuildinfo.

Deviations from the accepted architecture (with evidence)

  • In 0.3.289 the subagent tool is Agent (sdk.d.ts:1562). Both Task and Agent are treated as subagent tools.
  • TS FISpanKind has no CONVERSATION (fi-semantic-conventions SemanticConventions.ts:607), so the conversation span is CHAIN, as the spec required.
  • Tool spans are built from the message stream (tool_use / tool_result). The spec requires user hooks to pass through unchanged, so no hooks are injected.
  • Usage source: the spec mapped tokens to "result message usage". result.usage is main-loop only and per turn in streaming input (sdk.d.ts:5683); tokens now come from modelUsage (sdk.d.ts:5687) with per-session deltas (sdk.d.ts:5679).
  • Streaming input stays one conversation span per query(); Python's ClaudeSDKClient path makes one per user turn. Totals cover every turn.
  • The ESM build config is copied from traceai_anthropic, because the traceai_openai_agents config emits dist/esm/src/index.js while its package.json points at dist/esm/index.js.
  • Found and filed separately (TH-8383), not fixed here: in fi-core 1.0.0, register({ batch: true }) does not batch on sdk-trace-base 2.x. This package flushes explicitly.

Tests (run on 4cf737f, clean tree)

  • pnpm test (jest --no-cache) in the package: 101 passed (5 suites: wrapper, usage, parity, collector, package); tsc --noEmit clean.
  • PYTHONPATH=python/tests uv run --no-project --python 3.11 --with pytest --with opentelemetry-api --with opentelemetry-sdk --with opentelemetry-exporter-otlp-proto-http pytest typescript/packages/traceai_claude_agent_sdk/contract -q --noconftest -o addopts='': 12 passed
    • Shared-harness contract tests through the Receiver from test(harness): shared OTLP harness for TH-8103 contract tests #203: collector path, X-Api-Key/X-Secret-Key, project_name/project_type=observe.
    • Seven run the real SDK 0.3.289 CLI against a loopback Messages-API mock (options.env.ANTHROPIC_BASE_URL): one Read tool; close() + asyncDispose mid-tool; resume; continue + fork; resume after a process restart (shared HOME); streaming input with two user turns; streaming input ended by close() after its result. The usage scenarios assert that promoted cost/tokens summed over all traces equal the final cumulative totals.
    • Pack test (no SDK file, native binary or tsbuildinfo) and ESM/CJS import.
  • Mutation check: putting tokens on every assistant turn now fails 2 tests (it passed all 40 on e14df89).
  • Not run on this head: the Node 18/20/22 matrix (only Node 26.8.1 on the test host; e14df89 was checked on 18/20/22 by the implementer).

Review

Independent fresh-context review of e14df89: CHANGES_REQUESTED, findings R1-R7. Fixed test-first on top of e14df89; every new test failed on the old code first.

ID Finding Resolution Commit RED before fix
R1 (blocking) Resumed sessions double-count cost; options.continue unhandled Promote cumulative minus a per-session baseline (LRU, 1000 sessions; fork uses the parent's); unknown baseline promotes nothing; continue labelled resumed c6fe543 real CLI resume/continue-fork/restart scenarios failed (0.00054 promoted for 0.00036 spent)
R2 (blocking) Tokens from result.usage (main loop, per turn) Sum the latest modelUsage across models, same delta handling c6fe543 streaming scenario: 20 tokens promoted vs 40
R3 close() / Symbol.asyncDispose skip the tracer Intercepted; open spans end cancelled, then forward 14e3adf real CLI exported 0 spans after close/dispose
R4 FI_HIDE_* fails open on 1/yes/ true Only an explicit false opts in 5d0efff 7 jest cases failed
R5 Streaming-input turn 2 starts at the query start Next turn starts after the previous result 378acc5 turn 2 overlapped turn 1 by 28 ms
R6 task_updated / backgroundTasks() backgrounding ignored Tracked; subagent ends at task_notification; failed/stopped status is ERROR 8b1ab15 6 jest cases failed
R7 (nits) Wrong repo URLs; providers held forever; tsbuildinfo in tarball Repo metadata = traceai_anthropic's; WeakRef registry; files excludes tsbuildinfo 06a3a20 3 failed (incl. 20 providers held after GC)

Round 2 on 06a3a20, run as separate installed cards (Opus 5.5):

  • pr-reviewer t_9581bf6d: APPROVE. R1-R7 resolved. New findings N1-N4.
  • pr-verifier t_61027f62: CHANGES_REQUESTED. It confirmed N1-N4 and R1-R7, and added V1.

Fixed in cc1a581 (test first: 8 jest cases and a new real-CLI contract test failed with the fix stashed) and in a8831d6 (docs):

ID Finding Resolution
N1 close()/asyncDispose after a successful result marked the conversation ERROR "cancelled" After a result, with no open turn, tool or subagent, it ends as returned (OK). Work still in flight stays cancelled. New real-CLI streaming_close scenario
N2 task_notification without the optional tool_use_id (sdk.d.ts:5997) left a background subagent open Falls back to the task_started task_id mapping
N3 An unknown baseline suppressed every later result in the query The first result's totals are the baseline for later results; the flag stays true
N4 resumeSessionAt used the latest saved totals; /clear edge resume/fork with resumeSessionAt has no known baseline; the /clear undercount is documented
V1 The process-local baseline can promote spend twice when one session is resumed from several processes Documented: exact while a session stays on one worker; claude_agent.cumulative.* is exact for every deployment. A cross-process store (shared state, new dependency) and dropping promotion on every resume (loses resumed spend in the common case) were both rejected

Round 3 on a8831d6 (installed pr-reviewer t_2411725e): CHANGES_REQUESTED.

  • N2, N3, N4 and V1 resolved. V1's documentation was judged adequate; a shared store would be a follow-up feature.
  • M1 (introduced by N1): the N1 shortcut keyed off resultSeen, which never resets. An abort in a later streaming turn, between model steps, was therefore recorded as OK. Fixed in a1dfe62: an idleAfterResult flag is set by a result and cleared by any later assistant or user message. Test first: 3 cases (close, asyncDispose, abort) failed on a8831d6. The README names the one remaining gap: an abort after a new prompt is sent, before the SDK yields anything for it.
  • M2 (optional): reset the in-query baseline on conversation_reset. Not taken: whether the CLI's running totals restart there is unverified. The documented /clear undercount stays.

pr-verifier t_46d2f7cf on a8831d6: CHANGES_REQUESTED. It confirmed M1 and M2, and found N2-N4 and V1 resolved. Its fix note: clear the idle state only on main-loop messages. A background subagent's messages after the result must not turn a later close() into a cancellation. Applied in 005a70f; the new test failed on a1dfe62.

Round 4 (pr-verifier t_2915cba8 on a8831d6..005a70f): M1 fixed for close, asyncDispose and AbortController. M2 is acceptable to defer. Two P3 findings were fixed test first in 4cf737f (3 new cases failed on 005a70f):

  • R4-1: a background task_notification runs its own main-thread turn (sdk.d.ts:5440), so close() after one is a cancellation again. system/session_state_changed is authoritative when the CLI sends it.
  • R4-2: a main-loop stream_event after a result starts a turn.

Round 5 (pr-verifier t_2cc44bc8 on 005a70f..4cf737f): VERIFIED with no regressions and no blocking findings. Its two P3 follow-ups are listed under Limits.

Review is complete at 4cf737f.

Video demo

A narrated terminal demo (5:46) was recorded at the reviewed head 4cf737f9d2d100118da08cda40c3242ea70ee51b. It is a real terminal recording, not a slideshow.

What it runs:

  • It drives the real Claude Code CLI from @anthropic-ai/claude-agent-sdk 0.3.289 against a local loopback mock of the Messages API. There is no live Anthropic call.
  • Spans are exported through the Future AGI exporter into the shared harness Receiver.

Where to find it: the video, captions, transcript and media check are attached privately to Linear issue TH-8235 as th-8235-demo.mp4, th-8235-demo.srt, transcript.md and media-verification.md. The mp4's sha256 is ebec6535…467d63d99.

Chapters:

  • 00:00 The problem
  • 00:36 The reviewed head (branch, SHA, 0 uncommitted files, SDK 0.3.289)
  • 01:10 A real agent run: conversation CHAIN, 2 assistant_turn LLM spans and tool.Read TOOL, exported to /tracer/v1/traces with both auth headers; usage keys on the conversation only
  • 02:14 Cost counted once across resume: 0.000180 + 0.000180 = 0.000360, which equals total_cost_usd
  • 03:07 Ending a query: close() and asyncDispose mid-tool end all 6 spans ERROR cancelled; a streaming close() right after the result ends OK
  • 03:38 The change: recordUsage
  • 04:38 Tests on camera: jest 5 suites, 101 passed; contract suite 12 passed; 0 uncommitted files after
  • 05:10 Review and limits

The demo driver is run_demo.sh + demo_spans.py. It is read-only and reuses the contract module's real-SDK runner.

Limits

  • No live Anthropic call, and no real fi-collector: the contract tests use the loopback Receiver.
  • startup() is not wrapped.
  • The usage baseline is per process. A resume after a restart writes no promoted usage keys, and a session resumed from several processes can be promoted twice; both are documented, and claude_agent.cumulative.* is exact in every case.
  • Open P3 follow-ups from verify-r5:
    • Every post-result task_notification is treated as starting a main-loop turn. The SDK types don't say whether ambient or skip_transcript tasks do; the error is on the conservative side (cancelled).
    • session_state_changed is only sent when the CLI opts in (CLAUDE_CODE_EMIT_SESSION_STATE_EVENTS), so the authoritative idle path is usually off.
  • M2 (reset the in-query baseline on conversation_reset) is deferred: it would only ever undercount, and the CLI behaviour is unverified.
  • The docs page lives in company-brain PR Dev #62. A future-agi/docs PR is not opened here.

Stacked on #203 (shared harness, includes the chunked-body fix 3eaadc8). Not merged.

Linear: TH-8235. approved: Nikhil 2026-10-03 blanket

nik13 added 4 commits October 3, 2026 18:58
The Node OTLP exporter streams with Transfer-Encoding: chunked and sends
no Content-Length. The Receiver read Content-Length only, so every Node
export came back 400 with zero spans; two TH-8103 children had to put a
de-chunking relay in front of it. Receiver now accepts both framings.

It also records one entry per accepted export (path, lower-cased
headers, flattened resource attributes) via requests(), so a contract
test can assert X-Api-Key/X-Secret-Key and project_name through the
shared harness instead of a private recorder.

Verified: 6 harness tests pass, and the TanStack example's real Node
exporter delivered 4 chunked exports (4 spans, collector path, both
auth headers, project_name/project_type) with no relay.

Refs: TH-8339, TH-8103
New package typescript/packages/traceai_claude_agent_sdk. wrapQuery()
wraps the SDK's exported query() and turns its message stream into
conversation, assistant_turn, tool_execution, mcp_tool and subagent
spans. Each span keeps the Python claude_agent.span_kind string and sets
gen_ai.span.kind and fi.span.kind. The conversation span falls back to
CHAIN because TS FISpanKind has no CONVERSATION.

The attribute names are copied from the Python _attributes.py, and a
test fails when one is missing. Tokens and cost come from the result
message. Cost is also written to gen_ai.cost.total, the key the
collector promotes. The session id goes to session.id. The wrapper also
records fork origin, provider (custom when ANTHROPIC_BASE_URL is not an
anthropic.com host), abort and errors.

Prompts, tool inputs and outputs, and assistant text are hidden unless
hideInputs/hideOutputs or FI_HIDE_* opt in. This is stricter than the
Python package.

@anthropic-ai/claude-agent-sdk (Anthropic Commercial Terms) is a peer
dependency, ^0.3.142. It is never imported at runtime or bundled. The
contract test packs the tarball and checks it holds no SDK file or
native binary.

Tests: 40 jest tests, plus shared-harness contract tests. The contract
tests send fixture journeys and the real SDK 0.3.289 / 0.3.142 CLI,
pointed at a loopback Messages mock through options.env
.ANTHROPIC_BASE_URL, through fi-core to the harness Receiver. No
Anthropic call.

pnpm-lock.yaml gets only this importer and the 15 packages/snapshots
it needs. pnpm 10.30.1 also re-resolved other importers' ts-node peers,
and those changes were left out.

Refs: TH-8235
nik13 added 11 commits October 4, 2026 16:39
…y false

resolveContentPolicy() used fi-core's rule (value.toLowerCase() === "true"),
so FI_HIDE_INPUTS=1, =yes or =" true" turned prompt and tool capture ON.
Only an explicit "false" (trimmed, any case) now opts in; every other value,
including the empty string, keeps the hidden default. The traceConfig option
still takes precedence. README documents the rule.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
The Query proxy only intercepted next/return/throw/asyncIterator.
Query.close() (sdk.d.ts:3236-3243, the SDK's abort path) and
Symbol.asyncDispose (`await using`) went straight to the original, so a
query stopped mid-run never finished its spans and the trace was never
exported. Both are now intercepted: the open spans are ended as
cancelled (ERROR, claude_agent.cancelled=true), then the call is
forwarded. After a completed stream the call changes nothing.

Tests: jest close()/asyncDispose cases on the fake Query, and a real-SDK
contract scenario (run_real_sdk.mjs SCENARIO=close) that stops two real
CLI queries mid-tool. That scenario exported 0 spans before the fix. The
real-SDK contract test now also asserts the x-secret-key header. The fake
Query gains close/asyncDispose/backgroundTasks and a per-message delay.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
… result

In a streaming-input session one query() yields a result per user turn.
The next main-loop turn fell back to the query start time
(spans.ts startTurn: nextTurnStartMs ?? this.startTimeMs), so the second
turn overlapped the first. onResult now sets the main scope's
nextTurnStartMs, the same way a user message does.

Test: a two-result streaming-input journey (new fixture, numbers taken
from the real CLI against the mock) checks turn 2 starts at or after
turn 1's end; it failed by 28 ms before the fix.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
…sk_notification

A foreground subagent moved to the background arrives as system/task_updated
with patch.is_backgrounded (sdk.d.ts:6059; task_id only), or follows an app
call to Query.backgroundTasks() (sdk.d.ts:3234). Both were ignored: the
subagent span ended at the "running in the background" tool_result and the
subagent's later turns were parented to an ended span.

- task_started records task_id -> tool_use_id; task_updated with
  patch.is_backgrounded marks that subagent background.
- The Query proxy intercepts backgroundTasks(toolUseId?): matching open
  foreground subagents are marked before the call is forwarded, and
  unmarked if it rejects or resolves false. The caller gets the original
  promise.
- A task_notification with status failed/stopped now makes the subagent
  span ERROR in the foreground case too (it ended OK before when the
  tool_result itself was not an error).

README documents background subagents.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
R1: total_cost_usd is a running total that a resumed, continued or forked
session carries forward (sdk.d.ts:5679), and the collector promotes
gen_ai.cost.total / gen_ai.usage.* on any span and Observe sums them per
trace and per session.id. The wrapper copied the cumulative value, so each
resume re-counted earlier spend (real CLI: 0.00018 then 0.00036 promoted
for 0.00036 spent).
- A process-local LRU (1000 sessions) maps session id -> last running totals.
  The promoted keys carry the cumulative value minus that baseline: zero for
  a new session, the session's own totals for resume/continue, the parent's
  for a fork. A drop within one query (/clear) counts both sides.
- No known baseline (resume after restart, continue+forkSession, or a first
  result below the saved totals): no promoted key and no
  claude_agent.cost.total_usd; claude_agent.usage.baseline_unknown=true.
- Running totals always go on unpromoted claude_agent.cumulative.* keys.
- options.continue is now labelled session.is_resumed=true / is_new=false.
- A result with nothing usable, or a zeroed crash result, writes no usage
  key and leaves the baseline alone.

R2: tokens came from result.usage, which is main-loop only and per turn in
streaming input (sdk.d.ts:5683). They now come from the latest result's
modelUsage summed over models (input, output, cache read, cache creation),
with the same delta handling. Streaming input stays one conversation span
per query() (recorded in the README; Python's ClaudeSDKClient makes one per
user turn).

Tests: usage.test.ts (21 jest: streaming totals, multi-model modelUsage,
no usage on turn/tool/subagent spans, absent usage, resume/continue/fork
deltas, unknown baselines, /clear, crash result, LRU). Real-CLI contract
scenarios resume, continue_fork, restart (two processes, shared HOME) and
streaming assert the promoted keys sum to the final cumulative totals; all
four failed before the fix. Fixture results now carry modelUsage.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
…op tsbuildinfo

R7 review nits:
- package.json repository/bugs/homepage pointed at github.com/futureagi/traceai
  (copied from strands). They now match traceai_anthropic:
  future-agi/traceAI, with this package's directory.
- The module-level knownProviders Set kept every provider passed to an
  instrumentation alive forever. It now holds WeakRefs, pruned on read and
  add, and dedupes a provider set twice. shutdown() with no argument still
  flushes every live provider.
- The tarball shipped dist/**/*.tsbuildinfo; files now excludes it.

Tests: package.test.ts (metadata against the sibling; 20 dropped providers
are released after a forced GC, failed with 20 held before; a live
provider is tracked once). The pack contract test rejects *.tsbuildinfo
(failed on dist/esm/tsconfig.esm.tsbuildinfo before).

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
…ask_id notifications, partial baselines

N1: close()/asyncDispose/abort after a result with no open turn, tool or
subagent ends the query as returned (conversation OK), the normal end of a
streaming-input session. A close with work still in flight stays cancelled.
N2: task_notification without tool_use_id (optional, sdk.d.ts:5997) is
matched through the task_started task_id mapping.
N3: an unknown usage baseline covers only the first result; later results
in the same query promote their delta against it (flag stays true).
N4: resume/fork with resumeSessionAt has no known baseline. README notes
the /clear undercount edge.

Tests first: 8 jest cases and the new real-CLI streaming_close contract
test failed on 06a3a20 and pass now.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
…ion stays in one process

The usage baseline is process-local. Resuming one session from several
processes can promote spend made in another process again. Documented, with
claude_agent.cumulative.* as the exact source for those deployments.
Rejected: a cross-process store (shared state and a new dependency) and
dropping promotion on every resume (loses resumed spend in the common
single-process case).

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
… result ends OK on close

The N1 shortcut keyed off resultSeen, which never resets, so an abort in a
later streaming turn between model steps (turn span ended, next assistant
message not yet out) was recorded as OK. Track idleAfterResult instead: a
result sets it, any later assistant/user message clears it. close(),
asyncDispose and AbortController aborts mid-turn are cancelled again.
README names the residual gap (abort after a new prompt is sent, before the
SDK yields anything for it). Comments in queryWrapper.ts refreshed.

M2 (reset the in-query baseline on conversation_reset) not taken: whether the
CLI's running totals restart there is unverified; the README keeps the
documented /clear undercount.

Test first: 3 new jest cases (close, asyncDispose, abort) failed on a8831d6.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
… the idle-after-result state

Background subagent messages (parent_tool_use_id set) that arrive after the
main loop's result no longer clear idleAfterResult, so close() after that
subagent's task_notification ends the conversation OK; a subagent still
open stays cancelled through nothingInFlight().

Test first: the new post-result background subagent case failed on a1dfe62.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
…ssion state drive the idle-after-result state

R4-1: a background task_notification is queued as a main-thread command
that runs its own turn (sdk.d.ts:5440, 4603), so close() after one is a
cancellation again (supersedes 005a70f's expectation). R4-2: a main-loop
stream_event (includePartialMessages) after a result also starts a turn.
system/session_state_changed is authoritative when sent (sdk.d.ts:5858):
'idle' after a result ends OK, 'running' cancelled. README updated.

Test first: 3 new jest cases failed on 005a70f.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8235
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant