Skip to content

feat(ag2): traceAI-ag2 for AG2 1.x via native TelemetryMiddleware (TH-8236) - #210

Open
nik13 wants to merge 15 commits into
devfrom
feat/th-8236-ag2
Open

nik13 wants to merge 15 commits into
devfrom
feat/th-8236-ag2

Conversation

@nik13

@nik13 nik13 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

New Python package traceAI-ag2 for AG2 1.x (TH-8236). It is native-first: it attaches AG2's own TelemetryMiddleware with capture_content=False and adds AG2SpanProcessor ahead of the Future AGI exporter. It does not wrap Agent, and it does not depend on autogen or autogen-agentchat.

  • Kinds: invoke_agent, chat and execute_tool map to AGENT, LLM and TOOL. await_human_input and record_usage spans get no kind rather than a guessed one.
  • The three non-semconv AG2 usage keys are copied to their dotted GenAI names; the originals are kept.
  • Token double-count fixed. record_usage model_call repeats the chat span's tokens (ag2 1.1.2 telemetry.py 437-440 vs 501-504).
    • fi-collector promotes gen_ai.usage.* into the token columns on any span (fi-collector/pkg/adapter/adapter.go), and Observe sums total_tokens over every span of a trace (tracer/services/clickhouse/v2/span_reader.py, tracer/tasks/session.py).
    • The processor moves the promoted keys to ag2.usage.* only where they would be counted twice. On model_call spans the chat span already carries the same tokens. A subtask rollup is moved only when that worker is instrumented in the same trace (agent names tracked per trace, bounded). aggregation keeps its tokens. A real-AG2 test checks that the trace total equals AG2's UsageReport.
  • With hide_inputs / hide_outputs, tool arguments, tool results and human-input text are also dropped (TraceConfig.mask does not cover them). hide_input_text / hide_output_text drop AG2's JSON message attributes; the README lists which flags apply.
  • A later setup(config=...) replaces the installed config and logs a warning. on_start copies traceAI context (using_session etc.) onto AG2 spans.
  • AG2SpanProcessor stays first after a later add_span_processor call: fi's call resets the chain, so the processor is lifted out and put back first. A ConcurrentMultiSpanProcessor gets a warning. Normalized spans keep BoundedAttributes and the dropped-attribute count. At max_span_attributes (128 by default), an added key pushes out the oldest one and the drop is counted (README, CHANGELOG).

Deviations from the accepted architecture (with evidence)

  • setup(*agents, ...) takes the agents. AG2 has no process-wide middleware registry: middleware attaches per agent (plugin.py:176 add_middleware) or per call. create_telemetry_middleware() covers agents built later.
  • Version floor is ag2>=1.0.0, not >=1.1.2. TelemetryMiddleware is present from 1.0.0 (checked across every 1.x wheel). max_tool_result_chars is forwarded only when set (it exists from 1.0.4).
  • Session: AG2 emits no conversation id. The app can set span_attributes={"session.id": ...}, and that path is tested.
  • Requires fi-instrumentation-otel>=1.1.0. 0.1.x lacks the span-kind constants, and published 1.0.0 crashes in _headers (5 tests fail on it).

Tests (run on d64de18, clean tree, by the coordinator)

PYTHONPATH="python/frameworks/ag2:python:python/tests" uv run --no-project --python <py> --with pytest --with opentelemetry-api --with opentelemetry-sdk --with opentelemetry-exporter-otlp-proto-http --with wrapt --with requests --with jsonschema --with "ag2==<v>" pytest python/frameworks/ag2/tests -q -p no:cacheprovider --noconftest -o addopts= -rs
ag2 1.1.2, Python 3.10: 97 passed
ag2 1.1.2, Python 3.11: 97 passed
ag2 1.1.2, Python 3.12: 97 passed
ag2 1.1.2, Python 3.13: 97 passed
ag2 1.0.3, Python 3.11: 94 passed, 3 skipped
ag2 1.0.0, Python 3.11: 88 passed, 9 skipped

No file changed outside python/frameworks/ag2, and no secret-like lines were added.

Review

Round 1, an independent fresh-context review of ec9fc07 (a subagent, not an installed card): CHANGES_REQUESTED. Fixed test first on top of ec9fc07. All final tests run against ec9fc07 give 23 failed / 65 passed; 21 of those are real assertion failures.

ID Finding Resolution Commit
R1 (blocking) aggregation tokens wrongly moved off Kept; the real-AG2 trace total equals UsageReport 00802f3
R2 (blocking) subtask rollup counted twice when the worker is instrumented Moved only when the worker is instrumented in the same trace 00802f3
R3 fi floor 1.0.0 crashes fi-instrumentation-otel>=1.1.0 30479c4
R4 A later setup() config ignored Replaces the config, with a warning fbaac8e
R5 traceAI context not applied on_start copies it; context session.id wins fbaac8e
R6 Text-hide flags leave AG2 JSON content Whole content attributes dropped (7 failed before) 4a9922e
R7 Processor order and attribute limits Stays first after a later add_span_processor (2 failed before); BoundedAttributes kept (2 failed before) c1196a4, 9c2c1bf

Round 2: installed pr-reviewer, then pr-verifier, on 9c2c1bf. Verdicts will be added here.

Round 2 follow-ups (9c2c1bf..d64de18)

The installed review passed this PR at 9c2c1bf: pr-reviewer t_0d642613 APPROVE, pr-verifier t_62f6dbc6 VERIFIED. Its non-blocking findings were still fixed before the demo, test-first. Run against 9c2c1bf, the new tests give 4 failures on ag2 1.1.2 and 3 on 1.0.0.

ID Finding Resolution
N1 A second provider-less setup() registered a second provider and counted subtask spend twice setup() registers once and reuses that provider until it is shut down (2d7cd80); the failing case was (91, 15) vs (51, 11) tokens
N3 session.id was the first key evicted at the attribute limit traceAI context keys are moved after AG2's keys before re-bounding (2ffef0a)
N5 PRD J4 (middleware short-circuit) neither covered nor disclosed README Limits and CHANGELOG: an outer short-circuit has no span. AG2's own halt gives an UNSET chat span when passed per call through Plugin (1.1.2), and no span with Agent(assembly=...); both are tested (c0a2015)
N6 README version/attribute claims not matched by tests Only versions that ran are listed; gen_ai.request.parameters is documented as never emitted by ag2 (d64de18)
M1, N2, N4 Namesake dedupe, unknown agent name, concurrent first-call race README Limits

Round 3 (installed pr-verifier t_3f913b48 at d64de18) returned VERIFIED, ready to merge.

  • N1 and N3 are fixed in code; N5, M1, N2 and N4 are disclosed in the README.
  • N6 was partly fixed: the README's tested-version line named runs that weren't in the supplied evidence. The coordinator has now run them at d64de18 on a clean tree: ag2 1.0.3 on Python 3.11 gave 94 passed, 3 skipped; ag2 1.1.2 on Python 3.10 and 3.12 gave 97 passed each. The line now stands as written.
  • No regressions.

Review is complete at d64de18. One P3 follow-up remains, for after merge:

  • Provider reuse covers only a provider that setup() registered itself. setup(a, tracer_provider=P) followed by setup(b) still registers a second provider. Agents set up before a provider shutdown keep that provider's middleware on a later setup(). One README sentence would cover both.

Video demo

Narrated terminal demo, 6:28, recorded at head d64de18 (the reviewed head), 1080p H.264/AAC, 9 chapters. sha256 41a51914bff49cee8abc49ecdecc7305782dcb22f18b9ce08551e4e09086664b.

The video, captions, transcript, chapter list, preview and media check are attached privately to Linear TH-8236 (Future AGI workspace access needed). Every run uses AG2 1.1.2 with a fake model, placeholder keys and the shared harness receiver: no vendor call and no live fi-collector.

Time Chapter
00:00 The problem
00:49 Head d64de18: branch, 0 uncommitted files, 12 commits on 3eaadc8
01:26 Run: the example in a separate process; 1 export to /tracer/v1/traces with both auth headers (values not shown); 5 spans in 1 trace with AGENT / LLM / TOOL kinds; promoted tokens 11/7 over the trace = one model call counted once; no content keys and no prompt, answer or tool text by default
02:14 Session: inside using_attributes, session.id=demo-session-42 and user.id=demo-user-7 on all 5 spans
02:42 Subtask before/after: at 9c2c1bf, 2 providers and 2 exports, trace tokens 91/15 against an AG2 UsageReport of 51/11 (the worker's spend counted twice); at the head, one provider, 1 export, 51/11
03:31 Limits before/after: at 2d7cd80, with max_span_attributes=14 the LLM span loses session.id and user.id; at the head the 4 drops fall on AG2's own keys and both are kept
04:13 The change: provider reuse in _setup.py (2d7cd80) and the eviction order in _processor.py (2ffef0a)
05:19 Tests: 97 passed on camera (Python 3.11, ag2 1.1.2), 0 uncommitted files after, plus the logged matrix at this head
05:56 Review and limits

Under recording load VHS dropped frames, so each shot holds its final verified frame under the narration (no retiming). Checked by the coordinator: sha256, full decode, chapter markers, pauses under 3 s, frames per chapter against the driver output, no keys or header values in frames, captions or transcript.

Limits

  • No live model call; the model is AG2's TestConfig.
  • The AG2 network hub was not run. The traceparent hand-off is simulated the way network/client/handlers.py:255 does it.
  • The docs draft is in company-brain PR Dev #62; the corrected setup() call is in the README.

Stacked on #203. Not merged.

Linear: TH-8236. approved: Nikhil 2026-10-03 blanket

nik13 added 5 commits October 3, 2026 18:58
The Node OTLP exporter streams with Transfer-Encoding: chunked and sends
no Content-Length. The Receiver read Content-Length only, so every Node
export came back 400 with zero spans; two TH-8103 children had to put a
de-chunking relay in front of it. Receiver now accepts both framings.

It also records one entry per accepted export (path, lower-cased
headers, flattened resource attributes) via requests(), so a contract
test can assert X-Api-Key/X-Secret-Key and project_name through the
shared harness instead of a private recorder.

Verified: 6 harness tests pass, and the TanStack example's real Node
exporter delivered 4 chunked exports (4 spans, collector path, both
auth headers, project_name/project_type) with no relay.

Refs: TH-8339, TH-8103
New package python/frameworks/ag2 (traceAI-ag2, import traceai_ag2) for
AG2 1.x (PyPI ag2, import ag2). Native-first per D-8236:

- setup(*agents, tracer_provider=None, capture_content=False, ...) registers
  the Future AGI exporter (register(project_type=OBSERVE) when no provider is
  passed) and attaches AG2's own TelemetryMiddleware to each agent through the
  public Agent.add_middleware, overriding the upstream capture_content=True
  default. Idempotent per agent. Agent is not wrapped or patched.
- create_telemetry_middleware() for agents built later or per-call middleware.
- AG2SpanProcessor, prepended ahead of the exporter, on the
  opentelemetry.instrumentation.ag2 scope only: sets gen_ai.span.kind from
  gen_ai.operation.name (chat/execute_tool/invoke_agent), copies
  cache_creation_input_tokens, cache_read_input_tokens and thinking_tokens to
  the dotted GenAI semconv keys (originals kept), and applies TraceConfig as a
  second content gate.
- ag2>=1.0.0 (first stable 1.x with TelemetryMiddleware), no dependency on
  autogen or autogen-agentchat; fi-instrumentation-otel>=1.0.0.

Tests: processor unit tests, setup tests against real ag2 with its scripted
TestConfig model, import isolation (no autogen), wire-level resource/header
test, and shared-harness Receiver contract tests over real OTLP/HTTP.

Refs: TH-8236
record_usage model_call repeats the chat span's input/output tokens
(ag2 1.1.2 telemetry.py 437-440 vs 501-504). fi-collector promotes
gen_ai.usage.* into the token columns on any span (adapter.go
inputTokenKeys/outputTokenKeys/totalTokenKeys) and Observe sums
total_tokens over every span of a trace (span_reader v2, tasks/session.py),
so the trace total doubled.

On model_call and aggregation usage spans the processor now moves the
promoted keys to ag2.usage.*. Values are kept. subtask and compaction
usage spans keep their tokens: no chat span repeats them.

Tests first: the contract test now asserts the summed promoted input
tokens equal one model call; 2 new processor tests. 52 passed on ag2
1.1.2, 50 passed + 2 skipped on 1.0.3.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
nik13 added 10 commits October 4, 2026 16:51
Review of #210 (reproduced with real ag2 1.1.2) found two token-total
errors against AG2's own UsageReport:

- aggregation: ag2 aggregate.py 97-99 / 193-195 call the model client on a
  throwaway Context(MemoryStream()) that never reaches on_llm_call, so the
  record_usage aggregation span is the only copy of that spend. Demoting it
  dropped it from the trace (11 promoted vs 111 actual input tokens).
  'aggregation' is removed from the demoted kinds; compaction stays kept.
- subtask: the rollup (run_task.py 63-93, label = sub-agent name) repeats
  the sub-agent's chat spans when that sub-agent is itself instrumented
  (91 vs 51). AG2SpanProcessor now records agent names from ended
  invoke_agent spans per trace in a bounded LRU map (1024 traces) and moves
  a subtask rollup's promoted tokens to ag2.usage.* only when its
  ag2.usage.label names an agent seen in the same trace. An uninstrumented
  worker's rollup keeps its tokens (the only copy).

Comment, module docstring, README usage table and CHANGELOG describe the
moves per usage kind.

Tests first, red on ec9fc07 (7 failed): new tests/test_usage_totals.py runs
real AG2 with the scripted TestConfig model and asserts summed promoted
input and output tokens equal UsageReport for aggregation, instrumented and
uninstrumented subagent_tool workers, a namesake in another trace, and one
run covering main loop + compaction + aggregation + subtask; synthetic
processor tests for aggregation/compaction kept, subtask demoted/kept and
the bounded map. Mutants "demote every subtask" and "agent names not
scoped per trace" each turn guard tests red. 63 passed on ag2 1.1.2,
61 passed + 2 skipped on 1.0.3.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
Published fi-instrumentation-otel 1.0.0 crashes in register() with
opentelemetry-exporter-otlp-proto-http 1.45.0 (otel.py:289
AttributeError: 'HTTPSpanExporter' object has no attribute '_headers').

Evidence, this branch's suite with PYTHONPATH=python/frameworks/ag2:
python/tests (no in-repo fi_instrumentation) and ag2 1.1.2:
- --with fi-instrumentation-otel==1.0.0: 5 failed, 58 passed
  (3x the _headers AttributeError, 1x register() got an unexpected
  keyword argument 'timeout')
- --with fi-instrumentation-otel==1.1.0: 63 passed

Floor raised in pyproject.toml with the reason; README and CHANGELOG note it.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
R4: a second setup(..., config=TraceConfig(hide_inputs=True,
hide_outputs=True)) on a provider that already had an AG2SpanProcessor
was silently ignored, so content was exported despite the hide request
(reviewer probe: gen_ai.input/output.messages still on the chat span).
install_span_processor now applies an explicit config to the installed
processor (AG2SpanProcessor.update_config; most recent call wins) and
logs a warning when it changes. config=None keeps the installed config,
so a later setup() never re-exposes content.

R5: on_start did nothing, so using_session / using_user / using_metadata /
using_attributes were ignored on AG2 spans. on_start now copies
fi_instrumentation get_attributes_from_context() onto AG2-scope spans
without overriding keys AG2 already set (later AG2 set_attribute calls
still win), except session.id: the context session is re-applied in
on_end so it beats a static span_attributes value. Pending sessions are
kept in a bounded map keyed by (trace_id, span_id) and released on end.
Foreign-scope spans are untouched.

README (setup, sessions, conformance table), setup()/processor docstrings
and CHANGELOG updated.

Tests first, red on ec9fc07: test_later_setup_config_replaces_installed_
processor_config, test_using_session_sets_session_on_every_ag2_span,
test_using_attributes_context_reaches_ag2_spans,
test_context_session_wins_but_other_ag2_keys_are_not_overridden,
test_pending_context_sessions_are_bounded_and_released_on_end (also red
against a no-eviction mutant: 40 == 16). Guards that stay green:
test_later_setup_without_config_keeps_installed_config,
test_context_attributes_skip_foreign_scope_spans. 70 passed on ag2 1.1.2.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
R6: AG2 records message content as one JSON string per span
(gen_ai.input.messages / gen_ai.output.messages, telemetry.py 477/517 at
1.1.2), never as the flattened gen_ai.input.messages.{i}.message.content
keys that TraceConfig.mask's hide_input_text / hide_output_text /
hide_input_images rules match. With capture_content=True, those flags
left every prompt and completion in place, while the module docstring
and README claimed "the other TraceConfig.mask rules" applied.

The processor now drops whole attributes (privacy-safe; the text inside
the JSON cannot be redacted on its own):
- hide_inputs / hide_input_text: gen_ai.input.messages,
  gen_ai.system_instructions, gen_ai.tool.call.arguments,
  ag2.human_input.prompt
- hide_outputs / hide_output_text: gen_ai.output.messages,
  gen_ai.tool.call.result, ag2.human_input.response
- hide_input_messages: gen_ai.input.messages, gen_ai.system_instructions
- hide_output_messages: gen_ai.output.messages
hide_input_images and hide_embedding_vectors have nothing to act on (AG2
serialises only text parts and records no embeddings); a real-AG2 test
pins that image bytes never reach a span. Module docstring, README
(privacy table of which flags apply) and CHANGELOG say so.

Tests first, red on fbaac8e (7 failed, 75 passed; final test files copied
into a detached fbaac8e worktree): test_trace_config_flags_on_ag2_json_
content[hide_input_text, hide_output_text, hide_inputs,
hide_input_messages, hide_input_text+hide_output_text],
test_text_hide_flags_remove_ag2_message_json and
test_hide_input_text_keeps_outputs (real AG2). Passing on fbaac8e by
design: the hide_outputs / hide_output_messages params,
test_flags_with_nothing_to_act_on_leave_ag2_content,
test_ag2_never_records_image_bytes. Green: 82 passed on ag2 1.1.2;
75 passed, 7 skipped on ag2 1.0.0.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
R7: the provider from fi_instrumentation.register() keeps
_default_processor=True, so the application's first
provider.add_span_processor() call after setup() shut down and cleared
the whole chain (fi_instrumentation/otel.py add_span_processor,
lines 329-340), AG2SpanProcessor included. The new exporter then got
un-normalized AG2 spans (no gen_ai.span.kind, no usage moves, no
TraceConfig content gate).

install_span_processor now wraps the provider instance's
add_span_processor once (marker attribute): the AG2 processor is lifted
out of the chain for the call, so fi's reset neither shuts it down nor
drops it, and is put back first afterwards. provider.shutdown() still
shuts it down. Installing on a provider whose active processor is a
ConcurrentMultiSpanProcessor logs a warning, since parallel on_end calls
do not guarantee normalization before export. Docstring, README (set-up
note, troubleshooting) and CHANGELOG updated.

Tests first, red on 4a9922e (2 failed, 2 passed):
test_later_add_span_processor_on_fi_provider_keeps_ag2_processor_first
(chain was just the new SimpleSpanProcessor) and
test_install_warns_when_processors_run_concurrently. Passing on
4a9922e by design: test_later_add_span_processor_on_sdk_provider_keeps_
ag2_processor_first and test_provider_shutdown_still_shuts_the_processor_
down; mutant "do not put the processor back" turns both later_add tests
red. Green: 86 passed on ag2 1.1.2; 79 passed, 7 skipped on ag2 1.0.0.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
…spans

R7: AG2SpanProcessor.on_end replaced span._attributes with a plain dict.
ReadableSpan.dropped_attributes (exported as dropped_attributes_count)
reads BoundedAttributes.dropped and returns 0 for anything else, so
every normalized AG2 span reported 0 dropped attributes, and the span's
max_span_attributes / max_attribute_length limits and its ended-span
immutability were lost.

on_end now rebuilds a BoundedAttributes with the previous maxlen,
max_value_len and immutability, and carries the previous dropped count
over. Keys the processor adds count against the same limit: on a span
already full, the oldest key is evicted and counted, as the SDK does
(only reachable at the limit, 128 by default). Docstring, README note
and CHANGELOG updated.

Tests first, red on c1196a4 (2 failed):
test_normalizing_keeps_the_sdk_dropped_attribute_count (assert 0 == 1
on span.dropped_attributes) and
test_keys_the_processor_adds_count_against_span_limits (plain dict kept
3 keys over a limit of 2). Green: 88 passed on ag2 1.1.2.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
…g again

setup() without tracer_provider called fi register() on every call, so
setup(planner) then setup(worker) put each agent on its own provider with
its own AG2SpanProcessor. The planner's processor never saw the worker's
invoke_agent span, kept the subtask rollup's promoted tokens, and the trace
counted the worker's spend twice (real AG2: 91/15 input/output tokens
against a UsageReport of 51/11).

setup() now remembers the provider it registered and reuses it on later
calls without tracer_provider until that provider is shut down. A later
config= still replaces the installed TraceConfig; a different project_name
is ignored with a warning. README and CHANGELOG updated.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
…bute limits

on_start writes the traceAI context keys (session.id, user.id, metadata, ...)
before AG2 sets any attribute, so they are the oldest keys on every AG2 span.
When the keys the processor adds (span kind, usage aliases) overflow
max_span_attributes, BoundedAttributes evicts oldest-first, so session.id
and user.id went first and the span dropped out of its session's totals.

on_end now moves the context keys behind AG2's keys before re-bounding, so
AG2's oldest keys are evicted instead. The eviction count is unchanged.
Test: real AG2 chat span with using_attributes at a lowered limit.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
…tured

setup() adds TelemetryMiddleware as the innermost middleware. Observed with
real AG2 (1.1.2 and 1.0.0) and now pinned by tests:
- a middleware registered before setup() that answers on_llm_call without
  calling call_next leaves no chat span (invoke_agent ends UNSET);
- one that answers on_turn leaves no span at all;
- AG2's halt (HaltEvent from AlertPolicy) with Agent(assembly=...) policies
  leaves no chat span, because the halt check is added before telemetry;
- with per-call Plugin policies (ag2 1.1.2 only) the halted call gets a chat
  span with status UNSET and no attribute carrying the halt reason.

No capture is added. README gains a Limits section; CHANGELOG notes it.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
- Tested-version line names only what the suite ran: ag2 1.0.0, 1.0.3 and
  1.1.2 on Python 3.11, and ag2 1.1.2 on Python 3.10, 3.12 and 3.13.
- hide_llm_invocation_parameters has nothing to act on: AG2 1.0.0 and 1.1.2
  set no gen_ai.request.parameters. A drift guard now checks AG2's
  telemetry source for that key; the processor docstring says the same.
- Limits: the subtask dedupe matches by agent name within a trace (M1);
  create_telemetry_middleware() defaults agent_name=None ("unknown") and
  tracer_provider=None (OTel global provider) (N2); concurrent first
  add_span_processor calls can race with fi's unlocked reset (N4).

approved: Nikhil 2026-10-03 blanket
Refs: TH-8236
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant