Skip to content

[FEATURE] Workflow executor emits no OpenTelemetry spans (workflow.run/workflow.node) #3871

Description

@afiodorov

Problem Statement

The workflow executor (src/workflow/executor/step-executor.js) never opens an OTel span for a run or for individual steps — confirmed via grep -rl "withSpan(" src/workflow/ returning nothing. Meanwhile the agent runtime (src/agent/runtime/index.js) does instrument itself (agent.generate, agent.tool_execute, etc.), so with @veryfront/ext-observability-opentelemetry registered you get a scatter of agent-internal spans with no workflow structure above them — no way to see which node a tool call belongs to, how long the whole run took, or a step's resolved input/output.

The only hook exposed for this is onStepStart(nodeId, input) / onStepComplete(nodeId, output) on StepExecutorConfig, and it's not enough to build proper tracing without hacks:

  • onStepStart fires synchronously inside the step's async execution scope, but onStepComplete fires outside it (after the awaited body returns) — so a caller can't hold "current span" via normal context.with()/context-propagation; we had to reach into the OTel context manager's private _asyncLocalStorage field and call enterWith() directly just to get agent spans nested under the right workflow node.
  • Neither hook passes a run id, span, or context — just a bare nodeId. See to see how we can achieve things now - and it's very hacky
    workflow-client.ts

Ask: instrument the workflow executor itself — a workflow.run span around the whole run and a workflow.node span per step, using the tracer already installed by the OTel extension — so this is available out of the box instead of requiring a userland AsyncLocalStorage/enterWith workaround.

Proposed Solution

Workflow should have OpenTelemetry spans so that it can be traces with the unified mechanism as agent with tools.

Component

Workflows

Alternatives Considered

No response

Example Use Case

I want to spin up jaeger ui and trace entire workflow execution not just the agent with tools.

Verification

  • I have searched the existing issues to make sure this is not a duplicate
  • I have provided sufficient context for the team to understand the request

Activity

kwakayama commented on Aug 20, 2026

@kwakayama
Contributor

This was generated by AI during triage.

Thanks for this one — it was an unusually good report. The hook asymmetry you described was real (we verified it), and the workflow-client.ts you attached is what made the root cause obvious: reaching into the context manager's private _asyncLocalStorage and calling enterWith() is only necessary because nothing in the executor was opening a span for the agent call to attach to. That pointed straight at the fix.

This is shipped. The workflow executor now emits OpenTelemetry spans via the tracer the OTel extension installs — no userland wiring.

What you get

  • A workflow.run span per execution.
  • A workflow.node <id> span per node.
  • Agent spans nest beneath the node that produced them through ordinary context propagation. Your enterWith() workaround should be deletable — that was the actual goal here, so please let us know if anything in your setup still needs it.

The span opens at the DAG dispatch layer rather than in the step path, so every node type gets one: step, parallel, map, branch, wait, subWorkflow, loop.

One note on your hook suggestion: onStepStart / onStepComplete signatures are unchanged. The asymmetry you found is real, but the span makes it moot rather than requiring a signature change, and we would rather not break that contract for everyone if the tracing need is met another way.

Attributes

Attribute Notes
workflow.run_id Always the root persisted run id, on every span
workflow.id
workflow.node.id / .type / .status / .attempts / .skipped
workflow.sub_run_id / workflow.sub_workflow_id On subWorkflow nodes

A failed node or run sets span status to ERROR, so the standard errored-span filters in Jaeger, Tempo and Datadog work. Retries add a workflow.node.retry span event carrying workflow.node.attempt, workflow.node.retry_delay_ms and workflow.node.error_type — a bounded classification, not a message.

Things we did differently than you asked, and why

No resolved input/output on spans. You asked for a step's resolved input and output as attributes. We deliberately did not do this. Workflow inputs routinely carry customer data, and spans get exported to third-party backends, so no span in this codebase carries payload content — sizes and identifiers only. If you need the payloads, they belong in your own sink where you control retention and residency, not on a span.

One trace per execution attempt, not per run. A run that pauses at a wait node or a pending approval and later resumes produces a separate trace from its first execution. Correlate with workflow.run_id, which is on every span in every attempt. Span links across resumes are a plausible future upgrade, but they are not there today.

workflow.run is a trace root only when nothing traces the caller. Started from an instrumented HTTP handler or webhook it becomes a child of that request's span and joins its trace — usually what you want. Note the request span typically finishes before the run does, since execution is dispatched without being awaited.

Cardinality is on you to bound. Node ids are used verbatim in span names, and generated children get generated ids (<map>_0, <map>_1, …), so a map over a large collection produces one span and one distinct span name per item. Backends that aggregate by span name (Tempo's metrics generator, for one) will see a series per item. There is no framework-level cap; extensions/ext-observability-opentelemetry/README.md covers this with pointers to OTEL_BSP_MAX_QUEUE_SIZE / OTEL_BSP_MAX_EXPORT_BATCH_SIZE and collector-side tail sampling.

Status

Both parts are merged. The instrumentation landed first; a follow-up tightened what a failed
span reports, because marking failures as ERROR initially forwarded the thrown error's own
message — a raw error string on the wire, the same class of thing we avoided with inputs and
outputs. Failed spans now carry a bounded classification instead: the node span reports which
node failed, the run span reports a classification such as VeryfrontError:500. The status is
still ERROR and workflow.node.status is still failed, so errored-span queries are
unaffected; only the free text is gone.

That one was found by exporting a real workflow over OTLP to a local collector rather than
asserting in-process — the attribute sanitiser redacts key-shaped tokens but passes ordinary
prose through, so an email in an error message reached the wire while an API key in the same
string did not.

Closing this as implemented. If the workaround is still needed anywhere in your workflows after this, reopen or file a new issue with the shape of the graph and we will look.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions