Skip to content

fix(core): Attribute errors to the span they escaped - #23666

Open
logaretm wants to merge 7 commits into
developfrom
awad/error-span-attribution
Open

fix(core): Attribute errors to the span they escaped#23666
logaretm wants to merge 7 commits into
developfrom
awad/error-span-attribution

Conversation

@logaretm

@logaretm logaretm commented Aug 27, 2026

Copy link
Copy Markdown
Member

Errors are now attributed to the span they were thrown in, not whichever span happened to be active when captureException ran. startSpan records the escaping error's span in a WeakMap and we prefer that when building the event.

It only applies inside the error's own trace, this was flagged as an unlikely edge case by the clanker so I decided to add a sanity check for it regardless of how unlikely it is.

We seemed to have tests asserting the old behavior, so those needed to change to match the new one.

closes #16206

@github-actions

github-actions Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

size-limit report 📦

⚠️ Warning: Base artifact is not the latest one, because the latest workflow run is not done yet. This may lead to incorrect results. Try to re-run all tests to get up to date results.

Path Size % Change Change
@sentry/browser 28.88 kB +0.29% +82 B 🔺
@sentry/browser - with treeshaking flags 27.18 kB +0.26% +69 B 🔺
@sentry/browser - with treeshaking flags tracing without tracing 27.07 kB +0.24% +64 B 🔺
@sentry/browser (incl. Tracing) 49.28 kB +0.13% +60 B 🔺
@sentry/browser (incl. Tracing + Span Streaming) 49.29 kB +0.14% +68 B 🔺
@sentry/browser (incl. Tracing, Profiling) 52.21 kB +0.17% +84 B 🔺
@sentry/browser (incl. Tracing, Replay) 88.83 kB +0.08% +71 B 🔺
@sentry/browser (incl. Tracing, Replay) - with treeshaking flags 78.01 kB +0.08% +58 B 🔺
@sentry/browser (incl. Tracing, Replay with Canvas) 93.51 kB +0.08% +69 B 🔺
@sentry/browser (incl. Tracing, Replay, Feedback) 106.46 kB +0.08% +84 B 🔺
@sentry/browser (incl. Feedback) 46.38 kB +0.19% +85 B 🔺
@sentry/browser (incl. sendFeedback) 33.94 kB +0.22% +74 B 🔺
@sentry/browser (incl. FeedbackAsync) 39.04 kB +0.18% +68 B 🔺
@sentry/browser (incl. Metrics) 29.9 kB +0.26% +76 B 🔺
@sentry/browser (incl. Logs) 30.16 kB +0.24% +70 B 🔺
@sentry/browser (incl. Metrics & Logs) 30.83 kB +0.27% +82 B 🔺
@sentry/react 30.63 kB +0.24% +73 B 🔺
@sentry/react (incl. Tracing) 51.63 kB +0.14% +70 B 🔺
@sentry/vue 36.13 kB +0.22% +77 B 🔺
@sentry/vue (incl. Tracing) 51.56 kB +0.17% +83 B 🔺
@sentry/svelte 28.9 kB +0.23% +66 B 🔺
CDN Bundle 30.62 kB +0.22% +67 B 🔺
CDN Bundle (incl. Tracing) 49.85 kB +0.23% +114 B 🔺
CDN Bundle (incl. Logs, Metrics) 32.89 kB +0.21% +67 B 🔺
CDN Bundle (incl. Tracing, Logs, Metrics) 51.81 kB +0.23% +115 B 🔺
CDN Bundle (incl. Replay, Logs, Metrics) 73.55 kB +0.1% +70 B 🔺
CDN Bundle (incl. Tracing, Replay) 87.37 kB +0.1% +84 B 🔺
CDN Bundle (incl. Tracing, Replay, Logs, Metrics) 89.3 kB +0.14% +120 B 🔺
CDN Bundle (incl. Tracing, Replay, Feedback) 93.31 kB +0.11% +96 B 🔺
CDN Bundle (incl. Tracing, Replay, Feedback, Logs, Metrics) 95.32 kB +0.12% +114 B 🔺
CDN Bundle - uncompressed 90.71 kB +0.28% +253 B 🔺
CDN Bundle (incl. Tracing) - uncompressed 148.51 kB +0.24% +343 B 🔺
CDN Bundle (incl. Logs, Metrics) - uncompressed 97.29 kB +0.28% +262 B 🔺
CDN Bundle (incl. Tracing, Logs, Metrics) - uncompressed 154.48 kB +0.23% +343 B 🔺
CDN Bundle (incl. Replay, Logs, Metrics) - uncompressed 226.55 kB +0.12% +262 B 🔺
CDN Bundle (incl. Tracing, Replay) - uncompressed 268.1 kB +0.13% +343 B 🔺
CDN Bundle (incl. Tracing, Replay, Logs, Metrics) - uncompressed 274.05 kB +0.13% +343 B 🔺
CDN Bundle (incl. Tracing, Replay, Feedback) - uncompressed 281.81 kB +0.13% +343 B 🔺
CDN Bundle (incl. Tracing, Replay, Feedback, Logs, Metrics) - uncompressed 287.75 kB +0.12% +343 B 🔺
@sentry/nextjs (client) 54.07 kB +0.13% +68 B 🔺
@sentry/sveltekit (client) 49.74 kB +0.17% +82 B 🔺
@sentry/core/server 37.08 kB +0.25% +89 B 🔺
@sentry/core/browser 13.66 kB +0.82% +111 B 🔺
@sentry/node 127.96 kB +0.14% +170 B 🔺
@sentry/node/import (ESM hook with diagnostics-channel injection) 81.61 kB - -
@sentry/node - without tracing 88.86 kB +0.17% +149 B 🔺
@sentry/node - without channel injection 107.21 kB +0.17% +181 B 🔺
@sentry/aws-serverless 97.24 kB +0.16% +148 B 🔺
@sentry/cloudflare (withSentry) - minified 202.42 kB +0.23% +445 B 🔺
@sentry/cloudflare (withSentry) 503.91 kB +0.25% +1.23 kB 🔺

View base workflow run

@logaretm
logaretm marked this pull request as ready for review August 27, 2026 02:47
@logaretm
logaretm requested review from Lms24 and mydea August 27, 2026 02:47
*/
const escapedSpanTraceContexts = new WeakMap<object, TraceContext>();

function toKey(error: unknown): object | undefined {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

l: maybe add a comment here why we do this

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

without a bit more context, from the outside it seems as if we're converting the error to something else, which I was confused about initially 😅

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I renamed the fn to be clearer and added a comment

/**
* The trace context of the span an error escaped, keyed by the error itself.
*
* We store the plain trace context rather than the span, so that an error object cannot keep a

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

l/m: is this necessary? Since this is a weakmap, should this be garbage collected anyhow?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think the implementation is fine but I would change the comment - it is not necessary to keep this as trace context for gc reasons, but we apply the trace context directly later so this is fine.

@logaretm logaretm Aug 27, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-worded it a bit, initially I was only storing the span id but noticed for the whole thing to be correct-ish I needed the rest of the trace context.

trace: {
...eventTraceContext,
span_id: traceContext.span_id,
parent_span_id: traceContext.parent_span_id,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

m: IMHO we should invert this, if a span is already set on this, likely we should not overwrite it I think? 🤔 but this also makes the logic a bit trickier, because we cannot just do:

trace: {
  span_id: traceContext.span_id,
  parent_span_id: traceContext.parent_span_id
  ...eventTraceContext
}

because that could lead to a case where eventTraceContext has a span_id but no parent_span_id and then they would be incorrectly in sync.

I guess we generally do have a traceContext here already set, right, even if we do not actually have the span? 🤔

Would it work if we move the invocation of applyEscapedErrorSpanToEvent to prepareEvent.ts like this:

if (span) {
    applySpanToEvent(prepared, span);
  } else {
   applyEscapedErrorSpanToEvent(prepared, hint);
}

or something along these lines? 🤔

@logaretm logaretm Aug 27, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So there are three cases here:

  • The span id is wrong: this is the main fix here because we cannot trust the span id already set because we know for a fact it is wrong, so we have to overwrite. The concern here would be, could the already set span id be more accurate than the span id that caught and re-thrown the error?
  • The span id matches the span the error escaped from: We write the same values so it is idempotent here.
  • No trace/span_id: This is the tricky part, if we change placement we won't have a trace_id to compare against and so we could have a mismatching dsc.

I clarified these in a comment just now if it makes sense, WDYT?

@logaretm
logaretm force-pushed the awad/error-span-attribution branch from 99fac5c to 07c56d2 Compare August 27, 2026 15:33

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/core/src/utils/errorSpanAttribution.ts
@logaretm
logaretm requested a review from mydea August 27, 2026 18:54

@Lms24 Lms24 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Q: When we call recordEscapedErrorSpan, could we instead of the weakmap approach, throw a _span non-enumberable property onto the error object? Which we'd later on pick up and adjust the trace context accordingly...

Just wondering if there's smaller alternative to the weakmap.

(fwiw, I'm probably biased because my idea a few years to keep scope data was basically this: Throw the scope onto the error ein every startSpan/withScope callback and merge scopes. Never tried this though so I'm probably missing reasons, so don't feel forced to rewrite!)

@logaretm

Copy link
Copy Markdown
Member Author

When we call recordEscapedErrorSpan, could we instead of the weakmap approach, throw a _span non-enumberable property onto the error object?

I thought we generally preferred weakmaps in recent PRs I observed. Doesn't seem like a big difference to me other than it avoids any funny business with runtimes that serialize stuff like cloudlfare (not saying they do in this case, but they were sneakily serializing tracing channel payloads).

Happy to switch over since it's not a strong opinion, what do you think?

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

👋 @mydea — Please review this PR when you get a chance!

@mydea

mydea commented Sep 2, 2026

Copy link
Copy Markdown
Member

When we call recordEscapedErrorSpan, could we instead of the weakmap approach, throw a _span non-enumberable property onto the error object?

I thought we generally preferred weakmaps in recent PRs I observed. Doesn't seem like a big difference to me other than it avoids any funny business with runtimes that serialize stuff like cloudlfare (not saying they do in this case, but they were sneakily serializing tracing channel payloads).

Happy to switch over since it's not a strong opinion, what do you think?

I think this makes sense to me, then we can possibly just pick this up if it exists in the code where we set the trace context based on the active span? 🤔

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

👋 @mydea — Please review this PR when you get a chance!

Adds failing coverage for #16206: an error is attributed to whatever span
is active at captureException time rather than the span it was thrown in.

Three cases fail on develop (sync nested span, concurrent group, deepest
escaped span). The fourth asserts the cross-trace bail-out, which passes
today and must keep passing.
An error was attributed to whichever span happened to be active when
captureException ran, not to the span that actually failed. Record the
span's trace context in a WeakMap keyed on the error as it unwinds, and
prefer that when building the error event.

The first (deepest) span wins, non-recording spans are skipped, and the
attribution only applies within the error's own trace so the envelope
header and body can never name different traces.
The error thrown in a hapi route handler escapes the router span, so it is
now attributed to that span rather than to the request span. Assert the new
relationship (error span is a child of the transaction's span, and is the
router span) instead of the old identity.
Rename `toKey` to `toWeakMapKey` and explain that thrown primitives cannot be
keyed. The previous WeakMap comment claimed a GC reason that does not hold.
… merge

The trace context of an event captured with no active span only exists as of
that merge, and without its trace id we cannot check the recorded span belongs
to the same trace.
@logaretm
logaretm force-pushed the awad/error-span-attribution branch from 07c56d2 to b1d8b5d Compare September 8, 2026 14:48
Comment thread packages/core/src/utils/errorSpanAttribution.ts
@logaretm

logaretm commented Sep 8, 2026

Copy link
Copy Markdown
Member Author

I think this makes sense to me [...]

@mydea which approach do u mean?

I moved the lookup to preapreEvent so that the event processors also pick up the values changed rather than the incorrect values that would be patched later.

@logaretm
logaretm requested a review from Lms24 September 8, 2026 16:09
Attribution ran once the event was fully assembled, so scope and client
event processors, along with the `postprocessEvent` hook, still saw the
span that happened to be active at capture time.

Move the lookup next to `applySpanToEvent` in `prepareEvent`. The same
trace guard stays, but the trace id now comes from the scope when the
event has no trace context of its own, which is what pushed the call
downstream in the first place. The stored context is applied whole
rather than spliced in: at this point the event may have no trace
context at all, and the merge in `_prepareEvent` lets an existing one
win outright, so a partial context would never get its trace id filled
in.
`recordEscapedErrorSpan` skipped spans that were not recording, which
bundles two conditions: the span is sampled, and it has not ended. Only
sampling is relevant, since an unsampled span is never sent and its id
would point at nothing.

`handleCallbackErrors` runs its error handler after the callback's own
`finally` block, and `startSpanManual` leaves ending the span to the
caller. So the usual manual span idiom, ending the span in a `finally`
and letting the error propagate, ended the span before the error
unwound past it and was skipped entirely, despite being sampled and
sent. Check `spanIsSampled` instead.
() => callback(),
() => {
error => {
recordEscapedErrorSpan(error, span);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: The error span attribution fix is incomplete. Next.js route handlers call handleCallbackErrors directly, bypassing the startSpan logic, so errors are not correctly attributed to the handler's span.
Severity: MEDIUM

Suggested Fix

To ensure consistent error attribution, wrap the route handler execution in wrapRouteHandlerWithSentry.ts with a startSpan call. This will ensure that runCallback is executed, which in turn calls recordEscapedErrorSpan upon an error, correctly associating the error with the active span before it's captured. This pattern should be applied to any other framework integrations that currently call handleCallbackErrors directly.

Prompt for AI Agent
Review the code at the location below. A potential bug has been identified by an AI
agent. Verify if this is a real issue. If it is, propose a fix; if not, explain why it's
not valid.

Location: packages/core/src/tracing/trace.ts#L675

Potential issue: The fix for error span attribution is incomplete. In certain framework
integrations, such as the Next.js route handler wrapper
(`wrapRouteHandlerWithSentry.ts`), errors are handled by calling `handleCallbackErrors`
directly, without being wrapped in a `startSpan`. The new attribution logic, which uses
`recordEscapedErrorSpan`, is only called from within the `runCallback` function used by
`startSpan`. As a result, errors thrown from these specific handlers will not be
correctly associated with their originating span, reverting to the old behavior where
attribution depends on the active span at the time of capture.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4099fde. Configure here.

const eventTraceContext = event.contexts?.trace;
const eventTraceId = eventTraceContext?.trace_id ?? (scope && getTraceContextFromScope(scope).trace_id);

if (eventTraceId !== traceContext.trace_id) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Trace check uses the wrong scope

Medium Severity

The same-trace fallback reads finalScope, but the later merge and DSC still come from currentScope. Passing a Scope as captureContext can skip attribution, or write the escaped span's trace_id onto an event whose envelope DSC names a different trace.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 4099fde. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Experiment: Thrown errors passing through startSpan should be associated with first passed through span

3 participants