Skip to content

fix(vmcp): end sessions whose stored credentials backends reject on rebuild - #6

Open
CiraciNicolo wants to merge 1 commit into
dp-stablefrom
fix/vmcp-end-sessions-with-stale-credentials
Open

CiraciNicolo wants to merge 1 commit into
dp-stablefrom
fix/vmcp-end-sessions-with-stale-credentials

Conversation

@CiraciNicolo

@CiraciNicolo CiraciNicolo commented Sep 24, 2026 •

Copy link
Copy Markdown

Summary

  • Background rebuilds reuse the identity captured when the session was created. When its token has expired, or is missing because the session was restored from storage, the backends that need it reject the rebuild. The session loses them, and the lacking-backend reconcile (e755493) retries them every 5 minutes with the same credentials.
  • This PR ends such a session instead. The client's next request gets 404 and, as the MCP transport requires, it starts a new session with its current credentials, which brings those backends back.
  • It changes the refresh path from e755493, 35680e2 and c9ab3cc.

Type of change

  • Bug fix

What we saw

  • Hundreds of sessions lacking one token-injected remote backend were rebuilt every 5 minutes. Each rebuild sends one initialize with the token captured at creation, which had expired hours before.
  • Those rebuilds made up most of the /mcp requests that backend received, all 401. vMCP logs them as Failed to initialise backend for session; continuing without it … unauthorized (401).
  • The loop runs until mcp-go's idle sweep ends the session: 24-36 h after the client's last request with a 24 h sessionTTL, since the sweep runs every sessionTTL/2. The reconcile's GetMultiSession refreshes the Redis TTL, not the sweep's activity clock.

Change

  • What counts as a rejection (backend.IsCredentialRejection, matched with errors.Is):
    • a 401 from the backend (mcptransport.ErrUnauthorized);
    • no token for the outgoing auth strategy: ErrUpstreamTokenNotFound, and a new ErrCallerTokenEmpty, returned by upstream_inject with the unchanged message.
    • Not counted: a 403, which a fresh token for the same caller would get too, and failures of the auth strategy itself, such as a token exchange that cannot be reached.
    • vmcp.IsAuthenticationError is not used. It matches every strategy failure, through authentication failed for backend …, and it misses mcp-go's 403 (request failed with status 403).
    • A test through the real connector against an httptest backend checks that the sentinels survive the auth round tripper, net/http and mcp-go.
  • Two metadata keys:
    • vmcp.backend.rejected_ids: the backends that rejected this build, written by the factory on every build (absent when none did);
    • vmcp.backend.rejected_ids_at_creation: the same list from the creation build. CreateSession records it, RefreshSession carries it (never from the rebuild), RestoreSession copies it. An absent key reads as empty, so sessions created before this change heal too.
  • Decision, in refreshSessionCapabilities, for every background rebuild whatever triggered it (the scheduler merges lacking and containing triggers into refreshAllSessions). If the session has a subject and a backend rejected the rebuild without being on the creation list:
    • Terminate the session (storage delete, the rebuilt session is closed), drop its activeClientSessions entry and unregister it from the SDK server, then close the previous session after the usual grace period;
    • if Terminate fails, roll back as for any other refresh failure; the next refresh tries again.
    • Request-time rebuilds (refreshSessionForStaleBackend) are unchanged.
  • Logs: runBackendRefresh reports ended next to failed. Each ended session logs ended session: backends rejected its stored credentials with backend_ids.

Why the creation list

  • Without it, a backend that rejects a caller's fresh token would end every new session of that caller at its first rebuild, every few minutes. Examples: a misconfigured audience, or a token-injected backend in anonymous mode. With the list, those sessions keep today's behaviour.
  • The subject comes from the session before the rebuild. A restored session has a subject but no token, and its rebuild writes an empty token hash. So neither ShouldAllowAnonymous nor the token hash can tell that it has an owner.

Behaviour changes

  • The trade. A session whose credentials went stale now ends at its next background rebuild.
    • Before, it lost only its token-injected backends and kept every other one.
    • Now the whole session ends. That is a win only if the client re-initializes on 404, as the MCP transport requires. A client that does not now loses every backend instead of a few.
    • Not verified for any client yet, neither interactive IDE and CLI agents nor automation clients that forward their own token.
  • If a backend answers 401 to valid tokens during an outage, each rebuilt session ends once, instead of losing that backend. Sessions created during the outage have the backend on their creation list and do not end again.
  • After the rollout:
    • abandoned sessions go with the old pods: nothing registers them again, and their Redis records expire;
    • live sessions restored from Redis still lose their token-injected backends at restore (no token), as today. They now end at their next rebuild, within about one reconcile pass.
  • Server-side termination leaves mcp-go's per-session transport maps for the old ID until the idle sweep removes them.

Before merging

  • Merge condition: 404 recovery verified for an interactive client and for an automation client.
  • A quick way to check it on a staging vMCP, without waiting for a token to expire:
    1. keep a client connected with a session that includes an upstream_inject backend;
    2. restart the vMCP pod. The session is restored without a token, so the next reconcile pass (about 5 minutes, plus the 15 s debounce) ends it;
    3. the client's next call shows whether it recovers.

Verification

  • go test ./pkg/vmcp/...: the new tests pass. Two failures also happen on dp-stable and are not touched here:
    • TestRestoreHijackPrevention_AuthenticatedRoundTrip expects a different token with the same subject to be rejected; sessions are bound to the subject since e492cbf;
    • under -race, a data race in TestDefaultAggregator_QueryAllCapabilities.
  • golangci-lint v2.13.2 on the changed packages: no findings beyond those dp-stable already has.
  • Mutation check: 23 hand-made mutations, all caught by the tests. Among them:
    • dropping the carry-forward, the creation record or the restore copy;
    • counting 403 or strategy failures as rejections;
    • breaking the %w chain;
    • exempting sessions via ShouldAllowAnonymous;
    • skipping the rollback;
    • keeping the registrations.

Not in this PR

  • Refreshing a restored session binds it as anonymous (pre-existing, separate fix).
    • RefreshSession passes ShouldAllowAnonymous(CreatorIdentity()), which is true for a restored identity (subject, no token);
    • PreventSessionHijacking then writes an empty token hash;
    • the owner's token-bearing calls are then rejected as a session upgrade.
  • Rebuilding with the caller's live identity instead of the stored one. It would heal sessions without the 404, but it is a bigger change.

…ebuild

Background rebuilds (health or capability triggers, and the 5-minute
lacking-backend reconcile) reuse the identity captured when the session
was created. Once its token expires, or is missing because the session
was restored from storage, every backend that needs it rejects the
rebuild. The session loses those backends and the reconcile retries them
every 5 minutes, one 401 each time, until the idle sweep ends the session
24-36 h after the client's last request. In one deployment this reached
hundreds of sessions and most of the traffic one backend received.

Record, per build, the backends that rejected the credentials: a 401 from
the backend, or no token for its upstream_inject strategy
(ErrUpstreamTokenNotFound, and a new ErrCallerTokenEmpty sentinel with the
same message as before). The session manager keeps the list from the
creation build and carries it through refreshes and restores.

After a background rebuild, a backend that rejects the credentials but
did not reject them at creation means they went stale. Terminate the
session instead of publishing the degraded rebuild, and drop its
registrations. The client's next request gets 404 and it starts a new
session with its current credentials. If terminating fails, roll back to
the previous session as for any other refresh failure.

Sessions without a subject, and backends that already rejected the
credentials at creation, keep today's behaviour: a backend that never
accepts a caller cannot end every new session of that caller. A 403 and
failures of the auth strategy itself (a token exchange that cannot be
reached) do not count as rejections.

Signed-off-by: Nicolò Ciraci <nicolo.ciraci@docplanner.com>
@CiraciNicolo
CiraciNicolo force-pushed the fix/vmcp-end-sessions-with-stale-credentials branch from c47f2ba to 33579d1 Compare September 24, 2026 12:15
@github-actions github-actions Bot added size/L and removed size/L labels Sep 24, 2026
@CiraciNicolo
CiraciNicolo marked this pull request as ready for review September 24, 2026 12:15
@github-actions github-actions Bot added size/L and removed size/L labels Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants