Skip to content

fix(vmcp): record the owner URL in the session placeholder - #7

Draft
andresrsanchez wants to merge 1 commit into
dp-stablefrom
fix/vmcp-owner-url-in-placeholder
Draft

andresrsanchez wants to merge 1 commit into
dp-stablefrom
fix/vmcp-owner-url-in-placeholder

Conversation

@andresrsanchez

@andresrsanchez andresrsanchez commented Sep 25, 2026 •

Copy link
Copy Markdown

The problem

  • Each spoke vMCP runs as 3 pods. A hub session opens with initialize on one pod, which becomes its owner.
  • The ingress sends each later request (tools/list, resources/list, DELETE) to any of the 3 pods. A pod that gets a request for a session it doesn't own is supposed to forward it to the owner.
  • To forward, it needs the owner's address from Redis. But the owner only saved its address after it finished connecting to all its backends, which takes a moment.
  • A request that arrived in that window found no address, so the pod built its own full copy of the session instead of forwarding.
  • The DELETE at the end only reaches the owner, so the copy on the other pod stays for 24–36h until the idle sweep clears it.
  • During a hub storm, thousands of these copies pile up and the spoke runs out of memory. That's where the 108 GB and the OOM restarts come from.

The fix

  • The owner now saves its address the instant the session is created, before connecting to backends. Every other pod can forward from the very first request, so no copies get built.

Smaller fixes for when the owner pod has died

  • Other pods give up trying to reach a dead owner after 3s instead of 30s.

  • Taking over a dead owner's session no longer fails just because the client stopped waiting.

  • A DELETE (or a GET) for a dead owner's session no longer rebuilds the whole session first.

  • A tool call the client has abandoned is not run a second time.

  • Forwarding exists only in our fork. We added it in April ("forward live session to owning replica", "session handover on pod churn"). Upstream Generate() still saves an empty placeholder, and nothing in upstream stores an owner URL.

  • Upstream avoids copies by keeping each client on one pod. Any pod that gets a request for a session it doesn't hold rebuilds it from Redis. Upstream counts on the Service's ClientIP stickiness to prevent that. Our remote spokes are reached through the ingress, which balances every request separately and skips the Service stickiness. So upstream would build copies on remote spokes too, probably more than today, because nothing forwards.

  • Upstream does cap the copies. Since about v0.45.0 (08-26), each pod keeps at most 1,000 live sessions (CacheCapacity) and drops the least recently used. At p2's session size that's roughly 350 MB per pod, so it would stop the OOMs. It doesn't stop the extra builds, though: dropped sessions just get rebuilt, so backend load stays high.

  • Our fork doesn't have that cap. It's behind upstream, and the v0.39.0 sync (fork PR chore: sync fork with upstream v0.39.0 (fixes ToolHive restart churn, #5064 + #5300) #3) is still open.

Expected result: each session is built once instead of about twice, memory stays flat during hub storms, and backends get about 44% fewer connections.

The owner URL was written to storage only after CreateSession finished
connecting every backend, so a tools/list or resources/list that the
ingress sent to another replica in that window was not forwarded. That
replica restored a full copy of the session and kept it until the idle
sweep, which is what grows remote spoke vMCP memory during hub storms.

Generate() now stores this pod's owner URL in the placeholder and
CreateSession keeps it, so early requests are forwarded to the owner.
Every session-scoped request except initialize is forwarded, and records
already marked terminated are not.

Takeover of a dead owner is hardened in the same change: the forwarding
client gives up dialing after 3s, the ownership claim survives a
cancelled caller context, DELETE and GET to a dead owner no longer
restore the session, DELETE is forwarded on a detached context, and an
abandoned tools/call is not re-run locally.
@github-actions github-actions Bot added size/L and removed size/L labels Sep 25, 2026
@andresrsanchez
andresrsanchez marked this pull request as draft September 25, 2026 10:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant