Repository navigation
fix(vmcp): record the owner URL in the session placeholder - #7
Draft
andresrsanchez wants to merge 1 commit into
Draft
andresrsanchez wants to merge 1 commit into
andresrsanchez wants to merge 1 commit into
Conversation
The owner URL was written to storage only after CreateSession finished connecting every backend, so a tools/list or resources/list that the ingress sent to another replica in that window was not forwarded. That replica restored a full copy of the session and kept it until the idle sweep, which is what grows remote spoke vMCP memory during hub storms. Generate() now stores this pod's owner URL in the placeholder and CreateSession keeps it, so early requests are forwarded to the owner. Every session-scoped request except initialize is forwarded, and records already marked terminated are not. Takeover of a dead owner is hardened in the same change: the forwarding client gives up dialing after 3s, the ownership claim survives a cancelled caller context, DELETE and GET to a dead owner no longer restore the session, DELETE is forwarded on a detached context, and an abandoned tools/call is not re-run locally.
andresrsanchez
marked this pull request as draft
September 25, 2026 10:42
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
The fix
Smaller fixes for when the owner pod has died
Other pods give up trying to reach a dead owner after 3s instead of 30s.
Taking over a dead owner's session no longer fails just because the client stopped waiting.
A DELETE (or a GET) for a dead owner's session no longer rebuilds the whole session first.
A tool call the client has abandoned is not run a second time.
Forwarding exists only in our fork. We added it in April ("forward live session to owning replica", "session handover on pod churn"). Upstream Generate() still saves an empty placeholder, and nothing in upstream stores an owner URL.
Upstream avoids copies by keeping each client on one pod. Any pod that gets a request for a session it doesn't hold rebuilds it from Redis. Upstream counts on the Service's ClientIP stickiness to prevent that. Our remote spokes are reached through the ingress, which balances every request separately and skips the Service stickiness. So upstream would build copies on remote spokes too, probably more than today, because nothing forwards.
Upstream does cap the copies. Since about v0.45.0 (08-26), each pod keeps at most 1,000 live sessions (CacheCapacity) and drops the least recently used. At p2's session size that's roughly 350 MB per pod, so it would stop the OOMs. It doesn't stop the extra builds, though: dropped sessions just get rebuilt, so backend load stays high.
Our fork doesn't have that cap. It's behind upstream, and the v0.39.0 sync (fork PR chore: sync fork with upstream v0.39.0 (fixes ToolHive restart churn, #5064 + #5300) #3) is still open.
Expected result: each session is built once instead of about twice, memory stays flat during hub storms, and backends get about 44% fewer connections.