Skip to content

fix(ibkr): contain superseded connect attempts - #1681

Open
mehdijamshidian90 wants to merge 2 commits into
TraderAlice:devfrom
mehdijamshidian90:fix/ibkr-superseded-connect
Open

mehdijamshidian90 wants to merge 2 commits into
TraderAlice:devfrom
mehdijamshidian90:fix/ibkr-superseded-connect

Conversation

@mehdijamshidian90

@mehdijamshidian90 mehdijamshidian90 commented Sep 30, 2026 •

Copy link
Copy Markdown

Problem

When IB Gateway restarts while a UTA is already in IBKR auto-recovery, recovery can start a new waitForConnect() on the shared EClient before the previous attempt has finished opening its socket. The two attempts then interfere through shared state.

  1. Teardown hits the successor. The superseded attempt's teardown (client.disconnect() in RequestBridge.waitForConnect) runs after its successor has already replaced EClient.conn. It destroys the successor's socket, and its connectionClosed() rejects the successor with Connection to TWS/Gateway closed during handshake. The successor's Connection.connect() never settles, so that attempt hangs.
  2. Unobserved rejection kills the process. The superseded attempt's own TCP connect then completes. EClient.connect() resumes with this.conn === null, waitForHandshake() rejects immediately, and this.conn.sendMsg() throws before await handshake. The rejected handshake promise is never observed, and Node terminates the UTA process:
Error: Connection closed before handshake started
    at EClient.waitForHandshake (packages/ibkr/src/client/base.ts:345:16)
    at EClient.connect (packages/ibkr/src/client/base.ts:198:30)
    at async Promise.allSettled (index 0)
    at async RequestBridge.waitForConnect (.../request-bridge.ts:166:48)
    at async IbkrBroker.init (.../IbkrBroker.ts:197:7)
    at async UnifiedTradingAccount._attemptReach (.../UnifiedTradingAccount.ts:292:7)

In a real session (Windows 11, IB Gateway paper, pnpm dev), the UTA log showed a run of superseded / closed during handshake recovery failures, then this crash. Guardian logged exited (code=1) — optional service offline, continuing, and the account stayed offline until the UTA was restarted by hand. IB Gateway's daily auto-restart can set up the same sequence without anyone at the machine.

Closes #1679

Approach

Make every connect attempt own the resources it created, in both layers:

  • EClient.connect() keeps the Connection it opened in a local variable and checks, after each await, that the client still points at it. If a newer connect() or a disconnect()/reset() replaced it, the attempt closes only its own socket, with the wrapper detached, and returns. waitForHandshake() now takes that connection explicitly. Its promise is observed as soon as it is created, so no later synchronous throw can leave a rejection unobserved.
  • EClient wrapper ownership. Each Connection reports through a wrapper that forwards error()/connectionClosed() until a newer connect() opens a replacement. So a superseded socket that answers, closes, or times out late cannot mark its successor dead. Teardown of the current connection still reports, including handleReaderError(), which resets the client before closing the socket.
  • RequestBridge.waitForConnect() numbers its attempts. Only the latest attempt may disconnect() the shared client or clear the nextValidId waiter. A superseded attempt still rejects with Previous TWS/Gateway connection attempt was superseded, as before.

Tradeoffs:

Out of scope:

  • UTA respawn. Guardian treats UTA as an optional service. After an unexpected UTA exit it logs optional service offline, continuing and, unlike the Connector (armConnectorRecovery), does not respawn it. With this fix the crash above no longer happens, but automatic UTA respawn could still be worth a separate change.
  • The overlap source. UnifiedTradingAccount.nudgeRecovery() can start a second concurrent broker.init() while a recovery attempt is in flight (details in the review thread). This PR contains the effect at the IBKR boundary and leaves that behaviour alone.

Evidence

Regression tests

  • tests/integration/ibkr-request-bridge/request-bridge.spec.ts:
    • lets a superseding connect finish while the superseded attempt only closes its own socket. Two overlapping waitForConnect() calls against a local TCP server. Without the fix, an unhandled Connection closed before handshake started is raised and the successor never settles. With the fix:
      • the first attempt rejects as superseded;
      • the second connects (nextValidId 700);
      • exactly one server socket remains;
      • nothing is unhandled.
    • Suite superseded connect that settles late: the stale attempt is held in its handshake while its successor connects. Then its socket answers late, is closed from the gateway side, or hits the 10s handshake timeout. The successor must stay connected and not connectionDead, and the wrapper must receive no connectionClosed()/error(). On dev all three fail, because the overlap itself breaks there. The gateway-side close case also fails without the wrapper-ownership change.
  • packages/ibkr/src/client/base.connect-replaced-attempt.spec.ts:
    • closes only its own socket when the client is reset while the socket opens. Without the fix: the same unhandled rejection, plus a leaked socket. With the fix: the attempt resolves, the wrapper receives no error/connectionClosed, and its socket is closed.
    • still reports connectionClosed when a reader failure tears down the current connection. A real socket delivers a malformed frame; the current connection's close must still reach the wrapper.

Live checks (read-only, IB Gateway paper on port 4002, separate client ids, no orders)

Overlapping connects. The dev column is from the original submission; the branch column is re-run on the updated head.

scenario dev this branch
normal connect connected (9 ms, serverVersion 222) connected (8 ms, serverVersion 222)
two overlapping waitForConnect() calls first rejects as superseded; second never settles (waited 12 s), client not connected first rejects as superseded; second connects
unhandled rejections Connection closed before handshake started none

Gateway restart on the updated head. A read-only UnifiedTradingAccount around IbkrBroker was watched while IB Gateway was shut down and started again as a new process, with a manual login.

  • It went offline, then four recovery attempts reached down while the Gateway was down.
  • After the login it recovered to readable, and an account read succeeded.
  • 0 unhandled rejections, and the process stayed up.

Commands run on the updated head (Windows 11, Node 24.19.0, pnpm 11.7.0)

  • Typechecks pass: packages/ibkr, services/uta, root tsc --noEmit, and tests/tsconfig.json.
  • pnpm exec vitest run packages/ibkr/src/client tests/integration/ibkr-request-bridge tests/integration/ibkr-reader-recovery: 30 passed.
  • pnpm test:integration:uta: pass.
  • pnpm test:select --package @traderalice/ibkr: 98 passed.
  • pnpm test: 7552 passed, 17 failed in 14 files. None of the failures are in packages/ibkr or the UTA broker code. The same 14 files also fail on a clean origin/dev worktree on this machine (17 failed there too).

@vercel

vercel Bot commented Sep 30, 2026

Copy link
Copy Markdown

@mehdijamshidian90 is attempting to deploy a commit to the luokerenx4's Team Team on Vercel.

A member of the Team first needs to authorize it.

When IB Gateway restarts during IBKR auto-recovery, a new waitForConnect()
can start on the shared EClient before the previous attempt has opened its
socket. The superseded attempt then tore down its successor's connection,
and once its own TCP connect completed it resumed with a null connection and
left a rejected handshake promise unobserved. Node treats that as fatal, so
the UTA process exited with "Connection closed before handshake started".

EClient.connect() now owns the Connection it opens: after each await it
checks that the client still points at it, and if not, it closes only that
socket with the wrapper detached. The handshake promise is observed as soon
as it exists. RequestBridge.waitForConnect() lets only the latest attempt
disconnect the shared client or clear the nextValidId waiter.
…rapper

A superseded connect attempt's socket can still answer, close, or time out
after its successor connects. Connection reported a late close straight to
the shared wrapper, so RequestBridge marked the healthy successor dead.

EClient now hands each Connection a wrapper that forwards error() and
connectionClosed() until a newer connect() opens a replacement. Teardown of
the current connection still reports, including handleReaderError(), which
resets the client before closing the socket.

Adds RequestBridge cases for a stale attempt whose handshake succeeds,
closes, or times out after the successor connected, an EClient case for a
reader failure on the current connection, and moves the EClient
replaced-attempt spec next to base.ts per the current test layout.
@mehdijamshidian90
mehdijamshidian90 force-pushed the fix/ibkr-superseded-connect branch from 6ca1944 to 5901ab3 Compare October 2, 2026 21:51
@mehdijamshidian90

Copy link
Copy Markdown
Author

Updated per the review on #1679. The branch is rebased on current dev. The first commit is the original change; the rebase carried its RequestBridge cases into the relocated spec. A second commit adds the following.

Test layout

  • The EClient replaced-attempt spec moved to packages/ibkr/src/client/base.connect-replaced-attempt.spec.ts.
  • The RequestBridge cases now live in tests/integration/ibkr-request-bridge/request-bridge.spec.ts.

Late success / reject / timeout
New suite RequestBridge — superseded connect that settles late. The stale attempt is held in its protocol handshake while its successor connects. Then the stale socket:

  • answers the handshake late,
  • is closed from the gateway side, or
  • hits EClient's 10s handshake timeout (fake timers).

Each case asserts that:

  • the stale call rejects as superseded and its socket closes;
  • the successor stays connected and not connectionDead, with nextValidId intact;
  • the wrapper receives no connectionClosed()/error();
  • there are no unhandled rejections.

Against dev's base.ts and request-bridge.ts, all three cases fail, because the overlap itself already breaks there. The late-close case also failed on the previous head of this PR. Connection's own close handler reported the stale socket's close straight to the shared wrapper, before the replaced attempt could detach it, so RequestBridge marked the healthy successor dead. A late socket error would also have reported CONNECT_FAIL (502), which marks the connection dead as well.

Fix: EClient gives each Connection a wrapper that forwards error() and connectionClosed() until a newer connect() opens a replacement. Teardown of the current connection still reports. That includes handleReaderError(), which resets the client before closing the socket. A new EClient case drives that path over a real socket; the existing reader-recovery spec injects the wrapper into a fake Connection directly, so it would not catch a regression there. docs/ibkr-wire-protocol.md documents the rule.

Where overlapping connects come from (observation only, not changed here)
UnifiedTradingAccount.nudgeRecovery() schedules an immediate attempt even while a recovery attempt is in flight, because _attemptReach() has no in-flight guard. It is called:

  • by reads against an unhealthy account: getAggregatedEquity() and contract search in uta-manager.ts, and the HTTP queryAccount() helper;
  • on a 1101/1102 restored event.

So during a Gateway outage, a read while an attempt is in flight starts a second concurrent broker.init() on the same EClient. I kept this PR to the IBKR boundary; I can open a separate issue if that is useful.

Local verification on the updated head (Windows 11, Node 24.19.0)

  • Typechecks pass: packages/ibkr, services/uta, root, and tests/tsconfig.json.
  • These targeted runs pass:
    • the ibkr client specs plus the request-bridge and reader-recovery integration specs (30 tests);
    • pnpm test:integration:uta;
    • pnpm test:select --package @traderalice/ibkr (98 tests).
  • pnpm test: 7552 passed, 17 failed in 14 files. The failing files are cli lifecycle/uninstall, inbox, workspace routes/upgrades, supervisor TUI, the headless watchdog, and similar. The same 14 files also fail on a clean origin/dev worktree on this machine (17 failed there too), so I read them as environment-specific and unrelated to this change.

Live paper check on the updated head (connection only, no orders)
I ran a read-only UnifiedTradingAccount around IbkrBroker from this branch: paper account, IB Gateway on 4002, its own client id.

  • IB Gateway was shut down. The account went offline with "Connection to TWS/Gateway lost".
  • While the Gateway was down (no Gateway process and nothing listening on 4002 when checked), four recovery attempts reached down with "Connection to TWS/Gateway closed during handshake". That is the bridge's message for a socket that closes before nextValidId, which also covers a refused connect. The process stayed up.
  • The Gateway was started again as a new process, with a manual login. The next attempt reconnected and reached readable, and a follow-up account read succeeded.
  • 0 unhandled rejections throughout.

The restart run's attempts were sequential, so it does not by itself prove that the overlapping-connect race was hit live. That race is covered by the deterministic specs above and by the read-only overlapping-connect check from the PR description, re-run on this head: the first waitForConnect() rejects as superseded, the second connects, and nothing is unhandled.

Automatic UTA respawn is untouched, as scoped. I also updated the PR description to match the current head.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant