fix: cancel selected-workspace PTC across processes - #196
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a937bffbaf
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
a937bff to
a75dfad
Compare
a75dfad to
e45a407
Compare
|
Codex Review: Didn't find any major issues. Already looking forward to the next diff. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
beb5ad7 to
503de80
Compare
|
Codex Review: Didn't find any major issues. More of your lovely PRs please. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 503de80c55
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
503de80 to
09180cd
Compare
09180cd to
aa5882a
Compare
|
Codex Review: Didn't find any major issues. More of your lovely PRs please. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
aa5882a to
01749d1
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d017f3ed06
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
d017f3e to
dd3d293
Compare
|
@codex review exact head dd3d293968c748e704c2abfe9dca5f63e9fb555d |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dd3d293e41
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review exact head 9adede8. Please verify the four prior distributed-cancellation findings: external cancellation now wakes the original waiter through the process-wide registry; retryable cancellation bodies stay abortable; cancellation targets are attached before enqueue so failed admission cannot orphan runnable work; and reconnect reconciliation retries with bounded backoff. Also verify the bounded credential-refresh settlement drain. |
9adede8 to
c8876d1
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c8876d1306
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review Please review exact head a5b528f and confirm the reviewed commit. Both latest findings are addressed by deepening the shared outcome interface. Fencing now returns a discriminated cancelled/expired/completed outcome, with the winning result in the same atomic Redis response. Enqueue, subscription, completion and disconnect recovery consume that snapshot without follow-up GETs. The normal Stop route requests only the small decision, not the result body. Initial worker result recovery uses one MGET snapshot as well. Jobs carry the producer cancellation retention requirement; completion uses the larger local/producer retention. Late Stop atomically renews completion evidence to cover the renewed request tombstone, using millisecond precision to avoid rounded-TTL gaps. Focused real Redis tests cover connection loss after the decision, immediate recovery without a lost queue event, timeout configuration drift, and subsecond retention renewal. Local end-to-end verification will be repeated on this exact head. The previous immutable outcome/deadline and lost-reply invariants remain covered. Please audit any remaining lifecycle boundary as a whole rather than suggesting independent signal snapshots after a durable outcome has already won. |
|
Final local acceptance passed again on exact head a5b528f. Real LibreChat + locally linked Agents SDK + Code API + Redis/Mongo + native SRT workers, isolated non-default ports, deterministic model responses only. Physical create/read/edit persistence, approvals, workspace-bound PTC/private skills, worker isolation and offline fail-closed checks passed. Both command and PTC Stop verified a live shell PID, killed it within five seconds, prevented the delayed write after waiting past its scheduled time, and allowed workspace reuse. One focused E2E scenario passed in 1.3 minutes; temporary services were cleaned up. No hosted deployment or existing workspaces changed. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a5b528f8ab
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review Please review exact head b75bbe1. Confirm that commit in your result. All three findings are fixed together at the outcome boundary. Selected native workspace jobs now map and durably commit their result inside the existing bridge sessionResultFinalizer, before the bridge clears the mutation fence, regardless of whether egress restoration is needed. Failure or ambiguity there leaves the root quarantined. A successful committed handoff remains successful through later Stop, egress revocation, or bridge cleanup failure. No separate post-handoff signal check can overwrite that result. Attached request tombstones no longer renew on Stop at all, removing the split renewal window; admission retention bounds the mapping. Deterministic result decoding/validation is outside the Redis transport retry loop. Real Redis regressions cover handoff completion versus cancellation/quarantine, missing and corrupt results with exactly one attempt, and interrupted split renewal. 26 focused Redis/bridge tests plus 23 registry tests pass. The actual local LibreChat + Code API + native SRT worker acceptance just passed again, including physical command and PTC Stop/no-late-write checks, files, approvals and private skills. Please audit these shared invariants and remaining alternate transitions. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b75bbe11d3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Final post-commit local verification passed on b75bbe1 after rebuilding Code API. The real LibreChat + SDK integration + native SRT worker acceptance passed again in 1.3 minutes, including verified live PIDs before UI Stop, PID termination within five seconds, no delayed file writes, and subsequent workspace reuse for both commands and PTC. Files, approvals, private skills, isolation, reload and offline fail-closed checks passed. Separate real-SRT speculative-network regression passed too. All services used disposable identities and non-default ports; existing deployments/workspaces were untouched. |
|
@codex review Please review exact head ad14f28 and confirm the reviewed commit. The stalled-job redelivery finding is fixed at admission as well as settlement. Before any sandbox work, one atomic Redis operation returns the cached committed result or acquires a once-only execution claim. Duplicate processors and ambiguous/lost claim replies cannot authorize another execution; the bounded claim is retained through the producer/local cancellation horizon. The claim shares the cancellation decision transaction, so cancel-before-start fails closed. commitJobResult now distinguishes committed, cancelled and already_completed. An existing completion cannot be attributed to a second mutation handoff: the native bridge finalizer throws and quarantines that root. The original durable result is never overwritten. Real Redis tests cover concurrent claims (exactly one execution), lost claim ACK, pre-cancellation, missing cached payload, duplicate handoff quarantine and preserving the first result. 23 outcome tests plus 30 registry/request tests pass; service build passes with only existing unrelated warnings. Final local LibreChat/native-SRT acceptance is being repeated on this head. Please inspect remaining lifecycle invariants. |
|
Final local acceptance and real-SRT speculative replay canary both pass on ad14f28. Real LibreChat, locally linked SDK, Code API, Redis/Mongo and native SRT workers on disposable non-default ports. Both command and PTC Stop verified live shell PIDs, termination within five seconds, absence of delayed writes after their scheduled time, and workspace reuse. File persistence, approval modes, private skills, worker isolation, reload and offline failure checks pass. Temporary services cleaned up; no hosted deployment or existing workspace changed. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ad14f28ad0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review Please re-review exact unchanged head ad14f28. All valid findings are fixed, CI is fully green, and final real local LibreChat/Code API/native-SRT acceptance passes. The last finding is disproved by source and real Redis failure injection. In enqueueForActiveIncarnation, workspaceQuarantineKey is KEYS[7], created by SET key assignmentId with NO TTL. Assignment/queue/deadline/receipt expiry does not expire that fence. A failed SET that would relabel it quarantined leaves the original persistent pending fence blocking reuse. The test acknowledged and fulfilled a native assignment, injected failure into the quarantine eval, waited beyond the deadline, and obtained PTTL=-1 both before and after; a new dispatch was rejected WORKSPACE_QUARANTINED. Detailed evidence is in discussion_r4006669482. No code change was needed. Please confirm the exact-head result after accounting for that persistent fence, and report only additional actionable findings with a demonstrated lifecycle path. |
|
Codex Review: Didn't find any major issues. 👍 Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
ad14f28 to
e8b7b34
Compare
* fix: fail denied input downloads once per batch (LibreChat-AI#177) * fix: preserve retries for transient egress ledger conflicts (LibreChat-AI#179) * fix: distinguish retryable ledger contention from scope denials * fix: preserve error classification through marker discovery and relay * test: exercise classified denials through gateway configuration * fix: honor bounded gateway retry hints during object downloads * perf: reuse authorized input versions and make egress accounting atomic (LibreChat-AI#180) * perf: reuse authorized input versions and make egress accounting atomic * perf: resolve authorized input manifests once per execution * test: preserve fetch signature in revocation fixture * fix: isolate shared-grant failures and prevent ledger replay * 🧹 fix: Evict Stale File-Object Index Entries (LibreChat-AI#182) * fix: evict stale file-object index entries Forget cached locators after successful deletion and missing-object downloads so replacement keys resolve immediately. Reuse exact resolver matching in the delete route to avoid prefix collisions. Fixes LibreChat-AI#181 * fix: keep file-object deletion storage-authoritative * fix: retire superseded upload objects * fix: canonicalize replacement object keys * fix: collapse legacy object-key siblings * fix: recover reads from stale locators * fix: namespace canonical object identities * perf: enable authorized input reuse by default (LibreChat-AI#183) * fix: honor requested input destinations (LibreChat-AI#184) * fix: disambiguate legacy dotted object identities (LibreChat-AI#186) * fix: helm egress deployment getting stuck on install (LibreChat-AI#176) * fix: helm egress deployment getting stuck on install * only wait for redis if the ledger is required * Update helm/codeapi/templates/egress-gateway-deployment.yaml Co-authored-by: Danny Avila <danacordially@gmail.com> --------- Co-authored-by: Danny Avila <danacordially@gmail.com> * feat: Add Trusted VM Command Policy (LibreChat-AI#187) * 🛰️ feat: Add trusted VM command policy * docs: clarify trusted VM socket boundary * fix: Retry Clean Cancelled BYOM Settlements Through Stop Grace (LibreChat-AI#188) A clean atomic workspace mutation rejection was settled with retries cut off at the original execution deadline. When Stop arrived near that deadline, process-tree termination finished after it, so the first settlement attempt was aborted immediately while Code API was still draining the cancellation. The client received ASSIGNMENT_EXPIRED and the durable mutation guard stayed armed. Route clean workspace mutation rejections through the known-clean rejection recovery path: a transient-retrying heartbeat and settlement retries through the rejection acknowledgement grace, floored at the bridge cancellation settlement grace. Worker shutdown still fails closed. Share the grace constant from the protocol module so the bridge and worker stay aligned. Closes LibreChat-AI#173 * fix: Release Unassigned Workspace Slots When Dispatch Cleanup Fails (LibreChat-AI#189) Closes LibreChat-AI#170 * fix: Exit Cleanly When Native Executor Shuts Down Concurrently (LibreChat-AI#191) * fix: Exit Cleanly When Native Executor Shuts Down Concurrently A native BYOM worker under systemd KillMode=control-group receives SIGTERM at the same time as its forked SRT executor. The child ignores IPC once it is shutting down, so the parent's close handshake is left pending until the child exits, which rejects it with 'Native executor is unavailable'. That rejection escaped the CLI finally block and turned an idle administrative stop into exit status 1. Treat the close handshake as best-effort: the executor is terminated in finally regardless, and the active command has already drained, so a lost or stalled reply carries no mutation risk. Also make the child report exit status 0 when its own SRT teardown succeeded. Closes LibreChat-AI#190 * fix: Surface Explicit Executor Cleanup Failures During Close Only a lost, refused, or stalled close handshake is benign at shutdown. A negative close reply from the executor is a real cleanup failure and still rejects so pool shutdown can aggregate it. * fix: authenticate GitHub App Git operations (LibreChat-AI#192) * fix: isolate file deletion rate limits (LibreChat-AI#193) * feat: broker GitHub CLI authentication (LibreChat-AI#194) * feat: Run PTC in Selected BYOM Workspaces (LibreChat-AI#195) * feat: run PTC in selected BYOM workspaces * fix: harden native workspace PTC replay * fix: preserve replay isolation and bridge limits * test: tolerate hosts without filesystem cloning * test: surface copy-on-write clone faults * fix: harden native workspace PTC admission * fix: close native replay effect and finalization boundaries * feat: Report Truncated Output Artifacts (LibreChat-AI#199) * feat: report truncated output artifacts * fix: classify omitted artifacts precisely * fix: preserve artifact scan invariants * fix: bound depth truncation probes * fix: bound capped directory enumeration * fix: constrain truncation probes across the job * fix: stop exhausted artifact probes * fix: cancel selected-workspace PTC across processes (LibreChat-AI#196) * fix: cancel replay jobs across API and worker processes * fix: drain worker cancellation watches promptly * fix: close programmatic cancellation races * fix: preserve cancellation response ordering * fix: close distributed cancellation races * fix: harden cancellation under concurrent load * fix: make cancellation ownership durable through completion * fix: recover durable replay outcomes across lost replies * fix: return atomic cancellation outcomes with aligned retention * fix: commit native results inside the workspace mutation fence * fix: claim programmatic execution before stalled-job redelivery * fix: classify capped artifact probe candidates (LibreChat-AI#206) * feat: Report Deleted Code Session Files (LibreChat-AI#200) * fix: report deleted persisted files * fix: reconcile deletions across code runtimes * fix: preserve protected session inputs * fix: classify reserved runtime paths * fix: preserve trusted jq path for PTC (LibreChat-AI#207) * fix: isolate native PTC readiness and watchdog phases (LibreChat-AI#208) * feat: Declare Named Worker Project Environments (LibreChat-AI#209) * feat: declare named worker project environments * fix: preserve environment trust and negotiated action boundaries * Harden environment loading and executor identity * Protect environment root traversal and exact config bytes * Reject self-controlled environment root aliases * Check filesystem identities at environment trust boundaries * Validate environment containment across Linux mount aliases * Handle stacked mounts conservatively without blocking unrelated paths * fix: Allow Trusted Own-Root Environment Symlinks (LibreChat-AI#212) * fix: Allow Trusted Own-Root Environment Symlinks * fix: Check Alias Parent Ownership by Filesystem Identity * fix: Enforce Parent Ownership Across Every Environment Path * fix: Retain Quarantine After Failed Environment Setup (LibreChat-AI#213) * fix: Retain Quarantine After Failed Environment Setup * test: Run Native Environment Setup Lifecycle in CI * fix: Document and Verify Local Setup Quarantine Recovery * fix: Allow Lambda MicroVM metadata in hardened mode (LibreChat-AI#215) * docs: add self-hosted worker setup runbook (LibreChat-AI#219) * fix: bound memory buffering for streamed file uploads (LibreChat-AI#218) * ci: Automate Auditable Main Releases (LibreChat-AI#216) * fix: decouple repository release versions * ci: automate releases after successful main builds * fix: forward input-file limit into sandbox guests (LibreChat-AI#217) * feat: Inspect Local Coding Projects (LibreChat-AI#221) * feat: add bounded local project inventory * fix: report incomplete Git metadata reads * fix: preserve incomplete discovery and remote identities * fix: finalize discovery budgets and nested remote identities * fix(code): stop project traversal at filesystem budget boundaries * fix: make automatic release tip check read-only (LibreChat-AI#225) * feat: Discover Bounded Repository Instructions for Attached Workspaces (LibreChat-AI#226) * Discover bounded repository instructions for opted-in workspaces * Verify snapshot digests and cross-platform confinement * feat: Register Selected Coding Projects (LibreChat-AI#222) * feat(code): register explicitly selected project roots * fix(code): reject shared Git metadata for selected projects * fix(code): pin selected project identity through executor admission * fix(code): preserve full filesystem identity precision * Check selected project identity before replay staging * Anchor replay copies to the verified working directory * fix: Bind Selected Project Operations to Held Directory Descriptors * test: Cover Selected Project PTC and Load Native Fixtures Before Platform Simulation * fix: Keep Native Root Bindings Worker-Local and Verify Directory Ancestry * fix: Anchor Project Admission and Preserve Search Permissions * fix: Report workspace admission capacity without ambiguous timeout errors (LibreChat-AI#227) * fix: Distinguish workspace admission capacity from execution expiry * test: Preserve execution uncertainty while classifying blocked follow-ups * fix: Classify admission expiry at the enqueue boundary * ci: Fix Release Version Resolution for Untagged and Resumed Runs (LibreChat-AI#233) * ci: Fix Release Version Resolution for Untagged and Resumed Runs The release workflow resolved its version in one inline shell block under `set -euo pipefail`, where two paths could not succeed. Filtering tags through `grep` made a no-match fatal. On the ordinary untagged tip of `main`, `git tag --points-at HEAD | grep -E '^v[0-9]+...'` exits 1, and the step died before reaching its skip handling or `next-release-version.sh`, so a deployable commit could not obtain a release version (LibreChat-AI#228). Selecting stable tags now reads exit 1 as an empty answer while exit 2 and above still fail the release, which also lets the missing-previous-tag case report its own error. The rerun-resume path then rejected the tag it had itself chosen. With a stable tag already pointing at `HEAD` and no release published, the version comes from that tag, and the following existence check failed merely because the ref existed (LibreChat-AI#229). It now compares the tag's commit against the release commit, so only a tag on some other commit is a collision; `Create tag` already tolerates a tag that exists. The block moved into `.github/scripts/resolve-release-version.sh`, beside the `next-release-version.sh` it calls, so `tests/release-version-resolution.sh` can cover every path: automatic, resumed, skipped, dispatched, pushed-tag, and the runs that must be refused, each against a throwaway repository with a stubbed `gh`. * fix: harden release resolver execution --------- Co-authored-by: Lia <lia@librechat.ai> Co-authored-by: Danny Avila <danny@librechat.ai> * feat: Route GitHub App credentials per repository (LibreChat-AI#236) * feat: Route GitHub App credentials per repository * test: Make repository routing assertion deterministic * fix: Harden repository credential routing * fix: Bound shared GitHub credential refreshes * fix: Bind GitHub credentials to admitted workspaces * fix: Authenticate GitHub App bot identity lookup (LibreChat-AI#237) * feat: Advertise Workspace Command Timeout Ceiling (LibreChat-AI#238) * 🌳 feat: Provision Conversation-Scoped Code Worktrees (LibreChat-AI#239) * feat: provision conversation-scoped code worktrees * fix: isolate conversation checkout metadata * docs: clarify isolated conversation checkouts * fix: revalidate conversation checkout sources * fix: preserve synchronous legacy execution startup * fix: harden conversation worktree lifecycle * test: use canonical workspace isolation keys * fix: secure conversation worktree provisioning * fix: preserve isolated workspace lifecycle * fix: harden conversation worktree provisioning * fix: fence worktree setup and credential routing * fix: use kernel-backed provisioning locks * fix: load worktree locking only when provisioned * fix: retain conversation provisioning ownership through recovery * fix: reserve provisioning before launching checkout writers * fix: pin Git provisioning inputs and close instance admission gaps * fix: Close native scratch directory streams (LibreChat-AI#240) * feat(code): route trusted VM GitHub App tokens by checkout (LibreChat-AI#248) * feat(code): route trusted VM GitHub App tokens by checkout * fix(code): resolve checkout credentials within canonical root * fix(code): reject cross-root credential aliases --------- Co-authored-by: Danny Avila <danny@librechat.ai> Co-authored-by: Ignaz "Ian" Kraft <ignaz.k@live.de> Co-authored-by: Danny Avila <danacordially@gmail.com> Co-authored-by: Jackson Riding <99007683+jacksonriding@users.noreply.github.com> Co-authored-by: lia-by-librechat[bot] <328778573+lia-by-librechat[bot]@users.noreply.github.com> Co-authored-by: Lia <lia@librechat.ai> Co-authored-by: busla <3162968+busla@users.noreply.github.com>
Summary
Scale and reliability
Verification
Depends on #195.