fix(orchestrator): recover lost ancestor entries when persisting headers - #3489
ValentaTomas wants to merge 5 commits into
Conversation
A pause clones its source header's Builds map, so a mapping-referenced build whose entry was lost (e.g. the header was peer-served/incomplete when an earlier pause persisted it) stays absent in every descendant header. Since #3447 the read path recovers from such gaps, but each descendant then pays a header refresh per gap forever. Close the write path: when appendAncestorBuilds resolves nothing locally and the entry is still missing, load the referenced build's own header and copy its self entry before persisting. Builds without a header file (legacy uncompressed) stay absent and keep resolving on the read path.
PR SummaryMedium Risk Overview Reviewed by Cursor Bugbot for commit a198284. Bugbot is set up for automated code reviews on this repo. Configure here. |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Older releases fabricated a zero BuildData at load time for every build a header's mapping referenced but Builds lacked, and a pause serialized those into descendant headers as real entries — on the wire indistinguishable from the legit V3 uncompressed sentinel. For a compressed ancestor the claimed suffix-less object does not exist, so every fault failed permanently, even after the load-time backfill fix. Instead of failing the createDiff size lookup, treat a zero entry whose basic-name object is missing as a persisted gap and resolve the build's own header, mirroring the no-entry branch. Legit V3 sentinels are unaffected: their object exists and no extra roundtrip happens.
When the zero-entry fallback cannot load the build's own header, the load error was dropped and only the original object miss surfaced, so a pruned genuine V3 build was indistinguishable from an unreadable header. Join both into the returned size-lookup error.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
Bugbot Autofix prepared a fix for the issue found in the latest run.
- ✅ Fixed: Zero-entry recovery skips peer transition
- Added PeerTransitionedError handling to the zero-entry recovery logic so it recovers via refreshHeader like the no-entry branch does.
Or push these changes by commenting:
@cursor push d39bdce632
Preview (d39bdce632)
diff --git a/packages/orchestrator/pkg/sandbox/build/storage_diff.go b/packages/orchestrator/pkg/sandbox/build/storage_diff.go
--- a/packages/orchestrator/pkg/sandbox/build/storage_diff.go
+++ b/packages/orchestrator/pkg/sandbox/build/storage_diff.go
@@ -209,7 +209,15 @@
if size == 0 {
size, err = upstream.Size(ctx)
- if hasEntry && initialFT == storage.UncompressedFullFrameTable && errors.Is(err, storage.ErrObjectNotExist) {
+ var shouldRecover bool
+ if hasEntry && initialFT == storage.UncompressedFullFrameTable {
+ shouldRecover = errors.Is(err, storage.ErrObjectNotExist)
+ if !shouldRecover {
+ var transErr *storage.PeerTransitionedError
+ shouldRecover = errors.As(err, &transErr)
+ }
+ }
+ if shouldRecover {
// The zero entry claimed an uncompressed V3-era ancestor, but the
// suffix-less object does not exist: the entry is a gap persisted
// as zero by older releases. Resolve the build's own header likeYou can send follow-ups to the cloud agent here.
Reviewed by Cursor Bugbot for commit f26f8a6. Configure here.
This comment has been minimized.
This comment has been minimized.
f26f8a6 to
ac5603c
Compare
The zero-entry fallback only matched ErrObjectNotExist, but the ancestor is opened through the peer-routing provider, whose Size reports a peer miss as PeerTransitionedError. A stale zero entry on a peer-served build therefore failed the read for the length of the transition window instead of refreshing the build's own header, which is what the no-entry branch already does for the same signal. Refresh on either error, attributing the peer case to the existing peer_transitioned cause so zero_entry_miss keeps counting only baked-zero recoveries.
Recovering a lost ancestor entry is an optimization: an absent entry is what every release before the recovery persisted, and createDiff still resolves the build's own header per fault-in. Aborting the header store when that load fails turns an unreadable ancestor header into a failed snapshot — fatal on the Checkpoint path, which has no retry and tears the sandbox down, and permanent for a deterministic read error, which burns the whole retry budget. Log and leave the gap instead, keeping the failure fatal only when the context is already done, where continuing would bury the cause under a store-header error.


Two complementary fixes for headers whose mapping references builds their
Buildsmap lost. #3447 stopped load-time fabrication of zero entries for such gaps; this PR closes the two paths it left open.Write path: stop persisting inherited gaps
A pause carries its source header's
Buildsmap forward, so an entry lost once (e.g. the live header was peer-served/incomplete when an earlier pause persisted it) propagates to every descendant header, and each descendant diff pays a proactive header refresh per gap forever. Now, whenappendAncestorBuildsresolves nothing locally and the entry is still missing, it recovers the entry from the referenced build's own stored header before the diff header is persisted — gaps heal on the next pause instead of being inherited, at the cost of one header GET per still-missing entry (normally zero). Builds with no stored header (legacy uncompressed) stay absent and keep resolving on the read path. The heal lives inappendAncestorBuildsrather thanStoreHeaderbecause that's where the ancestor barrier, storage handle, and file type already are.Read path: recover stale zero entries
Before #3447, the fabricated zero entries were also serialized whenever a pause ran on a header loaded from storage — baked into descendant headers as real entries, on the wire indistinguishable from the legit V3 uncompressed sentinel. For a compressed ancestor the claimed suffix-less object does not exist, and
createDifffailed the size lookup permanently; #3447 only helps when the entry is absent. Now a zero entry whose basic-name object is missing falls back to the build's own header (same recovery as the no-entry branch, telemetry causezero_entry_miss). Legit V3 sentinels are unaffected — their object exists, so no extra roundtrip.Stored zero entries are re-persisted by later pauses (indistinguishable from V3 sentinels without a per-pause probe), but reads now always recover; rewriting them is fleet hygiene, not orchestrator logic.
orchestrator.storage.diff.frame_table_refresh(cause=proactive) should trend down as affected lineages pause past the write-path fix; cause=zero_entry_miss counts the baked-zero recoveries.Tests