Repository navigation
fix(ssr): recover custom error pages after transient failures - #4436
Conversation
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
📦 Client bundle boundary
A server module in a client graph aborts hydration in the browser. New leaks fail CI; known leaks are tracked in |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (3)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe fallback now distinguishes confirmed missing pages from transient failures. New tests verify that discovery and loading outages do not prevent a later successful 500 error-page response. ChangesError-page fallback recovery
Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to Custom SSR error pages now recover after temporary failures instead of remaining stuck on the generic fallback. The supplied validation reports broad passing coverage, with no actionable merge blocker. Sequence Diagram(s)sequenceDiagram
participant tryErrorPageFallback
participant tryLoadErrorPage
participant ErrorPageLoader
tryErrorPageFallback->>tryLoadErrorPage: discover and load error page
tryLoadErrorPage->>ErrorPageLoader: resolve or probe page
ErrorPageLoader-->>tryLoadErrorPage: transient failure
tryLoadErrorPage-->>tryErrorPageFallback: return null without caching absence
tryErrorPageFallback->>tryLoadErrorPage: retry after recovery
tryLoadErrorPage-->>tryErrorPageFallback: return rendered 500 response
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Code Review — Score: 88/100 (Good)Solid, narrowly-scoped bug fix: it stops caching a "no custom error page" miss when the underlying failure is transient (a stat/read/import hiccup) rather than a confirmed absence (ENOENT/ENOTDIR across every extension). The fix is easy to reason about, adds no new dependency or public API surface, and comes with strong regression coverage. Strengths
Concerns (non-blocking, worth a look)
Nothing here blocks merge — the fix is correct, well-tested, and appropriately scoped. The two edge-case notes above are suggestions for a fast follow-up or a comment, not requirements. Generated by Claude Code |
|
Note Automatic reviews are paused because your trial's included automatic processing has been used for this period. Upgrade now, or comment "Gitar review" to run a review anytime. Code Review ✅ ApprovedFixes custom error pages being unavailable after transient failures by caching only confirmed absence instead of failed discovery attempts. The fix preserves the generic fallback during outages and retries loading on the next request, with no new retry loop, dependency, or public API change. Comprehensive test coverage includes recovery cases across cold and warm caches, both resolution paths, and missing dependencies. No issues found. OptionsDisplay: compact → Showing less information. Comment with these commands to change the behavior for this request:
Important Your trial ends in 1 day — upgrade now to keep code review, CI analysis, auto-apply, custom automations, and more. Was this helpful? React with 👍 / 👎 | Gitar |
|
@codex review |
|
Codex Review: Didn't find any major issues. Keep it up! Reviewed commit: ℹ️ About Codex in GitHubCodex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback". |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
|
Reviewed the non-blocking feedback against the queued head:
The PR head remains |
Staging verification resultsVerified framework The disposable fixture used the deployed image, read-only application sources, no application credentials, and no mounted service-account token.
Cleanup is complete: the temporary Deployment, Service, ConfigMap, NetworkPolicy, pods, and local port-forward are gone. Manifests and results are retained locally for reproduction. Remaining rollout gateTwo-replica failover was not tested. A fresh quota check still shows 76 GiB of 80 GiB allocated to memory limits, leaving only 4 GiB. Two temporary 4 GiB replicas need 8 GiB, plus whatever rollout headroom the capacity owner requires. Resource sizing is being handled separately; no quotas or existing application workloads were modified for this verification. The single-pod replacement check is not a substitute for multi-replica availability verification. These results do not approve production rollout, and this verification made no production changes. |
Capacity mission checkpoint — 2026-09-07PR head The previously documented two-replica staging failover verification remains open. A fresh read shows staging at 80/80 GiB memory limits, 41.4/42 CPU limits, 33.25/40 GiB memory requests, and 44/55 pods, with the API rolling and still 4/4 available. Two additional 4 GiB verification replicas do not fit the current memory-limit quota. The active capacity lane in veryfront/veryfront-issue-inbox#1041 now explicitly owns the fixture budget and source-of-truth quota/rollout proposal needed to unblock this verification. Renderer memory/root-cause work proceeds separately in veryfront/veryfront-issue-inbox#1040 and #1035. No quota, workload, HPA or replica changes were made for this checkpoint. Existing passing single-pod checks are preserved; multi-replica availability and production rollout are not claimed verified. |
Staging failover unblock preparedThe capacity change is now in veryfront-infrastructure#336, head It proposes staging admission limits of 96 GiB memory / 48 CPU cores, preserving memory/CPU requests, pod quota and availability/isolation settings. With the recorded 76 GiB / 39.9 CPU steady allocation, the two 4 GiB / 1-CPU verification replicas plus one API surge require 86 GiB / 42.4 CPU. The fixture's retained resource metrics show 2 GiB / 250m requests per replica. The recovered 12-worker cluster has sufficient aggregate reservation headroom for this bounded scenario including loss of one large worker, but aggregate arithmetic does not prove node-local placement or peak safety. Run the fixture on separate eligible workers and without unrelated rollout overlap; recheck live quota and usage immediately before execution. Independent review and normal release/apply gates are pending for the capacity PR. The multi-replica failover check has not yet run; no quota was changed. Track execution and results in veryfront/veryfront-issue-inbox#1041. PR #4436 itself remains merged with its previously recorded staging checks preserved. |
Capacity dependency reviewedveryfront-infrastructure#336 is now independently reviewed at head The source proposal is ready for the normal merge/apply gate. Live staging is still 80 GiB / 42 CPU; the proposed 96 GiB / 48 CPU has not been applied. The remaining two-replica staging failover test is therefore still pending approved activation, a fresh quota/usage check, distinct-worker placement and a release window without unrelated rollouts. Existing single-pod/browser evidence remains valid for its recorded scope. Execution and final verification are tracked in veryfront/veryfront-issue-inbox#1041. This update does not claim the multi-replica test or production rollout is complete. |
Staging gate unblocked and verifiedApplied the merged veryfront-infrastructure#336 quota to only the staging ResourceQuota, using source from merge
The remaining two-replica pod-failover gate is now verified using the exact deployed staging renderer image ( Results:
Fixture setup corrections are recorded: projected ConfigMap symlinks were replaced with regular files copied by an init container, preserving the renderer's file-safety policy; the init request was raised to the namespace's existing 64 MiB minimum. During replacement, scheduling briefly waited for old-pod reservations/required affinity to clear, then recovered. These were resolved setup/transition events, not an unresolved quota blocker. The team can continue #4436 staging work. This validates bounded two-replica pod replacement at the sampled request cadence, not arbitrary peak load, a real node outage, production rollout, or activation of the separate memory-recycle PR #4438. Continue avoiding overlapping release waves and retain the existing availability/isolation controls. Tracking: veryfront/veryfront-issue-inbox#1041. Reproduction evidence is retained locally in |
Two-replica staging verification completedThe capacity blocker from the earlier verification is cleared. Retested the currently deployed framework Results:
All temporary resources have been removed: Deployment, Service, ConfigMap, NetworkPolicy, probe pod, renderer pods, and local port-forwards. The manifests and probe logs are retained locally for reproduction. No quota or existing application workload changes were made by this test. This completes the remaining controlled two-replica recovery check for this fix. It is a short verification of graceful pod replacement, not an abrupt node-loss, network-partition, sustained-load, or deferred isolated-executor activation test. Production promotion remains a separate approval and release action; this test made no production changes. |
Strict staging release gate passedRemote E2E Health run 34141011437 completed successfully for staging only:
No source changes, disabled pinning, increased retries, or weakened test expectations were used to obtain this result. The earlier transitive dependency-snapshot 409 is recorded in issue #1042; its cause remains unconfirmed and must not be described as fixed by this rerun. Draft promotion PRs are prepared for the exact staging-verified artifacts: Operator #240 and Server #340. Their CI and artifact-availability checks pass. Both remain drafts with auto-merge disabled, pending capacity-owner sign-off, release-owner risk review, and second-person Code Owner approval. No production deployment was triggered. |



Description
Keep custom error pages available after temporary filesystem or dependency failures.
Previously, a failed discovery, source read, or module import could cache an existing error page as absent. Later requests continued to use the generic fallback even after the underlying failure recovered.
This fix caches only confirmed absence. It preserves the existing generic fallback during an outage and retries loading on the next request. It adds no retry loop, dependency, or public API change.
Related issues
Discovered during verification following #4423. Broader renderer integration remains tracked separately in veryfront/veryfront-issue-inbox#1035.
Type of change
Verification
deno task test:file src/server/handlers/request/ssr: 163 steps pass.deno task test:file src/platform/compat/fs.test.ts: 58 steps pass.codex review --uncommittedandcodex review --base origin/main: no actionable findings for head8af1c6389b43cc1d790316d834115bc5be68d58a.Local verification does not replace staging rollout verification. This PR does not change resource sizing or deploy to production.
Checklist
Summary by CodeRabbit