Skip to content

feat(updates): orchestrate package and system updates on-device - #379

Open
andrescera wants to merge 158 commits into
mainfrom
feat/update-system-overhaul
Open

andrescera wants to merge 158 commits into
mainfrom
feat/update-system-overhaul

Conversation

@andrescera

Copy link
Copy Markdown
Member

Affected repo & language: CeraUI — TypeScript, Svelte, Bun

What

Add one persisted update orchestrator for package updates, signed system-image staging, slot-sync eligibility, stream admission and update transport. Expose settings, capability-gated update details, notifications and a device-local Updates dialog. Prepare the next stable CalVer in package.json without publishing a release.

Why

The legacy package-only flow cannot chain system updates or safely coordinate downloads and streaming. This brings the device control plane up to the update-system-overhaul contract while keeping legacy images on their exact-name package roster.

How to verify

Run bun install --frozen-lockfile, bun run test, the isolated netns tests, bun run check:tech-debt, bun run build:federation, bun run test:release-package-contracts, the workspace typechecks and Biome. Run the functional Playwright update-system spec on desktop and mobile. The PR must remain unmerged until PR-head package deployment and subsequent board drills complete.

Risks

The capable-image paths are not active on released images. On a legacy image, only the 15-name app roster remains actionable. Package self-upgrade and startup must be checked on both bench boards; any regression is fixed on this PR head and the prior installed package version is the rollback target.


Checklist

  • Documentation and affected contracts updated
  • Rebased on current main
  • Local package and unit checks run; browser load failures from the full local run are under hosted CI review
  • Both-board legacy-mode PR-head deployment and exact-head independent review

@coderabbitai

coderabbitai Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 65e4cf39-3be9-4c5f-ae2d-679023271ead

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…ve it

recoverSoftwareUpdate() reattached to a detached apt unit and returned
as soon as onAttached() fired, leaving the actual drain/cleanup and
lastUpdateSucceeded/lastUpdateFailure assignment to run in a detached,
fire-and-forget continuation. Boot's own startup recovery call (before
startUpdateOrchestrator()) could therefore resolve before the outcome
was recorded, and its own crash-to-restart tail (invariant() on a
plain success) could fire before the 3s-later orchestrator tick ever
got a chance to observe the settled state — so a backend restart could
lose the evidence entirely: the unit already gone, and the in-memory
success flag reset by the very crash that was supposed to report it.
resumeOrchestratorState()'s own (separate) recovery call then found
nothing to reattach to and misclassified a successful transaction as
commit_unit_absent_on_resume.

recoverDetachedAptUpgrade() now reports whether the unit was already
finished at inspect time (wasAlreadyFinished). createSoftwareUpdateProcessMonitor()
splits its completion handling into an awaitable settle() (bounded:
drain + cleanup + the getUpdateState()-relevant assignments) and a
still-detached crash/log tail, so a caller can await settle() without
risking that the intentional restart throw gets caught upstream.
recoverSoftwareUpdate() awaits settle() only when the unit was already
finished (bounded); a still-running unit stays fire-and-forget, since
that can take minutes and must never block boot.
An independent review of the previous commit found the settle-await
split alone did not guarantee ordering: the orchestrator's own resume
call to recoverSoftwareUpdateIfRunning() ran only after
startUpdateOrchestrator()'s own async file read (loadOrchestratorState),
which is a real I/O yield point the runtime can use to detect and act
on the crash-to-restart tail's unhandled rejection BEFORE resume ever
reads the settled outcome and persists the resulting phase.

Reordering main.ts so the orchestrator starts (and its resume runs)
BEFORE the general standalone recovery call closes that gap: resume's
own recovery call becomes the first real observation of a finished
unit, and reduceOrchestrator()+deps.persist() (a synchronous,
non-awaited write) run in the same microtask wave the crash tail's
scheduling belongs to, ahead of the yield point where the runtime
would next check for an unhandled rejection. The standalone call after
it is then a harmless no-op whenever the orchestrator already handled
the unit, and remains the only recovery path for a detached transaction
the orchestrator never tracked.

Also hardens settle() and onAttached() so a broadcast/notification
failure can never surface as a thrown exception: the getUpdateState()
-relevant assignments happen first and unconditionally, and the
broadcast/notification calls are wrapped so their failure is logged,
never propagated. Before this, such a failure would reject the whole
recovery attempt via runRecovery()'s catch, losing the very evidence
the recovery path exists to observe (proven by a new test that mocks
broadcastMsg to throw).
A second independent review of a99ae5f traced the fix by inspection and
concluded the boot reorder does close the persist-before-crash race for
the success path, but asked for a test that actually drives
startUpdateOrchestrator() (not just resumeOrchestratorState() directly)
to prove it, rather than relying on inspection alone.

Adds exactly that: seeds a persisted "committing" agent.json, mocks
invariant() to record rather than throw, and overrides only
recoverSoftwareUpdateIfRunning/getPackageInstallWireState/persist on the
real defaultOrchestratorRuntimeDeps. Confirms the crash-tail's .then()
reaction (registered first on the shared settle() promise) fires before
the sibling reaction that unblocks resumeCommitting()'s own
continuation -- expected, since that's promise registration order, not
something boot ordering changes -- and, critically, that this does NOT
prevent the orchestrator's synchronous persist() from running to
completion in the same microtask wave. A real invariant() throw only
becomes a process-terminating unhandled rejection once the engine's
microtask queue fully drains, which is after persist() has already run;
the durable write survives even though the crash attempt is logged
first.

Also independently re-confirms (via the same broadcastMsg spy already
in this file, now cross-checked against a fresh reviewer's claim that
compat.broadcastMsg is a different function than the one
software-updates.ts imports) that the spy target is correct:
apps/backend/src/modules/ui/websocket-server.ts re-exports broadcastMsg
directly from rpc/compat.ts, so spyOn(compat, "broadcastMsg") does
intercept the real call path -- confirmed by the stack trace of the
thrown mock surfacing at the exact onAttached()/settle() call sites.
…dy-applied download resumes

Covers the three round-17 traces end to end: updates disabled keeps the
download and D8 still reaches the commit-stage probe; a throwing probe is
re-asked until conclusive (recovery happens once, or the live unit is
polled); a proven absence with nothing actionable ends idle with no failure,
no installed notice and no quarantine. Also restart idempotence and that a
deferral never judges a newer download.
…nd rediscovered

Replaces the retry-is-a-no-op claim: an empty plan fails the launcher.
Documents the proof requirement, the per-tick re-adjudication of an
undecided resume, rediscovery instead of replay, and what a skipped probe
(updates disabled) does.
…German

"zurückgesetzt" can read as a device reset.
…e generation

The deferred re-adjudication awaited a pending-plan delete before checking
whether the download was still the one it judged, so a stream abort plus a
replacement install during that await could lose the replacement's plan;
enteredAt was used as identity although two attempts can share a
millisecond; and the deferral marker was dropped before the delete, so one
failed delete wedged the download again.

A generation counter now advances on every state change. A probe answer is
applied only if the generation is unchanged, in the same synchronous step as
the transition; the boot probe runs with the persisted download visible so
a stream start meanwhile wins. The adjudication no longer deletes the plan:
every install start now replaces it (rewrite, or clear when discovery holds
no plan), so a dropped download's plan can never reach a reader.
…t failures with the deferred resume

Slow probe while D8 aborts and a replacement install starts, the same-
millisecond replacement, a stream start during the boot probe, a persist
failure on the recovery, and an install that starts without a discovered
plan. The X1 replays now expect the old plan to stay until the next install
rewrites it.
… the deferral is generation-fenced

Absence is not identity-checked; only a loaded unit is. Documents that the
plan file is replaced by the next install rather than deleted on resume and
that a late probe answer never touches a newer attempt.
…s from a non-available wire

Revert maybeStartPackageInstall() to its a6b8210 body: the else-branch that
cleared pending-packages.json deleted the plan of an install restored in
awaiting-idle after a restart, so its commit failure quarantined nothing and
its success named no packages. A dropped download is always followed by
discovery, and an install started from an available wire overwrites the
record before its unit exists, so no cleanup is needed for X1.
The test that asserted the removed deletion now asserts the opposite: an
install started from a non-available wire leaves the record as it was.
Adds the round-19 repro (restart in awaiting-idle, manual install from an
idle wire: failure quarantines the plan, success names it) and the X1
end-to-end proof (proven-absent drop, discovery, the new plan overwrites the
stale one and is the only one quarantined).
…esume suite

Comment-only: the plan is left on disk by the drop and overwritten by the
next install start, which comes through discovery; it is not cleared and
not rewritten by every start.
After a proven-absent drop the next install start comes through discovery
and overwrites pending-packages.json from an available wire; no start clears
it. The restart-in-awaiting-idle start from a non-available wire keeps the
record as before, now listed as an existing limitation. Also removes the
stale DEVICE-UPDATES.md paragraph that said an absent unit returns to
awaiting-idle and retries the install.
…load's plan

A start from an available wire rewrites pending-packages.json, but discovery does not keep the wire available until launch. A wire reset or a restart in awaiting-idle starts from a non-available wire, which keeps the earlier record as at a6b8210. Say so in the recovery docs, the backend AGENTS contract and the comments, and list it as an inherited limitation.
Make the on-device update orchestrator recover safely from interrupted
OS staging and package updates, and prove that behaviour against
captured real-device command output.

- Preserve the validated recovery wire vocabulary and retain the update
  pin identity through recovery.
- Fence the singleton backend writer's ownership and lifetime. Source
  development runs (NODE_ENV=development) are exempt so CI end-to-end
  and `bun run dev` start without the packaged guardian.
- Bind guarded OS staging to physical ownership proof and ship the
  production-only stage guardian in the package.
- Compose durable recovery with startup authority, and publish recovery
  state across the operator surfaces.
- Add normalized real-device captures and replay-test foundations, and
  document the recovery contracts and the limits of the evidence.
…oting competitors

What: Reserve main-table protocol 242 for a single elected-uplink preference. Preserve every foreign default, reconcile both families, serialize applies and releases, sweep crash residue before the first election, and drain/release preferences during termination.

Why: Competitor demotion rewrote NetworkManager/DHCP-owned routes, left residue after failback or removal, and fed prior demoted metrics into later ceilings. Runner reproductions fail for all three defects on cfde40f.

Tests replaced: The demotion-specific assertions in gateway-route-repair and gateways-migration now assert equivalent NIC routability, argv isolation, family selection and exact partial-failure rollback under owned-route semantics. Keeping their old metric/delete expectations would require the defect. No unrelated tests were removed, skipped or weakened; the existing CLI-drift diagnostic test is unchanged.

How to verify: bun install --frozen-lockfile; bun run --filter backend check; bunx biome check on changed TypeScript files; bun run --filter backend test; BUILD_ARCH=amd64 bun run --filter backend build:backend-only. Full backend result: 7915 pass, 8 existing skips, 0 fail across 710 files. Focused route suites: 39 pass. Gateway/connectivity/update-transport/boot regression subset: 129 pass, 2 existing skips. Added lifecycle, rollback, ownership and 512 seeded flip-sequence coverage.

Risks: Protocol 242 is reserved for these main-table routes. IPv4 metric-0 and IPv6 metric-1 competitors cannot be beaten with a lower representable metric; refuse without mutations rather than claiming success. Switches briefly use baseline routing between owned delete/add. Rollback failure stays visible. Prior unmarked DHCP-looking residue cannot be safely swept. Kernel and both-board qualification remain owed: repeat portal/failback/removal/restore, shutdown/restart/crash sweep, family selection and APT/OS transport drills on a rebuilt head. No board contact, deployment, push or publication is part of this change.
… a metric-0 competitor

What: Build owned delete and rollback argv from parsed supported attributes, separate cleanup inventory from candidate validation, and prepend an owned IPv4 metric-0 default when a foreign competitor is already at the floor. Verify the exact FIB winner before accepting acquisition or retaining a floor preference. Split inventory, acquisition and transaction responsibilities into small modules.

Why: Carrier-loss display flags made ip reject cleanup, unrelated foreign source-specific/multipath defaults blocked startup and release, and strict metric subtraction refused the board-confirmed no-metric HiLink case. An identical owned row is not evidence that equal-metric ordering still selects it.

How to verify: The new fixture regressions and all six real-kernel namespace scenarios failed before the correction. Run backend check, changed-file Biome, focused default-route/gateway suites, surrounding gateway/connectivity/update-transport/software-updates/shutdown suites, and the full backend test script. Measured full backend: 7932 pass, 8 unchanged skips, 0 fail across 712 files (baseline 7915/8/0 across 710). BUILD_ARCH=amd64 backend build passed. Kernel tests run the Bun test process itself under unshare -Urn, with dummy/veth links and the real setDefaultRoute/lifecycle over the injected runner; NEW kernel cases skip with the stated probe stderr reason only when user/network namespaces are unavailable. All six ran on Linux 7.2.9 / iproute2 7.2.0 / Bun 1.4.2 without sudo. The 512-transition test checks the intended winner and exact preference presence/family/metric on every transition; gateway argv/family assertions are exact again.

Risks: Protocol 242 must be exclusively reserved for CeraUI main-table defaults; collisions are treated as owned and swept. IPv6 metric-1 competitors still refuse metric-exhausted: never use IPv6 prepend, which can merge a foreign nexthop. IPv4 fibmatch uses a documentation-range destination without sending traffic; more-specific/policy routes or unreadable FIB fail closed. Mutations are not crash-atomic and NM/DHCP may reorder after verification. No board, deployment, publication or network qualification is claimed. Both boards still owe C2a failback/C9 restoration, flapping/carrier-loss, shutdown/restart sweep, v4/v6 selection, HiLink no-metric recovery, APT admission, UID-pinned OS transport coexistence and Wi-Fi second-uplink re-drills from an owner-established clean foreign baseline.
Relocate update-system contracts into main's docs/agents structure while retaining slim routing manifests. Preserve PR paragraphs and both route fixes; non-Markdown changes are exactly main's documentation-reference and documentation-check updates.
…le winner

Stage distinct preferences before retiring the old route. Use nonzero realm generations for same-winner IPv4 floor repairs, ordered endpoint restoration for legacy rows, and pre-mutation refusal when rollback order cannot be preserved. Reject non-main FIB matches and cover both kernel counterexamples and nested rollback failures.
Merges the final route-fix commits (2ac0d51, 90f639f, 4b7d5c6) into the PR branch so the deb drilled on the boards is built from the merged PR head.
Keep fake hrtime in the imported scheduler deadline domain.
Late-loaded parallel workers now exercise the skip retry consistently.
Document the fixture and assert that both declining runner calls occur.
Poll the route table within a two-second monotonic deadline.
Carrier-loss fixtures no longer race asynchronous linkdown propagation.
Keep the original regex assertion and show the last table on timeout.
Source-only sockets follow the current FIB winner even with a unique address. Bind ordinary candidates to the device like duplicate-IP twins, ignoring curlrc and proxies. Keep probes read-only.

Update the old ordinary-roster source-binding assertions because they encoded the misattribution; preserve candidate, exclusion, target, family and twin-independence assertions. Real user/netns packet counters fail on b789561 and kill restored source binding under both 241/0 and 242/49 competitors. The separately device-bound Rock repository timeout remains unproven, not a claimed fix.
Probe every tied foreign winner while an owned preference masks it. Require three healthy completed sweeps spanning ten seconds, resetting after impairment, identity changes or gaps over thirty seconds. Re-read route identities and the final monotonic clock before forward-only release.

This trades a short dwell on Wi-Fi or a metered winner for resistance to one-off recovery and rapid Check clicks. Keep maintenance queued while ownership is held; no foreign edits, metric ratchet, new rule/table or durable recovery state. Seven real-kernel cases cover recovery, interruption, topology, stale gaps, rapid checks and final-clock/topology changes; sticky and no-dwell mutants are killed.
Await independent candidate observations concurrently, then retain record-order repository ranking and HTTP fallback. All candidate results stay bound to their named interface; no timer returns while unseen batch work continues.

Real five-uplink curl/kernel timeout regression takes 30103ms on exact b789561 and 6028ms with this change. Completion barriers catch serialization and completion-order ranking mistakes; a serial-loop mutant is killed. Existing subprocess budgets remain; DNS, local I/O and APT mean this is not a hard UI Check deadline.
… election

A source-specific foreign IPv6 default cannot be modeled for automatic failback, but it must not turn an otherwise healthy IPv4 preference into a refused Check. Withhold recovery on an unsupported foreign shape while preserving fail-closed inventory reads and owned-selector validation.

The real-kernel counterexample is retained RED against 57d5522e, GREEN after the guard, and kills renewed error propagation. Existing acquisition, forward-only release and foreign routes remain unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant