restore/upgrade: stop running from the flash before rewriting it - #224
Conversation
restore and upgrade rewrote the whole flash while the system kept executing from it. umount_all() only detached mounts lazily and left "/" alone, so every process kept paging from the root squashfs, and the first fault on a rewritten block took the system down with the flash half written. isolate_from_flash() now runs between the dry run and the first erase: - mlockall() so ipctool itself is never paged in from flash; - SIGKILL every userspace process but init (not kill(-1), which also hits signal-accepting kernel threads; not SIGTERM, since `udhcpc -R` drops the address on a graceful exit); - if one of them held /dev/watchdog, take it over and feed it through the whole flash. The HiSilicon SDK driver ignores write() and the standard keepalive, so every form is sent; the HISINEW_* ioctls move to watchdog.h; - remount every writable flash filesystem read-only (stopping the jffs2 garbage collector), then unmount; - no global sync() from then on, as it would wait on a dead network mount; reboot fsync()s just the log. xm_disable_watchdog() failed whenever one of the two EV300 watchdog modules was not loaded, which aborted every restore on such a board; a module that is absent is now fine. Tested on a Hi3516EV300 XM board with 16M NOR, both ways, with the log on NFS: XM firmware -> OpenIPC and OpenIPC -> XM (the #223 case). All partitions read back identical to the image, apart from jffs2 blocks written after boot.
PR Summary by QodoIsolate restore and upgrade from flash before rewriting it
AI Description
Diagram
High-Level Assessment
Files changed (4)
|
Code Review by Qodo
1.
|
Review follow-up: - mlockall() failing aborts before anything is stopped (pin_to_ram()), instead of flashing a process that can still fault on the flash. - init is the one process left running and wakes to reap; its executable and libraries are mapped under MCL_FUTURE, which locks the page-cache pages it faults on. - The watchdog feeders are killed and the device taken over first, with no unattended gap behind a full kill pass; a takeover that fails is fatal. - A flash filesystem that will not go read-only is fatal: it can write over the image. - A failure after the isolation can no longer return to a system with no services: it reboots, cleanly when nothing was erased yet. restore now exits nonzero when it does not flash, and umount_fs() no longer exit()s. - The log flush before reboot runs in a child with 5 s to finish, so a stalled share cannot hold the reboot. Re-tested both ways on the Hi3516EV300 XM board (XM -> OpenIPC from telnet, OpenIPC -> XM from ssh with majestic holding the watchdog).
|
Addressed the review in eee2707:
Re-tested both directions on the Hi3516EV300 XM board with this code. The partitions read back identical and both firmwares boot. |
Fixes #223.
restoreandupgraderewrote the whole flash while the system was still executing from it.umount_all()only detached mounts lazily and skipped/, so every process kept paging from the root squashfs. The first page fault on a rewritten block killed the system and left the flash half written.What changes
A new step,
isolate_from_flash(), runs after the dry run (Analyzing) and before the first erase:mlockall), so it is never paged in from flash.kill(-1): that also reaches kernel threads that accept signals, such asjffs2_gcd_*.udhcpc -Rreleases its lease on a graceful exit, which takes down a network share holding the log.write()returns 0 and the board resets 10 s after the open, so every keepalive form is sent, including the HiSilicon one fromwatchdog.c.HISINEW_*defines move towatchdog.hso both files share them.sync()from then on, since it could hang on an unreachable NFS mount.reboot_with_msg()fsyncs only the log, so a log on a share keeps its ending.A terminal on stdout/stderr is redirected to
/dev/console; a redirect to a file is left alone.Also fixed:
xm_disable_watchdog()returned failure whenever one of the two EV300 watchdog modules was not loaded (open_wdt→ ENOENT). That aborted every XM restore on such a board.Testing
Tested on a Hi3516EV300 XM board (XT25F128B, 16M NOR), UPX-packed arm32 build, log redirected to NFS, recovery through U-Boot TFTP on standby:
/dev/watchdog, overlay rootmtddiffers only in jffs2 blocks Sofia wrote after boot; Sofia upEach direction passed twice, the last round with the final binary.
ipctool backupon OpenIPC produces a file whose SHA1s all match.Host build and unit tests pass; the arm32 and arm64 cross builds are clean.