Summary
ipctool restore <file> overwrites the whole flash in offset order while the running OS is still executing from that flash. Once the write reaches the partition that backs the live root filesystem, the system dies under ipctool and the camera is left half-flashed: part vendor firmware, part the old firmware. This happened on a lab camera restoring its own vendor backup over OpenIPC. It now needs serial (UART) recovery.
Setup
- Hi3516EV300, XM board
HI3516EV300_85H50AI, 16M NOR (XT25F128B), original XM U-Boot still in place.
- Running: OpenIPC (2021 build, kernel 4.9.37). Flash layout
boot 256k, kernel 1984k, rootfs 5056k, rootfs_data. / is an overlay with /rom (squashfs on mtdblock2) as the lower layer and /overlay (jffs2 on mtdblock3) as the upper layer.
- Backup: an ipctool backup of the same device, made in 2021 on the XM firmware (
boot, romfs, usr, web, custom, mtd = 0x1000000). All SHA1s were verified and the total size matches the current flash.
- ipctool: master @ 429e8c3, cross-built for arm32 and UPX-packed.
- Precautions taken: ipctool and the backup file were copied to
/tmp (tmpfs), majestic was killed, the watchdog is fed by the [hidog] kernel thread, and the restore ran detached (setsid nohup ... restore -f /tmp/backup.bin).
What happened
Unmounting /rom
Unmounting /overlay
Checking boot... [0] 0x00000000 0x00040000 4876308a 4876308a
... (all SHA1 OK)
Backups were checked
Analyzing boot ... Analyzing mtd
Restoring boot
Restoring romfs
Flashing [wwwwwwwwwwwwwwwwwwwwwwwwwwwwwwwwwwwe................................................]
The log stops at romfs block 35/84, which is absolute flash offset 0x270000, i.e. mtd2 (OpenIPC rootfs) + 0x40000. The ssh session timed out and the camera never answered ping or ARP again. boot had been written and verified (XM U-Boot + XM env), so U-Boot still works. Everything from 0x40000 to 0x270000 is XM data and the rest is OpenIPC data, so the device cannot boot either firmware.
Root cause
src/backup.c:
umount_all() detaches squashfs/cramfs/ubifs and rw mtdblock mounts with MNT_FORCE | MNT_DETACH, and umount_fs() explicitly skips /. A lazy detach only removes the mount from the namespace. The overlay root, and every running process (init, busybox applets, dropbear, syslogd, the shell), still pages text and data from the squashfs on mtdblock2.
do_flash() then erases and writes that partition block by block. The next page fault into squashfs reads erased or foreign data, the root filesystem falls apart, and the system hangs or panics mid-write. mtd_write() gives no retry or rollback, so the flash is left in a mixed state.
- The pre-flash checks (SHA1, total size, the simulated
do_flash("Analyzing")) all pass, because none of them asks whether the running system depends on the blocks about to be overwritten.
On vendor firmware the same thing applies to whatever / is mounted from. The XM path xm_kill_stuff() reduces the risk but does not remove it.
Proposed fix
- Relocate before flashing. Before the first erase: copy a minimal busybox (or rely on ipctool alone, since it is static) into a tmpfs,
pivot_root/chroot into it, kill every process that is not a kernel thread and whose executable or mappings live on flash (keeping the watchdog fed), then echo 3 > /proc/sys/vm/drop_caches. Only then erase anything.
- Otherwise refuse. Resolve the device behind
/ (and behind the overlay lower layer) to an mtd number and flash range. If the restore would overwrite it and step 1 was not done, abort before the first erase with a clear message, and allow an explicit override.
- Cheaper partial mitigation: order the writes so that the region backing the live root is written last. The boot partition and everything else would then be complete before the system is at risk. This is still not safe on its own.
- Optional: add a
--dry-run that runs everything up to and including the "Analyzing" pass and prints the mapping of each backup block onto the current mtd devices, highlighting blocks that overlap the live root.
The same concern applies to upgrade (do_upgrade()), which also calls free_resources() and then rewrites the kernel and rootfs regions.
Recovery (for reference)
U-Boot is intact, so: TFTP the flat 16M image (the backup with the YAML header and the per-block length prefixes stripped), then sf probe 0; sf erase 0x40000 0xfc0000; sf write <ram+0x40000> 0x40000 0xfc0000.
Summary
ipctool restore <file>overwrites the whole flash in offset order while the running OS is still executing from that flash. Once the write reaches the partition that backs the live root filesystem, the system dies under ipctool and the camera is left half-flashed: part vendor firmware, part the old firmware. This happened on a lab camera restoring its own vendor backup over OpenIPC. It now needs serial (UART) recovery.Setup
HI3516EV300_85H50AI, 16M NOR (XT25F128B), original XM U-Boot still in place.boot 256k, kernel 1984k, rootfs 5056k, rootfs_data./is an overlay with/rom(squashfs on mtdblock2) as the lower layer and/overlay(jffs2 on mtdblock3) as the upper layer.boot, romfs, usr, web, custom, mtd= 0x1000000). All SHA1s were verified and the total size matches the current flash./tmp(tmpfs),majesticwas killed, the watchdog is fed by the[hidog]kernel thread, and the restore ran detached (setsid nohup ... restore -f /tmp/backup.bin).What happened
The log stops at romfs block 35/84, which is absolute flash offset 0x270000, i.e.
mtd2(OpenIPC rootfs) + 0x40000. The ssh session timed out and the camera never answered ping or ARP again.boothad been written and verified (XM U-Boot + XM env), so U-Boot still works. Everything from 0x40000 to 0x270000 is XM data and the rest is OpenIPC data, so the device cannot boot either firmware.Root cause
src/backup.c:umount_all()detaches squashfs/cramfs/ubifs and rwmtdblockmounts withMNT_FORCE | MNT_DETACH, andumount_fs()explicitly skips/. A lazy detach only removes the mount from the namespace. The overlay root, and every running process (init, busybox applets, dropbear, syslogd, the shell), still pages text and data from the squashfs onmtdblock2.do_flash()then erases and writes that partition block by block. The next page fault into squashfs reads erased or foreign data, the root filesystem falls apart, and the system hangs or panics mid-write.mtd_write()gives no retry or rollback, so the flash is left in a mixed state.do_flash("Analyzing")) all pass, because none of them asks whether the running system depends on the blocks about to be overwritten.On vendor firmware the same thing applies to whatever
/is mounted from. The XM pathxm_kill_stuff()reduces the risk but does not remove it.Proposed fix
pivot_root/chrootinto it, kill every process that is not a kernel thread and whose executable or mappings live on flash (keeping the watchdog fed), thenecho 3 > /proc/sys/vm/drop_caches. Only then erase anything./(and behind the overlay lower layer) to an mtd number and flash range. If the restore would overwrite it and step 1 was not done, abort before the first erase with a clear message, and allow an explicit override.--dry-runthat runs everything up to and including the "Analyzing" pass and prints the mapping of each backup block onto the current mtd devices, highlighting blocks that overlap the live root.The same concern applies to
upgrade(do_upgrade()), which also callsfree_resources()and then rewrites the kernel and rootfs regions.Recovery (for reference)
U-Boot is intact, so: TFTP the flat 16M image (the backup with the YAML header and the per-block length prefixes stripped), then
sf probe 0; sf erase 0x40000 0xfc0000; sf write <ram+0x40000> 0x40000 0xfc0000.