fix(orch): pin nixpkgs to a supported release for the NixOS base image - #3478
Conversation
The premade NixOS base image built from `channel:nixos-24.05`. That branch stopped receiving commits on 2024-12-30 and is long past its ~7 month support window, so the image shipped an EOL userland that no rebuild could refresh. Publishing it as the first tag in a public repo would have made an unpatched base the default for every NixOS template. Pin an exact channel release instead of a channel name. `channel:nixos-XX.YY` resolves at build time, so the same commit evaluated to a different closure on every run — unreproducible, unbisectable, and with no answer to "what changed". The release URL is immutable, so this is a real pin while still being the artifact the channel serves. Bumping NIXOS_SERIES/NIXPKGS_RELEASE is now the whole maintenance story, and is how the image picks up security updates. Also derive the pre-activation /etc/os-release from NIXOS_SERIES rather than hardcoding it: it still read 24.05, and nothing would have caught that, since the only field read from it is ID. system.stateVersion tracks the pin, which is safe here because the image holds no state across the bump — it is rebuilt from scratch on every publish and each sandbox boots it fresh. NOT READY TO MERGE: the 26.05 image builds and imports, but a template built from it fails the base layer — envd never binds 49983 and the build ends with "failed to init envd ... syncing took too long". A control build from the previous 24.05 image on the same dev slot succeeds, so this is a real 26.05 boot-path regression and not the environment. Root cause is still open.
PR SummaryMedium Risk Overview Reviewed by Cursor Bugbot for commit 16fcbc9. Bugbot is set up for automated code reviews on this repo. Configure here. |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Pinning the base image to 26.05 (previous commit) is not enough on its own: the template build fails at the base layer, envd never binds 49983, and the guest console shows systemd freezing with "Unit default.target not found". Up to NixOS 24.11, $toplevel/init was the stage-2 shell script: it mounted /proc and /sys, ran $toplevel/activate to populate /etc from the store, and only then exec'd systemd. From 25.05 that file IS the systemd binary (md5-identical to systemd-260.2/lib/systemd/systemd) and activation moved into the systemd stage-1 initrd. We boot the rootfs directly with no initrd, so pointing init at it lands in PID 1 against an empty /etc — no units at all. Ship an /sbin/e2b-nixos-init shim in the image that does what stage 2 did, and point the nixos profile's InitBinary at it. The shim needs an explicit PATH: there is no FHS userland before activation, so mount, install and ln are not found without one, and the failure is silent. It also has to mount /proc before activating — nix-store reads /proc/self/exe, so the e2bNixDb snippet that loads the store registration errored out and, being deliberately non-fatal, left the store DB unloaded. Verified on real KVM: C1-C11 all pass on 26.05, including C9 (check_validity OK, 568 valid paths) and C11 (raw envd exec resolves git/jq/sh, reclaim and guestSync replicas rc=0). os-release reports NixOS 26.05 (Yarara).
#3477's rootfs check asserted sbin/init -> $toplevel/init, which this branch changes: that path is the systemd binary from NixOS 25.05 on and runs no activation. Point the assertion at the shim and check the two properties whose absence is silent — that it runs the activation script, and that it sets a PATH (without one, mount/install/ln resolve to nothing before activation).
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit ad586f2. Configure here.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ad586f22f9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Review follow-ups, all stale text this branch's own changes created: the workflow's boot-chain comment and the README's 'three pieces of glue' still described /sbin/init -> $toplevel/init, and the README pin table named a NIXPKGS_REV that no longer exists (the pin moved to NIXOS_SERIES/ NIXPKGS_RELEASE when it switched to the release tarball). Also records why the shim deliberately does not 'set -e': NixOS's activate exits non-zero even on a healthy boot (observed: activate_exit=1 in a sandbox passing C1-C11), so aborting PID 1 on it panics the kernel instead of booting a working sandbox — tried, and it fails the template build outright.
🤖 I have created a release *beep* *boop* --- ## [0.0.2](orchestrator-v0.0.1...orchestrator-v0.0.2) (2026-07-31) ### Bug Fixes * **orch:** choose the chrony time source at boot instead of at build time ([#3440](#3440)) ([10b1bae](10b1bae)) * **orch:** pin nixpkgs to a supported release for the NixOS base image ([#3478](#3478)) ([6e8e9d3](6e8e9d3)) * **template-build:** create trailing-slash COPY targets before moving files ([#3458](#3458)) ([dc53fa7](dc53fa7)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com>
…atrix (#3498) Adds the premade NixOS image to `TestTemplateBuildDistroFamilies`, so CI catches a base image that doesn't boot. The tar-only publish gate in `nixos-base-image.yml` provably can't: the pre-shim 26.05 image passed every tar assertion and still froze at boot with no units. The profile installs nothing, so reaching ready proves exactly the parts that are ours — the busybox `Bootstrap`, the `/sbin/e2b-nixos-init` shim (#3478), and the `rm -rf /etc/systemd/system` that lets `setup-etc` take over. **Immutable tag, never `:latest`** — `phases/base/hash.go` keys the base layer on `Config.FromImage` as written, so a republished `:latest` would keep building from the stale cached layer and a failure wouldn't reproduce. **Cost, measured on this run.** Lands in `templates-builds-2` (`Distro` misses `TEMPLATE_BUILDS_SHARD1_RE`). The subtest is **156s** — 515 MB compressed / 1.05 GiB unpacked, cold pull every time. Package wall 166s → 255s, job **5.7 → 7.2 min** against a 30 min cap. It grows the wall by more than 156/4 because the pull and unpack also slow its three concurrent neighbours (utilisation 3.7× → 3.0×). That makes shard 2 the long pole against shard 1's 4m52s; leaving `TEMPLATE_BUILDS_SHARD1_RE` alone — that 2.3 min gap wants its own measuring run, the way #3479 did it. **`compression-tests.tsv`: no change.** The test is already absent; the documented criterion is reading a snapshot back, and provisioning is compression-invariant like the other four cases. Entries are per top-level test anyway, so a subtest couldn't be listed on its own.
#3478) The premade NixOS base image built from `channel:nixos-24.05`, whose nixpkgs branch stopped receiving commits on **2024-12-30** — long EOL. Same defect class as E2B#1625 (`from_fedora_image` defaulting to an EOL Fedora), one layer down, and it would otherwise have been the first tag ever pushed to `e2bdev/nixos`. This pins an exact channel release instead of a channel name: a channel resolves at build time, so the same commit evaluated to a different closure every run. The release URL is immutable, so it is a real pin. Bumping `NIXOS_SERIES`/`NIXPKGS_RELEASE` is now the whole maintenance story. Pinning alone doesn't boot. Up to 24.11 `$toplevel/init` was the stage-2 script — it mounted `/proc`, ran `activate` to populate `/etc`, then exec'd systemd. From 25.05 that file **is** the systemd binary (md5-identical to `systemd-260.2/lib/systemd/systemd`); activation moved into the stage-1 initrd. We boot the rootfs directly with no initrd, so PID 1 came up against an empty `/etc` and froze on `Unit default.target not found`. The image now ships `/sbin/e2b-nixos-init`, which does what stage 2 did, and the `nixos` profile points `InitBinary` at it. Two silent traps: it needs an explicit `PATH` (no FHS userland pre-activation, so `mount`/`install`/`ln` aren't found), and it must mount `/proc` before activating — `nix-store` reads `/proc/self/exe`, so the non-fatal `e2bNixDb` snippet was failing and leaving the store DB unloaded. Also fixes the pre-activation `/etc/os-release`, which still hardcoded `24.05` — invisible to every check, since `ID` is the only field read. **Verified on real KVM (x86_64, KVM slot): C1–C11 all pass on 26.05** — envd active and 49983 connected, unit parity + GOMEMLIMIT, chrony, sshd, CA bundle + `https_code=200`, default user + NOPASSWD sudo, hostname, firewall clear, **C9 `check_validity=OK` / 568 valid paths**, `os_release=nixos/NixOS 26.05 (Yarara)`, and C11 raw-envd exec resolving git/jq/sh with reclaim and guestSync replicas `rc=0`. The container build reproduced the host preflight toplevel byte-for-byte. > Includes a merge of `main` to pick up #3477, whose rootfs check asserted the old `sbin/init` target; the assertion now covers the shim.
🤖 I have created a release *beep* *boop* --- ## [0.0.2](orchestrator-v0.0.1...orchestrator-v0.0.2) (2026-07-31) ### Bug Fixes * **orch:** choose the chrony time source at boot instead of at build time ([#3440](#3440)) ([7fa567c](7fa567c)) * **orch:** pin nixpkgs to a supported release for the NixOS base image ([#3478](#3478)) ([bcdb4aa](bcdb4aa)) * **template-build:** create trailing-slash COPY targets before moving files ([#3458](#3458)) ([a19b99b](a19b99b)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com>
…atrix (#3498) Adds the premade NixOS image to `TestTemplateBuildDistroFamilies`, so CI catches a base image that doesn't boot. The tar-only publish gate in `nixos-base-image.yml` provably can't: the pre-shim 26.05 image passed every tar assertion and still froze at boot with no units. The profile installs nothing, so reaching ready proves exactly the parts that are ours — the busybox `Bootstrap`, the `/sbin/e2b-nixos-init` shim (#3478), and the `rm -rf /etc/systemd/system` that lets `setup-etc` take over. **Immutable tag, never `:latest`** — `phases/base/hash.go` keys the base layer on `Config.FromImage` as written, so a republished `:latest` would keep building from the stale cached layer and a failure wouldn't reproduce. **Cost, measured on this run.** Lands in `templates-builds-2` (`Distro` misses `TEMPLATE_BUILDS_SHARD1_RE`). The subtest is **156s** — 515 MB compressed / 1.05 GiB unpacked, cold pull every time. Package wall 166s → 255s, job **5.7 → 7.2 min** against a 30 min cap. It grows the wall by more than 156/4 because the pull and unpack also slow its three concurrent neighbours (utilisation 3.7× → 3.0×). That makes shard 2 the long pole against shard 1's 4m52s; leaving `TEMPLATE_BUILDS_SHARD1_RE` alone — that 2.3 min gap wants its own measuring run, the way #3479 did it. **`compression-tests.tsv`: no change.** The test is already absent; the documented criterion is reading a snapshot back, and provisioning is compression-invariant like the other four cases. Entries are per top-level test anyway, so a subtest couldn't be listed on its own.

The premade NixOS base image built from
channel:nixos-24.05, whose nixpkgs branch stopped receiving commits on 2024-12-30 — long EOL. Same defect class as E2B#1625 (from_fedora_imagedefaulting to an EOL Fedora), one layer down, and it would otherwise have been the first tag ever pushed toe2bdev/nixos. This pins an exact channel release instead of a channel name: a channel resolves at build time, so the same commit evaluated to a different closure every run. The release URL is immutable, so it is a real pin. BumpingNIXOS_SERIES/NIXPKGS_RELEASEis now the whole maintenance story.Pinning alone doesn't boot. Up to 24.11
$toplevel/initwas the stage-2 script — it mounted/proc, ranactivateto populate/etc, then exec'd systemd. From 25.05 that file is the systemd binary (md5-identical tosystemd-260.2/lib/systemd/systemd); activation moved into the stage-1 initrd. We boot the rootfs directly with no initrd, so PID 1 came up against an empty/etcand froze onUnit default.target not found. The image now ships/sbin/e2b-nixos-init, which does what stage 2 did, and thenixosprofile pointsInitBinaryat it. Two silent traps: it needs an explicitPATH(no FHS userland pre-activation, somount/install/lnaren't found), and it must mount/procbefore activating —nix-storereads/proc/self/exe, so the non-fatale2bNixDbsnippet was failing and leaving the store DB unloaded.Also fixes the pre-activation
/etc/os-release, which still hardcoded24.05— invisible to every check, sinceIDis the only field read.Verified on real KVM (x86_64, KVM slot): C1–C11 all pass on 26.05 — envd active and 49983 connected, unit parity + GOMEMLIMIT, chrony, sshd, CA bundle +
https_code=200, default user + NOPASSWD sudo, hostname, firewall clear, C9check_validity=OK/ 568 valid paths,os_release=nixos/NixOS 26.05 (Yarara), and C11 raw-envd exec resolving git/jq/sh with reclaim and guestSync replicasrc=0. The container build reproduced the host preflight toplevel byte-for-byte.