Repository navigation
ci: add a manual workflow to publish the NixOS base image to Docker Hub - #3477
Conversation
PR SummaryMedium Risk Overview Reviewed by Cursor Bugbot for commit 9f19cab. Bugbot is set up for automated code reviews on this repo. Configure here. |
❌ 2 Tests Failed:
View the top 3 failed test(s) by shortest run time
View the full list of 2 ❄️ flaky test(s)
To view more test analytics, go to the Test Analytics Dashboard |
8ea5415 to
211fc90
Compare
211fc90 to
8d59b48
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 8d59b48. Configure here.
build.sh has been able to build the premade NixOS base image since NixOS support landed, but it defaults to the local dev registry, so no pullable image exists: templates cannot be built `fromImage` a NixOS base, and neither the distro test matrix nor the docs can reference one. Add a workflow_dispatch job that runs build.sh on a hosted runner (nix runs inside the nixos/nix container, so no KVM and no self-hosted pool) and pushes to Docker Hub as `docker.io/<namespace>/nixos:<tag>`, namespace defaulting to e2bdev. The namespace is its own input rather than being derived from DOCKERHUB_USERNAME, so the credential and the image path stay independent; the job fails up front, naming DOCKERHUB_USERNAME/DOCKERHUB_TOKEN, if either secret is unset. The tag input is required and must be unused — the base-layer cache key includes the image reference, so a rebuild under a reused tag is silently ignored in favour of the cached layer, which also means a broken tag cannot be fixed in place. The job therefore refuses an existing tag up front and verifies the rootfs tar before it pushes: ID=nixos in /etc/os-release, the /sbin/init and /nix/var/nix/profiles/system symlinks with their expected targets, the system toplevel those resolve to actually packed, and a non-empty nix/var/nix/db-registration. The checks read the tar rather than running the image because it has no FHS userland before its first activation. `latest` is tagged and pushed only in the push step, after those checks pass, so it always names a verified build. It is a convenience pointer, not a reproducible reference: hash.go keys the base layer on the literal image string rather than the resolved digest, so a template built FROM `nixos:latest` keeps its cache key when latest moves and keeps using the stale base layer until it is force-rebuilt. The README and the job summary both say so, and the immutable per-publish tag stays the reference for reproducible builds. The concurrency group drops the tag so two publishes cannot race to move latest. The unused-tag guard fails closed: `imagetools inspect` exits non-zero for "no such tag" and for 429/5xx/DNS alike, so the exit status alone cannot be read as "tag is free" without risking an overwrite of a published tag. It now classifies the registry's answer, retries an inconclusive one, and aborts if it stays inconclusive. build.sh grows one hook, E2B_NIXOS_SKIP_PUSH, so the workflow can import and inspect before publishing; the default path is unchanged.
8d59b48 to
c774d61
Compare
There was a problem hiding this comment.
LGTM — this is a manual CI workflow (not part of any request-serving path), and the tag-existence race flagged earlier in review has already been fixed and verified by the author in c774d61.
What was reviewed: the tag-existence guard (now classifies inspect output instead of trusting exit status, retries transient errors, fails closed), the rootfs verification gate before push (os-release/init symlinks/closure/db-registration), and the namespace/registry input handling in build.sh — no injection risk since these are quoted shell variables set only by a manual, permissioned workflow_dispatch.
Extended reasoning...
Overview
Adds a workflow_dispatch-only GitHub Actions job that builds the premade NixOS base image via the existing build.sh and publishes it to Docker Hub, plus doc updates. No production service code is touched.
Security risks
The workflow uses DOCKERHUB_USERNAME/DOCKERHUB_TOKEN secrets via docker/login-action, which is standard. Namespace and tag are workflow_dispatch inputs but the workflow itself is manual-trigger and gated by repo permissions (only users with actions:write can dispatch it), so there's no untrusted-input injection surface beyond what a trusted operator already controls. The bug hunting system's candidate issues (no-repo vs no-permission conflation, TOCTOU on tag check, missing namespace validation) were raised and refuted — the classification logic distinguishes these cases explicitly and namespace lacking validation is low-risk given the trigger is manual and permissioned.
Level of scrutiny
This is CI/release tooling for a slow-moving, manually-published artifact, not a runtime code path — appropriate scrutiny is moderate rather than the level given to auth or sandbox-execution code. The one substantive bug found by Cursor bugbot (tag check failing open on Hub errors, allowing a silent overwrite) was already caught, fixed, and verified by the author in commit c774d61, which matches the current diff.
Other factors
Codecov reports full coverage of changed lines and passing tests. Docs were updated in the same change to reflect the new publish flow and the latest-vs-immutable-tag caching caveat, consistent with the project's documentation expectations. No outstanding unaddressed review comments remain on the thread.
There was a problem hiding this comment.
No new issues found on this run. This is a first-time review from me on this PR — the fail-open tag-existence bug that Cursor's bot flagged earlier was already confirmed and fixed by the author in c774d61 (retry + fail-closed classification, verified against the real registry per their comment).
What was reviewed:
- Tag validation/existence-check logic in the workflow (state machine for exists/free/no-repo/unknown, retry with backoff, fail-closed abort on inconclusive answers)
- build.sh changes (image renamed to
nixos,E2B_NIXOS_SKIP_PUSHgating the push behind CI's tar checks, workdir override) - The rootfs tar verification step (ID=nixos, boot symlinks and their targets, closure presence, non-empty db-registration)
- Checked the IMAGE-then-LATEST push ordering for partial-failure risk — a mid-sequence failure leaves
lateststale but doesn't corrupt the immutable tag, which remains the reproducible reference
Given this publishes public, customer-facing base images with mutable latest and a non-trivial amount of new bash logic, I'd still suggest a human pass over the workflow file even though no concrete bugs turned up.
Extended reasoning...
Overview
This PR adds a new workflow_dispatch GitHub Actions workflow (.github/workflows/nixos-base-image.yml) that builds and publishes the premade NixOS base image to Docker Hub, plus small supporting changes to build.sh (image renamed from e2b-nixos to nixos, a new E2B_NIXOS_SKIP_PUSH env var to stop after docker import) and an updated README. The workflow validates secrets and tag shape up front, checks the target tag is unused via docker buildx imagetools inspect with a retry/backoff loop that fails closed on inconclusive registry responses, verifies boot-path invariants in the built rootfs tar before any push, and only then pushes both the immutable <tag> and the moving latest reference.
Security risks
No auth/crypto/permission logic is introduced — Docker Hub credentials are handled via the standard docker/login-action with repository secrets, and the workflow has contents: read permissions only. The main risk class here isn't injection or auth bypass but supply-chain/availability: a bad publish under latest (or an overwritten immutable tag) could silently serve a broken or stale base layer to every template built from it, since phases/base/hash.go caches on the literal image string rather than resolved digest. This PR's own stated purpose is mitigating that risk, and the fail-open version of the tag-existence check (found by Cursor's bot) was already fixed and verified by the author before this run.
Level of scrutiny
This is CI/ops tooling rather than a production request path, but it publishes public, customer-facing artifacts that other templates depend on, and the bash logic (retry/backoff state classification, tar member verification) is non-trivial enough that subtle edge cases in shell parsing or registry error-message matching are plausible. That combination — public blast radius plus first-time, moderately complex script logic — puts this above 'skim and approve' even though the visible logic reads correctly and the author has already demonstrated real-registry testing of the fixed tag-check.
Other factors
No existing test coverage applies to GitHub Actions workflow files (Codecov reports 100% coverage of the Go/testable changes only, none of which is central here). The author's response to Cursor's bug report was thorough and included real-registry verification of all six failure modes, which increases my confidence in that specific fix, but I have not independently exercised the workflow end-to-end (e.g., dry-running the tag-check state machine against a live registry) and it has not yet received a human approval in the timeline.
#3477's rootfs check asserted sbin/init -> $toplevel/init, which this branch changes: that path is the systemd binary from NixOS 25.05 on and runs no activation. Point the assertion at the shim and check the two properties whose absence is silent — that it runs the activation script, and that it sets a PATH (without one, mount/install/ln resolve to nothing before activation).
#3478) The premade NixOS base image built from `channel:nixos-24.05`, whose nixpkgs branch stopped receiving commits on **2024-12-30** — long EOL. Same defect class as E2B#1625 (`from_fedora_image` defaulting to an EOL Fedora), one layer down, and it would otherwise have been the first tag ever pushed to `e2bdev/nixos`. This pins an exact channel release instead of a channel name: a channel resolves at build time, so the same commit evaluated to a different closure every run. The release URL is immutable, so it is a real pin. Bumping `NIXOS_SERIES`/`NIXPKGS_RELEASE` is now the whole maintenance story. Pinning alone doesn't boot. Up to 24.11 `$toplevel/init` was the stage-2 script — it mounted `/proc`, ran `activate` to populate `/etc`, then exec'd systemd. From 25.05 that file **is** the systemd binary (md5-identical to `systemd-260.2/lib/systemd/systemd`); activation moved into the stage-1 initrd. We boot the rootfs directly with no initrd, so PID 1 came up against an empty `/etc` and froze on `Unit default.target not found`. The image now ships `/sbin/e2b-nixos-init`, which does what stage 2 did, and the `nixos` profile points `InitBinary` at it. Two silent traps: it needs an explicit `PATH` (no FHS userland pre-activation, so `mount`/`install`/`ln` aren't found), and it must mount `/proc` before activating — `nix-store` reads `/proc/self/exe`, so the non-fatal `e2bNixDb` snippet was failing and leaving the store DB unloaded. Also fixes the pre-activation `/etc/os-release`, which still hardcoded `24.05` — invisible to every check, since `ID` is the only field read. **Verified on real KVM (x86_64, KVM slot): C1–C11 all pass on 26.05** — envd active and 49983 connected, unit parity + GOMEMLIMIT, chrony, sshd, CA bundle + `https_code=200`, default user + NOPASSWD sudo, hostname, firewall clear, **C9 `check_validity=OK` / 568 valid paths**, `os_release=nixos/NixOS 26.05 (Yarara)`, and C11 raw-envd exec resolving git/jq/sh with reclaim and guestSync replicas `rc=0`. The container build reproduced the host preflight toplevel byte-for-byte. > Includes a merge of `main` to pick up #3477, whose rootfs check asserted the old `sbin/init` target; the assertion now covers the shim.
…ub (#3477) `build.sh` can already build the premade NixOS base image, but it defaults to the local dev registry, so nothing pullable exists. This adds a `workflow_dispatch` job that runs the same script on a hosted runner (nix runs inside the `nixos/nix` container, so no KVM) and pushes `docker.io/e2bdev/nixos:<tag>`. The namespace is its own input rather than derived from `DOCKERHUB_USERNAME`, keeping the credential and the image path independent; the job fails up front naming a missing secret. `tag` is required and must be unused, and the rootfs tar is verified before any push (`ID=nixos`, the boot symlinks and their targets, the toplevel closure, a non-empty `db-registration`). The unused-tag guard fails closed: a 429/5xx/DNS error is retried and then aborts rather than being read as "tag is free". `:latest` is tagged only after those checks, so it always names a verified build — but it is a convenience pointer, not a reproducible reference: `phases/base/hash.go` keys the base layer on the literal image string, not the resolved digest, so a template built `FROM e2bdev/nixos:latest` keeps its cache key when `latest` moves and goes on using the stale base layer until force-rebuilt. The immutable `:<tag>` remains the reproducible reference. **Ordering — this merges, but don't publish from it yet.** `build.sh` currently builds from `channel:nixos-24.05`, whose nixpkgs branch last moved 2024-12-30; it is long EOL and receives no security updates. Merging here is harmless (dispatch-only — it publishes nothing until someone runs it). The 26.05 bump lands separately, and **the first tag ever pushed to `e2bdev/nixos` should be built from a supported release.** Publishing an EOL base as that repo's first image would repeat E2B#1625 (`from_fedora_image` defaulting to an EOL Fedora 42) one layer down.
#3478) The premade NixOS base image built from `channel:nixos-24.05`, whose nixpkgs branch stopped receiving commits on **2024-12-30** — long EOL. Same defect class as E2B#1625 (`from_fedora_image` defaulting to an EOL Fedora), one layer down, and it would otherwise have been the first tag ever pushed to `e2bdev/nixos`. This pins an exact channel release instead of a channel name: a channel resolves at build time, so the same commit evaluated to a different closure every run. The release URL is immutable, so it is a real pin. Bumping `NIXOS_SERIES`/`NIXPKGS_RELEASE` is now the whole maintenance story. Pinning alone doesn't boot. Up to 24.11 `$toplevel/init` was the stage-2 script — it mounted `/proc`, ran `activate` to populate `/etc`, then exec'd systemd. From 25.05 that file **is** the systemd binary (md5-identical to `systemd-260.2/lib/systemd/systemd`); activation moved into the stage-1 initrd. We boot the rootfs directly with no initrd, so PID 1 came up against an empty `/etc` and froze on `Unit default.target not found`. The image now ships `/sbin/e2b-nixos-init`, which does what stage 2 did, and the `nixos` profile points `InitBinary` at it. Two silent traps: it needs an explicit `PATH` (no FHS userland pre-activation, so `mount`/`install`/`ln` aren't found), and it must mount `/proc` before activating — `nix-store` reads `/proc/self/exe`, so the non-fatal `e2bNixDb` snippet was failing and leaving the store DB unloaded. Also fixes the pre-activation `/etc/os-release`, which still hardcoded `24.05` — invisible to every check, since `ID` is the only field read. **Verified on real KVM (x86_64, KVM slot): C1–C11 all pass on 26.05** — envd active and 49983 connected, unit parity + GOMEMLIMIT, chrony, sshd, CA bundle + `https_code=200`, default user + NOPASSWD sudo, hostname, firewall clear, **C9 `check_validity=OK` / 568 valid paths**, `os_release=nixos/NixOS 26.05 (Yarara)`, and C11 raw-envd exec resolving git/jq/sh with reclaim and guestSync replicas `rc=0`. The container build reproduced the host preflight toplevel byte-for-byte. > Includes a merge of `main` to pick up #3477, whose rootfs check asserted the old `sbin/init` target; the assertion now covers the shim.

build.shcan already build the premade NixOS base image, but it defaults to the local dev registry, so nothing pullable exists. This adds aworkflow_dispatchjob that runs the same script on a hosted runner (nix runs inside thenixos/nixcontainer, so no KVM) and pushesdocker.io/e2bdev/nixos:<tag>. The namespace is its own input rather than derived fromDOCKERHUB_USERNAME, keeping the credential and the image path independent; the job fails up front naming a missing secret.tagis required and must be unused, and the rootfs tar is verified before any push (ID=nixos, the boot symlinks and their targets, the toplevel closure, a non-emptydb-registration). The unused-tag guard fails closed: a 429/5xx/DNS error is retried and then aborts rather than being read as "tag is free".:latestis tagged only after those checks, so it always names a verified build — but it is a convenience pointer, not a reproducible reference:phases/base/hash.gokeys the base layer on the literal image string, not the resolved digest, so a template builtFROM e2bdev/nixos:latestkeeps its cache key whenlatestmoves and goes on using the stale base layer until force-rebuilt. The immutable:<tag>remains the reproducible reference.Ordering — this merges, but don't publish from it yet.
build.shcurrently builds fromchannel:nixos-24.05, whose nixpkgs branch last moved 2024-12-30; it is long EOL and receives no security updates. Merging here is harmless (dispatch-only — it publishes nothing until someone runs it). The 26.05 bump lands separately, and the first tag ever pushed toe2bdev/nixosshould be built from a supported release. Publishing an EOL base as that repo's first image would repeat E2B#1625 (from_fedora_imagedefaulting to an EOL Fedora 42) one layer down.