Build and install the capture extensions with the default native build - #144
Merged
Merged
Conversation
zaoxing
force-pushed
the
feat/capture-build-install
branch
2 times, most recently
from
September 24, 2026 14:07
bf2e908 to
782200a
Compare
zaoxing
force-pushed
the
feat/native-capture-service
branch
from
September 24, 2026 14:46
d14ca30 to
b3dfb0d
Compare
The documented rebuild is `make -C native clean && make -C native`. clean removes build/, where _dmi_native_sink and _dmi_native_store lived, and `all` built only _native_backend, so the documented command deleted both capture extensions and nothing put them back until someone remembered their separate targets. `all` now depends on a `capture` target that links both extensions in build/ and installs them next to _native_backend in src/dmi/, the first directory the loader searches; clean removes the installed copies too. It is unconditional rather than "when libcurl is found", because skipping silently would strand them again and surface only at runtime; CAPTURE=0 builds the backend alone. The targets are now the files they produce, so an up-to-date extension is not relinked on every make, with their headers listed by directory so the file targets cannot go stale; build/_dmi_native_sink and build/_dmi_native_store stay as aliases for CI and the loaders' error messages. `capture` is a CPU-only goal. CURL_INCDIR/CURL_LIBDIR: explicit values still win (CI passes the system paths), then `pkg-config libcurl`, then the extracted dev package under CURL_SYSROOT (default unchanged, which is how this host builds). A check-libcurl step, order-only before every curl target, compiles and links one libcurl call and on failure names libcurl4-openssl-dev instead of a page of compiler output. No torch RUNPATH on the sink: a bare import without torch fails on libtorch_cpu.so, but every loader route imports torch first (dmi.storage.capture does), and that route loads it. Evidence: 7 new cpu tests in tests/test_cpu_native_build.py, 6 failing before the change (the CAPTURE=0 guard passes on both); 13 cpu passed after. Real run: make -C native clean, then make -C native capture rebuilt both into native/build and src/dmi, a second make relinked nothing, and both loaded through the normal loaders from src/dmi. Full build (C++20 Makefile copy, CUDA 13.0, sm_89): clean then all produced _native_backend, _dmi_native_sink and _dmi_native_store; all three load through the loader, the sink bound to the real RecordSink (no stand-ins). The package-layout wheel still contains no .so. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
install.md, huggingface.md, vllm.md and megatron.md tell users to run `make -C native clean && make -C native`, which now also rebuilds _dmi_native_sink and _dmi_native_store. Say so, name the capture target and the build/ aliases, add libcurl4-openssl-dev (and pkg-config, which the build uses to find it) to the system prerequisites, document the CURL_SYSROOT / CURL_INCDIR / CURL_LIBDIR fallbacks and CAPTURE=0, and add troubleshooting entries for the libcurl error and a missing capture extension. code-organization.md names where the extensions are installed. Evidence: the new install.md smoke check loads _native_backend, _dmi_native_sink and _dmi_native_store from src/dmi through the loader after `make -C native clean` and the full build on this host. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
The wheel guard looked only for dmi/_native_backend*.so. The native build now installs _dmi_native_sink and _dmi_native_store into src/dmi beside it (and `make host` puts _host_backend there), so a packaging change that started shipping one of those would have passed. Check every .so under dmi/ instead. Evidence: _validate_archive on a synthetic wheel carrying dmi/_dmi_native_sink*.so was accepted before and is refused after; the real wheel built from a tree holding all three installed extensions has no .so members and is still accepted. `make test-package` itself cannot run on this host (ensurepip in the temporary venv exits 127), so the full check was not run locally. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
… only if it is there The documented no-root fallback extracted only libcurl4-openssl-dev. That package's libcurl.so links to libcurl.so.4.6.0, which ships in the runtime package libcurl4, so the sysroot held a dangling link; ld skipped it for libcurl.a, and check-libcurl failed telling the user to install the very package they had just extracted (without the check, the store linked the static libcurl and failed at import on GSS_C_NT_HOSTBASED_SERVICE). It worked on this host only because /tmp/opencode/sysroot also holds the runtime library. install.md and the Makefile now extract both packages, and check-libcurl names a dangling libcurl.so as the missing runtime package. The default build also started searching /tmp/opencode/sysroot on every host where pkg-config did not report libcurl: a world-writable /tmp path, passed as -I/-L ahead of the system directories. The sysroot default now applies only when that directory exists and belongs to the user running make (CURL_SYSROOT_DEFAULT, so the rule is testable); otherwise no -I/-L is passed and the compiler's default paths apply. Explicit CURL_INCDIR, CURL_LIBDIR or CURL_SYSROOT still win, so CI's invocations are unchanged. Evidence: two new cpu tests in tests/test_cpu_native_build.py failed first (generic message; sysroot flags present with no sysroot), then pass; the file's cpu tests: 15 passed. From an empty scratch sysroot, the documented recipe (apt-get download libcurl4-openssl-dev libcurl4, dpkg-deb -x each) passes check-libcurl, builds _dmi_native_store against it, and the loader imports it; with the dev package alone the new message names libcurl4. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
…changes Naming the capture targets after their files made them up to date whenever their sources and headers were, and nothing about torch, pybind11 or the Python ABI was a prerequisite. A torch upgrade in the venv therefore left the sink linked against the old libtorch, and the command the loader's ImportError recommends (make -C native build/_dmi_native_sink) printed "Nothing to be done"; before the rename it always relinked. Both extensions now depend on build/torch-extension.config, a record of the torch version and git revision, its C++ ABI flag, include and library paths, the Python include directory and EXT_SUFFIX, written only when it changes by the same write-stamp helper as the CUDA toolkit stamp. Evidence: the new cpu test (prerequisite in make's rule database; a real stamp run under a temporary BUILD_DIR that keeps its mtime when unchanged and rewrites a stale record) failed first, then passes. By hand: a second `make capture` relinks nothing, and after the stamp's torch version is edited, `make build/_dmi_native_sink` relinks the sink. A full clean + `all` (CUDA 13.0, sm_89, C++20 scratch copy of the Makefile) builds all three extensions, a second `all` runs no compiler, and all three load through the loader. make test-cpu with CUDA_VISIBLE_DEVICES="": 2224 passed, 1 skipped (no CUDA device), 324 deselected. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
The default build now includes the capture storage extensions, so on a host without libcurl's dev package `make -C native -j` fails at check-libcurl before it builds anything, _native_backend included. The reviewer reproduced it with `make -C native -j16 SM_ARCH=sm_89 CURL_SYSROOT= PKG_CONFIG=false` after a clean: exit 2, no .o, no .so. The message named libcurl4-openssl-dev and CURL_INCDIR/CURL_LIBDIR but not the opt-out, so a user of the in-memory path, which needs no capture extensions, had no hint how to get the backend back. The message now ends by saying to pass CAPTURE=0 to build without the capture storage extensions, and docs/install.md carries the same hint beside the `make -C native -j` step. A dry run of the default goal with CURL_SYSROOT= PKG_CONFIG=false plans the libcurl check under CAPTURE=1 and not under CAPTURE=0, while both still plan _native_backend. test_missing_libcurl_names_the_package asserts the message names CAPTURE=0 (it failed before the Makefile change), and now runs with BUILD_DIR under tmp_path, as the torch-relink test does, so its real make no longer writes the torch stamp into the checkout's native/build. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
pkg-config assigned CURL_INCDIR and CURL_LIBDIR with := before the sysroot's ?= lines were read, so a CURL_SYSROOT the user passed was ignored whenever `pkg-config --exists libcurl` succeeded. The review showed it with a fake pkg-config: `make CURL_SYSROOT=/sysroot` compiled with -I/-L from pkg-config and never searched /sysroot. Naming a sysroot is the only reason to pass one, so a non-empty CURL_SYSROOT from the command line or the environment now skips pkg-config; the default sysroot still ranks behind it, and an empty CURL_SYSROOT= still leaves pkg-config in charge. test_libcurl_directories_resolve_in_documented_order gains the command-line and exported cases; both failed before this change (the link line carried the pkg-config directories) and pass after it. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
`ifeq ($(CAPTURE),1)` treated every other value as CAPTURE=0, so CAPTURE=yes, CAPTURE=true, or an unrelated CAPTURE exported by some other tool built the backend without the capture extensions and said nothing -- the silent stranding the default build exists to prevent. The value is now stripped and must be exactly 0 or 1; anything else stops make with an error naming the variable, where it came from, and both accepted values: Makefile:211: *** CAPTURE must be 0 or 1, not 'yes' (from the command line): 1 builds the capture storage extensions with the default build, 0 leaves them out. Stop. The check runs at parse time, so it applies to every goal, clean included: a CAPTURE that means something else is worth a stop wherever it shows up. New tests: test_capture_accepts_only_zero_or_one (CAPTURE=yes, =true, and an exported CAPTURE, each through `make -n clean`) and test_capture_setting_ignores_surrounding_whitespace. All four failed before this change -- the first three exited 0, and CAPTURE=' 1 ' left the extensions out because of its trailing space -- and pass after it. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
The loader (dmi/transport/native.py, _search_dirs) tries src/dmi/ before native/build/. With src/dmi holding copies, anything that rebuilt native/build alone left the copy stale and in front: the review reproduced it with a branch without this change, whose `make build/_dmi_native_sink` writes only native/build, and the loader kept importing the copy this branch had installed. src/dmi now gets a relative symlink to the build/ file, one inode under two names, so a rebuild in either place is what gets imported. A copy an earlier build installed is replaced even when it is newer than build/, which the file-named target alone would have called up to date. Checked before switching, in a scratch build of both extensions: - the loader's spec_from_file_location path imports each through the link, and __file__ is the link; - in one process, the loader route and the torch suites' route (native/build on sys.path) give the same module object in either order, as they did with copies; - `rm -rf` in `clean` removes the links, not what they name, and a dangling link (build/ removed) is skipped by the loader's exists(); - the package-layout wheel built from the tree has no .so with valid or dangling links present, and _validate_archive still refuses a wheel with a symlinked extension added. (check_package.py's venv smoke step fails on this host for every tree, with or without extensions: EnvBuilder copies the uv Python binary, whose $ORIGIN/../lib libpython is then missing, exit 127.) test_capture_extensions_install_as_links_to_the_build installs into a stand-in checkout holding newer copies, then rebuilds build/ by rename the way ld does. It failed before this change (the copy stayed) and passes after it; a second run redoes nothing. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
The torch stamp recorded torch, its ABI and include/library paths, and the Python build, but not CXX or the libcurl flags. Pointing CURL_INCDIR or CURL_LIBDIR at another libcurl, or switching compilers, changed no source file and no recorded field, so an up-to-date _dmi_native_store kept its old link. The record now ends with CXX, CURL_CPPFLAGS and CURL_LDFLAGS, and is still rewritten only when it changes. It stays one stamp for both extensions, so the sink (no libcurl) also relinks on a libcurl change; that costs seconds and avoids a second stamp. test_capture_extensions_relink_when_compiler_or_libcurl_changes builds the store three times for real in a stand-in checkout, with a compiler that logs its link lines and writes an empty file: first build links, an unchanged rerun does not, and changing CURL_INCDIR (or CXX) relinks once with the new value on the line. Both cases failed before this change (0 relinks) and pass after it. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
install.md said the build took libcurl from pkg-config and otherwise from the compiler's default paths. The Makefile also falls back to CURL_SYSROOT_DEFAULT=/tmp/opencode/sysroot -- a default that predates this branch and that one development host's builds rely on -- and, since the CURL_SYSROOT fix, ranks an explicit CURL_SYSROOT ahead of pkg-config. A reader of the docs could not know either, nor why a /tmp path shows up in a link line. The curl paragraph now lists the five steps in the Makefile's order: explicit CURL_INCDIR/CURL_LIBDIR, an explicit non-empty CURL_SYSROOT, pkg-config, the default sysroot (used only if the directory exists and is owned by the user running make, and movable with CURL_SYSROOT_DEFAULT=), then the compiler's defaults -- each directory resolved on its own. The Makefile's comment names the default's origin the same way. The default itself stays. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
… command Rebasing onto #138 brought its C++20 guard, which checks every compile in `native all` that reads torch's include directory. This branch puts the capture extensions in `all`, and the guard then failed on "2 of 13 torch-including compile(s) below C++20", for two reasons: - The store extension compiled at -std=c++17 while passing torch's include directory for pybind11. It now builds as C++20, like every compile that reads that directory. Its sources are the same C++17-clean catalog and store code the drivers build; rebuilt as C++20 it compiles with 0 warnings, and the storage service and native reader parity suites pass on it (60 passed), with the wiring suite (37 passed). - The guard read the dry run line by line, so a recipe continued with backslashes had its -std= on one line and its torch -I flags on the next, and the continuation read as a compile with no standard. It now joins continued lines before judging them. With the store put back at C++17, the guard flags exactly that one compile. tests/test_cpu_native_build.py: 33 passed, the gpu-planned cases included. CPU tier: 2299 passed, 1 skipped (no CUDA device). Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing
force-pushed
the
feat/capture-build-install
branch
from
September 24, 2026 15:01
782200a to
162e0fb
Compare
zaoxing
added a commit
that referenced
this pull request
Sep 24, 2026
#144 landed - #138 is on main, so the temporary step that applied its C++20 flags with sed changed nothing and would only raise its "delete me" warning. Deleted, as its own comment asked. - #144 puts the capture extensions in `make -C native all`, and the store needs libcurl's headers; without them the build now stops before anything else, by design. The job installs libcurl4-openssl-dev and pkg-config and lets the Makefile find libcurl through pkg-config -- its default lookup, with no CURL_* override -- so it builds exactly what docs/install.md tells a user to run. It lists the three extensions it expects in src/dmi afterwards. Validated locally: the workflow parses and the job's steps are as listed; ruff 0.16.5 F821 passes on the merged src/dmi; with #144's Makefile the CPU targets build and install as links, and the chain test passes (2 passed). The CUDA build itself runs only in CI. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing
added a commit
that referenced
this pull request
Sep 24, 2026
…, backend compile, undefined-name lint (#145) * Drive the real native sink through to the catalog in the live job Every live suite staged its packs through the conformance_sink driver, which feeds the sink core JSON and base64, so the torch-facing NativePackSink that the ring hands envelopes to never reached a catalog in CI. tests/test_native_capture_chain_live.py drives it with torch CPU tensors (through its attach/submit_envelope test entry points) into a spool, runs CaptureStorageService in-process against ClickHouse and the signature-verifying fake S3, and reads the payloads back through NativeCaptureReader, requiring the bytes to match. No CUDA, no ring, no conformance driver. It also sends one 68 MiB pack up as a multipart upload, twice. The fake S3 accepts parts of any size, takes the part list on trust and never checks the completion ETags, so against it the test asserts the part sizes the client sent. Against MinIO, S3's 5 MiB part floor is enforced by the server; the MinIO case skips only when DMI_MINIO_ENDPOINT is unset. The clickhouse-live job now builds build/_dmi_native_sink and starts MinIO (quay.io image pinned by digest, via docker run because a service container takes no command) and sets that variable, so its skip gate turns a lost variable into a red job. Evidence, local ClickHouse 127.0.0.1 and a MinIO binary extracted from the pinned image on 127.0.0.1:19100: 3 passed; with the endpoint unset, 2 passed, 1 skipped. With the client's multipart chunk shrunk to 4 MiB (not committed), the fake-S3 case failed on the part count (18 != 5) and the MinIO case on CompleteMultipartUpload EntityTooSmall. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Compile the full native backend and the ring tests in CI Nothing in CI compiled _native_backend: the cpu job builds only the CPU-only goals and dry-runs `host`, so the .cu sources, bindings.cpp and the torch/CUDA link could break unnoticed, as the C++20 floor of torch 2.14 did (#138). The new native-backend-compile job builds clickhouse-cpp (docs/install.md step 5), runs `make -C native all`, imports the result, compiles all six tests/native/ring binaries and runs the two that need no device. It installs the CUDA 13.0 pieces from NVIDIA's apt repository on the runner instead of using an nvidia/cuda devel container. A container job pulls its image before any step can free disk, and that image is 3.68 GiB compressed before the ~4.4 GB that torch 2.14+cu130 takes installed. The package list (nvcc, cudart-dev, cublas/cusparse/cusolver dev, about 2.7 GB by Installed-Size) is what a real build reads: the backend's -MMD files and `nvcc -M` of the ring tests, mapped with dpkg -S. #138 is not in this base, so a step makes #138's three -std=c++17 -> c++20 edits with sed; after #138 merges it is a no-op and should be deleted. Local evidence: the sed step run against this Makefile changes exactly lines 125, 130 and 175 (the #138 diff) and against #138's head changes nothing; with those edits, CUDA_HOME=/usr/local/cuda-13.0, SM_ARCH=sm_89 and the libcuda stub on LIBRARY_PATH, `make all` built and passed check-link, and the backend imported with no GPU visible. The stock flags stop at bindings.cpp on "#error C++20 or later compatible compiler is required". Ring tests: all six built with CXX_STD=c++20, and test_record_consumer (14 passed) and test_clickhouse_record_sink (21 passed) ran with CUDA_VISIBLE_DEVICES empty. Not verified until CI runs: the apt install on ubuntu-24.04, disk headroom, gcc 13 (local is 11.4), and the job as a whole. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Make src/dmi pass ruff's undefined-name check ruff F821 over src/dmi reported 14 names, in two groups, and neither is a NameError today: - config.py annotates capture_sink_config and capture_storage_config with the strings "NativeSinkConfig" and "NativeCaptureStorageConfig" and imports neither. With `from __future__ import annotations` nothing evaluates them at runtime, but typing.get_type_hints(MonitoringConfig) raises NameError on the first one (reproduced; nothing in src/ calls it today). They are now imported under TYPE_CHECKING, so a static tool resolves them and the module still loads no storage package at runtime (checked: importing dmi.config loads no dmi.storage module). get_type_hints still raises, as before; B2's move of NativeSinkConfig into live code is where a runtime import can go. - hooks/specs.py compares against HOOK_TYPE_Q and eleven siblings that the module binds through a globals() loop over _HOOK_DEFS, which a linter cannot see. Each of those ten lines now carries its own `# noqa: F821` with a comment saying why: per line rather than per file, so the rest of the module stays checked. A typo inside one of those ten lines would now be suppressed too; that is the cost. No behaviour change. ruff 0.16.5 `check --select F821 src/dmi`: 14 errors before, "All checks passed!" after; tests/test_tp_shapes.py, test_hook_spec_flags.py and test_config.py: 63 passed. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Lint src/dmi for undefined names in the cpu job An undefined name is a NameError that waits for its line to run, and the cpu job had nothing that would see one first: `make check` compiles with `python -m compileall`, which accepts any name. The new step runs ruff's F821 rule, and only that rule, over src/dmi, pinned to ruff 0.16.5 with GitHub annotations. The 14 existing hits are handled in the previous commit, in the source, not excluded here. Local evidence with ruff 0.16.5: the step's command passes on this tree, and with an undefined name appended to src/dmi/config.py (not committed) it exits 1 naming the file and line. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Check every catalog metadata field in the capture chain test The chain test compared only capture_id, payload bytes, shape and dtype, so a metadata field the torch sink's parser mangled went unnoticed. A reviewer changed native/csrc/sink/record_row.cpp to store step_number + 7 and the chain test still passed, as did every CPU suite. This chain is the only CI test that reaches record_row.cpp, so it has to catch that. Each capture's read-back descriptor is now compared field by field against the CaptureMetadata the test submitted for it: all 21 fields, with shape as a tuple and a NULL adapter_revision as the empty string the native reader returns for it. The integer fields now also take distinct values (token_start and batch_position vary apart from step_number), so a parser that wires one field into another's slot fails too. producer_rank stays fixed, at 1, because the sink seals a pack when the rank changes. The fake-S3 multipart check counted every logged PUT carrying a partNumber, so one transient retry would have failed the part count. It now keys parts by part number, keeping the last attempt. Evidence, from a scratch copy of HEAD built with make build/_dmi_native_sink build/_dmi_native_store: - step_number + 7 mutant: 2 failed, 1 skipped (MinIO unset); both fail on {'step_number': 7} != {'step_number': 0} for chain-0000. - unmutated rebuild: 2 passed, 1 skipped. Same in this worktree. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Drop the unused E402 noqa from the capture chain test The fake-S3 import sits among the module's top-level imports, before any statement, so E402 never fires on it and the directive suppresses nothing. ruff 0.16.5 with --select E402,RUF100 reports it as an unused noqa (unused: E402); with it removed the same check passes. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Warn when the C++20 sed step has nothing left to change The step exists only until #138 moves the Makefile to -std=c++20, and it had no way to say when that happened: after the merge the sed matches nothing, and the closing grep still counts 3, because the three lines already read c++20. The step would pass silently forever. An unchanged native/Makefile after the sed is exactly that state (before #138 the sed always rewrites all three lines), so the step now emits a ::warning:: asking for its own deletion. Rehearsed on a scratch copy of the Makefile: at this branch's Makefile the diff is 3 lines and no warning; with those edits committed, as #138 would leave them, the warning prints and the count is still 3. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Check the capture sink binds the backend's RecordSink in CI _dmi_native_sink picks its RecordSink base once, when it loads: the _native_backend registration when that backend is loadable, module-local stand-ins otherwise. Only the first is attachable, since create_record_runtime checks isinstance against the backend's class. The cpu and live jobs build the sink with no backend, and native-backend-compile built the backend with no sink, so the backend branch of test_the_sink_derives_from_the_engines_record_sink ran in no job. native-backend-compile now builds build/_dmi_native_sink after the backend, with the same torch and PYTHON, and checks the binding in both import orders: through dmi's loaders (backend first, as the engine does) and as a bare import (the sink finds the backend itself). Each asserts RING_TYPES_ARE_STANDINS is false and NativePackSink subclasses the backend's RecordSink. Then it runs the adapter test and refuses a skip. The assertions are needed because the test alone is not enough. Rehearsed on a scratch copy of this branch with a copied torch 2.14+cu130 backend: with the backend present, both checks print an MRO through _native_backend.RecordSink and the test passes (1 passed). With the backend removed, both assertions fail on the stand-ins, and the test still passes, because its stand-in branch is a pass by design. The rehearsal also found that the bare import needs torch imported first, since the .so has no rpath to libtorch_cpu, so the check does that, as the suite does. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Bound the MinIO health probe and drop the last-release claim The health loop ran curl -sf with no timeout, so a socket that accepts and never answers would have held the first probe until the 30-minute job timeout, and the loop's 30 attempts would never have run. Each probe is now capped with --max-time 2. Checked locally against a listener that accepts and sleeps: curl gives up after 2.01 s with exit 28, which the loop treats as not up yet and retries. The comment also called the pinned release "the last release published as an image". Review found later tags on quay.io, for example RELEASE.2025-09-07T16-13-09Z.hotfix.7aa24e772. (I could not list the tags from here: quay's API wants a token.) The pin is still right, so the comment now says only what it does: tag and digest together keep an upstream retag from changing what the job tests against. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Describe the runner disk step as a precaution, with measured figures The job comment justified freeing disk with "the ~14 GB a hosted runner guarantees", but the runner this job actually got had far more. The df output from run 35933626430 (job 107425626899) shows 87 GB free before the step, 103 GB after it, 99 GB after the CUDA apt install and 94 GB after torch. Freeing the Android SDK and .NET took about two minutes of that run. The step stays, since the free space belongs to the runner image and can change, but both comments now call it a precaution and quote only the figures that log shows. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Pull MinIO from Bitnami's archive, and let a warning pass the sink check Two CI failures on the previous push, neither in the code under test: - The MinIO step's image could no longer be pulled. The MinIO project has withdrawn its public images: quay.io/minio/minio and docker.io/minio/minio both now answer anonymous pulls with 401, including `latest`, although this job pulled the pinned quay image fine a day earlier. The step now runs bitnamilegacy/minio:2025.7.23-debian-12-r5, pinned by digest -- a frozen archive, which suits a fixture whose S3 multipart rules do not change. The Bitnami image starts its own server from MINIO_ROOT_USER/_PASSWORD, so the `server /data` argument is gone, and the health loop allows about two minutes for its entrypoint's setup. The live skip gate then failed only because the suite never ran. - The sink-binding check passed ("1 passed, 1 warning in 1.50s") but the step greps for '^1 passed in ', and torch's NumPy warning changed the summary line. The step now accepts a summary with warnings and still fails on any skip. Validated locally: the workflow parses, and the new pass condition accepts "1 passed" and "1 passed, 1 warning" and refuses "1 skipped". The image's digest, amd64 manifest and entrypoint were read from Docker Hub's registry; running it needs docker, which only CI has. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Drop MinIO from the live job MinIO was there for one check: that a multipart pack follows real S3's rules (every part but the last at least 5 MiB, and a "-N" ETag), which the in-repo fake S3 does not enforce. Nothing in DMI runs against MinIO -- the object store it uses is Garage -- and the MinIO project has withdrawn its public images, so the job had come to depend on a frozen third-party archive for a store DMI does not use. The CI step, its DMI_MINIO_* settings and the MinIO test case are removed. The fake-S3 multipart case stays: it asserts the parts the client sends (every part but the last is exactly the client's chunk and at least 5 MiB) and reads the pack back byte-exact. Checks against a real store belong to the Garage suites. Chain test against local ClickHouse: 2 passed, 0 skipped; no tables left. The workflow parses, and the only step removed from any job is the MinIO one. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Build the documented native build in the compile job, now that #138 and #144 landed - #138 is on main, so the temporary step that applied its C++20 flags with sed changed nothing and would only raise its "delete me" warning. Deleted, as its own comment asked. - #144 puts the capture extensions in `make -C native all`, and the store needs libcurl's headers; without them the build now stops before anything else, by design. The job installs libcurl4-openssl-dev and pkg-config and lets the Makefile find libcurl through pkg-config -- its default lookup, with no CURL_* override -- so it builds exactly what docs/install.md tells a user to run. It lists the three extensions it expects in src/dmi afterwards. Validated locally: the workflow parses and the job's steps are as listed; ruff 0.16.5 F821 passes on the merged src/dmi; with #144's Makefile the CPU targets build and install as links, and the chain test passes (2 passed). The CUDA build itself runs only in CI. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y --------- Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #143. Merge #143 first; GitHub then retargets this PR to
main.Problem
make -C native clean && make -C nativedeleted_dmi_native_sinkand_dmi_native_storeand never rebuilt them.cleanremovesnative/build/, which was the only place they lived, andallbuilt only_native_backend. A user who followed the install docs and enabled the capture config got anImportErroratcreate_record_runtime.Changes
allincludes a newcapturetarget, unconditionally. Mode-1-only builds opt out withCAPTURE=0.captureis a CPU-only goal, so it never resolves a CUDA toolkit. Building the extensions only when libcurl happened to be present would strand them again silently, so a missing libcurl now fails the build immediately. Acheck-libcurlstep compiles one libcurl call and prints a single line naminglibcurl4-openssl-dev.Install location. Each extension is linked in
native/build/and installed insrc/dmi/next to_native_backendas a relative symlink to that file (see the update below).src/dmi/is the first directory_search_dirs()looks in.cleanremoves both copies. Loading through the loader and a direct import both return the same module, checked in both import orders.Targets are named after the files they produce, so a second
makerelinks nothing.build/_dmi_native_sinkandbuild/_dmi_native_storeremain as aliases, since CI and the docs call them. A stamp relinks both extensions when the torch, pybind11 or Python they were built against changes.libcurl lookup order:
CURL_INCDIR/CURL_LIBDIR, which win, so CI's invocations are unchanged;pkg-config libcurl;CURL_SYSROOT, used only if that directory exists.The earlier default, silent
-I/-Lflags under a world-writable/tmppath, is gone.No torch RUNPATH on the sink. Every loader route imports torch first, and a bare import without torch is the only thing that fails. The Makefile comment records why.
Docs.
install.mdcovers the prerequisites (libcurl4-openssl-dev,pkg-config), thecapturetarget and its aliases,CAPTURE=0, the curl variables, a loader smoke check and troubleshooting. The HF, vLLM and Megatron guides andcode-organization.mdget short notes.Package check.
tests/tools/check_package.pynow refuses any.sounderdmi/in the wheel, not only_native_backend.Evidence
Seven new cpu tests in
tests/test_cpu_native_build.py. They are make dry runs plus rule-database checks under acleangoal, because a dry run ofallwould need CUDA on CI's CPU runner. They cover:allincludescapture;_search_dirs()[0];cleanthencapturerebuilds whatcleanremoved;curl.hfails before the store compiles, and the error names the package.Six of the seven failed before the change.
By hand:
cleanthencapturerebuilt both extensions intonative/buildandsrc/dmi, a second run did nothing, and both load through the loaders. A full CUDA build (C++20, sm_89),cleanthenall, produced all three extensions, and the sink subclasses the real_native_backend.RecordSink.CPU tier: 2221 passed, 1 skipped (no CUDA device), 0 failed. The capture storage live suite passed 6.
An independent review found one major and two minors, all fixed:
libcurl.sosymlink;Not run locally: GitHub Actions. The no-sysroot path, where the compiler's default paths apply, was only dry-run, because this host has no system libcurl headers.
Note for C++20 copies of the Makefile: the stamp adds one line near the top, so #138's three
-std=c++17lines move down by one. #138 and the CI compile job in the A5 PR both match by pattern, not line number.Since the independent review (2026-09-24)
An independent review found one major and five minors; all are fixed (041b79c..782200a):
make -C native -jstopped before building anything,_native_backendincluded, and never mentioned the opt-out. The libcurl error now ends with "To build without the capture storage extensions, which the in-memory path does not need, pass CAPTURE=0", anddocs/install.mdsays so beside the build command.CURL_SYSROOTnow beats pkg-config.CAPTUREaccepts exactly0or1; anything else (yes,true, a stray exported variable) stops make with a message, instead of silently dropping the extensions.src/dmiare symlinks intonative/build, so a stale copy can no longer shadow a fresh build when switching branches./tmp/opencode/sysroot, used only if it exists and is owned by the user).Current evidence:
tests/test_cpu_native_build.py25 passed; scratchmake cleanthenmake capturerebuilds both extensions as links, loadable through the loader; the CI cpu-job invocation passes (2234 tests). CI green at 782200a.