Skip to content

Build the torch-including native targets as C++20 - #138

Merged
zaoxing merged 4 commits into
mainfrom
fix/native-cxx20
Sep 24, 2026
Merged

zaoxing merged 4 commits into
mainfrom
fix/native-cxx20

Conversation

@zaoxing

@zaoxing zaoxing commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

What

main can't build its own native backend against the torch it installs. PyTorch's headers refuse anything below C++20:

torch/include/ATen/ATen.h:5:2: error: #error C++20 or later compatible compiler is required to use ATen.

native/Makefile compiled all three torch-including flag sets as -std=c++17, so on a fresh install neither make -C native nor make -C native host compiles:

Flag set Target Line
CXXFLAGS full backend (_native_backend) 125
HOST_CXXFLAGS CPU host backend (_host_backend) 130
NVCCFLAGS CUDA ring kernels 175

CI installs torch>=2.8,<3, which currently resolves to a release that enforces the check. _dmi_native_sink was already C++20, which is why cpu-goals kept building.

This is a three-line change: all three move to -std=c++20.

Why CI never caught it

CI only dry-runs host (test_host_build_plan_has_no_cuda_toolchain_or_libraries). The targets it actually compiles are cpu-goals and the conformance drivers, and none of those include ATen under C++17. So the failure only appears when someone builds the backend for real. It surfaced while standing up a GPU end-to-end run of the native capture path.

The guard

test_every_torch_including_compile_requests_cxx20 checks the plan make would execute, not the Makefile's text. Every compile that carries TORCH_EXTENSION_NAME has to request C++20 or later, for both host and all. That makes it independent of how the flags get assembled.

  • Before the fix: 3 stale compiles in host and 8 in all.
  • After the fix: it passes.
  • Only host is a cpu test. Planning all needs a CUDA toolkit, so that case carries the gpu marker and the cpu job doesn't select it. An earlier revision marked both cases cpu. On the runner, all then skipped with "CUDA toolkit resolution failed", and the cpu gate correctly failed the build, because a missing build prerequisite isn't absent hardware. Rewording the skip reason to fit the gate's allow-list would have hidden the problem, not fixed it.

The dry runs also pass PYTHON=sys.executable. That interpreter's torch is the one the build targets, and the Makefile's default python may not exist on a given machine.

Verification

Built from a clean native/build with torch 2.14, CUDA 13 and an RTX 4090 (SM_ARCH=sm_89):

Result
make -C native host 0 errors (fails on main)
make -C native (full, CUDA) 0 errors (fails on main)
make -C native cpu-goals 0 errors, 0 warnings
pytest -m cpu 2183 passed, 0 unexplained skips (checked through the CI gate's own logic)
tests/test_native_sink_ring_e2e.py (GPU) 6 passed against the rebuilt backend

On the warnings. host and all report 24 and 40 respectively. None come from C++20 behaviour.

  • 52 are inside torch's own headers: its bundled pybind11.h (-Wattributes) and an nvcc constexpr note in BFloat16.h.
  • The 4 in DMI code are false positives:
    • bindings.cpp:549 compares a signed loop index that counts up from 0 against py::len().
    • batching_queue.hpp:272 is GCC's std::optional -Wmaybe-uninitialized, on a value that is only dereferenced under if (max_linger_ns_ && oldest).

Both are left alone.

Not in this PR

The same end-to-end run turned up three design points. Each changes behaviour, so they belong in their own discussions:

  • CaptureReader's default max_coalesce_gap_bytes=4096 is too narrow for captures interleaved across requests. A 256 KiB gap made a one-prompt, all-layers hydrate 15× faster (144.6 s → 9.4 s), at the cost of overfetch.
  • CatalogIndexer's 128 MiB estimated-work bound means a large run has to be indexed in chunks, and nothing between the uploader and the indexer does that for the caller.
  • CaptureQuery has no request_id or step filter, so per-prompt queries scan every descriptor in the run.

Since the independent review (2026-09-24)

An independent review found the guard missed half of what it claims to check. Both findings are fixed, test-first:

  • The guard now checks the nvcc compiles (d10d3da). It picked compile lines by -DTORCH_EXTENSION_NAME=, which NVCCFLAGS never sets, so setting only the CUDA line back to -std=c++17 still passed. It now selects every compile that includes torch's headers (11 in the all plan: 8 g++, 3 nvcc). With only the CUDA line at C++17 it fails ("3 of 11 torch-including compile(s) below C++20"); against the base Makefile it fails every case.
  • The ring test harness builds as C++20 (460e84d). tests/native/ring/Makefile defaulted to CXX_STD ?= c++17, and its torch-including binaries failed with ATen's #error. The default is now c++20, pinned by a third guard case. The stale Makefile comment ("the one target that links against torch") is corrected.

tests/test_cpu_native_build.py: 13 passed. CI green at 460e84d.

PyTorch's headers refuse anything older: ATen.h opens with
`#error C++20 or later compatible compiler is required to use ATen`, and the
range CI installs (torch>=2.8,<3) resolves to a release that enforces it. The
full backend (CXXFLAGS), the CPU host backend (HOST_CXXFLAGS) and the CUDA ring
(NVCCFLAGS) were all still -std=c++17, so on a fresh install neither
`make -C native` nor `make -C native host` could compile. _dmi_native_sink was
already C++20, which is why cpu-goals kept building.

CI never noticed. It only DRY-RUNS `host` (test_cpu_native_build), its real
builds are cpu-goals and the conformance drivers, and none of those reach ATen
at C++17.

The guard pins the rule on the plan make would EXECUTE rather than on the
Makefile's text: every compile carrying TORCH_EXTENSION_NAME must request C++20
or later, for `host` and for `all`. It failed on the old flags (3 stale
compiles in host, 8 in all) and passes now. `all` needs a CUDA toolkit to plan
at all and skips where there is none, so `host` is the case CI checks. The dry
runs also pass PYTHON=sys.executable: that interpreter's torch is the one the
build targets, and the Makefile's default `python` need not exist.

Verified from a clean native/build with torch 2.14 / CUDA 13 / RTX 4090 (sm_89):
host, all and cpu-goals build with 0 errors. The remaining warnings are not new
C++20 behaviour: 52 are inside torch's own headers (its bundled pybind11.h
-Wattributes, and a BFloat16.h nvcc constexpr note), and the 4 in DMI code are
false positives -- a signed loop index compared against py::len(), and GCC's
std::optional -Wmaybe-uninitialized on a value only read under
`if (max_linger_ns_ && oldest)`. cpu suite 2184 passed; the GPU
ring -> NativePackSink suite 6 passed against the rebuilt backend.
Copilot AI lite review requested due to automatic review settings September 23, 2026 00:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

The cpu gate fails any skip whose reason is not absent hardware, and it was
right to fail this one. Planning `all` resolves the CUDA toolkit and stops
without it, so on a cpu runner the case skipped with "CUDA toolkit resolution
failed" -- a missing build prerequisite, which is exactly the kind of skip the
gate exists to refuse. Rewording the reason to match its allow-list would have
hidden it rather than fixed it.

The case was never a cpu test. It now carries the gpu marker, so the cpu job
does not select it, while host -- which plans anywhere -- stays cpu and is what
CI checks. With C++17 restored on HOST_CXXFLAGS the cpu case still fails
(3 stale compiles), so the move did not weaken the guard.

Local run through the same gate: 2183 collected, 0 unexplained skips.
@zaoxing

zaoxing commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator Author

GPU-host evidence for this PR's full build, at head 3d21207 (RTX 4090, torch 2.14.0+cu130, nvcc from /usr/local/cuda):

  • Clean full build: make -C native clean && make -C native -j16 SM_ARCH=sm_89, plus build/_dmi_native_sink, from scratch. It finishes with 0 errors, and all 12 torch-including compile lines run at -std=c++20. It produces _native_backend and _host_backend in src/dmi/, and _native_backend and _dmi_native_sink in native/build/.
  • tests/test_cpu_native_build.py + tests/test_native_sink_ring_e2e.py, unfiltered: 18 passed. That includes the GPU-marked test_every_torch_including_compile_requests_cxx20[all] guard this PR adds (CI can't run it, having no GPU runner), plus the full ring → NativePackSink suite, CUDA-graph replays across ring wraps (f32/bf16/i64) among them.

Without this PR, the same make -C native fails on torch 2.14: ATen.h:4-5 has #error C++20 or later compatible compiler is required. Every other capture-path PR (#135, #143 and the follow-ups) was verified on GPU with a local C++20 build for that reason, so this PR is the first thing to land.

🤖 Generated with Claude Code

The guard picked its compile lines by -DTORCH_EXTENSION_NAME=, which only
CXXFLAGS and HOST_CXXFLAGS define. NVCCFLAGS never sets it, so the three
ring .cu compiles in the `all` plan -- which include torch's headers just
the same -- were never checked. Reverting only NVCCFLAGS to -std=c++17
left the guard green (2 passed) while a real
`make build/ring/ring_engine_py.o` failed with ATen's
`#error C++20 or later compatible compiler is required`.

The guard now selects every command whose tokens carry a torch include
path (torch.utils.cpp_extension.include_paths()), so g++ and nvcc compiles
are checked alike: the `all` plan goes from 8 checked compiles to 11.

Evidence, on scratch copies of the tree:
- PR head: 2 passed (host, all).
- NVCCFLAGS alone at c++17: old guard 2 passed; new guard fails `all`
  with "3 of 11 torch-including compile(s) below C++20".
- origin/main: fails host (3 of 3) and all (11 of 11).

Also reword the stale comment above the sink target: C++20 is the floor
for every torch-including target, and C++17 stays only on the torch-free
drivers.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
tests/native/ring/Makefile defaulted CXX_STD to c++17, and three of its
binaries -- test_ring_engine, test_record_consumer and
test_clickhouse_record_sink -- include torch's headers, so against torch
2.14 each one stops at ATen's `#error C++20 or later compatible compiler
is required`. The default is now c++20; CXX_STD can still be overridden.

The C++20 guard now also dry-runs this Makefile's `all` (gpu-marked, as
planning it runs the CUDA resolver), so the default cannot slip back.

Evidence:
- Guard with the ring case, Makefile still c++17: 1 failed, 12 passed
  ("3 of 3 torch-including compile(s) below C++20 in tests/native/ring").
  With c++20: 13 passed.
- `make -C tests/native/ring build/test_record_consumer SM_ARCH=sm_89`
  on a scratch copy of the tree: rc=2 with the ATen #error at c++17,
  rc=0 at c++20. `make all` at c++20 builds all six ring binaries.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 23, 2026
Nothing in CI compiled _native_backend: the cpu job builds only the
CPU-only goals and dry-runs `host`, so the .cu sources, bindings.cpp
and the torch/CUDA link could break unnoticed, as the C++20 floor of
torch 2.14 did (#138). The new native-backend-compile job builds
clickhouse-cpp (docs/install.md step 5), runs `make -C native all`,
imports the result, compiles all six tests/native/ring binaries and
runs the two that need no device.

It installs the CUDA 13.0 pieces from NVIDIA's apt repository on the
runner instead of using an nvidia/cuda devel container. A container job
pulls its image before any step can free disk, and that image is 3.68
GiB compressed before the ~4.4 GB that torch 2.14+cu130 takes
installed. The package list (nvcc, cudart-dev, cublas/cusparse/cusolver
dev, about 2.7 GB by Installed-Size) is what a real build reads: the
backend's -MMD files and `nvcc -M` of the ring tests, mapped with
dpkg -S. #138 is not in this base, so a step makes #138's three
-std=c++17 -> c++20 edits with sed; after #138 merges it is a no-op and
should be deleted.

Local evidence: the sed step run against this Makefile changes exactly
lines 125, 130 and 175 (the #138 diff) and against #138's head changes
nothing; with those edits, CUDA_HOME=/usr/local/cuda-13.0, SM_ARCH=sm_89
and the libcuda stub on LIBRARY_PATH, `make all` built and passed
check-link, and the backend imported with no GPU visible. The stock
flags stop at bindings.cpp on "#error C++20 or later compatible compiler
is required". Ring tests: all six built with CXX_STD=c++20, and
test_record_consumer (14 passed) and test_clickhouse_record_sink (21
passed) ran with CUDA_VISIBLE_DEVICES empty. Not verified until CI runs:
the apt install on ubuntu-24.04, disk headroom, gcc 13 (local is 11.4),
and the job as a whole.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
Nothing in CI compiled _native_backend: the cpu job builds only the
CPU-only goals and dry-runs `host`, so the .cu sources, bindings.cpp
and the torch/CUDA link could break unnoticed, as the C++20 floor of
torch 2.14 did (#138). The new native-backend-compile job builds
clickhouse-cpp (docs/install.md step 5), runs `make -C native all`,
imports the result, compiles all six tests/native/ring binaries and
runs the two that need no device.

It installs the CUDA 13.0 pieces from NVIDIA's apt repository on the
runner instead of using an nvidia/cuda devel container. A container job
pulls its image before any step can free disk, and that image is 3.68
GiB compressed before the ~4.4 GB that torch 2.14+cu130 takes
installed. The package list (nvcc, cudart-dev, cublas/cusparse/cusolver
dev, about 2.7 GB by Installed-Size) is what a real build reads: the
backend's -MMD files and `nvcc -M` of the ring tests, mapped with
dpkg -S. #138 is not in this base, so a step makes #138's three
-std=c++17 -> c++20 edits with sed; after #138 merges it is a no-op and
should be deleted.

Local evidence: the sed step run against this Makefile changes exactly
lines 125, 130 and 175 (the #138 diff) and against #138's head changes
nothing; with those edits, CUDA_HOME=/usr/local/cuda-13.0, SM_ARCH=sm_89
and the libcuda stub on LIBRARY_PATH, `make all` built and passed
check-link, and the backend imported with no GPU visible. The stock
flags stop at bindings.cpp on "#error C++20 or later compatible compiler
is required". Ring tests: all six built with CXX_STD=c++20, and
test_record_consumer (14 passed) and test_clickhouse_record_sink (21
passed) ran with CUDA_VISIBLE_DEVICES empty. Not verified until CI runs:
the apt install on ubuntu-24.04, disk headroom, gcc 13 (local is 11.4),
and the job as a whole.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
The step exists only until #138 moves the Makefile to -std=c++20, and it
had no way to say when that happened: after the merge the sed matches
nothing, and the closing grep still counts 3, because the three lines
already read c++20. The step would pass silently forever.

An unchanged native/Makefile after the sed is exactly that state (before
#138 the sed always rewrites all three lines), so the step now emits a
::warning:: asking for its own deletion. Rehearsed on a scratch copy of
the Makefile: at this branch's Makefile the diff is 3 lines and no
warning; with those edits committed, as #138 would leave them, the
warning prints and the count is still 3.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
@zaoxing
zaoxing merged commit 68f7de2 into main Sep 24, 2026
2 checks passed
zaoxing added a commit that referenced this pull request Sep 24, 2026
… command

Rebasing onto #138 brought its C++20 guard, which checks every compile in
`native all` that reads torch's include directory. This branch puts the
capture extensions in `all`, and the guard then failed on "2 of 13
torch-including compile(s) below C++20", for two reasons:

- The store extension compiled at -std=c++17 while passing torch's include
  directory for pybind11. It now builds as C++20, like every compile that
  reads that directory. Its sources are the same C++17-clean catalog and
  store code the drivers build; rebuilt as C++20 it compiles with 0
  warnings, and the storage service and native reader parity suites pass
  on it (60 passed), with the wiring suite (37 passed).
- The guard read the dry run line by line, so a recipe continued with
  backslashes had its -std= on one line and its torch -I flags on the next,
  and the continuation read as a compile with no standard. It now joins
  continued lines before judging them. With the store put back at C++17,
  the guard flags exactly that one compile.

tests/test_cpu_native_build.py: 33 passed, the gpu-planned cases included.
CPU tier: 2299 passed, 1 skipped (no CUDA device).

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
#144)

* Build and install the capture extensions with the default native build

The documented rebuild is `make -C native clean && make -C native`. clean
removes build/, where _dmi_native_sink and _dmi_native_store lived, and
`all` built only _native_backend, so the documented command deleted both
capture extensions and nothing put them back until someone remembered
their separate targets.

`all` now depends on a `capture` target that links both extensions in
build/ and installs them next to _native_backend in src/dmi/, the first
directory the loader searches; clean removes the installed copies too.
It is unconditional rather than "when libcurl is found", because skipping
silently would strand them again and surface only at runtime; CAPTURE=0
builds the backend alone. The targets are now the files they produce, so
an up-to-date extension is not relinked on every make, with their headers
listed by directory so the file targets cannot go stale;
build/_dmi_native_sink and build/_dmi_native_store stay as aliases for CI
and the loaders' error messages. `capture` is a CPU-only goal.

CURL_INCDIR/CURL_LIBDIR: explicit values still win (CI passes the system
paths), then `pkg-config libcurl`, then the extracted dev package under
CURL_SYSROOT (default unchanged, which is how this host builds). A
check-libcurl step, order-only before every curl target, compiles and
links one libcurl call and on failure names libcurl4-openssl-dev instead
of a page of compiler output.

No torch RUNPATH on the sink: a bare import without torch fails on
libtorch_cpu.so, but every loader route imports torch first
(dmi.storage.capture does), and that route loads it.

Evidence: 7 new cpu tests in tests/test_cpu_native_build.py, 6 failing
before the change (the CAPTURE=0 guard passes on both); 13 cpu passed
after. Real run: make -C native clean, then make -C native capture
rebuilt both into native/build and src/dmi, a second make relinked
nothing, and both loaded through the normal loaders from src/dmi. Full
build (C++20 Makefile copy, CUDA 13.0, sm_89): clean then all produced
_native_backend, _dmi_native_sink and _dmi_native_store; all three load
through the loader, the sink bound to the real RecordSink (no
stand-ins). The package-layout wheel still contains no .so.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Document the capture extensions in the native build instructions

install.md, huggingface.md, vllm.md and megatron.md tell users to run
`make -C native clean && make -C native`, which now also rebuilds
_dmi_native_sink and _dmi_native_store. Say so, name the capture target
and the build/ aliases, add libcurl4-openssl-dev (and pkg-config, which
the build uses to find it) to the system prerequisites, document the
CURL_SYSROOT / CURL_INCDIR / CURL_LIBDIR fallbacks and CAPTURE=0, and add
troubleshooting entries for the libcurl error and a missing capture
extension. code-organization.md names where the extensions are
installed.

Evidence: the new install.md smoke check loads _native_backend,
_dmi_native_sink and _dmi_native_store from src/dmi through the loader
after `make -C native clean` and the full build on this host.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Refuse any native extension in the package-layout wheel

The wheel guard looked only for dmi/_native_backend*.so. The native build
now installs _dmi_native_sink and _dmi_native_store into src/dmi beside
it (and `make host` puts _host_backend there), so a packaging change that
started shipping one of those would have passed. Check every .so under
dmi/ instead.

Evidence: _validate_archive on a synthetic wheel carrying
dmi/_dmi_native_sink*.so was accepted before and is refused after; the
real wheel built from a tree holding all three installed extensions has
no .so members and is still accepted. `make test-package` itself cannot
run on this host (ensurepip in the temporary venv exits 127), so the
full check was not run locally.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Extract libcurl's runtime package into the sysroot, and use a sysroot only if it is there

The documented no-root fallback extracted only libcurl4-openssl-dev. That
package's libcurl.so links to libcurl.so.4.6.0, which ships in the runtime
package libcurl4, so the sysroot held a dangling link; ld skipped it for
libcurl.a, and check-libcurl failed telling the user to install the very
package they had just extracted (without the check, the store linked the
static libcurl and failed at import on GSS_C_NT_HOSTBASED_SERVICE). It
worked on this host only because /tmp/opencode/sysroot also holds the
runtime library. install.md and the Makefile now extract both packages,
and check-libcurl names a dangling libcurl.so as the missing runtime
package.

The default build also started searching /tmp/opencode/sysroot on every
host where pkg-config did not report libcurl: a world-writable /tmp path,
passed as -I/-L ahead of the system directories. The sysroot default now
applies only when that directory exists and belongs to the user running
make (CURL_SYSROOT_DEFAULT, so the rule is testable); otherwise no -I/-L is
passed and the compiler's default paths apply. Explicit CURL_INCDIR,
CURL_LIBDIR or CURL_SYSROOT still win, so CI's invocations are unchanged.

Evidence: two new cpu tests in tests/test_cpu_native_build.py failed first
(generic message; sysroot flags present with no sysroot), then pass; the
file's cpu tests: 15 passed. From an empty scratch sysroot, the documented
recipe (apt-get download libcurl4-openssl-dev libcurl4, dpkg-deb -x each)
passes check-libcurl, builds _dmi_native_store against it, and the loader
imports it; with the dev package alone the new message names libcurl4.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Relink the capture extensions when the torch they were built against changes

Naming the capture targets after their files made them up to date whenever
their sources and headers were, and nothing about torch, pybind11 or the
Python ABI was a prerequisite. A torch upgrade in the venv therefore left
the sink linked against the old libtorch, and the command the loader's
ImportError recommends (make -C native build/_dmi_native_sink) printed
"Nothing to be done"; before the rename it always relinked. Both
extensions now depend on build/torch-extension.config, a record of the
torch version and git revision, its C++ ABI flag, include and library
paths, the Python include directory and EXT_SUFFIX, written only when it
changes by the same write-stamp helper as the CUDA toolkit stamp.

Evidence: the new cpu test (prerequisite in make's rule database; a real
stamp run under a temporary BUILD_DIR that keeps its mtime when unchanged
and rewrites a stale record) failed first, then passes. By hand: a second
`make capture` relinks nothing, and after the stamp's torch version is
edited, `make build/_dmi_native_sink` relinks the sink. A full clean +
`all` (CUDA 13.0, sm_89, C++20 scratch copy of the Makefile) builds all
three extensions, a second `all` runs no compiler, and all three load
through the loader. make test-cpu with CUDA_VISIBLE_DEVICES="": 2224
passed, 1 skipped (no CUDA device), 324 deselected.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Name CAPTURE=0 when the libcurl check stops the default build

The default build now includes the capture storage extensions, so on a
host without libcurl's dev package `make -C native -j` fails at
check-libcurl before it builds anything, _native_backend included. The
reviewer reproduced it with `make -C native -j16 SM_ARCH=sm_89
CURL_SYSROOT= PKG_CONFIG=false` after a clean: exit 2, no .o, no .so.
The message named libcurl4-openssl-dev and CURL_INCDIR/CURL_LIBDIR but
not the opt-out, so a user of the in-memory path, which needs no
capture extensions, had no hint how to get the backend back.

The message now ends by saying to pass CAPTURE=0 to build without the
capture storage extensions, and docs/install.md carries the same hint
beside the `make -C native -j` step. A dry run of the default goal with
CURL_SYSROOT= PKG_CONFIG=false plans the libcurl check under CAPTURE=1
and not under CAPTURE=0, while both still plan _native_backend.

test_missing_libcurl_names_the_package asserts the message names
CAPTURE=0 (it failed before the Makefile change), and now runs with
BUILD_DIR under tmp_path, as the torch-relink test does, so its real
make no longer writes the torch stamp into the checkout's native/build.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Let an explicit CURL_SYSROOT win over pkg-config

pkg-config assigned CURL_INCDIR and CURL_LIBDIR with := before the
sysroot's ?= lines were read, so a CURL_SYSROOT the user passed was
ignored whenever `pkg-config --exists libcurl` succeeded. The review
showed it with a fake pkg-config: `make CURL_SYSROOT=/sysroot` compiled
with -I/-L from pkg-config and never searched /sysroot. Naming a sysroot
is the only reason to pass one, so a non-empty CURL_SYSROOT from the
command line or the environment now skips pkg-config; the default
sysroot still ranks behind it, and an empty CURL_SYSROOT= still leaves
pkg-config in charge.

test_libcurl_directories_resolve_in_documented_order gains the
command-line and exported cases; both failed before this change (the
link line carried the pkg-config directories) and pass after it.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Refuse a CAPTURE value other than 0 or 1

`ifeq ($(CAPTURE),1)` treated every other value as CAPTURE=0, so
CAPTURE=yes, CAPTURE=true, or an unrelated CAPTURE exported by some other
tool built the backend without the capture extensions and said nothing
-- the silent stranding the default build exists to prevent. The value
is now stripped and must be exactly 0 or 1; anything else stops make
with an error naming the variable, where it came from, and both accepted
values:

  Makefile:211: *** CAPTURE must be 0 or 1, not 'yes' (from the command
  line): 1 builds the capture storage extensions with the default build,
  0 leaves them out.  Stop.

The check runs at parse time, so it applies to every goal, clean
included: a CAPTURE that means something else is worth a stop wherever
it shows up.

New tests: test_capture_accepts_only_zero_or_one (CAPTURE=yes, =true,
and an exported CAPTURE, each through `make -n clean`) and
test_capture_setting_ignores_surrounding_whitespace. All four failed
before this change -- the first three exited 0, and CAPTURE=' 1 ' left
the extensions out because of its trailing space -- and pass after it.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Install the capture extensions in src/dmi as links to the build

The loader (dmi/transport/native.py, _search_dirs) tries src/dmi/ before
native/build/. With src/dmi holding copies, anything that rebuilt
native/build alone left the copy stale and in front: the review
reproduced it with a branch without this change, whose
`make build/_dmi_native_sink` writes only native/build, and the loader
kept importing the copy this branch had installed. src/dmi now gets a
relative symlink to the build/ file, one inode under two names, so a
rebuild in either place is what gets imported. A copy an earlier build
installed is replaced even when it is newer than build/, which the
file-named target alone would have called up to date.

Checked before switching, in a scratch build of both extensions:
- the loader's spec_from_file_location path imports each through the
  link, and __file__ is the link;
- in one process, the loader route and the torch suites' route
  (native/build on sys.path) give the same module object in either
  order, as they did with copies;
- `rm -rf` in `clean` removes the links, not what they name, and a
  dangling link (build/ removed) is skipped by the loader's exists();
- the package-layout wheel built from the tree has no .so with valid
  or dangling links present, and _validate_archive still refuses a
  wheel with a symlinked extension added. (check_package.py's venv
  smoke step fails on this host for every tree, with or without
  extensions: EnvBuilder copies the uv Python binary, whose
  $ORIGIN/../lib libpython is then missing, exit 127.)

test_capture_extensions_install_as_links_to_the_build installs into a
stand-in checkout holding newer copies, then rebuilds build/ by rename
the way ld does. It failed before this change (the copy stayed) and
passes after it; a second run redoes nothing.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Relink the capture extensions when the compiler or libcurl changes

The torch stamp recorded torch, its ABI and include/library paths, and
the Python build, but not CXX or the libcurl flags. Pointing CURL_INCDIR
or CURL_LIBDIR at another libcurl, or switching compilers, changed no
source file and no recorded field, so an up-to-date _dmi_native_store
kept its old link. The record now ends with CXX, CURL_CPPFLAGS and
CURL_LDFLAGS, and is still rewritten only when it changes. It stays one
stamp for both extensions, so the sink (no libcurl) also relinks on a
libcurl change; that costs seconds and avoids a second stamp.

test_capture_extensions_relink_when_compiler_or_libcurl_changes builds
the store three times for real in a stand-in checkout, with a compiler
that logs its link lines and writes an empty file: first build links,
an unchanged rerun does not, and changing CURL_INCDIR (or CXX) relinks
once with the new value on the line. Both cases failed before this
change (0 relinks) and pass after it.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Document the libcurl lookup order exactly, default sysroot included

install.md said the build took libcurl from pkg-config and otherwise
from the compiler's default paths. The Makefile also falls back to
CURL_SYSROOT_DEFAULT=/tmp/opencode/sysroot -- a default that predates
this branch and that one development host's builds rely on -- and, since
the CURL_SYSROOT fix, ranks an explicit CURL_SYSROOT ahead of
pkg-config. A reader of the docs could not know either, nor why a /tmp
path shows up in a link line.

The curl paragraph now lists the five steps in the Makefile's order:
explicit CURL_INCDIR/CURL_LIBDIR, an explicit non-empty CURL_SYSROOT,
pkg-config, the default sysroot (used only if the directory exists and
is owned by the user running make, and movable with
CURL_SYSROOT_DEFAULT=), then the compiler's defaults -- each directory
resolved on its own. The Makefile's comment names the default's origin
the same way. The default itself stays.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Build the capture store as C++20, and judge a continued recipe as one command

Rebasing onto #138 brought its C++20 guard, which checks every compile in
`native all` that reads torch's include directory. This branch puts the
capture extensions in `all`, and the guard then failed on "2 of 13
torch-including compile(s) below C++20", for two reasons:

- The store extension compiled at -std=c++17 while passing torch's include
  directory for pybind11. It now builds as C++20, like every compile that
  reads that directory. Its sources are the same C++17-clean catalog and
  store code the drivers build; rebuilt as C++20 it compiles with 0
  warnings, and the storage service and native reader parity suites pass
  on it (60 passed), with the wiring suite (37 passed).
- The guard read the dry run line by line, so a recipe continued with
  backslashes had its -std= on one line and its torch -I flags on the next,
  and the continuation read as a compile with no standard. It now joins
  continued lines before judging them. With the store put back at C++17,
  the guard flags exactly that one compile.

tests/test_cpu_native_build.py: 33 passed, the gpu-planned cases included.
CPU tier: 2299 passed, 1 skipped (no CUDA device).

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

---------

Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
zaoxing added a commit that referenced this pull request Sep 24, 2026
Nothing in CI compiled _native_backend: the cpu job builds only the
CPU-only goals and dry-runs `host`, so the .cu sources, bindings.cpp
and the torch/CUDA link could break unnoticed, as the C++20 floor of
torch 2.14 did (#138). The new native-backend-compile job builds
clickhouse-cpp (docs/install.md step 5), runs `make -C native all`,
imports the result, compiles all six tests/native/ring binaries and
runs the two that need no device.

It installs the CUDA 13.0 pieces from NVIDIA's apt repository on the
runner instead of using an nvidia/cuda devel container. A container job
pulls its image before any step can free disk, and that image is 3.68
GiB compressed before the ~4.4 GB that torch 2.14+cu130 takes
installed. The package list (nvcc, cudart-dev, cublas/cusparse/cusolver
dev, about 2.7 GB by Installed-Size) is what a real build reads: the
backend's -MMD files and `nvcc -M` of the ring tests, mapped with
dpkg -S. #138 is not in this base, so a step makes #138's three
-std=c++17 -> c++20 edits with sed; after #138 merges it is a no-op and
should be deleted.

Local evidence: the sed step run against this Makefile changes exactly
lines 125, 130 and 175 (the #138 diff) and against #138's head changes
nothing; with those edits, CUDA_HOME=/usr/local/cuda-13.0, SM_ARCH=sm_89
and the libcuda stub on LIBRARY_PATH, `make all` built and passed
check-link, and the backend imported with no GPU visible. The stock
flags stop at bindings.cpp on "#error C++20 or later compatible compiler
is required". Ring tests: all six built with CXX_STD=c++20, and
test_record_consumer (14 passed) and test_clickhouse_record_sink (21
passed) ran with CUDA_VISIBLE_DEVICES empty. Not verified until CI runs:
the apt install on ubuntu-24.04, disk headroom, gcc 13 (local is 11.4),
and the job as a whole.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
The step exists only until #138 moves the Makefile to -std=c++20, and it
had no way to say when that happened: after the merge the sed matches
nothing, and the closing grep still counts 3, because the three lines
already read c++20. The step would pass silently forever.

An unchanged native/Makefile after the sed is exactly that state (before
#138 the sed always rewrites all three lines), so the step now emits a
::warning:: asking for its own deletion. Rehearsed on a scratch copy of
the Makefile: at this branch's Makefile the diff is 3 lines and no
warning; with those edits committed, as #138 would leave them, the
warning prints and the count is still 3.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
#144 landed

- #138 is on main, so the temporary step that applied its C++20 flags with
  sed changed nothing and would only raise its "delete me" warning. Deleted,
  as its own comment asked.
- #144 puts the capture extensions in `make -C native all`, and the store
  needs libcurl's headers; without them the build now stops before
  anything else, by design. The job installs libcurl4-openssl-dev and
  pkg-config and lets the Makefile find libcurl through pkg-config -- its
  default lookup, with no CURL_* override -- so it builds exactly what
  docs/install.md tells a user to run. It lists the three extensions it
  expects in src/dmi afterwards.

Validated locally: the workflow parses and the job's steps are as listed;
ruff 0.16.5 F821 passes on the merged src/dmi; with #144's Makefile the
CPU targets build and install as links, and the chain test passes (2
passed). The CUDA build itself runs only in CI.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
…, backend compile, undefined-name lint (#145)

* Drive the real native sink through to the catalog in the live job

Every live suite staged its packs through the conformance_sink driver,
which feeds the sink core JSON and base64, so the torch-facing
NativePackSink that the ring hands envelopes to never reached a catalog
in CI. tests/test_native_capture_chain_live.py drives it with torch CPU
tensors (through its attach/submit_envelope test entry points) into a
spool, runs CaptureStorageService in-process against ClickHouse and the
signature-verifying fake S3, and reads the payloads back through
NativeCaptureReader, requiring the bytes to match. No CUDA, no ring, no
conformance driver.

It also sends one 68 MiB pack up as a multipart upload, twice. The fake
S3 accepts parts of any size, takes the part list on trust and never
checks the completion ETags, so against it the test asserts the part
sizes the client sent. Against MinIO, S3's 5 MiB part floor is enforced
by the server; the MinIO case skips only when DMI_MINIO_ENDPOINT is
unset. The clickhouse-live job now builds build/_dmi_native_sink and
starts MinIO (quay.io image pinned by digest, via docker run because a
service container takes no command) and sets that variable, so its
skip gate turns a lost variable into a red job.

Evidence, local ClickHouse 127.0.0.1 and a MinIO binary extracted from
the pinned image on 127.0.0.1:19100: 3 passed; with the endpoint unset,
2 passed, 1 skipped. With the client's multipart chunk shrunk to 4 MiB
(not committed), the fake-S3 case failed on the part count (18 != 5)
and the MinIO case on CompleteMultipartUpload EntityTooSmall.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Compile the full native backend and the ring tests in CI

Nothing in CI compiled _native_backend: the cpu job builds only the
CPU-only goals and dry-runs `host`, so the .cu sources, bindings.cpp
and the torch/CUDA link could break unnoticed, as the C++20 floor of
torch 2.14 did (#138). The new native-backend-compile job builds
clickhouse-cpp (docs/install.md step 5), runs `make -C native all`,
imports the result, compiles all six tests/native/ring binaries and
runs the two that need no device.

It installs the CUDA 13.0 pieces from NVIDIA's apt repository on the
runner instead of using an nvidia/cuda devel container. A container job
pulls its image before any step can free disk, and that image is 3.68
GiB compressed before the ~4.4 GB that torch 2.14+cu130 takes
installed. The package list (nvcc, cudart-dev, cublas/cusparse/cusolver
dev, about 2.7 GB by Installed-Size) is what a real build reads: the
backend's -MMD files and `nvcc -M` of the ring tests, mapped with
dpkg -S. #138 is not in this base, so a step makes #138's three
-std=c++17 -> c++20 edits with sed; after #138 merges it is a no-op and
should be deleted.

Local evidence: the sed step run against this Makefile changes exactly
lines 125, 130 and 175 (the #138 diff) and against #138's head changes
nothing; with those edits, CUDA_HOME=/usr/local/cuda-13.0, SM_ARCH=sm_89
and the libcuda stub on LIBRARY_PATH, `make all` built and passed
check-link, and the backend imported with no GPU visible. The stock
flags stop at bindings.cpp on "#error C++20 or later compatible compiler
is required". Ring tests: all six built with CXX_STD=c++20, and
test_record_consumer (14 passed) and test_clickhouse_record_sink (21
passed) ran with CUDA_VISIBLE_DEVICES empty. Not verified until CI runs:
the apt install on ubuntu-24.04, disk headroom, gcc 13 (local is 11.4),
and the job as a whole.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Make src/dmi pass ruff's undefined-name check

ruff F821 over src/dmi reported 14 names, in two groups, and neither is
a NameError today:

- config.py annotates capture_sink_config and capture_storage_config
  with the strings "NativeSinkConfig" and "NativeCaptureStorageConfig"
  and imports neither. With `from __future__ import annotations` nothing
  evaluates them at runtime, but typing.get_type_hints(MonitoringConfig)
  raises NameError on the first one (reproduced; nothing in src/ calls
  it today). They are now imported under TYPE_CHECKING, so a static
  tool resolves them and the module still loads no storage package at
  runtime (checked: importing dmi.config loads no dmi.storage module).
  get_type_hints still raises, as before; B2's move of NativeSinkConfig
  into live code is where a runtime import can go.
- hooks/specs.py compares against HOOK_TYPE_Q and eleven siblings that
  the module binds through a globals() loop over _HOOK_DEFS, which a
  linter cannot see. Each of those ten lines now carries its own
  `# noqa: F821` with a comment saying why: per line rather than per
  file, so the rest of the module stays checked. A typo inside one of
  those ten lines would now be suppressed too; that is the cost.

No behaviour change. ruff 0.16.5 `check --select F821 src/dmi`: 14
errors before, "All checks passed!" after; tests/test_tp_shapes.py,
test_hook_spec_flags.py and test_config.py: 63 passed.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Lint src/dmi for undefined names in the cpu job

An undefined name is a NameError that waits for its line to run, and
the cpu job had nothing that would see one first: `make check` compiles
with `python -m compileall`, which accepts any name. The new step runs
ruff's F821 rule, and only that rule, over src/dmi, pinned to ruff
0.16.5 with GitHub annotations.

The 14 existing hits are handled in the previous commit, in the source,
not excluded here. Local evidence with ruff 0.16.5: the step's command
passes on this tree, and with an undefined name appended to
src/dmi/config.py (not committed) it exits 1 naming the file and line.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Check every catalog metadata field in the capture chain test

The chain test compared only capture_id, payload bytes, shape and dtype,
so a metadata field the torch sink's parser mangled went unnoticed. A
reviewer changed native/csrc/sink/record_row.cpp to store step_number + 7
and the chain test still passed, as did every CPU suite. This chain is the
only CI test that reaches record_row.cpp, so it has to catch that.

Each capture's read-back descriptor is now compared field by field against
the CaptureMetadata the test submitted for it: all 21 fields, with shape
as a tuple and a NULL adapter_revision as the empty string the native
reader returns for it. The integer fields now also take distinct values
(token_start and batch_position vary apart from step_number), so a parser
that wires one field into another's slot fails too. producer_rank stays
fixed, at 1, because the sink seals a pack when the rank changes.

The fake-S3 multipart check counted every logged PUT carrying a
partNumber, so one transient retry would have failed the part count. It
now keys parts by part number, keeping the last attempt.

Evidence, from a scratch copy of HEAD built with
make build/_dmi_native_sink build/_dmi_native_store:
- step_number + 7 mutant: 2 failed, 1 skipped (MinIO unset); both fail on
  {'step_number': 7} != {'step_number': 0} for chain-0000.
- unmutated rebuild: 2 passed, 1 skipped. Same in this worktree.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Drop the unused E402 noqa from the capture chain test

The fake-S3 import sits among the module's top-level imports, before any
statement, so E402 never fires on it and the directive suppresses
nothing. ruff 0.16.5 with --select E402,RUF100 reports it as an unused
noqa (unused: E402); with it removed the same check passes.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Warn when the C++20 sed step has nothing left to change

The step exists only until #138 moves the Makefile to -std=c++20, and it
had no way to say when that happened: after the merge the sed matches
nothing, and the closing grep still counts 3, because the three lines
already read c++20. The step would pass silently forever.

An unchanged native/Makefile after the sed is exactly that state (before
#138 the sed always rewrites all three lines), so the step now emits a
::warning:: asking for its own deletion. Rehearsed on a scratch copy of
the Makefile: at this branch's Makefile the diff is 3 lines and no
warning; with those edits committed, as #138 would leave them, the
warning prints and the count is still 3.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Check the capture sink binds the backend's RecordSink in CI

_dmi_native_sink picks its RecordSink base once, when it loads: the
_native_backend registration when that backend is loadable, module-local
stand-ins otherwise. Only the first is attachable, since
create_record_runtime checks isinstance against the backend's class. The
cpu and live jobs build the sink with no backend, and native-backend-compile
built the backend with no sink, so the backend branch of
test_the_sink_derives_from_the_engines_record_sink ran in no job.

native-backend-compile now builds build/_dmi_native_sink after the
backend, with the same torch and PYTHON, and checks the binding in both
import orders: through dmi's loaders (backend first, as the engine does)
and as a bare import (the sink finds the backend itself). Each asserts
RING_TYPES_ARE_STANDINS is false and NativePackSink subclasses the
backend's RecordSink. Then it runs the adapter test and refuses a skip.

The assertions are needed because the test alone is not enough. Rehearsed
on a scratch copy of this branch with a copied torch 2.14+cu130 backend:
with the backend present, both checks print an MRO through
_native_backend.RecordSink and the test passes (1 passed). With the
backend removed, both assertions fail on the stand-ins, and the test
still passes, because its stand-in branch is a pass by design. The rehearsal
also found that the bare import needs torch imported first, since the .so
has no rpath to libtorch_cpu, so the check does that, as the suite does.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Bound the MinIO health probe and drop the last-release claim

The health loop ran curl -sf with no timeout, so a socket that accepts
and never answers would have held the first probe until the 30-minute
job timeout, and the loop's 30 attempts would never have run. Each probe
is now capped with --max-time 2. Checked locally against a listener that
accepts and sleeps: curl gives up after 2.01 s with exit 28, which the
loop treats as not up yet and retries.

The comment also called the pinned release "the last release published
as an image". Review found later tags on quay.io, for example
RELEASE.2025-09-07T16-13-09Z.hotfix.7aa24e772. (I could not list the
tags from here: quay's API wants a token.) The pin is still right, so
the comment now says only what it does: tag and digest together keep an
upstream retag from changing what the job tests against.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Describe the runner disk step as a precaution, with measured figures

The job comment justified freeing disk with "the ~14 GB a hosted runner
guarantees", but the runner this job actually got had far more. The
df output from run 35933626430 (job 107425626899) shows 87 GB free
before the step, 103 GB after it, 99 GB after the CUDA apt install and
94 GB after torch. Freeing the Android SDK and .NET took about two
minutes of that run.

The step stays, since the free space belongs to the runner image and can
change, but both comments now call it a precaution and quote only the
figures that log shows.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Pull MinIO from Bitnami's archive, and let a warning pass the sink check

Two CI failures on the previous push, neither in the code under test:

- The MinIO step's image could no longer be pulled. The MinIO project has
  withdrawn its public images: quay.io/minio/minio and docker.io/minio/minio
  both now answer anonymous pulls with 401, including `latest`, although
  this job pulled the pinned quay image fine a day earlier. The step now runs
  bitnamilegacy/minio:2025.7.23-debian-12-r5, pinned by digest -- a frozen
  archive, which suits a fixture whose S3 multipart rules do not change. The
  Bitnami image starts its own server from MINIO_ROOT_USER/_PASSWORD, so the
  `server /data` argument is gone, and the health loop allows about two
  minutes for its entrypoint's setup. The live skip gate then failed only
  because the suite never ran.
- The sink-binding check passed ("1 passed, 1 warning in 1.50s") but the
  step greps for '^1 passed in ', and torch's NumPy warning changed the
  summary line. The step now accepts a summary with warnings and still
  fails on any skip.

Validated locally: the workflow parses, and the new pass condition accepts
"1 passed" and "1 passed, 1 warning" and refuses "1 skipped". The image's
digest, amd64 manifest and entrypoint were read from Docker Hub's registry;
running it needs docker, which only CI has.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Drop MinIO from the live job

MinIO was there for one check: that a multipart pack follows real S3's
rules (every part but the last at least 5 MiB, and a "-N" ETag), which the
in-repo fake S3 does not enforce. Nothing in DMI runs against MinIO -- the
object store it uses is Garage -- and the MinIO project has withdrawn its
public images, so the job had come to depend on a frozen third-party
archive for a store DMI does not use.

The CI step, its DMI_MINIO_* settings and the MinIO test case are removed.
The fake-S3 multipart case stays: it asserts the parts the client sends
(every part but the last is exactly the client's chunk and at least 5 MiB)
and reads the pack back byte-exact. Checks against a real store belong to
the Garage suites.

Chain test against local ClickHouse: 2 passed, 0 skipped; no tables left.
The workflow parses, and the only step removed from any job is the MinIO
one.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Build the documented native build in the compile job, now that #138 and #144 landed

- #138 is on main, so the temporary step that applied its C++20 flags with
  sed changed nothing and would only raise its "delete me" warning. Deleted,
  as its own comment asked.
- #144 puts the capture extensions in `make -C native all`, and the store
  needs libcurl's headers; without them the build now stops before
  anything else, by design. The job installs libcurl4-openssl-dev and
  pkg-config and lets the Makefile find libcurl through pkg-config -- its
  default lookup, with no CURL_* override -- so it builds exactly what
  docs/install.md tells a user to run. It lists the three extensions it
  expects in src/dmi afterwards.

Validated locally: the workflow parses and the job's steps are as listed;
ruff 0.16.5 F821 passes on the merged src/dmi; with #144's Makefile the
CPU targets build and install as links, and the chain test passes (2
passed). The CUDA build itself runs only in CI.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

---------

Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
@zaoxing
zaoxing deleted the fix/native-cxx20 branch September 24, 2026 15:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants