You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Since 2026-10-01, the kubernetes e2e suites fail their cleanup in the PR lane because one node's ZFS pool does not come back to its pre-suite level. The suite's own checks all pass; it goes red in cozy_cleanup, which compares each node with the baseline within a 512 MiB tolerance (cozy_linstor_pools_at_baseline in hack/e2e-chainsaw/_lib/run-kubernetes.sh). Each time, 450-830 MiB stays used on one or two nodes and is never freed.
36831242577 (renovate/indirect-go-modules, Go modules only): both kubernetes suites, srv1 541460 short, then srv3 793677 short. srv1's baseline for the second suite equals what the first one left, so the space is not released later.
Passing runs for comparison: 36785869560 on 2026-09-30 (under 1 MiB of drift) and 36845137883 on 2026-10-01 09:48 UTC. None of the failing PRs touches LINSTOR, the tenant teardown or run-kubernetes.sh, and the renovate branch failed before the Talos bump (#4598) merged.
The likely cause, not proven: a CDI scratch volume whose first CreateVolume failed leaves a partly created zvol behind. In 36844521765 the scratch PVC pvc-f1883e08 first failed with ResourceExhausted ... Not enough available nodes at 12:12:38 UTC, was then provisioned and attached on srv3, later logged VolumeFailedDelete ... still attached to node srv3, and is absent from the final LINSTOR resource list. srv3 is the node that came back short. For the other runs the events had already expired when cozyreport was collected.
What is missing to prove it: cozyreport has no zfs list of the pools, so the dataset holding the space can't be named, and a snapshot or ZFS metadata hasn't been ruled out.
Fix shape: first, on this failure print zfs list -t all -o name,used,refer,origin for every data-srvN pool together with the LINSTOR resources, so the next red names the dataset. Then fix the leak where it is, not by widening the 512 MiB tolerance, which would hide a real leak. #4651 changes only the replicated lane and does not affect this.
Since 2026-10-01, the kubernetes e2e suites fail their cleanup in the PR lane because one node's ZFS pool does not come back to its pre-suite level. The suite's own checks all pass; it goes red in
cozy_cleanup, which compares each node with the baseline within a 512 MiB tolerance (cozy_linstor_pools_at_baselineinhack/e2e-chainsaw/_lib/run-kubernetes.sh). Each time, 450-830 MiB stays used on one or two nodes and is never freed.Runs, in KiB (baseline → after cleanup):
kubernetes-latest: srv3 74348197 → 73715123, 633074 short. srv1 and srv2 came back exactly, andkubernetes-previousin the same run came back.kubernetes-latest: srv1 832247 short, srv3 463379 short.renovate/indirect-go-modules, Go modules only): both kubernetes suites, srv1 541460 short, then srv3 793677 short. srv1's baseline for the second suite equals what the first one left, so the space is not released later.Passing runs for comparison: 36785869560 on 2026-09-30 (under 1 MiB of drift) and 36845137883 on 2026-10-01 09:48 UTC. None of the failing PRs touches LINSTOR, the tenant teardown or
run-kubernetes.sh, and the renovate branch failed before the Talos bump (#4598) merged.The likely cause, not proven: a CDI scratch volume whose first
CreateVolumefailed leaves a partly created zvol behind. In 36844521765 the scratch PVCpvc-f1883e08first failed withResourceExhausted ... Not enough available nodesat 12:12:38 UTC, was then provisioned and attached on srv3, later loggedVolumeFailedDelete ... still attached to node srv3, and is absent from the final LINSTOR resource list. srv3 is the node that came back short. For the other runs the events had already expired when cozyreport was collected.What is missing to prove it: cozyreport has no
zfs listof the pools, so the dataset holding the space can't be named, and a snapshot or ZFS metadata hasn't been ruled out.Fix shape: first, on this failure print
zfs list -t all -o name,used,refer,originfor everydata-srvNpool together with the LINSTOR resources, so the next red names the dataset. Then fix the leak where it is, not by widening the 512 MiB tolerance, which would hide a real leak. #4651 changes only the replicated lane and does not affect this.