Current negative result: the usable historical detector evidence is DDP versus idle only. The primary PCIe-only baseline missed its sole training test run; reliable training-versus-inference detection and adversarial robustness are not established. See the indexed evidence and exact coverage analysis.
CommGuard is an installable research SDK for a controlled, single-host Kaggle experiment on two NVIDIA T4 GPUs. It collects content-agnostic GPU telemetry, calibrates accessible PCIe traffic readings against known NCCL workloads, runs bounded training/inference/control workloads, and evaluates leakage-resistant workload classifiers.
This is an independent, unofficial research prototype. It is not affiliated with or endorsed by SPAR, Kairos, ERA, UChicago XLab, William Fowler, or the authors and institutions cited in the related literature.
CommGuard’s Kaggle workflow is a single-node, dual-NVIDIA-T4 research prototype. It validates experimental methodology and software behavior on two local GPU ranks. It does not establish generalization to two physical 8-GPU nodes, NVLink/NVSwitch fabrics, RoCE or InfiniBand networks, large frontier-model workloads, or production treaty-verification deployments.
CommGuard is designed to answer a limited empirical question: whether short windows of signals available in one dual-T4 Kaggle session distinguish the included PyTorch DDP training workloads from included benign inference and control workloads. It does not validate frontier-scale, multi-node, NVLink, NVSwitch, RoCE, InfiniBand, privacy, security, production, or treaty claims. NVML PCIe TX/RX values are always described as PCIe traffic readings. They are not treated as NCCL byte counts or a complete measure of GPU communication.
The SDK distinguishes measured observations, evidence-supported inferences, untested hypotheses, and out-of-scope claims. A failed calibration or a classifier that does not generalize is a valid research result.
Modern calibration-v3 debugging evidence: the first attempted Kaggle T4 x2
run checked out the reviewed source and reached CUDA, NCCL, and the collective
runs, but every idle run failed with ValueError: unknown mode 'idle'. With zero
usable idle repetitions, the modern result was correctly not_supported. That
archive must not gate a benign run: it records a software dispatch defect, not
an empirical finding that PCIe telemetry is unusable. A clean calibration rerun
from this patched branch is required.
Observed limitation: the current pilot's feature coverage is DDP versus idle only. Although all 18 planned benign runs completed, the saved feature artifact contains 16 five-second windows from eight DDP/idle runs; inference, compute, and host-transfer runs were too short after warmup. Calibration-stage idle controls also entered the merged derived set E4 and E5.
Measured/derived negative result: PCIe-only detection missed the held-out training run (run-level balanced accuracy 0.5, training recall 0, false-negative rate 1.0) in a test containing only two runs. Combined and non-PCIe diagnostic features separated that tiny pilot, but current generalization evidence is insufficient. CommGuard has not established reliable training-versus-inference detection or adversarial robustness. See the evidence hash inventory and coverage-failure analysis before interpreting any historical notebook output E5.
CommGuard is distributed from GitHub and is not published to PyPI. For an interactive local checkout, install without changing the environment's dependency stack:
python -m pip install --no-build-isolation --no-deps -e .For the reproducible dual-T4 study, enable Kaggle Internet access, choose
GPU T4 x2, and run these canonical notebooks in order:
commguard_calibration_v3.ipynbcommguard_benign_corpus_v2.ipynbcommguard_detector_evaluation_v2.ipynbcommguard_adversarial_redteam_v1.ipynb
Each notebook installs, in order of preference, one checksum-pinned wheel, one checksum-pinned source archive, or a public 40-character Git commit. Editable source is restricted to explicitly labeled development smoke tests. Downstream notebooks require the exact SHA-256 printed by their predecessor.
The full calibration notebook records five idle repetitions and five AllReduce repetitions at 1, 4, 16, 64, and 128 MiB at the selected interval. A bounded idle collector study compares jitter and sampler duty at 1.0, 0.5, and 0.2 seconds. Smoke mode is explicitly ineligible for scientific acceptance. In the benign notebook, the restored calibration is explicitly prior-session input evidence; a fresh current-session calibration is the actual collection gate. Corpus, extraction, evaluation, and reporting paths verify the exact current calibration path, hash, session, environment, source commit, schema, and status rather than selecting a filename.
Calibration distinguishes supported, partially_supported, inconclusive,
not_supported, and failed result states while retaining compatibility with
older three-state artifacts.
Historical schema-1 support is reported only under its legacy contract and is
not treated as a pass of the new idle-aware gate.
Between Kaggle sessions, download each notebook's exported evidence archive and
upload it as a private Kaggle Dataset for the next notebook. See the
Kaggle prototype runbook for exact steps and
the experiment roadmap for the claim boundary.
git clone https://github.com/waqasm86/CommGuard.git
cd CommGuard
python -m pip install --no-build-isolation --no-deps -e .Equivalent shell commands:
python -m pip install --no-build-isolation --no-deps -e .
commguard preflight --strict --output artifacts
commguard run --profile smoke --output artifacts
commguard run --profile standard --repetitions 3 --output artifacts
# Inspect the matrix summary and continue only if primary_coverage_gate.passed is true.
commguard evaluate --input artifacts --output artifacts --minimum-runs-per-family 3
commguard report --input artifacts --output artifacts/report.md
commguard corpus plan --profile standard --repetitions 3
commguard verify-artifact evidence.tar.gz --sha256 <64-hex-digest>
commguard build-review-bundle --input artifacts --output /tmp/commguard-review.tar.gzThe standard and extended experiment profiles are calibration-gated. They stop when exactly two T4 GPUs, CUDA/NCCL, two distinct rank bindings, or responsive PCIe readings cannot be demonstrated. There is no CPU, Gloo, or one-GPU fallback for results labelled dual-GPU.
The standard profile is the bounded eight-family benign pilot. The extended profile is opt-in and adds configuration variants plus optional peer copy. Both write a corpus plan before execution and report per-family coverage before any detector evaluation. See the benign workload matrix.
Defensive adversarial strategies are not part of those profiles. They require an explicit human approval flag and a predeclared sealed holdout plan; no adversarial result is bundled. See bounded adversarial research.
from commguard import (
AdversarialHoldoutPlan,
check_environment,
evaluate_detector,
extract_features,
generate_report,
list_workloads,
load_artifact,
run_adversarial_matrix,
run_experiment,
run_matrix,
)Core schema and artifact tests are offline and CPU-safe:
pytest -m "not gpu and not multigpu and not network and not slow"See docs/ for architecture, artifact contracts, methodology, safety, and
reproducibility details:
- Kaggle dual-T4 instructions
- Kaggle prototype runbook
- Next Kaggle experiments
- Evidence index
- Current results and coverage failure
- Research report template
- Kaggle prototype report
- Artifact contracts
- Methodology
- Canonical versus executed notebooks
- Limitations and untested behavior
- Reproducibility
Source, releases, and issues are hosted at github.com/waqasm86/CommGuard.