Skip to content

Latest commit

 

History

47 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CommGuard

Current negative result: the usable historical detector evidence is DDP versus idle only. The primary PCIe-only baseline missed its sole training test run; reliable training-versus-inference detection and adversarial robustness are not established. See the indexed evidence and exact coverage analysis.

CommGuard is an installable research SDK for a controlled, single-host Kaggle experiment on two NVIDIA T4 GPUs. It collects content-agnostic GPU telemetry, calibrates accessible PCIe traffic readings against known NCCL workloads, runs bounded training/inference/control workloads, and evaluates leakage-resistant workload classifiers.

This is an independent, unofficial research prototype. It is not affiliated with or endorsed by SPAR, Kairos, ERA, UChicago XLab, William Fowler, or the authors and institutions cited in the related literature.

Scope

CommGuard’s Kaggle workflow is a single-node, dual-NVIDIA-T4 research prototype. It validates experimental methodology and software behavior on two local GPU ranks. It does not establish generalization to two physical 8-GPU nodes, NVLink/NVSwitch fabrics, RoCE or InfiniBand networks, large frontier-model workloads, or production treaty-verification deployments.

CommGuard is designed to answer a limited empirical question: whether short windows of signals available in one dual-T4 Kaggle session distinguish the included PyTorch DDP training workloads from included benign inference and control workloads. It does not validate frontier-scale, multi-node, NVLink, NVSwitch, RoCE, InfiniBand, privacy, security, production, or treaty claims. NVML PCIe TX/RX values are always described as PCIe traffic readings. They are not treated as NCCL byte counts or a complete measure of GPU communication.

The SDK distinguishes measured observations, evidence-supported inferences, untested hypotheses, and out-of-scope claims. A failed calibration or a classifier that does not generalize is a valid research result.

Current evidence status

Modern calibration-v3 debugging evidence: the first attempted Kaggle T4 x2 run checked out the reviewed source and reached CUDA, NCCL, and the collective runs, but every idle run failed with ValueError: unknown mode 'idle'. With zero usable idle repetitions, the modern result was correctly not_supported. That archive must not gate a benign run: it records a software dispatch defect, not an empirical finding that PCIe telemetry is unusable. A clean calibration rerun from this patched branch is required.

Observed limitation: the current pilot's feature coverage is DDP versus idle only. Although all 18 planned benign runs completed, the saved feature artifact contains 16 five-second windows from eight DDP/idle runs; inference, compute, and host-transfer runs were too short after warmup. Calibration-stage idle controls also entered the merged derived set E4 and E5.

Measured/derived negative result: PCIe-only detection missed the held-out training run (run-level balanced accuracy 0.5, training recall 0, false-negative rate 1.0) in a test containing only two runs. Combined and non-PCIe diagnostic features separated that tiny pilot, but current generalization evidence is insufficient. CommGuard has not established reliable training-versus-inference detection or adversarial robustness. See the evidence hash inventory and coverage-failure analysis before interpreting any historical notebook output E5.

Kaggle quick start

CommGuard is distributed from GitHub and is not published to PyPI. For an interactive local checkout, install without changing the environment's dependency stack:

python -m pip install --no-build-isolation --no-deps -e .

For the reproducible dual-T4 study, enable Kaggle Internet access, choose GPU T4 x2, and run these canonical notebooks in order:

  1. commguard_calibration_v3.ipynb
  2. commguard_benign_corpus_v2.ipynb
  3. commguard_detector_evaluation_v2.ipynb
  4. commguard_adversarial_redteam_v1.ipynb

Each notebook installs, in order of preference, one checksum-pinned wheel, one checksum-pinned source archive, or a public 40-character Git commit. Editable source is restricted to explicitly labeled development smoke tests. Downstream notebooks require the exact SHA-256 printed by their predecessor.

The full calibration notebook records five idle repetitions and five AllReduce repetitions at 1, 4, 16, 64, and 128 MiB at the selected interval. A bounded idle collector study compares jitter and sampler duty at 1.0, 0.5, and 0.2 seconds. Smoke mode is explicitly ineligible for scientific acceptance. In the benign notebook, the restored calibration is explicitly prior-session input evidence; a fresh current-session calibration is the actual collection gate. Corpus, extraction, evaluation, and reporting paths verify the exact current calibration path, hash, session, environment, source commit, schema, and status rather than selecting a filename.

Calibration distinguishes supported, partially_supported, inconclusive, not_supported, and failed result states while retaining compatibility with older three-state artifacts. Historical schema-1 support is reported only under its legacy contract and is not treated as a pass of the new idle-aware gate. Between Kaggle sessions, download each notebook's exported evidence archive and upload it as a private Kaggle Dataset for the next notebook. See the Kaggle prototype runbook for exact steps and the experiment roadmap for the claim boundary.

Local checkout

git clone https://github.com/waqasm86/CommGuard.git
cd CommGuard
python -m pip install --no-build-isolation --no-deps -e .

Equivalent shell commands:

python -m pip install --no-build-isolation --no-deps -e .
commguard preflight --strict --output artifacts
commguard run --profile smoke --output artifacts
commguard run --profile standard --repetitions 3 --output artifacts
# Inspect the matrix summary and continue only if primary_coverage_gate.passed is true.
commguard evaluate --input artifacts --output artifacts --minimum-runs-per-family 3
commguard report --input artifacts --output artifacts/report.md
commguard corpus plan --profile standard --repetitions 3
commguard verify-artifact evidence.tar.gz --sha256 <64-hex-digest>
commguard build-review-bundle --input artifacts --output /tmp/commguard-review.tar.gz

The standard and extended experiment profiles are calibration-gated. They stop when exactly two T4 GPUs, CUDA/NCCL, two distinct rank bindings, or responsive PCIe readings cannot be demonstrated. There is no CPU, Gloo, or one-GPU fallback for results labelled dual-GPU.

The standard profile is the bounded eight-family benign pilot. The extended profile is opt-in and adds configuration variants plus optional peer copy. Both write a corpus plan before execution and report per-family coverage before any detector evaluation. See the benign workload matrix.

Defensive adversarial strategies are not part of those profiles. They require an explicit human approval flag and a predeclared sealed holdout plan; no adversarial result is bundled. See bounded adversarial research.

Public API

from commguard import (
    AdversarialHoldoutPlan,
    check_environment,
    evaluate_detector,
    extract_features,
    generate_report,
    list_workloads,
    load_artifact,
    run_adversarial_matrix,
    run_experiment,
    run_matrix,
)

Core schema and artifact tests are offline and CPU-safe:

pytest -m "not gpu and not multigpu and not network and not slow"

See docs/ for architecture, artifact contracts, methodology, safety, and reproducibility details:

Source, releases, and issues are hosted at github.com/waqasm86/CommGuard.

About

CommGuard is an installable research SDK for a controlled, single-host Kaggle experiment on two NVIDIA T4 GPUs. It collects content-agnostic GPU telemetry, calibrates accessible PCIe traffic readings against known NCCL workloads, runs bounded training/inference/control workloads, and evaluates leakage-resistant workload classifiers.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages