Skip to content

Repository files navigation

Fleetsweeper: the fleet is the policy

The cluster that drifted is the one that pages you at 3am.
Fleetsweeper finds it before it does.

Release Go reference License: MIT

Fleetsweeper flags the one cluster that drifted from the fleet norm


You run twelve clusters. Or fifty. Or two hundred. They started identical and they did not stay that way. Versions skew. Admission policies drift. Service accounts get patched at 3am and nobody writes it down. Every cluster is "healthy" on its own, so every tool you already own stays quiet. Fleetsweeper finds the one cluster that wandered off the herd.

The fleet itself is the baseline. No rulebook. No thresholds to tune. The modified z-score across your own clusters does the work.

What you walk away with after one scan

  • A Fleet Score from 0 to 100 with a one-line headline you can put on a status TV.
  • The cluster that is most unlike the rest, plus the exact fields that flagged it.
  • Ranked, leverage-weighted recommendations. The fix that takes ten clusters from drifted to clean ranks ahead of the same fix on one.
  • An optional admission webhook that denies pods deviating from your fleet's actual norm, not from a static checklist.
  • One unified stream for inbound signal: scan findings, AlertManager, Falco, Trivy CVEs, Kyverno and Gatekeeper PolicyReports.

See it in 30 seconds

go install github.com/dcadolph/fleetsweeper@latest
fleetsweeper serve --demo --addr :8080

Open http://localhost:8080. A synthetic 26-cluster fleet renders across four continents with a 3D globe, findings, trends, outliers, capacity, and a guided tour. No kubeconfig required. The pulsing red dots are the cinematic part. The outlier detection under them is the real part.

Install for real

helm install fleetsweeper deploy/helm/fleetsweeper \
  --set auth.token=$(openssl rand -hex 32) \
  --set controller.enabled=true
kubectl apply -f deploy/examples/clusterscan-prod.yaml

The controller reconciles ClusterScan resources and writes outcomes back to .status. Full installation paths in docs/operator/helm.md. Scoped API keys for pipelines in docs/operator/rbac.md.

Why this and not what you already have

You already use What it tells you What Fleetsweeper adds
kubectl, k9s The state of one cluster, right now. A fleet-wide comparison across 24 dimensions. Names the outlier.
Argo CD, Flux Whether each cluster matches its manifest. Drift across clusters even when every cluster matches its own source of truth.
Prometheus, Grafana Time series for what you remembered to instrument. Statistical baselines derived from the fleet, with no rules to write.
Datadog Cluster Insights Per-cluster alerts scored by a vendor rulebook. The norm is your own fleet, not a vendor checklist.
OPA, Kyverno Violations against rules you authored. Detects drift you forgot to write a rule for. Complements, does not replace.

Production-ready out of the box

HA backends, leader election, scoped RBAC, audit log, declarative CRDs, Prometheus and OpenTelemetry, signed reports, backups, GitOps integrations, admission webhook, and supply-chain signed images. Full checklist: docs/production-readiness.md.

Where to go next

Start here

Concepts

Operator

Integrations

Reference

Contributing

Issues and PRs welcome. Start with CONTRIBUTING.md and the code of conduct. Security disclosures go through SECURITY.md.

License

MIT. See LICENSE.

Releases

Packages

Contributors

Languages