The cluster that drifted is the one that pages you at 3am.
Fleetsweeper finds it before it does.
You run twelve clusters. Or fifty. Or two hundred. They started identical and they did not stay that way. Versions skew. Admission policies drift. Service accounts get patched at 3am and nobody writes it down. Every cluster is "healthy" on its own, so every tool you already own stays quiet. Fleetsweeper finds the one cluster that wandered off the herd.
The fleet itself is the baseline. No rulebook. No thresholds to tune. The modified z-score across your own clusters does the work.
- A Fleet Score from 0 to 100 with a one-line headline you can put on a status TV.
- The cluster that is most unlike the rest, plus the exact fields that flagged it.
- Ranked, leverage-weighted recommendations. The fix that takes ten clusters from drifted to clean ranks ahead of the same fix on one.
- An optional admission webhook that denies pods deviating from your fleet's actual norm, not from a static checklist.
- One unified stream for inbound signal: scan findings, AlertManager, Falco, Trivy CVEs, Kyverno and Gatekeeper PolicyReports.
go install github.com/dcadolph/fleetsweeper@latest
fleetsweeper serve --demo --addr :8080
Open http://localhost:8080. A synthetic 26-cluster fleet renders across
four continents with a 3D globe, findings, trends, outliers, capacity, and
a guided tour. No kubeconfig required. The pulsing red dots are the
cinematic part. The outlier detection under them is the real part.
helm install fleetsweeper deploy/helm/fleetsweeper \
--set auth.token=$(openssl rand -hex 32) \
--set controller.enabled=true
kubectl apply -f deploy/examples/clusterscan-prod.yaml
The controller reconciles ClusterScan resources and writes outcomes back
to .status. Full installation paths in
docs/operator/helm.md. Scoped API keys for
pipelines in docs/operator/rbac.md.
| You already use | What it tells you | What Fleetsweeper adds |
|---|---|---|
kubectl, k9s |
The state of one cluster, right now. | A fleet-wide comparison across 24 dimensions. Names the outlier. |
| Argo CD, Flux | Whether each cluster matches its manifest. | Drift across clusters even when every cluster matches its own source of truth. |
| Prometheus, Grafana | Time series for what you remembered to instrument. | Statistical baselines derived from the fleet, with no rules to write. |
| Datadog Cluster Insights | Per-cluster alerts scored by a vendor rulebook. | The norm is your own fleet, not a vendor checklist. |
| OPA, Kyverno | Violations against rules you authored. | Detects drift you forgot to write a rule for. Complements, does not replace. |
HA backends, leader election, scoped RBAC, audit log, declarative CRDs,
Prometheus and OpenTelemetry, signed reports, backups, GitOps integrations,
admission webhook, and supply-chain signed images. Full checklist:
docs/production-readiness.md.
Start here
- Getting started. First scan, persistence, history, groups.
- Architecture. How the pipeline fits together.
- Scanners. The 24 dimensions Fleetsweeper checks.
Concepts
- The fleet is the policy. Why norm-based detection scales.
- Fleet Score. What the number means.
- Outliers. MAD-based statistical detection.
- Findings and remediation. Severity calibration and
kubectloutputs. - Globe view. Geolocation sources and overrides.
Operator
- Overview, Helm, Server mode
- ClusterScan CRD, Leader election, Backends
- RBAC and API keys, Audit log, OIDC
- Admission webhook, Recommend, What changed
Integrations
- Prometheus, Slack, Webhooks
- PolicyReport, FleetDriftReport
- AlertManager, Falco, Trivy
- Registry probing
Reference
Issues and PRs welcome. Start with CONTRIBUTING.md and
the code of conduct. Security disclosures go through
SECURITY.md.
MIT. See LICENSE.