Skip to content

fix(pod-health): surface high-restart and init-failure pods to the issues page - #37

Merged
hellodk merged 1 commit into
mainfrom
fix/pod-health-detection
Sep 10, 2026
Merged

hellodk merged 1 commit into
mainfrom
fix/pod-health-detection

Conversation

@hellodk

@hellodk hellodk commented Sep 10, 2026

Copy link
Copy Markdown
Owner

Live cluster check showed 117 pods where 113 were declared healthy while
70+ actually had elevated restart counts (jfrog: 742, kube-proxy: 72,
calico-node: 31) plus a CrashLoopBackOff and an OOMKilled. The scanner
only looked at Waiting.Reason and LastTerminationState, never at the
restart count, and ignored init container states entirely.

Backend (podhealth.go):

  • New 'high-restarts' category: Running pods with >5 restarts are no
    longer counted healthy
  • New 'init-failure' category: init containers in CrashLoopBackOff /
    ImagePullBackOff / non-zero exit are surfaced, not silently passed
  • PodHealthItem gains restartsPerHour so 742 restarts over 2 days and 6
    restarts overnight are distinguishable
  • categoryOrder is now a package var (also used by tests)

Frontend (issues page):

  • Pod health counts now come from /api/v1/pods/health (the scanner)
    instead of health.summary.unhealthyPods, which is computed from the
    rule engine's TopIssues and was 0 for pod-level problems
  • Pod Health breakdown adds 'High Restarts' and 'Init Failures' rows
  • Health ring now penalizes high-restart and init-failure pods

Closes #36

…sues page

Live cluster check showed 117 pods where 113 were declared healthy while
70+ actually had elevated restart counts (jfrog: 742, kube-proxy: 72,
calico-node: 31) plus a CrashLoopBackOff and an OOMKilled. The scanner
only looked at Waiting.Reason and LastTerminationState, never at the
restart count, and ignored init container states entirely.

Backend (podhealth.go):
- New 'high-restarts' category: Running pods with >5 restarts are no
  longer counted healthy
- New 'init-failure' category: init containers in CrashLoopBackOff /
  ImagePullBackOff / non-zero exit are surfaced, not silently passed
- PodHealthItem gains restartsPerHour so 742 restarts over 2 days and 6
  restarts overnight are distinguishable
- categoryOrder is now a package var (also used by tests)

Frontend (issues page):
- Pod health counts now come from /api/v1/pods/health (the scanner)
  instead of health.summary.unhealthyPods, which is computed from the
  rule engine's TopIssues and was 0 for pod-level problems
- Pod Health breakdown adds 'High Restarts' and 'Init Failures' rows
- Health ring now penalizes high-restart and init-failure pods

Closes #36
@hellodk
hellodk merged commit 6aae874 into main Sep 10, 2026
1 check passed
@hellodk
hellodk deleted the fix/pod-health-detection branch September 10, 2026 18:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(pod-health): detect high-restart pods, init failures, and wire scanner into issues page

1 participant