fix(pod-health): surface high-restart and init-failure pods to the issues page - #37
Merged
Merged
Conversation
…sues page Live cluster check showed 117 pods where 113 were declared healthy while 70+ actually had elevated restart counts (jfrog: 742, kube-proxy: 72, calico-node: 31) plus a CrashLoopBackOff and an OOMKilled. The scanner only looked at Waiting.Reason and LastTerminationState, never at the restart count, and ignored init container states entirely. Backend (podhealth.go): - New 'high-restarts' category: Running pods with >5 restarts are no longer counted healthy - New 'init-failure' category: init containers in CrashLoopBackOff / ImagePullBackOff / non-zero exit are surfaced, not silently passed - PodHealthItem gains restartsPerHour so 742 restarts over 2 days and 6 restarts overnight are distinguishable - categoryOrder is now a package var (also used by tests) Frontend (issues page): - Pod health counts now come from /api/v1/pods/health (the scanner) instead of health.summary.unhealthyPods, which is computed from the rule engine's TopIssues and was 0 for pod-level problems - Pod Health breakdown adds 'High Restarts' and 'Init Failures' rows - Health ring now penalizes high-restart and init-failure pods Closes #36
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Live cluster check showed 117 pods where 113 were declared healthy while
70+ actually had elevated restart counts (jfrog: 742, kube-proxy: 72,
calico-node: 31) plus a CrashLoopBackOff and an OOMKilled. The scanner
only looked at Waiting.Reason and LastTerminationState, never at the
restart count, and ignored init container states entirely.
Backend (podhealth.go):
longer counted healthy
ImagePullBackOff / non-zero exit are surfaced, not silently passed
restarts overnight are distinguishable
Frontend (issues page):
instead of health.summary.unhealthyPods, which is computed from the
rule engine's TopIssues and was 0 for pod-level problems
Closes #36