Skip to content

fix(chat): ground diagnosis in pod-level evidence - #43

Merged
hellodk merged 1 commit into
mainfrom
fix/chat-evidence-tools
Sep 19, 2026
Merged

hellodk merged 1 commit into
mainfrom
fix/chat-evidence-tools

Conversation

@hellodk

@hellodk hellodk commented Sep 19, 2026

Copy link
Copy Markdown
Owner

What

The Ask AI engine could only reach aggregate/summarized tools, so diagnosis of
failing workloads (CrashLoopBackOff, CreateContainerConfigError, stuck
rollouts, failed Jobs) was speculative rather than grounded in observed
evidence. This adds read-only, evidence-first tooling plus matching prompt
rules, mirroring the reality discovered in issue #42 (empty SealedSecrets
producing couldn't find key X in Secret Y container-config errors).

Tools added (read-only, typed k8s client, no shell, no writes)

  • describe_pod — container waiting/terminated state, exit codes, pod events
  • pod_logs — tail of current or --previous (last-terminated) container
  • rollout_status — deployment conditions (e.g. ProgressDeadlineExceeded)
    plus ReplicaSets by revision
  • job_status — Job conditions (e.g. BackoffLimitExceeded), failed count,
    and the last failed pod's log tail
  • secret_keys — Secret data keys only (never values)

get_pods now surfaces the container waiting reason (e.g.
CreateContainerConfigError with the missing-key message quoted by events).

Prompt changes

  • Planner: prefer direct-evidence tools over aggregate summaries for
    failing-workload questions
  • Synthesis: evidence-first diagnosis — map observed waiting/terminated
    state to cause classes; never attribute a workload failure to an aggregate
    health/security/cost score without observed evidence; state what could not
    be verified and which kubectl would reveal it

Tests

New chat_evidence_tools_test.go covers each evidence tool (waiting/terminated
state surfacing, previous-logs, rollout ReplicaSet attribution, Job backoff,
secret key-only output). Full analyzer + workspace suite green via make test.

Closes #42

The Ask AI planner could only pick aggregate in-memory tools, so failures
like CreateContainerConfigError (empty SealedSecret keys) and CrashLoopBackOff
(DB-auth crashes) were diagnosed by guesswork instead of observed state.

Add read-only evidence tools backed by the typed k8s client:
describe_pod (container waiting/terminated state, exit codes, pod events),
pod_logs (with previous=true), rollout_status (conditions + ReplicaSets),
job_status (conditions + failed count + last pod tail), and secret_keys
(metadata only, never values). get_pods now also surfaces the container
waiting reason on the summary line.

Tighten both prompts to evidence-first diagnosis: planners prefer the
evidence tools for failing-workload questions, and the synthesis prompt
maps status classes to causes while prohibiting attribution to aggregate
health scores without observed evidence.

Closes #42
@hellodk
hellodk merged commit 76dc6a4 into main Sep 19, 2026
2 checks passed
@hellodk
hellodk deleted the fix/chat-evidence-tools branch September 19, 2026 12:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: chat/AI analysis misses pod-level evidence (empty secrets, DB-auth crashes, stuck rollout)

1 participant