Skip to main content
Kai monitors your Amazon EKS cluster continuously, surfacing over-provisioned pods, underutilized nodes, and missing autoscaling policies before they cause outages.

The scenario

A platform team runs a production EKS cluster across several namespaces. CPU alerts are intermittent but investigation is slow—engineers run manual kubectl commands across hundreds of pods to correlate logs, metrics, and events.
Manual Kubernetes troubleshooting challenges across namespaces and resources

Manual Kubernetes troubleshooting challenges

The team asks Kai to assess the cluster end-to-end, identify resource waste, and recommend autoscaling policies where they are missing.

Walkthrough

What made this work

  • Kai queries the cluster API directly, replacing manual kubectl sessions and tool-switching.
  • Cross-layer correlation links pod utilization, node capacity, and scheduling patterns in a single analysis pass.
  • #report and #chart produce structured output Kai can reason over before surfacing findings.
  • #recommend generates actionable HPA policy changes rather than a raw metrics dump.
  • CloudKeepers can run this analysis on a schedule so findings arrive before on-call engineers are paged.

Try it yourself

Kai agent reference

Full capabilities of Kai, the Kubernetes Engineer agent

Kubernetes connection

Step-by-step guide to connecting CloudThinker to your EKS cluster

Topology Explorer

Map Kubernetes service dependencies for faster incident root cause analysis

CloudKeepers

Run continuous health checks across your Kubernetes workloads automatically