Quick and dirty query to find recent events indicating an unhealthy cluster (good one to run after maintenance):
kubectl get events -A --sort-by=.lastTimestamp | grep -iE 'preempt|evict|kill|fail|unhealthy|backoff'And here's a kind of elaborate one to get actual CPU/memory usage vs requests/limits for right-sizing: