Quick and dirty query to find recent events indicating an unhealthy cluster (good one to run after maintenance):
kubectl get events -A --sort-by=.lastTimestamp | grep -iE 'preempt|evict|kill|fail|unhealthy|backoff'And here's a kind of elaborate one to get actual CPU/memory usage vs requests/limits for right-sizing:
{
printf 'NAMESPACE/POD CPU CPU_REQ CPU_LIM MEM MEM_REQ MEM_LIM\n'
join <(kubectl get pods -A -o json | jq -r '
def q2m: if . == null then null else tostring |
if endswith("m") then (.[:-1]|tonumber)
elif endswith("n") then (.[:-1]|tonumber)/1e6
elif endswith("u") then (.[:-1]|tonumber)/1e3
else (tonumber*1000) end
end;
def q2Mi: if . == null then null else tostring |
if endswith("Ki") then (.[:-2]|tonumber)/1024
elif endswith("Mi") then (.[:-2]|tonumber)
elif endswith("Gi") then (.[:-2]|tonumber)*1024
elif endswith("Ti") then (.[:-2]|tonumber)*1048576
elif endswith("k") then (.[:-1]|tonumber)/1024
elif endswith("M") then (.[:-1]|tonumber)*1e6/1048576
elif endswith("G") then (.[:-1]|tonumber)*1e9/1048576
else (tonumber/1048576) end
end;
def total(f; unit): [.spec.containers[] | f] | map(select(. != null))
| if length == 0 then "-" else ((add|round|tostring) + unit) end;
.items[] | [
"\(.metadata.namespace)/\(.metadata.name)",
total(.resources.requests.cpu|q2m; "m"),
total(.resources.limits.cpu|q2m; "m"),
total(.resources.requests.memory|q2Mi; "Mi"),
total(.resources.limits.memory|q2Mi; "Mi")
] | u/tsv' | tr '\t' ' ' | sort) \
<(kubectl top pod -A --no-headers | awk '{print $1"/"$2, $3, $4}' | sort) \
| awk '{print $1, $6, $2, $3, $7, $4, $5}'
} | column -t
produces