Pod isn't serving — symptom-first diagnosis decision tree

Pod isn't serving — symptom-first diagnosis decision tree A workflow diagram generated by Archify. 01 / Symptom → first check 02 / STATUS · READY gates 03 / Matched ↓ diagnostic command Pod not serving · no service response · Symptom → first check Pod not serving no service response kubectl get pod -o wide · STATUS · READY · RESTARTS · Symptom → first check kubectl get pod -o wide STATUS · READY · RESTARTS No node assigned · PodScheduled=False · STATUS · READY gates No node assigned PodScheduled=False Image / setup · pull · CNI · mount · STATUS · READY gates Image / setup pull · CNI · mount CrashLoopBackOff · RESTARTS keeps rising · STATUS · READY gates CrashLoopBackOff RESTARTS keeps rising Pod Ready=False · Probe · readiness gate · STATUS · READY gates Pod Ready=False Probe · readiness gate Pod Ready=True · Verify response separately · STATUS · READY gates Pod Ready=True Verify response separately kubectl describe pod · Events: FailedScheduling · Matched ↓ diagnostic command kubectl describe pod Events: FailedScheduling kubectl describe pod · Events: pull · CNI · mounts · Matched ↓ diagnostic command kubectl describe pod Events: pull · CNI · mounts kubectl logs --previous · Reason · exit · events · Matched ↓ diagnostic command kubectl logs --previous Reason · exit · events kubectl describe pod · Events: Readiness probe failed · Matched ↓ diagnostic command kubectl describe pod Events: Readiness probe failed kubectl get endpointslices · Addresses + ready · selector · DNS · Matched ↓ diagnostic command kubectl get endpointslices Addresses + ready · selector · DNS no no no no Legend Diagnostic command STATUS · READY gate Network path check Symptom

Separate scheduling and initialization

  • • Pending pods can already have an assigned node
  • • Scheduling failure-reason counts can overlap
  • • After binding, inspect image, CNI, and mount states

Restart and image failures

  • • For OOMKilled distinguish limits from node pressure
  • • Probe failure can reflect a real application problem
  • • Check ECR identity, image existence, and network separately

Readiness and real reachability

  • • Excluding all Running pods can miss CrashLoops
  • • An EndpointSlice address can have ready=false
  • • Check policy ingress/egress, DNS, and LB health