IntermediateOrchestration & Kubernetes← All practice

Exit Code 137 at 03:00 Every Night

The report, in their words

One pod of the report-rendering service restarts most nights between 03:00 and 03:20. No alert has ever fired, because the Deployment always has other replicas ready and the restart is over in ten seconds. Someone finally looked at the pod: Restart Count: 46, last state Terminated, Exit Code: 137.

Pull evidence

One item at a time, and nothing here tells you which one matters. Deciding what is worth looking at is most of the diagnosis.

kubectl describe pod — last terminated state
The container's resources block
Container memory over 24 hours
What runs at 03:00
The node at the moment of the kill
Was anything evicted?
Why no alert ever fired
CPU throttling during the nightly run
Application log at 03:07

0 of 9 inspected. You are not required to open all of them — a real investigation is judged on how few you needed.

What is your diagnosis?

Commit to one. Nothing below is shown until you do.

Guessing wrong and being told exactly why is the point of this page. Reading the answer first is not practice.