During a deployment update, the rollout is stuck and new pods are not becoming ready. The developer checks the events and sees 'Back-off restarting failed container'. What is the most likely cause?
Trap 1: The liveness probe is failing
A failing liveness probe triggers a container restart, but that alone does not create a stuck rollout with back-off. After the kubelet restarts the container, subsequent successful probes keep the pod running, so no crash-related or image-related back-off occurs. A persistent liveness failure would eventually produce CrashLoopBackOff, but the observed reason would be 'Back-off restarting failed container' after repeated crashes, not an immediate ImagePullBackOff.
Trap 2: The image pull secret is missing
A missing image pull secret results in failed registry authentication, so the kubelet cannot download the container image from a private repository. The pod therefore transitions through phases of ErrImagePull and ImagePullBackOff, which are distinct from crash-loop back-off caused by command exits. Because the container never starts, the rollout cannot progress until the correct Secret is applied to the namespace or referenced in the podSpec.
Trap 3: The pod exceeds its memory limit and is OOMKilled
An OOMKilled status means the kernel terminated the container because it exceeded its memory limit, not because of a faulty entrypoint. The container state shows 'OOMKilled' and the exit reason is '137' (SIGKILL), which is not the same as the restart back-off seen when a command fails repeatedly. While the pod may be restarted, the underlying cause is memory pressure, so fixing it requires raising the limit or reducing workload memory usage, not debugging the entrypoint logic.
- A
The liveness probe is failing
Why it fails: A failing liveness probe triggers a container restart, but that alone does not create a stuck rollout with back-off. After the kubelet restarts the container, subsequent successful probes keep the pod running, so no crash-related or image-related back-off occurs. A persistent liveness failure would eventually produce CrashLoopBackOff, but the observed reason would be 'Back-off restarting failed container' after repeated crashes, not an immediate ImagePullBackOff.
- B
The container's entrypoint command fails immediately after start
An entrypoint that exits immediately with a non-zero code causes the container to crash right after startup. Kubernetes restarts the container, but each attempt fails equally quickly, so the restart loop triggers exponential back-off, producing CrashLoopBackOff. This directly stalls the deployment rollout because the new ReplicaSet can never reach the Ready state and the old ReplicaSet remains active.
- C
The image pull secret is missing
Why it fails: A missing image pull secret results in failed registry authentication, so the kubelet cannot download the container image from a private repository. The pod therefore transitions through phases of ErrImagePull and ImagePullBackOff, which are distinct from crash-loop back-off caused by command exits. Because the container never starts, the rollout cannot progress until the correct Secret is applied to the namespace or referenced in the podSpec.
- D
The pod exceeds its memory limit and is OOMKilled
Why it fails: An OOMKilled status means the kernel terminated the container because it exceeded its memory limit, not because of a faulty entrypoint. The container state shows 'OOMKilled' and the exit reason is '137' (SIGKILL), which is not the same as the restart back-off seen when a command fails repeatedly. While the pod may be restarted, the underlying cause is memory pressure, so fixing it requires raising the limit or reducing workload memory usage, not debugging the entrypoint logic.