Courseiva

CKA · domain

Troubleshooting

The Troubleshooting domain (30%) covers diagnosing broken workloads, services, and nodes on a live cluster. You are tested through scenario tasks: inspecting Pending pods, tracing Service endpoint failures, reading kubelet and control-plane logs, and repairing misconfigured networking, scheduling, or storage so the cluster recovers.

207 questions51 easy103 medium53 hard

Focused practice

Practice Troubleshooting questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Troubleshooting

You must systematically isolate faults across pods, Services, and nodes using kubectl describe, logs, events, and node-level systemctl/journalctl checks, then apply the fix. The single most important thing is reading events and status conditions before changing any configuration.

Diagnosing Pending pods via kubectl describe, events, taints, and resource requests

Tracing Service reachability using endpoints, selectors, kube-proxy, and CoreDNS

Reading kubelet state with systemctl status and journalctl -u kubelet

Resolving CNI and node NotReady conditions shown by kubectl get nodes

Watch out for

Common Troubleshooting exam traps

  • ▸Checking application logs first instead of running kubectl describe pod and kubectl get events to see scheduling or image errors
  • ▸Assuming a Service is broken when its selector does not match pod labels, leaving endpoints empty
  • ▸Ignoring node taints, cordons, or NotReady conditions that silently block scheduling and pod networking

Question index

All Troubleshooting questions (207)

Click any question to see the full explanation, or start a practice session above.

1

Which command should you use to view the logs of a container that has previously crashed in a Pod?

Easy
2

A user reports that their application cannot resolve DNS names for services in the cluster. The application runs in a pod with dnsPolicy: ClusterFirst. What is the most likely cause?

Hard
3

You need to see the startup logs of the kubelet service. Which command should you use?

Easy
4

A node in the cluster is reporting NotReady. You run 'journalctl -u kubelet' and see repeated errors: 'failed to get system container stats for "/system.slice/kubelet.service": failed to get cgroup stats'. What is the most likely cause?

Medium
5

You suspect a DNS issue within the cluster. Which TWO commands can you run from within a pod to test DNS resolution?

Hard
6

You need to test DNS resolution from within a pod. Which command should you run?

Medium
7

A pod is stuck in 'Pending' state. Which TWO of the following are common causes?

Medium
8

You are troubleshooting a pod that is failing to start due to a volume mount error. The pod spec references a PersistentVolumeClaim (PVC) named 'data-pvc'. You run 'kubectl get pvc data-pvc -n default' and see the status is 'Pending'. Which of the following is the MOST likely cause?

Medium
9

A pod is in 'ImagePullBackOff' state. Which TWO are valid first troubleshooting steps?

Easy
10

A pod is in ImagePullBackOff state. You run 'kubectl describe pod mypod' and see 'Failed to pull image "myapp:latest": rpc error: code = Unknown desc = Error response from daemon: manifest for myapp:latest not found: manifest unknown'. What is the most likely cause?

Medium
11

A Pod is running but cannot connect to a Service. You have verified that the Service endpoints are correct. Which of the following is the most likely cause if the Pod is using hostNetwork: true?

Hard
12

A pod is in 'ImagePullBackOff' state. You run 'kubectl describe pod' and see the event: 'Failed to pull image "nginx:latest": rpc error: code = Unknown desc = Error response from daemon: Get https://registry-1.docker.io/v2/: dial tcp: lookup registry-1.docker.io on 8.8.8.8:53: no such host'. What is the MOST likely cause?

Medium
13

You are tasked with troubleshooting a production Kubernetes cluster. A user reports that they cannot access a web application running in the cluster. The application is deployed as a Deployment named 'frontend' with 2 replicas, exposed via a Service of type LoadBalancer. You have kubectl access to the cluster. You run 'kubectl get pods -l app=frontend' and see both pods are Running and Ready. You run 'kubectl get svc frontend' and see the Service has an external IP of 192.168.1.100. However, when you curl http://192.168.1.100 from a machine outside the cluster, you get a connection timeout. You are able to curl the pod IPs directly from within the cluster and get a response. Which of the following is the most likely cause of the issue?

Easy
14

An application team reports that a Deployment's Pods in namespace 'web' become unreachable after a rolling update, even though 'kubectl get pods' shows them as Running and Ready. The Service 'web-svc' of type ClusterIP exists and has endpoints listed. You exec into a client Pod in the same namespace and run 'curl http://web-svc:8080'; the connection times out. You then run 'kubectl get networkpolicy -n web' and see a policy named 'deny-all-ingress' with podSelector matching the web Pods and policyTypes: Ingress, but no ingress rules. Which of the following is the MOST likely cause of the timeout?

Medium
15

You run 'kubectl get pods' and see a pod in 'ImagePullBackOff' state. What is the most likely cause?

Easy
16

You run 'kubectl top nodes' but get 'error: metrics not available yet'. What is the most likely cause?

Medium
17

You are troubleshooting a node that is in 'NotReady' state. Which command should you use to check the kubelet logs for errors?

Easy
18

Which TWO of the following are valid methods to diagnose why a node is in 'NotReady' state?

Medium
19

You see events like 'Failed to pull image' when running 'kubectl get events'. Which command can you use to get more detailed information about a specific pod's events?

Easy
20

A pod is stuck in Pending state. 'kubectl describe pod' shows '0/1 nodes are available: 1 node(s) had taint {key=value: NoSchedule}, 1 node(s) had taint {node.kubernetes.io/unreachable: }. What does this indicate?

Medium
21

You run 'kubectl get nodes' and one node shows 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. The output shows 'active (running)'. What should you check NEXT?

Medium
22

A pod is in CrashLoopBackOff state. Which command shows the last termination reason?

Easy
23

Which TWO of the following are valid commands to check resource usage of pods?

Easy
24

Which of the following are valid methods to troubleshoot a Node that is 'NotReady'? (Select all that apply)

Medium
25

A node in your cluster is reporting 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. The output shows 'Active: inactive (dead)'. Which command should you run FIRST to attempt to resolve this?

Hard
26

You are unable to resolve a Service DNS name from within a pod. To test DNS resolution, which command should you run inside the pod?

Medium
27

Which kubectl command is used to view the logs of a container that has previously crashed in a pod?

Easy
28

A pod is in 'Pending' state. 'kubectl describe pod' shows: '0/3 nodes are available: 1 Insufficient memory, 2 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate.' Which THREE actions would resolve the issue? (Choose three)

Medium
29

You need to check the logs of a container that previously crashed. The pod is currently running, but the previous instance of the container exited with an error. Which command will show you the logs from the crashed container?

Hard
30

You run 'kubectl top nodes' and it returns an error: 'error: metrics not available yet'. What does this indicate?

Easy
31

You have a multi-node cluster. One node shows 'NotReady'. You run 'journalctl -u kubelet' on that node and see 'network plugin is not ready'. What is the most likely cause?

Hard
32

Which TWO of the following are common causes of CrashLoopBackOff? (Choose two)

Medium
33

A pod is in 'Pending' state. After running 'kubectl describe pod', you see the event: '0/3 nodes are available: 3 PersistentVolumeClaim is not bound'. What is the most likely cause?

Medium
34

You need to investigate why a service is not reachable from within the cluster. Which of the following is the first step?

Medium
35

You need to check the CPU and memory usage of all pods in the 'production' namespace. Which command should you use?

Easy
36

A pod is failing to start with the error 'CrashLoopBackOff'. You check the logs with 'kubectl logs pod' and see nothing. What is the most likely reason?

Medium
37

You have a Deployment with the following resource limits for containers: memory: 256Mi. The pod is repeatedly killed with OOMKilled. You need to change the limit to 512Mi. Which field should you modify in the Deployment YAML?

Hard
38

You want to see the last 50 lines of logs from a pod named 'api-pod' for the container 'api-container'. Which command accomplishes this?

Easy
39

Which three are possible reasons for a pod being in Pending state? (Choose three.)

Medium
40

You run 'kubectl get pods' and see a pod in 'ImagePullBackOff' state. Which command would help you determine the exact reason for the image pull failure?

Medium
41

You have a Pod that is stuck in Pending state. Running 'kubectl describe pod' shows events: '0/4 nodes are available: 1 node(s) had taint {node-role.kubernetes.io/control-plane: }, 3 node(s) had taint {key: value}, that the pod didn't tolerate.' How can you resolve this issue?

Hard
42

A node in the cluster is showing NotReady status. Which steps should you take to diagnose the issue? (Select the BEST initial step.)

Medium
43

A pod runs but you cannot connect to its container port from another pod in the same namespace. 'kubectl exec' into the pod and 'curl localhost:8080' works. What is the MOST likely cause?

Hard
44

Which THREE of the following are valid steps to troubleshoot a Node in NotReady state? (Choose three)

Hard
45

You are troubleshooting a network connectivity issue between two pods in different namespaces. The pods have the following labels: pod-a in namespace 'foo' with labels {app: web}, pod-b in namespace 'bar' with labels {app: db}. You verify that both pods have IP addresses and can ping the Kubernetes service IP. However, pod-a cannot connect to pod-b on port 5432. What should you check first?

Hard
46

You run 'kubectl get pods' and see that a pod is in CrashLoopBackOff. Which THREE of the following are valid next steps? (Select 3)

Hard
47

A pod is stuck in Pending state. You describe the pod and see: '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate.' What is the most likely reason?

Hard
48

A worker node in your cluster has been marked NotReady for several minutes. You SSH into the node and find that the kubelet process is not running. You start the kubelet service, but it immediately exits. Running 'journalctl -u kubelet -n 50' shows repeated errors about being unable to load the kubeconfig file at /etc/kubernetes/kubelet.conf. Which of the following is the MOST likely cause?

Medium
49

Which THREE of the following are common causes for a Pod to remain in Pending state? (Select THREE.)

Medium
50

You run 'kubectl get nodes' and one node shows 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. Kubelet is active but 'journalctl -u kubelet -n 50' shows 'network plugin is not ready: cni config uninitialized'. What is the most likely cause?

Hard
51

Which command shows CPU and memory usage of nodes in the cluster?

Easy
52

A pod is in ImagePullBackOff. Which TWO of the following are possible causes? (Select 2)

Medium
53

You have a pod that is in 'CrashLoopBackOff' state. Which command should you use to view the logs from the previous instance of the container?

Easy
54

You are troubleshooting a node that is 'NotReady'. Which THREE of the following are possible causes? (Choose three.)

Hard
55

You have a pod that is CrashLoopBackOff. The logs show 'error: dial tcp: lookup service.default.svc.cluster.local: no such host'. What is the most likely cause?

Hard
56

You are troubleshooting a pod that cannot start. Running 'kubectl describe pod' shows the event: 'Failed to pull image "myregistry.io/myapp:1.0": rpc error: code = Unknown desc = Error response from daemon: manifest for myregistry.io/myapp:1.0 not found'. What is the MOST likely cause?

Hard
57

A node is 'NotReady'. Which THREE steps should you take to troubleshoot?

Medium
58

A pod is in ImagePullBackOff state. Which command is MOST useful to diagnose the issue?

Medium
59

You run 'kubectl get pods' and see a pod with status 'ImagePullBackOff'. Which of the following is a possible cause?

Hard
60

You try to run 'kubectl logs mypod' and get the error: 'Error from server (BadRequest): container "myapp" in pod "mypod" is waiting to start: PodInitializing'. What does this mean?

Medium
61

A pod in the 'production' namespace is in CrashLoopBackOff state. Running 'kubectl describe pod web-app -n production' shows the event 'OOMKilled'. What is the most appropriate action to resolve this issue?

Medium
62

You need to check the memory usage of all pods in the 'production' namespace. Which command fulfills this requirement?

Medium
63

Which TWO of the following kubectl commands can be used to view the logs of a container in a pod? (Choose two.)

Easy
64

A Pod is in 'CrashLoopBackOff' state. You run 'kubectl logs <pod> --previous' and see an error about a missing environment variable. The Pod spec defines the environment variable in a ConfigMap. What is the best next step to diagnose the issue?

Hard
65

A node in your cluster is reporting 'NotReady' status. You log into the node and run 'systemctl status kubelet'. The kubelet service is not running. Which command should you use to start the kubelet and enable it to start on boot?

Easy
66

Which THREE of the following are valid commands to troubleshoot network connectivity between pods? (Select 3)

Hard
67

After deploying a new Deployment, you run 'kubectl get events' and see 'FailedScheduling' events. What is a possible cause?

Medium
68

A Pod is stuck in Pending state. 'kubectl describe pod' shows the event: '0/4 nodes are available: 1 node had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate, 3 Insufficient cpu.' Which of the following is the most likely combination of issues?

Hard
69

A pod named 'app' is not starting. You run 'kubectl describe pod app' and see the event: 'MountVolume.SetUp failed for volume "pvc-volume" : rpc error: code = NotFound desc = volume not found'. What is the most likely issue?

Easy
70

You run 'kubectl get pods' and see a pod named 'db' in CrashLoopBackOff. 'kubectl logs db' shows nothing. 'kubectl logs db --previous' shows 'Error: database connection failed'. What is the most likely cause?

Medium
71

A Service of type ClusterIP is not reachable from within the cluster. Pods backing the Service are running and healthy. What is the most likely cause?

Hard
72

You need to check the resource usage of nodes in your cluster. Which command should you run?

Medium
73

Which of the following are valid methods to debug a failing CoreDNS pod? (Select TWO)

Medium
74

You have a Pod that is in CrashLoopBackOff. Which TWO of the following commands would be most helpful in diagnosing the issue?

Hard
75

A pod is stuck in Pending with event '0/4 nodes are available: 1 node(s) had taint "node.kubernetes.io/disk-pressure", and 3 node(s) had taint "node.kubernetes.io/memory-pressure", that the pod didn't tolerate'. What is the best approach to schedule the pod?

Hard
76

You run 'kubectl get pods' and see a pod with status 'Init:CrashLoopBackOff'. What does this indicate?

Medium
77

You run kubectl get nodes and see one node is NotReady. The kubelet is running on the node. What is the most likely cause?

Medium
78

A developer reports that a Pod named 'web-pod' in namespace 'frontend' is crashing repeatedly. You run 'kubectl logs web-pod -n frontend' but see no output. Which command should you run next to see the logs from the previous, crashed container instance?

Medium
79

A pod is in CrashLoopBackOff. The YAML for the initContainer is: apiVersion: v1 kind: Pod metadata: name: myapp spec: initContainers: - name: init image: busybox command: ['sh', '-c', 'sleep 5 && exit 1'] containers: - name: app image: nginx What is the most likely reason for the CrashLoopBackOff?

Hard
80

A node in the cluster is showing status 'NotReady'. You run 'kubectl describe node worker1' and see that the kubelet has not posted node status for more than 1 minute. Which command should you run on the node to check the kubelet logs?

Medium
81

You suspect the kubelet on a worker node has stopped. Which two commands should you run to confirm the kubelet status and check its logs?

Medium
82

You need to check the current resource usage of nodes in your cluster. Which command should you use?

Easy
83

Which command can you run to see the events related to a specific pod?

Easy
84

You need to check the status of control plane components. Which TWO commands are appropriate?

Medium
85

A node shows status NotReady. You SSH into the node and run 'systemctl status kubelet' which shows the kubelet is active (running). What is the next most likely step to diagnose the issue?

Medium
86

You run 'kubectl get pods' and see a pod with status 'CrashLoopBackOff'. You check the logs with 'kubectl logs <pod> --previous' and see: 'Error: unable to connect to database at db-svc:5432 (connection refused)'. What is the most likely cause?

Medium
87

A Pod is in CrashLoopBackOff. You run 'kubectl describe pod' and see that the container fails with 'Error: container command not found'. What is the most likely cause?

Medium
88

Which command shows events sorted by timestamp for troubleshooting recent issues?

Easy
89

After deploying a new Deployment, you notice that the pods are stuck in ImagePullBackOff. What is the most common cause?

Medium
90

You have a pod that is in 'Pending' state because it requires a PersistentVolumeClaim that is not bound. Which event would you see in 'kubectl describe pod'?

Easy
91

A pod is in Pending state. You see the event: '0/2 nodes are available: 2 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate'. What should you do to schedule the pod on one of the control-plane nodes?

Medium
92

A pod named 'web-app' is crashing repeatedly. You run 'kubectl describe pod web-app' and see that the container exited with code 137. What does this indicate?

Medium
93

A pod is in ImagePullBackOff state. Which command can you run to get more details about the underlying error?

Medium
94

A pod is in 'Pending' state. Which TWO of the following are possible causes? (Select 2)

Medium
95

A pod has been in Pending state for a long time. 'kubectl describe pod' shows the event: '0/3 nodes are available: 1 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate, 2 node(s) had taint {node.kubernetes.io/unreachable: }, that the pod didn't tolerate.' What is the most likely cause?

Medium
96

Which TWO of the following commands are useful for debugging network connectivity between pods?

Medium
97

You suspect a DNS issue inside a pod. Which command can you run to test DNS resolution from within a pod?

Medium
98

A pod is running but cannot be accessed via its ClusterIP service from another pod in the same namespace. The service endpoints list shows the pod's IP. What is the most likely cause?

Hard
99

A pod is stuck in Pending state. 'kubectl describe pod' shows the event: '0/3 nodes are available: 3 node(s) didn't match pod anti-affinity rules'. What is the most likely cause?

Hard
100

A pod is stuck in 'Pending' state. Which command would you run FIRST to diagnose the issue?

Medium
101

You deploy a pod with image 'nginx:1.21'. It stays in ImagePullBackOff. You run 'kubectl describe pod nginx-pod' and see the event: 'Failed to pull image "nginx:1.21": rpc error: code = Unknown desc = Error response from daemon: manifest for nginx:1.21 not found'. What is the most likely fix?

Medium
102

A Node is in NotReady state. Which action should be taken first to diagnose the issue?

Medium
103

Which THREE of the following are valid steps to troubleshoot DNS issues in a Kubernetes cluster?

Medium
104

You are trying to debug a network connectivity issue between two pods. Pod A can reach the internet but cannot reach Pod B's IP address. Which command should you use to test connectivity from within Pod A to Pod B's service?

Medium
105

You are debugging a DNS issue from within a pod. The pod is running 'busybox'. Which command would you use to test DNS resolution for 'kubernetes.default.svc.cluster.local'?

Medium
106

Based on the exhibit, the pod is in CrashLoopBackOff. Which command should you run NEXT to identify the root cause?

Medium
107

You are troubleshooting a pod that is in 'Pending' state. 'kubectl describe pod' shows '0/1 nodes are available: 1 Insufficient memory, 1 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate'. Which TWO actions can resolve the issue?

Hard
108

Which TWO of the following are valid reasons a pod might be stuck in 'Pending' state?

Easy
109

A pod is stuck in 'Pending' state. You run 'kubectl describe pod my-pod' and see the event: '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/unreachable: }, that the pod didn't tolerate'. What is the likely cause?

Easy
110

Which command shows resource usage (CPU and memory) for nodes in a cluster?

Easy
111

Which TWO actions can help troubleshoot a service that is not reachable from within a pod?

Medium
112

A pod is in 'ImagePullBackOff' state. Which of the following is NOT a common cause?

Medium
113

A pod is in ImagePullBackOff state. Which command would give you the most information about why the image pull failed?

Easy
114

Based on the exhibit, what is the most likely cause of the pod not running?

Hard
115

You run 'kubectl logs pod-name' and get no output. Which TWO steps should you take to troubleshoot further?

Medium
116

You are troubleshooting a scenario where a pod cannot communicate with another pod in the same namespace via service name. Which THREE steps would you take to diagnose the issue? (Select 3)

Hard
117

A pod is stuck in Pending state. You run 'kubectl describe pod my-pod' and see the event '0/1 nodes are available: 1 Insufficient cpu'. What is the most likely cause?

Easy
118

A pod is in ImagePullBackOff. Which command would help determine the exact reason?

Medium
119

Which TWO commands show cluster events that can help in troubleshooting?

Easy
120

A pod is in CrashLoopBackOff. 'kubectl logs my-pod --previous' shows: 'Error: failed to start: exec: "/app/start.sh": stat /app/start.sh: no such file or directory'. What is the most likely cause?

Hard
121

You suspect the kubelet on a worker node is not functioning correctly. Which command should you use to check the kubelet service status?

Easy
122

You have a pod that is in CrashLoopBackOff. Which two troubleshooting steps should you take first? (Choose two.)

Hard
123

A Deployment named 'web' is failing to schedule pods. You run 'kubectl describe pod web-xyz' and see the event: '0/3 nodes are available: 3 Insufficient cpu.' What is the most likely cause?

Medium
124

Which TWO of the following are common causes for a pod to be in the 'Pending' state?

Medium
125

You see a pod in 'Pending' state. 'kubectl describe pod' shows '0/4 nodes are available: 1 node(s) had taint(s) that the pod didn't tolerate, 3 Insufficient cpu'. What should you do?

Medium
126

Which THREE of the following are common causes for a pod to remain in 'Pending' state?

Hard
127

A pod is in CrashLoopBackOff. You run 'kubectl logs mypod --previous' and see 'Error: unable to connect to database'. What is the MOST likely cause?

Medium
128

A pod is stuck in 'Pending' state. You run 'kubectl describe pod mypod' and see: '0/4 nodes are available: 4 node(s) didn't match pod anti-affinity constraints'. What does this mean?

Medium
129

You want to check the logs of a container that previously crashed. Which command should you use?

Medium
130

A node is NotReady. Which THREE conditions could cause this?

Hard
131

Which THREE of the following are valid steps to troubleshoot a node that is in 'NotReady' state?

Hard
132

A ClusterIP Service is not reachable from within the cluster. You verify that the Service has endpoints. Which of the following could be the cause? (Select two.)

Hard
133

You run 'kubectl get events --sort-by=.lastTimestamp' and see the following events for a pod: 'Warning FailedScheduling 0/3 nodes are available: 3 Insufficient cpu'. What is the most likely solution?

Medium
134

You create a Deployment with the following YAML: apiVersion: apps/v1 kind: Deployment metadata: name: myapp spec: replicas: 3 selector: matchLabels: app: myapp template: metadata: labels: app: myapp spec: containers: - name: myapp image: myapp:latest resources: limits: cpu: "500m" memory: "256Mi" requests: cpu: "200m" memory: "128Mi" After applying it, the pods are in 'CrashLoopBackOff'. You check logs and see 'Error: container process exited with error'. What is the MOST likely cause?

Medium
135

A pod has status 'Init:Error'. What does this indicate?

Hard
136

You run 'kubectl get pods' and one pod shows 'ImagePullBackOff'. Which command would help you diagnose the issue?

Easy
137

Which THREE of the following are valid steps to troubleshoot a DNS issue within a Kubernetes cluster?

Hard
138

You have a pod that is in 'Pending' state. Which command would you run to get more information about why the pod cannot be scheduled?

Easy
139

You have a pod that is stuck in 'Pending' state. Running 'kubectl describe pod' shows the event: '0/3 nodes are available: 1 node(s) had taint {node-role.kubernetes.io/master: }, that the pod didn't tolerate, 2 node(s) didn't match pod anti-affinity rules.' What is the MOST likely solution?

Hard
140

You run 'kubectl logs my-pod -c my-container' and get no output, but you know the container produces logs. What should you do next to see previous container logs?

Medium
141

You have a Pod that is stuck in Pending state. Which command should you use to get detailed information about why the Pod is not running?

Easy
142

A Service of type LoadBalancer is created but the EXTERNAL-IP remains <pending>. The cluster is running on-premises without a cloud load balancer integration. Which of the following is the most likely reason?

Hard
143

You attempt to schedule a pod but it remains 'Pending'. 'kubectl describe pod' shows the event: '0/3 nodes are available: 3 node(s) didn't match node selector.' What is the MOST likely cause?

Medium
144

You are troubleshooting a connectivity issue between two pods in the same cluster. Pod A cannot reach Pod B's IP. Which command would you run from Pod A to test connectivity?

Hard
145

You are troubleshooting a DNS issue. From within a pod, you run 'nslookup kubernetes.default.svc.cluster.local' and get 'connection timed out; no servers could be reached'. What is the most likely cause?

Hard
146

A pod is in the 'Pending' state for a long time. You run 'kubectl describe pod pending-pod' and see the event: '0/4 nodes are available: 1 node(s) had taint {node.kubernetes.io/not-ready: }, 3 node(s) had taint {node-role.kubernetes.io/control-plane: } that the pod didn't tolerate.' What is the MOST likely solution?

Medium
147

A pod is in Pending state. You run 'kubectl describe pod pending-pod' and see an event: '0/3 nodes are available: 3 Insufficient memory'. However, you believe there is enough memory across the cluster. What could be the issue?

Hard
148

Which TWO of the following are correct methods to check the health of the kube-apiserver?

Medium
149

Which TWO of the following are valid commands to view cluster events sorted by timestamp?

Medium
150

A node in your cluster is in the 'NotReady' state. You SSH into the node and run 'systemctl status kubelet' which shows the kubelet is active but not functioning correctly. Which command should you use to get detailed logs to troubleshoot the kubelet?

Medium
151

A node in your cluster shows status 'NotReady'. You have SSH access to the node. What is the first command you should run to check the kubelet status?

Medium
152

You are troubleshooting a service named 'api' in the 'prod' namespace that is not reachable from other pods. You run 'kubectl get endpoints api -n prod' and see that the endpoints list is empty. The service selector is 'app=api', and there are three pods with that label running. Which of the following is the MOST likely cause?

Hard
153

You have a pod named 'web-pod' that is in a CrashLoopBackOff state. To examine the logs from the previous instance of the container, which command should you use?

Easy
154

A user reports that a pod is stuck in 'ContainerCreating' state. Which command would you run first to diagnose the issue?

Easy
155

You want to check the memory usage of nodes in your cluster. Which command should you use?

Easy
156

Which TWO of the following are valid commands to check the status of control plane components?

Easy
157

You deploy a pod with resource limits but no requests. The pod gets OOMKilled. What is the most likely reason?

Medium
158

A Deployment's pod is stuck in Pending state. 'kubectl describe pod' shows Events: '0/4 nodes are available: 1 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate, 3 Insufficient memory'. What is the likely fix?

Hard
159

You run 'kubectl get pods' and see that a pod is in 'CrashLoopBackOff'. You want to examine the container's previous exit code. Which command provides this information?

Medium
160

Which TWO of the following are valid steps to troubleshoot a pod that is in 'CrashLoopBackOff'?

Medium
161

The controller-manager logs show repeated errors: 'Failed to list *v1.Pod: connection refused'. What is the most likely cause?

Medium
162

You run 'kubectl get events --sort-by='.lastTimestamp'' and see repeated events: 'Failed to pull image "myimage:v2": rpc error: code = Unknown desc = Error response from daemon: manifest for myimage:v2 not found'. What is the issue?

Easy
163

Which command shows resource usage for pods and nodes in the cluster?

Easy
164

Drag and drop the steps to back up and restore etcd data for a Kubernetes cluster into the correct order.

Medium
165

You have a Deployment with 3 replicas. After updating the container image, the new pods are in 'ImagePullBackOff' state. You run 'kubectl describe pod <pod-name>' and see the event: 'Failed to pull image "myregistry/myapp:latest": rpc error: code = Unknown desc = Error response from daemon: manifest for myregistry/myapp:latest not found: manifest unknown: manifest unknown'. What is the MOST likely cause?

Medium
166

Which command can be used to view resource usage of nodes in a cluster?

Easy
167

You run 'kubectl get pods -n default' and see a pod named 'backend' in ImagePullBackOff state. What is the most likely cause?

Medium
168

You run 'kubectl top nodes' and get an error: 'error: metrics not available yet'. What is the most likely cause?

Medium
169

A Pod is stuck in CrashLoopBackOff. You run `kubectl logs <pod-name>` but see no output. What is the most likely cause?

Medium
170

You need to check the logs of a kubelet on a node. Which command should you run on the node?

Easy
171

You run 'kubectl get pods' and see some pods in 'ImagePullBackOff' state. Which command would best help identify the root cause?

Medium
172

You want to check the current resource usage (CPU and memory) of pods in the 'default' namespace. Which kubectl command should you use?

Hard
173

Which command retrieves logs from a container that has crashed and restarted?

Easy
174

A pod is in 'CrashLoopBackOff' state. Which THREE of the following are possible causes?

Hard
175

A pod remains in Pending state. You run 'kubectl describe pod mypod' and see the following event: '0/3 nodes are available: 2 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate, 1 node(s) didn't match pod anti-affinity rules.' What is the best action to schedule the pod?

Hard
176

A pod is stuck in Pending state. Running 'kubectl describe pod mypod' reveals the event '0/4 nodes are available: 3 Insufficient memory, 1 node(s) had taints that the pod didn't tolerate'. What is the most likely cause?

Medium
177

A pod is in CrashLoopBackOff. You check the logs with 'kubectl logs my-pod --previous' and see 'Error: cannot connect to database at 10.0.0.1:3306'. The database service is named 'mysql' and runs on port 3306. What is the most likely cause?

Medium
178

To view the logs of a specific container in a multi-container pod named 'web-pod', which command is correct?

Easy
179

You want to view the resource usage of all pods in the cluster. What command should you run?

Easy
180

You are troubleshooting DNS resolution from within a pod. You exec into the pod and run 'nslookup kubernetes.default.svc.cluster.local'. The command fails with 'connection timed out; no servers could be reached'. However, 'kubectl get svc -n kube-system' shows the kube-dns service with a ClusterIP. What is the MOST likely cause?

Medium
181

Which command shows the logs of a pod that has crashed and restarted?

Easy
182

Which command shows resource usage (CPU/memory) of all pods in the default namespace?

Easy
183

A node in your cluster is marked as NotReady. You SSH into the node and run 'systemctl status kubelet'. The output shows the kubelet is inactive (dead). What should you do FIRST to restore the node?

Medium
184

Which command lists all events in the cluster sorted by timestamp?

Easy
185

You run 'kubectl get nodes' and one node shows 'NotReady'. Which command should you run first to check the kubelet status on that node?

Medium
186

A pod is not able to communicate with another pod in the same namespace. Both pods are running and have IP addresses. Which command can you use to test connectivity from the first pod to the second pod's IP?

Hard
187

You run 'kubectl get pods' and see a pod in 'CrashLoopBackOff'. You want to see the logs of the last crashed instance. Which command should you run?

Medium
188

You run 'kubectl get pods' and see a pod in 'ImagePullBackOff' state. Which of the following is NOT a common cause?

Medium
189

You are troubleshooting a pod that is in 'CrashLoopBackOff' state. You run 'kubectl logs mypod' and get no output. You then run 'kubectl logs mypod --previous' and see an error: 'Error: failed to start container: context deadline exceeded'. What is the MOST likely cause?

Hard
190

A pod is unable to resolve DNS names. You exec into the pod and run 'nslookup kubernetes.default.svc.cluster.local'. The command hangs. What is the MOST likely cause?

Medium
191

You run 'kubectl get nodes' and see that a node is 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. The output shows 'Active: inactive (dead)'. What is the most likely cause?

Medium
192

Which command would you use to check the status of the kube-apiserver on a control plane node managed by systemd?

Easy
193

A node in your cluster is reporting a DiskPressure condition. Which kubectl command would you use to get details about the node's condition?

Easy
194

You have a Deployment with livenessProbe configured. The pod restarts every few minutes. 'kubectl describe pod' shows the liveness probe is failing with 'HTTP probe failed with statuscode: 503'. The application's /healthz endpoint returns 200 from within the pod using 'kubectl exec'. What could be the issue?

Hard
195

Which command can you use to view the logs of a container that has crashed and been restarted?

Easy
196

A node is NotReady. You ssh into the node and run 'systemctl status kubelet'. It shows 'Active: inactive (dead)'. What is the most appropriate next step?

Hard
197

You run 'kubectl get nodes' and see that one node is in the 'NotReady' state. Which command would you use FIRST to investigate the kubelet status on that node?

Medium
198

You have a pod that is in 'Pending' state. Which command would you use to view detailed information about the pod's status, including events that may indicate why it is not running?

Easy
199

You are a CKA managing a production cluster with 5 worker nodes. A developer reports that a new deployment 'payment-service' is not accessible from other pods via its Service 'payment-svc' in the 'default' namespace. The Service is of type ClusterIP with selector 'app: payment'. The deployment has 3 replicas, all showing 'Running' status. From a test pod, you run 'curl http://payment-svc:8080' and get 'Connection refused'. You verify that the pods are listening on port 8080 and the container's readiness probe passes. 'kubectl get endpoints payment-svc' shows no endpoints. 'kubectl describe svc payment-svc' shows the selector 'app=payment'. What is the most likely cause?

Hard
200

You run 'kubectl get pods' and see that a pod named 'db' is in 'CrashLoopBackOff'. Which TWO commands are most useful for diagnosing the issue? (Choose two)

Medium
201

A Node is reporting DiskPressure condition. Which action is most appropriate to resolve this without losing data?

Medium
202

You have a Deployment with 3 replicas. One of the pods is in 'Pending' state. 'kubectl describe pod' shows: 'Warning FailedScheduling 0/4 nodes are available: 1 node(s) had taint {key1: value1}, that the pod didn't tolerate, 3 node(s) didn't match pod anti-affinity rules.' Which two issues are preventing the pod from being scheduled?

Medium
203

A pod named 'web' in the 'default' namespace is in Running state, but users report that the application is not responding. You run 'kubectl exec web -- curl -v http://localhost:8080' and see 'Connection refused'. The pod's container is listening on port 8080. Which of the following is the MOST likely cause?

Medium
204

You are troubleshooting a node that shows 'NotReady' status. Which TWO commands can help you investigate the kubelet state?

Easy
205

You are debugging a DNS issue in the cluster. Which of the following tools is commonly used to test DNS resolution from within a Pod?

Medium
206

You deploy a pod with the following YAML: apiVersion: v1 kind: Pod metadata: name: test-pod spec: containers: - name: test image: nginx resources: requests: memory: "64Mi" cpu: "250m" limits: memory: "128Mi" cpu: "500m" The pod starts, but after a few minutes it is killed with OOMKilled. What is the MOST likely reason?

Medium
207

You run 'kubectl get pods' and see a pod in 'CrashLoopBackOff'. What command would you run to see the reason for the crash?

Medium

Frequently asked questions

What does the Troubleshooting domain cover on the CKA exam?
You must systematically isolate faults across pods, Services, and nodes using kubectl describe, logs, events, and node-level systemctl/journalctl checks, then apply the fix. The single most important thing is reading events and status conditions before changing any configuration.
How many questions are in this domain?
This page lists all 207 Troubleshooting questions in the CKA question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Troubleshooting questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
cncf-cka CNCF-CKA cka troubleshooting Practice Questions