Courseiva

CCNA Troubleshooting Questions

75 of 207 questions · Page 2/3 · Troubleshooting · Answers revealed

76
MCQmedium

You run 'kubectl get pods' and see a pod with status 'Init:CrashLoopBackOff'. What does this indicate?

A.An init container in the pod is failing and restarting
B.The pod's init container ran successfully but the main container has not started yet
C.The pod is still initializing but will eventually run
D.The main container is crashing and the pod is restarting
AnswerA

When a pod's status displays Init:CrashLoopBackOff, it indicates that one of its defined init containers has exited with a non-zero status code and is repeatedly failing during startup. Kubernetes will continuously attempt to restart this failing init container before it can proceed to the main application containers, blocking the pod from reaching the Running state.

Why this answer

The status 'Init:CrashLoopBackOff' indicates that an init container within the pod is failing and being repeatedly restarted by Kubernetes. Init containers run sequentially before any main containers start, and if one exits with a non-zero exit code, Kubernetes retries it with an exponential backoff delay, leading to the CrashLoopBackOff state. This is distinct from a main container crash, which would show 'CrashLoopBackOff' without the 'Init:' prefix.

Exam trap

The CKA exam often tests the distinction between init container failures and main container failures by using the 'Init:' prefix in the status, so candidates who overlook this prefix may mistakenly choose the main container crash option.

How to eliminate wrong answers

Option B is wrong because if an init container ran successfully, the pod would proceed to start the main container, not remain in an 'Init:' status; the 'Init:' prefix specifically indicates an init container is still running or failing. Option C is wrong because 'Init:CrashLoopBackOff' is not a transient initialization state—it signals a persistent failure with restarts, not eventual success without intervention. Option D is wrong because a crashing main container would show 'CrashLoopBackOff' (without 'Init:'), not 'Init:CrashLoopBackOff', which explicitly points to an init container issue.

77
MCQmedium

You run kubectl get nodes and see one node is NotReady. The kubelet is running on the node. What is the most likely cause?

A.The kubelet is not installed
B.Network connectivity issue between kubelet and API server
C.The node has been cordoned
D.The API server is down
AnswerB

This is the correct answer. The kubelet is responsible for reporting the node's status to the API server through periodic heartbeat updates. When there is a network connectivity issue between the kubelet and the API server, the API server's node controller does not receive these updates for longer than the node-monitor-grace-period (default 40 seconds). As a result, the node is marked as NotReady, even though the kubelet process itself continues to run locally and can still manage containers on the node.

Why this answer

When a node is NotReady but the kubelet is running, the most common cause is a network connectivity issue between the kubelet and the API server. The kubelet reports node status via periodic heartbeats (node-status-update-frequency, default 10s) and if the API server cannot receive these updates due to network problems (e.g., firewall rules, DNS resolution failure, or dropped packets), the node controller marks the node as NotReady after the node-monitor-grace-period (default 40s). The kubelet being running eliminates installation issues, and the API server being down would affect all nodes, not just one.

Exam trap

The trap here is that candidates confuse 'cordon' (which affects scheduling but not readiness) with 'NotReady' (which indicates a health or connectivity failure), leading them to pick Option C when the node is actually unreachable.

Why the other options are wrong

A

Question states kubelet is running.

C

Cordoning makes node Unschedulable, not NotReady.

D

If API server is down, kubectl get nodes would fail, not show NotReady.

78
MCQmedium

A developer reports that a Pod named 'web-pod' in namespace 'frontend' is crashing repeatedly. You run 'kubectl logs web-pod -n frontend' but see no output. Which command should you run next to see the logs from the previous, crashed container instance?

A.kubectl get events -n frontend --sort-by=.metadata.creationTimestamp
B.kubectl logs web-pod -n frontend --previous
C.kubectl logs web-pod -n frontend -c web-pod
D.kubectl exec -it web-pod -n frontend -- sh
AnswerB

This command is the correct approach because the --previous (or -p) flag instructs the kubelet to retrieve the stdout and stderr logs from the most recently terminated instance of the container. This is essential for diagnosing CrashLoopBackOff states where the current container has restarted and its active log buffer is empty or irrelevant.

Why this answer

The `kubectl logs --previous` flag retrieves logs from the previous instance of a container in a Pod, which is exactly what you need when the current container has crashed and restarted, leaving no logs from the current instance. Since `kubectl logs web-pod -n frontend` returned no output, the current container likely started fresh after a crash, and the logs from the crashed container are stored in the terminated container's log file. This flag accesses those logs without needing to specify a container name explicitly when there is only one container in the Pod.

Exam trap

The trap here is that candidates may think `kubectl logs` without flags is sufficient, or they may confuse `--previous` with `-c` (container name), not realizing that `--previous` is specifically designed to access logs from a terminated container instance, while `-c` only selects a container within a multi-container Pod.

How to eliminate wrong answers

Option A is wrong because `kubectl get events` shows cluster events (e.g., scheduling, pulling images) but does not provide container logs, which are needed to debug the crash. Option C is wrong because `-c web-pod` specifies a container name, but if the Pod has only one container (named 'web-pod'), this command is redundant and still fetches logs from the current (possibly empty) container, not the previous crashed instance. Option D is wrong because `kubectl exec` opens an interactive shell into the running container, but if the container is crashing repeatedly, it may not be running, and even if it were, this would not retrieve logs from the previous terminated instance.

79
MCQhard

A pod is in CrashLoopBackOff. The YAML for the initContainer is: apiVersion: v1 kind: Pod metadata: name: myapp spec: initContainers: - name: init image: busybox command: ['sh', '-c', 'sleep 5 && exit 1'] containers: - name: app image: nginx What is the most likely reason for the CrashLoopBackOff?

A.The main container image nginx is not pulled successfully
B.The init container command is misspelled
C.The init container exits with non-zero exit code
D.The pod has insufficient memory
AnswerC

The init container's role is to run to completion before the main container starts; if it fails, the pod cannot proceed. When an init container exits with a non-zero exit code, the kubelet restarts the pod according to the restartPolicy, causing the pod to enter CrashLoopBackOff after repeated failures. This is the standard mechanism for Init:CrashLoopBackOff, so a non-zero exit code is the correct explanation.

Why this answer

Init containers must complete successfully (exit 0) before the main container starts. An exit code 1 causes the pod to fail and restart.

80
MCQmedium

A node in the cluster is showing status 'NotReady'. You run 'kubectl describe node worker1' and see that the kubelet has not posted node status for more than 1 minute. Which command should you run on the node to check the kubelet logs?

A.cat /var/log/kubelet/kubelet.log
B.kubectl logs kubelet -n kube-system
C.systemctl restart kubelet
D.journalctl -u kubelet
AnswerD

journalctl -u kubelet is the correct command on systemd-based systems because the kubelet runs as a systemd service and writes its logs to the systemd journal. This command filters the journal to show only kubelet entries, revealing errors such as failed API server connectivity, CNI plugin failures, or node lease renewal problems. These logs are essential for diagnosing why the node is NotReady, as they contain the actual failure messages that file-based logs or kubectl cannot provide.

Why this answer

On systemd-based systems, 'journalctl -u kubelet' is the standard command to view kubelet logs. 'journalctl -f -u kubelet' follows the log, but the question asks for checking logs, not following.

81
MCQmedium

You suspect the kubelet on a worker node has stopped. Which two commands should you run to confirm the kubelet status and check its logs?

A.systemctl status docker and tail -f /var/log/syslog
B.kubectl get nodes and kubectl describe node <node>
C.systemctl restart kubelet and journalctl -u containerd
D.systemctl status kubelet and journalctl -u kubelet
AnswerD

This is the standard diagnostic pair for a systemd-managed kubelet. `systemctl status kubelet` shows whether the unit is active (running), failed, or inactive, and gives the last few log lines plus the main process PID; it also returns a non-zero exit code when the service is not running, which is useful for scripting. `journalctl -u kubelet` retrieves the full journal for the kubelet unit, allowing you to see startup errors, crashed state, or repeated restarts—critical for identifying why the kubelet stopped, such as a bad kubelet config, certificate expiry, or a missing CSI driver.

Why this answer

To check if the kubelet is running, use 'systemctl status kubelet'. To view logs, use 'journalctl -u kubelet'. The other options target the wrong service or are not applicable.

82
MCQeasy

You need to check the current resource usage of nodes in your cluster. Which command should you use?

A.kubectl top pods
B.kubectl get events
C.kubectl get nodes -o wide
D.kubectl top nodes
AnswerD

kubectl top nodes is the correct command because it queries the metrics.k8s.io API provided by metrics-server to return each node’s current total CPU and memory usage, along with the percentage relative to allocatable capacity. It aggregates pod usage plus node-level system reservations from cAdvisor and presents a concise per-node snapshot, making it the standard built-in way to assess current node resource consumption.

Why this answer

`kubectl top nodes` retrieves and displays real-time CPU and memory usage metrics for all nodes in the cluster, directly answering the question about current resource usage. This command relies on the metrics server being deployed in the cluster to collect resource utilization data from kubelets via the Summary API.

Exam trap

The trap here is that candidates confuse `kubectl top nodes` with `kubectl get nodes -o wide`, mistakenly thinking the latter shows resource usage when it only shows network and OS details, not utilization metrics.

How to eliminate wrong answers

Option A is wrong because `kubectl top pods` shows resource usage for pods, not nodes, so it does not meet the requirement to check node-level resource usage. Option B is wrong because `kubectl get events` lists cluster events (e.g., scheduling failures, pod lifecycle changes) and does not provide any resource utilization metrics. Option C is wrong because `kubectl get nodes -o wide` displays node metadata such as internal IP, external IP, and OS image, but not real-time CPU or memory usage.

83
MCQeasy

Which command can you run to see the events related to a specific pod?

A.kubectl logs pod-name
B.kubectl get pod pod-name
C.kubectl get events
D.kubectl describe pod pod-name
AnswerD

kubectl describe pod pod-name retrieves the Pod's full configuration along with a dedicated 'Events' section that records timestamped, sequential notifications from the kubelet and controller-manager about that Pod. These events describe actions like scheduling decisions, container creation, image pulling, probe failures, and restarts, which are exactly what you need when debugging why a Pod is stuck or repeatedly crashing. The describe command aggregates just the events for that specific Pod, making it the direct answer to the question.

Why this answer

`kubectl describe pod pod-name` includes a dedicated 'Events' section that lists all lifecycle events for that specific pod, such as scheduling, container pulls, and restarts. This command filters events to only those relevant to the pod, making it the most direct way to view pod-specific events without needing to parse all cluster events.

Exam trap

The trap here is that candidates often confuse `kubectl logs` (application output) with `kubectl describe` (cluster events), or assume `kubectl get events` is the only way to view events, missing that `kubectl describe` automatically filters events for the specified resource.

How to eliminate wrong answers

Option A is wrong because `kubectl logs pod-name` retrieves the container's stdout/stderr logs, not Kubernetes events; logs show application output, not cluster-level scheduling or lifecycle events. Option B is wrong because `kubectl get pod pod-name` only displays the pod's current status and metadata in a summary table, omitting the detailed event history. Option C is wrong because `kubectl get events` lists all events across the entire namespace or cluster, requiring manual filtering to find those related to a specific pod, which is less efficient and not targeted.

84
Multi-Selectmedium

You need to check the status of control plane components. Which TWO commands are appropriate?

Select 2 answers
A.kubectl get pods -n kube-system
B.systemctl status kube-apiserver
C.kubectl get componentstatuses
D.top -u kube
E.systemctl list-units --type=service
AnswersA, C

Control plane components run as static pods in the kube-system namespace, so listing pods there reveals their phase, restart counts and readiness. This works on clusters where the componentstatuses API is deprecated or disabled, making it the reliable fallback for verifying scheduler, controller-manager and etcd health.

Why this answer

To check the status of control plane components in a kubeadm-established cluster, use 'kubectl get pods -n kube-system' to inspect the static pods. 'kubectl get componentstatuses' (deprecated but still a valid status check) reports health of the control plane components. 'systemctl status kube-apiserver' is not appropriate because the API server runs as a static pod, not a systemd service.

Exam trap

The CKA exam environment is built using kubeadm. Do not look for systemd services for the apiserver, controller-manager, or scheduler, as they run as static pods. Only the kubelet and the container runtime (e.g., containerd) run as systemd services on the nodes.

85
MCQmedium

A node shows status NotReady. You SSH into the node and run 'systemctl status kubelet' which shows the kubelet is active (running). What is the next most likely step to diagnose the issue?

A.Restart the kubelet with systemctl restart kubelet
B.Check the container runtime status
C.Check kubelet logs with journalctl -u kubelet
D.Reboot the node
AnswerC

Checking kubelet logs with `journalctl -u kubelet` is the correct first diagnostic step because it provides the authoritative record of why the kubelet is reporting NotReady. The logs will show concrete errors such as failed CNI plugin execution, repeated API server connection timeouts, disk or memory pressure conditions (`eviction signals`), or Webhook errors, all of which directly explain the node's status. This is the most efficient way to pinpoint the failure rather than guessing, and it also gives you the timestamps to correlate with events from the control plane. Only after analyzing these logs should you decide whether a restart or other intervention is necessary.

Why this answer

Even if the kubelet is running, it may be unhealthy. Checking the kubelet logs with 'journalctl -u kubelet' can reveal errors such as network plugin failures or node pressure.

86
MCQmedium

You run 'kubectl get pods' and see a pod with status 'CrashLoopBackOff'. You check the logs with 'kubectl logs <pod> --previous' and see: 'Error: unable to connect to database at db-svc:5432 (connection refused)'. What is the most likely cause?

A.The pod's liveness probe is misconfigured
B.The pod's container image is missing
C.The database service is not running or is unreachable
D.The pod has a memory limit that is too low
AnswerC

A connection refused error—specifically ECONNREFUSED—indicates that the application's TCP handshake reached the target host but nothing was listening on that port, or the service endpoints are empty because the backing database pods are not ready. This commonly occurs when the database Deployment has zero ready replicas, the Service selector does not match any pods, or the pod is using an incorrect service name or port. The container's main process exits after failing to initialize its database connection, and the kubelet restarts it, cycling into CrashLoopBackOff.

Why this answer

The error message 'connection refused' indicates that the pod is attempting to connect to the database at 'db-svc:5432' but the target service is not accepting TCP connections on port 5432. This typically means the database pod or service is not running, or a network policy is blocking the connection. The 'CrashLoopBackOff' status confirms the application container repeatedly fails due to this startup dependency.

Exam trap

The CKA exam often tests the distinction between application-level errors (like 'connection refused') and infrastructure-level errors (like OOM or image pull failures), so candidates must read the exact error message in the logs rather than assuming a generic pod failure cause.

How to eliminate wrong answers

Option A is wrong because a misconfigured liveness probe would cause the pod to be restarted after it had started, not produce a 'connection refused' error in the application logs; liveness probes check container health after startup, not database connectivity. Option B is wrong because a missing container image would result in an 'ImagePullBackOff' or 'ErrImagePull' status, not a 'CrashLoopBackOff' with a database connection error in the logs. Option D is wrong because a memory limit that is too low would cause an 'OOMKilled' status or 'OutOfMemory' error in the logs, not a TCP connection refused error.

87
MCQmedium

A Pod is in CrashLoopBackOff. You run 'kubectl describe pod' and see that the container fails with 'Error: container command not found'. What is the most likely cause?

A.The container is running out of memory
B.The image pull is failing
C.The container image does not contain the specified command
D.The container command is not in the PATH
AnswerC

If the `command` or `args` defined in the Pod specification, or the `ENTRYPOINT` instruction within the container image's Dockerfile, refers to an executable that simply does not exist within the image's filesystem, the container runtime cannot execute it. This fundamental failure causes the container to exit immediately with a non-zero status. Kubernetes then detects this repeated failure and places the pod into a `CrashLoopBackOff` state, with `kubectl describe pod` typically showing an error like "executable file not found" or "no such file or directory" in the container's last termination message.

Why this answer

The error indicating a command is not found (typically presented as 'executable file not found in $PATH' in the container runtime events) most commonly occurs because the specified command or binary does not exist inside the container image at all (for example, trying to run 'bash' or 'curl' in a minimal distroless or scratch-based image, or due to a typo in the command name).

Exam trap

Do not confuse this with a system-level PATH misconfiguration on the Kubernetes host. The PATH being referred to is internal to the container image itself. If a command is not absolute (e.g., 'my-script' instead of '/usr/local/bin/my-script'), the container runtime will search the directories listed in the container's internal PATH variable.

If the binary is missing from the image entirely, it will fail with this error.

How to eliminate wrong answers

Option A is wrong because running out of memory typically causes an OOMKill, which appears as 'Exit Code 137' or 'Error: OOMKilled' in `kubectl describe pod`, not 'container command not found'. Option B is wrong because image pull failures result in events like 'ErrImagePull' or 'ImagePullBackOff', not a container exit error; the container never starts in that case. Option D is wrong because the 'command not found' error occurs when the command itself is missing from the image, not when it's simply not in the PATH; the container runtime (e.g., containerd) uses an absolute path or searches the image's default PATH, but if the binary is absent entirely, it fails regardless of PATH settings.

88
MCQeasy

Which command shows events sorted by timestamp for troubleshooting recent issues?

A.kubectl logs --events
B.kubectl get events --sort-by=.lastTimestamp
C.kubectl describe events
D.kubectl get events
AnswerB

kubectl get events --sort-by=.lastTimestamp is correct because it lists cluster Event resources and applies a JSONPath sort on the lastTimestamp field, which records the most recent time the event was observed. This produces a clean chronological ordering of events—oldest to newest—making it easier to correlate system activity and identify the root cause sequence during troubleshooting.

Why this answer

The correct option is B, `kubectl get events --sort-by=.lastTimestamp`, because `kubectl get events` lists cluster events and the `--sort-by=.lastTimestamp` flag sorts them by the time each event last occurred, making it easy to spot the most recent events for troubleshooting. The `.lastTimestamp` field is the exact sortable field on Event objects that reflects when the event was last observed. Option A is invalid because `kubectl logs` retrieves container logs and has no `--events` flag.

Option C is invalid because `kubectl describe` operates on a specific resource and does not accept `events` as a resource type in that form. Option D is incomplete: `kubectl get events` lists events but does not sort them by timestamp, so recent issues may not appear first.

89
MCQmedium

After deploying a new Deployment, you notice that the pods are stuck in ImagePullBackOff. What is the most common cause?

A.The liveness probe is misconfigured
B.The node has insufficient resources
C.The container image name or tag is incorrect
D.The container command fails on startup
AnswerC

Providing an invalid image name or an unavailable tag causes the container registry to return a 404 error to the kubelet. Consequently, the pod transitions into `ErrImagePull` and then `ImagePullBackOff` because the container runtime cannot locate or download the specified image layers.

Why this answer

The ImagePullBackOff status indicates that the kubelet is unable to pull the container image from the registry. The most common cause is an incorrect image name or tag, which results in a manifest not found error. This triggers an exponential backoff retry loop, leading to the ImagePullBackOff state.

Exam trap

The trap here is that candidates confuse ImagePullBackOff with CrashLoopBackOff, but ImagePullBackOff specifically relates to image retrieval failures, not container runtime errors.

How to eliminate wrong answers

Option A is wrong because a misconfigured liveness probe causes the container to be restarted or killed (CrashLoopBackOff), not an image pull failure. Option B is wrong because insufficient node resources result in a PodPending state with events like 'FailedScheduling' or 'OutOfMemory', not ImagePullBackOff. Option D is wrong because a container command that fails on startup leads to a CrashLoopBackOff state, as the container exits immediately after starting, not an image pull issue.

90
MCQeasy

You have a pod that is in 'Pending' state because it requires a PersistentVolumeClaim that is not bound. Which event would you see in 'kubectl describe pod'?

A.0/1 nodes are available: 1 node(s) didn't match node selector
B.Failed to pull image "myimage:latest"
C.0/1 nodes are available: 1 Insufficient memory
D.0/1 nodes are available: 1 pod has unbound immediate PersistentVolumeClaims
AnswerD

This is the exact event message emitted by the kube-scheduler when a Pod references a PersistentVolumeClaim that uses the Immediate volume binding mode but has not yet been bound to a PersistentVolume. The scheduler cannot place the Pod on a node until the underlying storage is successfully provisioned and bound to satisfy the Pod's volume requirements.

Why this answer

When a pod is in 'Pending' state due to an unbound PersistentVolumeClaim (PVC), the scheduler cannot place the pod until the PVC is bound to a PersistentVolume (PV). The event message '0/1 nodes are available: 1 pod has unbound immediate PersistentVolumeClaims' is generated by the Kubernetes scheduler when it evaluates the pod's PVC requirement and finds no matching PV that satisfies the claim's storage class, access modes, and size, causing the pod to remain unscheduled.

Exam trap

The trap here is that candidates confuse 'Pending' state causes—such as resource shortages or node selector mismatches—with the specific PVC binding failure, which produces a unique scheduler event message that is explicitly listed in 'kubectl describe pod' output.

How to eliminate wrong answers

Option A is wrong because '0/1 nodes are available: 1 node(s) didn't match node selector' indicates a node selector mismatch (e.g., nodeSelector or nodeAffinity), not a PVC binding issue. Option B is wrong because 'Failed to pull image' is an image pull error that occurs after scheduling, typically resulting in a 'ImagePullBackOff' or 'ErrImagePull' status, not a 'Pending' state caused by an unbound PVC. Option C is wrong because '0/1 nodes are available: 1 Insufficient memory' is a resource shortage error (memory pressure) that prevents scheduling, but it does not relate to PVC binding; the scheduler would report a different message for unbound claims.

91
MCQmedium

A pod is in Pending state. You see the event: '0/2 nodes are available: 2 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate'. What should you do to schedule the pod on one of the control-plane nodes?

A.Increase the pod's resource requests
B.Remove the taint from the control-plane node
C.Use a different namespace
D.Add a toleration to the pod spec matching the taint
AnswerD

Adding a toleration to the pod spec that matches the node's taint is the correct solution because tolerations explicitly opt a pod into scheduling on tainted nodes. The taint on the node uses key, value, and effect (e.g., node-role.kubernetes.io/control-plane:NoSchedule), and the toleration must mirror that key, value, and effect before the scheduler will place the pod there. This is the standard, least-privilege way to run a specific workload on a dedicated or control-plane node without weakening cluster-wide policies.

Why this answer

The pod is in Pending state because the control-plane nodes have a taint (node-role.kubernetes.io/control-plane) that the pod does not tolerate. By default, pods are not scheduled on control-plane nodes unless they explicitly tolerate that taint. Adding a toleration to the pod spec that matches the taint's key, effect, and optionally value allows the scheduler to place the pod on a control-plane node.

Exam trap

The trap here is that candidates may think removing the taint (Option B) is the correct fix, but the CKA exam expects you to use tolerations to selectively schedule pods on tainted nodes without altering node configuration.

How to eliminate wrong answers

Option A is wrong because increasing resource requests does not address taints or tolerations; it may even make scheduling harder by requiring more resources. Option B is wrong because removing the taint from the control-plane node would allow all pods to schedule there, which is not the intended solution for a specific pod and could compromise node isolation. Option C is wrong because namespaces are a logical isolation boundary and have no effect on taint/toleration mechanics or scheduling decisions.

92
MCQmedium

A pod named 'web-app' is crashing repeatedly. You run 'kubectl describe pod web-app' and see that the container exited with code 137. What does this indicate?

A.The container's entrypoint command failed
B.The container's readiness probe failed
C.The container image was not found
D.The container was killed because it exceeded its memory limit
AnswerD

Exit code 137 is calculated as 128 plus the signal number 9 (SIGKILL), which indicates that the operating system or container runtime forcefully terminated the process. In Kubernetes, this most commonly occurs when a container exceeds its configured memory limit, triggering the Linux Out-Of-Memory (OOM) killer to terminate the process to safeguard node stability.

Why this answer

Exit code 137 (128 + 9) means the container was killed by SIGKILL (signal 9). In Kubernetes, this typically occurs when the container exceeds its memory limit (specified in `resources.limits.memory`), causing the OOM (Out-Of-Memory) killer to terminate the process. The `kubectl describe pod` output will also show `OOMKilled` in the `State` field under container status.

Exam trap

The trap here is that candidates confuse exit code 137 with a generic application crash (exit code 1) or a probe failure, but 137 specifically indicates an OOM kill due to memory limit violation.

How to eliminate wrong answers

Option A is wrong because an entrypoint command failure would produce a non-zero exit code like 1 or 127, not 137, which is specifically tied to signal termination. Option B is wrong because a readiness probe failure does not cause the container to exit; it only removes the pod from service endpoints while the container continues running. Option C is wrong because an image-not-found error results in `ErrImagePull` or `ImagePullBackOff` status, not a container exit code.

93
MCQmedium

A pod is in ImagePullBackOff state. Which command can you run to get more details about the underlying error?

A.kubectl logs pod
B.kubectl get events --field-selector involvedObject.name=pod
C.kubectl describe pod
D.kubectl top pod
AnswerC

The kubectl describe pod command retrieves the pod's events and status conditions, which include the specific image pull failure reason and message. This satisfies the stem's need for underlying error detail that kubectl get pod alone does not expose.

Why this answer

The `kubectl describe pod` command provides detailed information about the pod, including its status, conditions, events, and container states. For an `ImagePullBackOff` error, the output will include the exact error message from the container runtime (e.g., 'Failed to pull image', 'manifest not found', or 'unauthorized'), which is essential for diagnosing the root cause.

Exam trap

The trap here is that candidates often confuse `kubectl logs` (which shows application output) with `kubectl describe` (which shows pod lifecycle events and container runtime errors), leading them to choose A when the container never started to produce logs.

How to eliminate wrong answers

Option A is wrong because `kubectl logs pod` retrieves container logs, which are generated by the application inside the container; if the container never started due to ImagePullBackOff, there are no logs to fetch. Option B is wrong because `kubectl get events` with a field selector filters events by the pod's name, but the output may not include the detailed pull error from the kubelet or container runtime; `kubectl describe pod` consolidates those events alongside other critical status fields. Option D is wrong because `kubectl top pod` shows resource usage (CPU/memory) of running pods, which is irrelevant when the pod is in a non-running state like ImagePullBackOff.

94
Multi-Selectmedium

A pod is in 'Pending' state. Which TWO of the following are possible causes? (Select 2)

Select 2 answers
A.Node has insufficient CPU or memory resources
B.Container exited with non-zero exit code
C.PersistentVolumeClaim is not bound
D.Container was killed due to OOM
E.Image name is misspelled
AnswersA, C

A Pod can only stay Pending when the scheduler is unable to place it on a node. If every node lacks sufficient allocatable CPU and/or memory to satisfy the Pod's `resources.requests`, kube-scheduler marks the Pod unschedulable and leaves it in Pending while continuously retrying scheduling. This is purely a pre-scheduling condition, so it remains until a node is scaled up or the requests are reduced.

Why this answer

Option A is correct because when a node lacks sufficient allocatable CPU or memory, the kube-scheduler cannot find a feasible node and leaves the pod in Pending with a FailedScheduling event. Option C is correct because a pod that references a PersistentVolumeClaim that is not yet Bound (for example, waiting on a dynamic provisioner or a matching PersistentVolume) will remain Pending until the PVC binds. Option B is incorrect because a non-zero container exit code indicates a runtime failure after the container started, producing CrashLoopBackOff or Error, not Pending.

Option D is incorrect because OOMKilled occurs when a running container exceeds its memory limit, resulting in a terminated container rather than a Pending pod. Option E is incorrect because a misspelled image name causes an ImagePullBackOff/ErrImagePull status once the kubelet attempts to pull the image, not a Pending phase.

Exam trap

CKA often tests the distinction between Pending (scheduling phase) and other failure states like CrashLoopBackOff or ImagePullBackOff, causing candidates to select runtime errors as causes of Pending.

95
MCQmedium

A pod has been in Pending state for a long time. 'kubectl describe pod' shows the event: '0/3 nodes are available: 1 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate, 2 node(s) had taint {node.kubernetes.io/unreachable: }, that the pod didn't tolerate.' What is the most likely cause?

A.The pod's image is incorrect
B.The kubelet on each node is not running
C.The pod has resource requests that exceed node capacity
D.The nodes are all cordoned
AnswerB

When kubelet is not running on a node, the node's heartbeat to the control plane is absent; after the node-monitor-grace-period, the Node controller marks the node NotReady and applies the node.kubernetes.io/not-ready:NoSchedule taint. Since no nodes are schedulable, the kube-scheduler cannot find a match for the pod, leaving it in Pending indefinitely. This is often the systemic cause when all nodes are unreachable or their kubelets have crashed.

Why this answer

The taints `node.kubernetes.io/not-ready` and `node.kubernetes.io/unreachable` are automatically added by the node controller when a node's kubelet stops reporting its status (the `node-monitor-grace-period`, default 40s, is exceeded). Since all three nodes exhibit these taints, the kubelet is not running on any of them, preventing the node from being marked `Ready` and causing the scheduler to find no suitable node for the pod.

Exam trap

A common trap is confusing taints added automatically by the node controller (like `node.kubernetes.io/not-ready` and `node.kubernetes.io/unreachable`) with taints added manually by an administrator (like `node.kubernetes.io/unschedulable` from `kubectl cordon`). In this scenario, the presence of these automatic taints on all nodes indicates that the kubelet is not running, not that nodes are cordoned.

How to eliminate wrong answers

Option A is wrong because an incorrect image would cause a `ErrImagePull` or `ImagePullBackOff` event, not a `Pending` state with taint-based scheduling failures. Option C is wrong because resource requests exceeding node capacity would produce events like `Insufficient cpu` or `Insufficient memory`, not taints related to node readiness or reachability. Option D is wrong because cordoned nodes have the `node.kubernetes.io/unschedulable:NoSchedule` taint (added by `kubectl cordon`), not the `not-ready` or `unreachable` taints; additionally, cordoning does not affect all nodes simultaneously unless explicitly done.

96
Multi-Selectmedium

Which TWO of the following commands are useful for debugging network connectivity between pods?

Select 2 answers
A.kubectl top pod <pod-name>
B.kubectl run test-pod --image=busybox --rm -it -- wget -O- http://service:port
C.kubectl edit deployment <deployment-name>
D.kubectl logs <pod-name>
E.kubectl exec <pod-name> -- ping <target-ip>
AnswersB, E

kubectl run test-pod --image=busybox --rm -it -- wget -O- http://service:port launches a temporary, interactive busybox Pod that performs an HTTP request against the target Service from within the cluster network. The --rm flag ensures the Pod is deleted after the command finishes, and -it lets you see the output immediately. This is a classic debugging pattern for validating DNS resolution, Service routing, and application response without altering any existing workloads.

Why this answer

Option B is correct because running an ephemeral busybox pod with `kubectl run test-pod --image=busybox --rm -it -- wget -O- http://service:port` lets you test DNS resolution and HTTP connectivity from a separate pod to a target service, which is a standard way to isolate whether a connectivity problem is pod-specific or service-wide. Option E is correct because `kubectl exec <pod-name> -- ping <target-ip>` executes a network utility inside an existing pod, directly verifying ICMP reachability and basic IP-level connectivity between that pod and the target. Option A is not useful here because `kubectl top pod` only reports CPU and memory usage metrics, not network reachability.

Option C is not useful because `kubectl edit deployment` modifies the deployment manifest and does not test connectivity. Option D is not useful because `kubectl logs` only retrieves container stdout/stderr output and does not actively probe network paths.

Exam trap

The trap here is that candidates confuse resource monitoring (`kubectl top`) or log inspection (`kubectl logs`) with active network probing, and may overlook that `ping` uses ICMP which is often filtered, while `wget` uses TCP which is more reliable for connectivity tests.

97
MCQmedium

You suspect a DNS issue inside a pod. Which command can you run to test DNS resolution from within a pod?

A.kubectl logs coredns -n kube-system
B.kubectl describe svc kubernetes
C.kubectl run test --image=busybox -- nslookup kubernetes.default
D.kubectl exec <pod-name> -- nslookup kubernetes.default
AnswerD

`kubectl exec <pod-name> -- nslookup kubernetes.default` runs the `nslookup` binary directly inside the target pod's network namespace. This uses the pod's own `/etc/resolv.conf`, including its `nameserver` (typically the kube-dns ClusterIP) and search domains (such as `default.svc.cluster.local`), to perform a real DNS query. It is the most direct way to verify that the pod can resolve a service name, because it replicates exactly what an application in that pod would experience.

Why this answer

The correct command is `kubectl exec <pod-name> -- nslookup kubernetes.default`, which runs the DNS lookup from inside the existing pod's network namespace, using the pod's own /etc/resolv.conf and DNS policy. This directly tests whether the pod can resolve cluster DNS names through kube-dns/CoreDNS, which is exactly what a DNS issue inside a pod requires. Running nslookup from within the affected pod isolates the problem to that pod's DNS configuration rather than the cluster DNS service as a whole.

Exam trap

CKA often tests whether candidates confuse observing DNS components (logs, Service descriptions) with actually exercising name resolution from the affected pod's network namespace.

How to eliminate wrong answers

Option A is wrong because `kubectl logs coredns -n kube-system` only shows CoreDNS server logs; it does not test resolution from the pod's perspective and may show no errors even when the pod's resolv.conf is misconfigured. Option B is wrong because `kubectl describe svc kubernetes` only displays the Service object's endpoints and metadata; it does not perform any DNS query or validate name resolution. Option C is wrong because `kubectl run test --image=busybox -- nslookup kubernetes.default` creates a brand-new pod with its own default DNS configuration, so it tests cluster DNS generally but not the suspect pod's DNS behavior.

98
MCQhard

A pod is running but cannot be accessed via its ClusterIP service from another pod in the same namespace. The service endpoints list shows the pod's IP. What is the most likely cause?

A.The kube-proxy is not running on the node
B.A NetworkPolicy is blocking the traffic
C.The service's targetPort is incorrect
D.The pod is running on a different node without proper routing
AnswerB

NetworkPolicy is a namespace-scoped firewall that can restrict egress traffic from a specific pod (via podSelector) to a destination service's backing pod IP or CIDR. Even if the Service object and endpoints are intact, a NetworkPolicy denying egress from the source pod to the backend pod's IP or port will silently drop the packets, making the Service unreachable only for the affected pod(s).

Why this answer

A NetworkPolicy can explicitly deny ingress traffic to a pod even when the service endpoints are correctly populated. Since the endpoints list shows the pod's IP, the service and pod are communicating at the network layer, but a NetworkPolicy with an ingress rule that does not allow traffic from the source pod's labels or CIDR will cause the packet to be dropped by the node's iptables or eBPF rules, resulting in a connection timeout or reset from the client pod.

Exam trap

The trap here is that candidates assume a populated endpoints list guarantees connectivity, but they overlook that NetworkPolicies operate at a lower layer (L3/L4) and can block traffic even when the service and pod are correctly configured.

Why the other options are wrong

A

kube-proxy issues would affect all services cluster-wide, not just one service with correct endpoints.

C

If targetPort were wrong, endpoints might still show but traffic would not reach the container; but endpoints are based on the container port, so if endpoints exist, targetPort matches.

D

ClusterIP services work across nodes; no extra routing needed.

99
MCQhard

A pod is stuck in Pending state. 'kubectl describe pod' shows the event: '0/3 nodes are available: 3 node(s) didn't match pod anti-affinity rules'. What is the most likely cause?

A.The nodes have insufficient resources
B.The pod has a requiredDuringSchedulingIgnoredDuringExecution anti-affinity rule that is too restrictive
C.The pod has a taint tolerance issue
D.The nodes are all cordoned
AnswerB

A requiredDuringSchedulingIgnoredDuringExecution anti-affinity rule is a hard constraint: the scheduler will only place the pod on a node that satisfies every term of the rule. If the rule's label selector and topologyKey match labels on pods running on every available node, no node passes the check. The resulting event is '0/3 nodes are available: 3 node(s) didn't match pod anti-affinity rules,' and the pod stays Pending until a node no longer runs a conflicting pod or the rule is updated. This is the only option where the pod's own scheduling constraints, not cluster conditions, make all nodes ineligible.

Why this answer

The event '0/3 nodes are available: 3 node(s) didn't match pod anti-affinity rules' directly indicates that the pod's scheduling is being blocked by anti-affinity constraints. Option B is correct because a `requiredDuringSchedulingIgnoredDuringExecution` anti-affinity rule is a hard constraint that must be satisfied at scheduling time; if no node meets the rule (e.g., the rule prevents co-location with other pods that are present on all nodes), the pod remains Pending.

Exam trap

CNCF often tests the distinction between hard and soft scheduling constraints; the trap here is that candidates may confuse anti-affinity errors with resource insufficiency or taint issues, but the specific event message directly points to anti-affinity rules.

How to eliminate wrong answers

Option A is wrong because insufficient resources would produce events like 'Insufficient cpu' or 'Insufficient memory', not a message about anti-affinity rules. Option C is wrong because taint/toleration issues generate events such as 'node(s) had taints that the pod didn't tolerate', not anti-affinity mismatches. Option D is wrong because cordoned nodes produce events like 'node(s) were cordoned' or 'node(s) were unschedulable', not a message about pod anti-affinity rules.

100
MCQmedium

A pod is stuck in 'Pending' state. Which command would you run FIRST to diagnose the issue?

A.kubectl logs <pod-name>
B.kubectl describe pod <pod-name>
C.kubectl top pod <pod-name>
D.kubectl exec -it <pod-name> -- sh
AnswerB

kubectl describe pod <pod-name> is the correct diagnostic command because it aggregates the pod's object metadata, current status, conditions, and most importantly, recent Events from the API server, scheduler, and kubelet. For a Pending pod, the Events section reveals whether the scheduler failed due to insufficient resources, taints/tolerations, node selector mismatches, or whether a PersistentVolume claim is awaiting binding — the precise reason for the stuck state.

Why this answer

A pod stuck in 'Pending' state means it has not been scheduled to a node yet. The `kubectl describe pod` command provides detailed event logs, scheduler decisions, and resource constraints (e.g., insufficient CPU/memory, persistent volume claims not bound, node selector mismatches) that reveal why scheduling failed. This is the first diagnostic step because it surfaces the root cause without requiring the pod to be running.

Exam trap

The trap here is that candidates often jump to `kubectl logs` or `kubectl exec` out of habit, forgetting that these commands only work for running pods, while 'Pending' indicates a pre-scheduling failure that requires inspecting events and conditions via `kubectl describe`.

How to eliminate wrong answers

Option A is wrong because `kubectl logs` retrieves container logs, but a pod in 'Pending' has no running containers yet, so there are no logs to fetch. Option C is wrong because `kubectl top pod` shows real-time resource usage metrics, which require the pod to be running on a node; a pending pod has no metrics. Option D is wrong because `kubectl exec` requires a running container to execute commands, which is impossible when the pod is still pending.

101
MCQmedium

You deploy a pod with image 'nginx:1.21'. It stays in ImagePullBackOff. You run 'kubectl describe pod nginx-pod' and see the event: 'Failed to pull image "nginx:1.21": rpc error: code = Unknown desc = Error response from daemon: manifest for nginx:1.21 not found'. What is the most likely fix?

A.Use a different container runtime
B.Add imagePullSecrets to the pod
C.Restart the kubelet on the node
D.Change the image tag to a valid one, e.g., nginx:1.21.6
AnswerD

The tag `nginx:1.21` does not exist in the Docker Hub repository; valid tags for the 1.21 series include specific patch versions like `1.21.6`. Changing the image reference to a known valid tag allows the runtime to pull the correct manifest and start the container, resolving the ImagePullBackOff. Always verify available tags in the registry when encountering pull errors.

Why this answer

The tag '1.21' does not exist in the registry. Use a valid tag like '1.21.6' or 'latest'.

102
MCQmedium

A Node is in NotReady state. Which action should be taken first to diagnose the issue?

A.kubectl describe node <node>
B.Check kubelet logs on the node
C.Restart kubelet
D.Check API server logs
AnswerA

kubectl describe node <node> is the correct first diagnostic because it displays the Node object's Status.Conditions (Ready, MemoryPressure, DiskPressure, PIDPressure, NetworkUnavailable) and a rolling list of recent Events that record taints, kubelet restarts, and CNI failures. This single command aggregates the node's current condition, capacity, allocatable resources, and the kubelet's reported heartbeats from the control plane, letting you immediately see which signal flipped the node to NotReady without requiring SSH access to the node.

Why this answer

When a node is in NotReady state, the first diagnostic step is to gather information about the node's current status, conditions, and recent events using `kubectl describe node <node>`. This command reveals the node's conditions (e.g., Ready, DiskPressure, MemoryPressure), the last heartbeat timestamp, and any relevant events that may indicate the root cause, such as network issues or kubelet failures. It provides a high-level overview without requiring direct node access, making it the most efficient initial action.

Exam trap

The trap here is that candidates often jump to checking kubelet logs or restarting the kubelet, forgetting that `kubectl describe node` provides immediate visibility into node conditions and events from the control plane, which is the standard first step in the Kubernetes troubleshooting workflow.

Why the other options are wrong

B

Useful but not the first step; describe node gives overview.

C

Recovery action, not diagnosis.

D

Unlikely to show node conditions.

103
Multi-Selectmedium

Which THREE of the following are valid steps to troubleshoot DNS issues in a Kubernetes cluster?

Select 3 answers
A.Run 'kubectl exec <pod> -- nslookup kubernetes.default'
B.Check the /etc/resolv.conf on the host
C.Restart all nodes in the cluster
D.Verify the kube-dns Service has endpoints
E.Check the logs of the CoreDNS pods
AnswersA, D, E

Running `kubectl exec <pod> -- nslookup kubernetes.default` is a valid troubleshooting step because it tests DNS resolution from inside the pod's network namespace, exactly where application failures occur. If this command fails, it confirms the issue is with DNS resolution rather than with the application itself. It uses the pod's configured DNS settings (resolv.conf and search domains) to query the cluster's DNS service. Success indicates the cluster DNS is reachable and resolving service names, while failure isolates the problem to DNS configuration or CoreDNS.

Why this answer

To troubleshoot DNS, you can check the DNS pod logs, test resolution from a pod, and verify the DNS service endpoints.

104
MCQmedium

You are trying to debug a network connectivity issue between two pods. Pod A can reach the internet but cannot reach Pod B's IP address. Which command should you use to test connectivity from within Pod A to Pod B's service?

A.kubectl exec pod-a -- nslookup service-b
B.curl http://<node-ip>:<nodeport>
C.ssh node-ip 'curl http://<pod-b-ip>:80'
D.kubectl exec pod-a -- curl http://service-b:80
AnswerD

This is the most effective command because `kubectl exec` runs `curl` directly within the network namespace of `pod-a`, simulating the exact origin of the communication. By targeting `http://service-b:80`, it simultaneously tests DNS resolution of the service name, the ability to establish a TCP connection to the service's ClusterIP on port 80, and the application's responsiveness. This provides a comprehensive end-to-end test from the perspective of the source pod.

Why this answer

It uses `kubectl exec` to run a command inside Pod A, then uses `curl` to reach Pod B's service by its DNS name (`service-b`) and port 80. This tests connectivity from Pod A's network namespace to the ClusterIP service, which is the correct way to verify pod-to-service communication within the cluster. Using the service name leverages Kubernetes internal DNS (CoreDNS) to resolve to the service's virtual IP, and `curl` sends an HTTP request to confirm reachability.

Exam trap

The trap here is that candidates confuse testing pod-to-service connectivity (which requires using the service DNS name from within the pod) with testing node-to-pod or DNS-only checks, leading them to pick options that bypass the pod's network namespace or only test DNS resolution.

How to eliminate wrong answers

Option A is wrong because `nslookup` only tests DNS resolution of the service name, not actual network connectivity to the service IP or pod. Option B is wrong because it tests connectivity from the node to the NodePort, not from within Pod A to the service; this bypasses Pod A's network namespace and does not verify pod-to-service communication. Option C is wrong because it runs `curl` from the node (via SSH) to Pod B's IP, which tests node-to-pod connectivity, not pod-to-service connectivity from within Pod A.

105
MCQmedium

You are debugging a DNS issue from within a pod. The pod is running 'busybox'. Which command would you use to test DNS resolution for 'kubernetes.default.svc.cluster.local'?

A.kubectl describe svc kubernetes -n default
B.kubectl exec -it my-pod -- curl kubernetes.default.svc.cluster.local
C.kubectl run test --image=busybox -- nslookup kubernetes.default.svc.cluster.local
D.kubectl exec -it my-pod -- nslookup kubernetes.default.svc.cluster.local
AnswerD

Executing `kubectl exec -it my-pod -- nslookup kubernetes.default.svc.cluster.local` is the most direct and effective method for debugging DNS resolution issues from within a specific pod. This command leverages `kubectl exec` to run `nslookup` directly inside `my-pod`, utilizing that pod's `/etc/resolv.conf` and its configured DNS server. It precisely tests whether the pod itself can successfully resolve the fully qualified domain name (FQDN) of the `kubernetes` service, providing immediate insight into its DNS capabilities.

Why this answer

`kubectl exec -it my-pod -- nslookup kubernetes.default.svc.cluster.local` runs the `nslookup` command directly inside the running pod, which uses the pod's configured DNS resolver (typically CoreDNS) to resolve the Kubernetes service FQDN. This is the standard method to test DNS resolution from within a pod, as it bypasses any external DNS and validates the cluster's internal DNS chain.

Exam trap

The trap here is that candidates often choose Option B (curl) thinking it tests DNS, but curl tests HTTP connectivity, not resolution; or they choose Option C (kubectl run) which creates a new pod with default DNS settings, missing the specific pod's DNS configuration that may be the root cause of the issue.

How to eliminate wrong answers

Option A is wrong because `kubectl describe svc kubernetes -n default` only shows the service's metadata and endpoints, not DNS resolution; it does not test the pod's ability to resolve the name. Option B is wrong because `curl` tests HTTP connectivity, not DNS resolution; a successful curl could still hide a DNS failure if the IP is cached or resolved via other means, and busybox may not include curl by default. Option C is wrong because `kubectl run test --image=busybox -- nslookup ...` creates a new ephemeral pod, which is unnecessary and slower; it also does not test DNS from the existing pod that is experiencing the issue, missing the specific pod's DNS configuration (e.g., dnsPolicy, resolv.conf).

106
MCQmedium

Based on the exhibit, the pod is in CrashLoopBackOff. Which command should you run NEXT to identify the root cause?

A.kubectl describe node node-1
B.kubectl top pod api-6f4d7b9d4c-abcde -n production
C.kubectl get deployment api -n production -o yaml
D.kubectl logs api-6f4d7b9d4c-abcde -n production --previous
AnswerD

kubectl logs api-6f4d7b9d4c-abcde -n production --previous is the correct command because it fetches the stdout/stderr from the previous, now-terminated container instance in the pod. In a CrashLoopBackOff, the currently restarted container usually has no useful logs — it may not have started, or it immediately restarted before writing anything — while the last crashed instance carries the actual error that triggered the restart. This gives you the application-level failure message (e.g., uncaught exception, missing config, listen EADDRINUSE) needed to fix the root cause; pair it with kubectl describe pod to see the last exit code and restart count.

Why this answer

The pod is in CrashLoopBackOff, which means the container starts, crashes, and restarts repeatedly. The `kubectl logs --previous` command retrieves the logs from the previous (crashed) container instance, which is the fastest way to see the error that caused the crash. This directly reveals the root cause, such as a missing dependency, configuration error, or application panic.

Exam trap

The trap here is that candidates may think `kubectl describe pod` or `kubectl get deployment` is needed to check the pod's status or configuration, but the fastest way to see the crash reason is the previous container's logs, not the current (restarted) container's logs which may be empty.

How to eliminate wrong answers

Option A is wrong because `kubectl describe node` shows node-level conditions and resource usage, not the application error causing the container to crash. Option B is wrong because `kubectl top pod` shows current CPU/memory metrics, which are irrelevant to a crash loop caused by an application error. Option C is wrong because `kubectl get deployment -o yaml` shows the desired state and pod template, but not the runtime logs or crash reason from the container.

107
Multi-Selecthard

You are troubleshooting a pod that is in 'Pending' state. 'kubectl describe pod' shows '0/1 nodes are available: 1 Insufficient memory, 1 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate'. Which TWO actions can resolve the issue?

Select 2 answers
A.Reduce the memory request in the container spec to fit available memory
B.Add a node selector to the pod spec to target a specific node
C.Increase the memory request to prioritize scheduling
D.Add resource limits without changing requests
E.Add a toleration for the control-plane taint to the pod spec
AnswersA, E

The scheduler reports Insufficient memory, meaning no node has enough allocatable memory for the pod's request. Lowering the container's memory request brings it within available capacity, allowing the scheduler to bind the pod to a suitable node.

Why this answer

The pod is pending because the single node in the cluster (0/1 nodes available) has two blocking issues: 1) Insufficient memory to satisfy the pod's request, and 2) a control-plane taint that the pod does not tolerate. To resolve this and allow the pod to schedule on this node, both issues must be addressed: you must reduce the memory request in the container spec to fit the available memory (Option A) AND add a toleration for the control-plane taint to the pod spec (Option E).

Exam trap

In a single-node cluster (indicated by '0/1 nodes are available'), any scheduling failure message lists all reasons why that single node failed. You must resolve all listed constraints (both the taint and the resource insufficiency) for the pod to schedule.

108
Multi-Selecteasy

Which TWO of the following are valid reasons a pod might be stuck in 'Pending' state?

Select 2 answers
A.Container is killed due to OOM
B.Container image pull fails because of authentication error
C.Not enough CPU or memory available on any node
D.Node is rebooted
E.Node has a taint that the pod does not tolerate
AnswersC, E

The scheduler cannot bind the pod to any node because every candidate lacks sufficient allocatable CPU or memory. Unschedulable resource requests leave the pod without a node assignment, so it remains Pending until capacity frees up or the pod's requests are reduced.

Why this answer

Option C is correct because the Kubernetes scheduler cannot bind a pod to any node when no node has sufficient allocatable CPU or memory to satisfy the pod's resource requests, leaving the pod in Pending with a FailedScheduling event. Option E is correct because taints on nodes repel pods that lack a matching toleration, so the scheduler finds no feasible node and the pod remains Pending. Option A is wrong because OOM-killed containers occur after scheduling, producing CrashLoopBackOff or OOMKilled statuses on a running pod, not Pending.

Option B is wrong because an image pull authentication failure happens on a node after the pod is already scheduled, resulting in ImagePullBackOff or ErrImagePull. Option D is wrong because rebooting a node affects already-running pods (e.g., eviction/rescheduling), not a pod that has never been scheduled.

Exam trap

The CKA exam often tests the distinction between pod scheduling failures (Pending) and runtime failures (CrashLoopBackOff, ImagePullBackOff) — candidates mistakenly associate any container startup issue with 'Pending', but 'Pending' strictly means the pod has not been assigned to a node yet.

109
MCQeasy

A pod is stuck in 'Pending' state. You run 'kubectl describe pod my-pod' and see the event: '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/unreachable: }, that the pod didn't tolerate'. What is the likely cause?

A.The pod's container image is not found
B.The pod has a resource request that cannot be met by any node
C.The nodes are unreachable or have network issues
D.The PersistentVolumeClaim is not bound
AnswerC

When a node becomes unreachable or experiences network partition issues, the node controller automatically applies the node.kubernetes.io/unreachable taint to it. Because the pending pod does not possess a matching toleration for this specific taint, the Kubernetes scheduler cannot assign the pod to any of these affected nodes, leaving it stuck in the Pending state.

Why this answer

The event '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/unreachable: }' indicates that all nodes in the cluster have the 'node.kubernetes.io/unreachable' taint, which is automatically applied by the node controller when a node becomes unreachable (e.g., due to network partition, node failure, or kubelet not reporting). Since the pod does not have a toleration for this taint, it cannot be scheduled on any node, resulting in a 'Pending' state. This is a classic scheduling failure caused by node unreachability, not by resource constraints or image issues.

Exam trap

The CKA exam tests the distinction between taint-based scheduling failures and resource-based failures; the trap here is that candidates may confuse the 'unreachable' taint with resource constraints or PVC issues, but the event message explicitly names the taint key, which directly points to node reachability problems.

How to eliminate wrong answers

Option A is wrong because an image-not-found error would produce a different event, such as 'Failed to pull image' or 'ErrImageNeverPull', and would not cause a taint-based scheduling failure. Option B is wrong because resource request issues would generate events like 'Insufficient cpu' or 'Insufficient memory', not a taint-based message referencing 'node.kubernetes.io/unreachable'. Option D is wrong because an unbound PersistentVolumeClaim would produce an event like 'persistentvolumeclaim not found' or 'failed to mount volume', and would not result in a taint-based scheduling failure.

110
MCQeasy

Which command shows resource usage (CPU and memory) for nodes in a cluster?

A.kubectl describe nodes
B.kubectl logs nodes
C.kubectl get nodes -o wide
D.kubectl top nodes
AnswerD

The `kubectl top nodes` command retrieves resource usage data from the Metrics API, which is typically backed by a metrics-server that collects node-level metrics from kubelets. It reports current CPU and memory consumption in both absolute units (e.g., millicores and MiB) and as percentages of node allocatable resources. This is exactly the command needed to answer the question about node resource usage.

Why this answer

kubectl top nodes displays resource usage if metrics-server is installed.

111
Multi-Selectmedium

Which TWO actions can help troubleshoot a service that is not reachable from within a pod?

Select 2 answers
A.Use kubectl delete svc and recreate
B.Use kubectl exec to curl the service IP from a pod
C.Use kubectl logs on the service pod to see application logs
D.Use kubectl top nodes to check resource usage
E.Use kubectl describe svc to check endpoints
AnswersB, E

Running `kubectl exec` into a pod and curling the Service's ClusterIP directly tests the Service's virtual IP from inside the cluster network namespace. This validates the entire data path: the pod's routing, kube-proxy or iptables/ipvs rules, and backend endpoint selection. It is especially useful because the ClusterIP is only reachable from within the cluster, so a successful curl proves the kube-proxy rules are installing correctly and that at least one healthy backend is receiving traffic. A failed curl with a timeout, rather than a connection refused, also gives a specific hint about whether packets are being silently dropped.

Why this answer

Checking endpoints and using curl from a pod are direct troubleshooting steps.

112
MCQmedium

A pod is in 'ImagePullBackOff' state. Which of the following is NOT a common cause?

A.The image registry requires authentication and no imagePullSecrets are configured
B.The image tag does not exist
C.The image name is misspelled
D.The container requires more memory than the limit allows
AnswerD

If a container demands more memory than its limit allows, the image has already been successfully pulled and the container has started, so the failure mode is OOMKilled or an overloaded kubelet eviction, not ImagePullBackOff. ImagePullBackOff is exclusively a pre-start image acquisition failure, making this the one option that could never produce that state.

Why this answer

ImagePullBackOff means the kubelet repeatedly failed to pull the container image and is backing off before retrying. Memory limits are enforced at container runtime after the image is pulled and the container starts, so an insufficient memory limit would cause OOMKilled or scheduling failures — not an image pull error. Therefore, 'requires more memory than the limit allows' is NOT a common cause of ImagePullBackOff.

Exam trap

The trap is conflating image-pull failures with runtime resource failures; candidates see 'memory' and think of pod failures generally, but ImagePullBackOff is strictly about fetching the image, not running it.

How to eliminate wrong answers

Option A is a genuine cause: if the registry requires authentication and no `imagePullSecrets` are configured on the pod or service account, the kubelet receives a 401/403 and enters ImagePullBackOff. Option B is a genuine cause: a non-existent tag causes the registry to return a manifest-unknown error, which the kubelet surfaces as ImagePullBackOff. Option C is a genuine cause: a misspelled image name results in a repository-not-found error, again producing ImagePullBackOff.

Only option D describes a runtime resource issue that manifests as OOMKilled or Pending scheduling, not an image pull failure.

113
MCQeasy

A pod is in ImagePullBackOff state. Which command would give you the most information about why the image pull failed?

A.kubectl get pod
B.kubectl logs <pod-name>
C.kubectl edit pod <pod-name>
D.kubectl describe pod <pod-name>
AnswerD

The `kubectl describe pod` command is the correct diagnostic because it aggregates the pod's status conditions, container states, and, crucially, the recent Events list from the kubelet and the image puller. For an `ImagePullBackOff`, the events will contain a specific reason such as `ErrImagePull`, `ImagePullBackOff`, `Failed to pull image`, or a registry authentication/not found error with the exact HTTP status. This detailed output is exactly what you need to pinpoint whether the problem is a typo in the image tag, missing credentials, or network connectivity to the registry.

Why this answer

`kubectl describe pod <pod-name>` provides detailed event logs, including the exact error message from the kubelet when it failed to pull the container image. This output includes the reason for the ImagePullBackOff state, such as authentication failures, image not found, or network issues, which is the most comprehensive information for troubleshooting.

Exam trap

The trap here is that candidates often think `kubectl logs` will show the error, but since the container never started, there are no logs; the real diagnostic data is in the pod's events and status conditions, which only `kubectl describe` reveals.

How to eliminate wrong answers

Option A is wrong because `kubectl get pod` only shows the current status (e.g., ImagePullBackOff) without any details about why the pull failed. Option B is wrong because `kubectl logs <pod-name>` retrieves container logs, but if the container never started due to an image pull failure, there are no logs to display. Option C is wrong because `kubectl edit pod <pod-name>` opens the pod specification for editing, which does not show the pull failure reason; it only allows you to modify the pod definition, which is not diagnostic.

114
MCQhard

Based on the exhibit, what is the most likely cause of the pod not running?

A.The volume driver is not installed on node-1.
B.The pod has exceeded its resource limits.
C.The node 'node-1' is experiencing disk pressure.
D.The Secret 'my-secret' does not exist in the namespace.
AnswerD

The exhibit's event message contains the exact Kubernetes error string: the secret `my-secret` could not be found in the pod's namespace, so the kubelet is unable to inject the environment variable or volume content required by the container spec. Every Secret reference is namespaced, and the kubelet queries the API server for the secret exactly as it appears in the pod manifest; any typo, wrong namespace, or omitted resource will immediately produce this failure. Because the error is explicit and points to a missing API object, the most likely cause is that `my-secret` simply does not exist in the namespace where the Pod is running.

Why this answer

The pod's status indicates it is waiting for a secret to be mounted, and the error message 'secret "my-secret" not found' directly points to the missing Secret resource. Without the Secret existing in the same namespace as the pod, the volume mount fails, preventing the pod from starting.

Exam trap

The trap here is that candidates may assume the issue is node-level (disk pressure or driver) or resource-related, overlooking the specific error message about the missing Secret, which is a common misdirection in CKA troubleshooting questions.

How to eliminate wrong answers

Option A is wrong because a missing volume driver would typically result in a different error, such as 'failed to mount volume' or 'driver not supported', not a secret not found error. Option B is wrong because exceeding resource limits would cause the pod to be in a CrashLoopBackOff or OOMKilled state, not a waiting state for a secret. Option C is wrong because disk pressure on node-1 would manifest as pod eviction or scheduling failures, not a secret mount error.

115
Multi-Selectmedium

You run 'kubectl logs pod-name' and get no output. Which TWO steps should you take to troubleshoot further?

Select 2 answers
A.Run 'kubectl get events --all-namespaces'
B.Run 'kubectl top pod pod-name' to check resource usage
C.Run 'kubectl describe pod pod-name' to check container state and events
D.Run 'kubectl logs --previous pod-name'
E.Run 'kubectl exec pod-name -- cat /var/log/container.log'
AnswersC, D

`kubectl describe pod pod-name` is a correct first step because it shows container states (Waiting, Running, Terminated) with detailed reason and message fields—for example, `CrashLoopBackOff`, `ImagePullBackOff`, or `OOMKilled`. It also lists recent events specific to that pod, such as failed volume mounts or failed liveness probes, which directly explain why a container may have never produced logs or why its log stream was cut short. When `kubectl logs` returns nothing, this command reveals whether the container even started, and if it did, what caused it to terminate or restart, making it an essential troubleshooting action.

Why this answer

Option C is correct because 'kubectl describe pod pod-name' surfaces the pod's container states (Waiting, Running, Terminated), restart counts, and the recent event stream, which reveals whether the container is crash-looping, stuck in ImagePullBackOff, or never started — all reasons 'kubectl logs' would return nothing. Option D is correct because 'kubectl logs --previous pod-name' retrieves the logs from the prior container instance, which is essential when the current container has restarted and its fresh instance has not yet produced output. Option A is not the right step because cluster-wide events are noisy and not scoped to the specific pod, making it far less targeted than 'kubectl describe pod'.

Option B is incorrect because 'kubectl top pod' only reports CPU/memory metrics and does not explain why logs are empty. Option E is incorrect because it assumes an in-container log file path that may not exist and bypasses the standard container log stream that kubectl already exposes.

Exam trap

CKA often tests whether candidates know the difference between current and previous container logs — the trap is running 'kubectl logs' repeatedly on a restarted container and missing that '--previous' is needed to see the crash output.

116
Multi-Selecthard

You are troubleshooting a scenario where a pod cannot communicate with another pod in the same namespace via service name. Which THREE steps would you take to diagnose the issue? (Select 3)

Select 3 answers
A.Run 'kubectl get nodes' to check node status
B.Run 'kubectl get endpoints' to verify the service has healthy endpoints
C.Exec into the pod and use curl to test connectivity to the service's cluster IP
D.Run 'kubectl logs' on the target pod to check application logs
E.Exec into the pod and run nslookup to verify DNS resolution of the service name
AnswersB, C, E

A Kubernetes Service only forwards traffic to Pod IPs listed in its Endpoints object, which are populated by the controller based on matching selectors and the readiness status of pods. If the selector matches no pods, or the pods are not Ready (e.g., failing readiness probes or CrashLoopBackOff), the Endpoints object is empty, so connections to the Service's ClusterIP are dropped or refused. Running 'kubectl get endpoints' is the quickest way to confirm whether the Service actually has healthy, Ready backends, directly exposing the most common cause of communication failure.

Why this answer

Options B, C, and E are correct. Checking endpoints (B) verifies the service has healthy pods. Exec into the pod and using curl (C) tests connectivity to the service's cluster IP.

Exec into the pod and running nslookup (E) checks DNS resolution of the service name. Option A checks node status, which is not directly related to pod-to-pod communication via service name. Option D checks logs of the target pod, which may not reveal network issues.

117
MCQeasy

A pod is stuck in Pending state. You run 'kubectl describe pod my-pod' and see the event '0/1 nodes are available: 1 Insufficient cpu'. What is the most likely cause?

A.The pod's memory limit is too low
B.The pod is trying to use a GPU that is not available
C.The pod's CPU request exceeds available node CPU
D.The kubelet on the node is not running
AnswerC

When a pod requests more CPU than any node can allocate, the kube-scheduler cannot find a feasible host and sets the pod to Pending, recording an event such as `Insufficient cpu`. CPU requests are guaranteed amounts that the scheduler sums across all existing pods and compares against a node's `allocatable` CPU (node capacity minus reserved system components). Until some running pods are terminated or a node with more free CPU joins the cluster, the pod will remain Pending, which exactly matches the given event.

Why this answer

The event indicates that no node has enough free CPU to satisfy the pod's CPU request.

118
MCQmedium

A pod is in ImagePullBackOff. Which command would help determine the exact reason?

A.kubectl logs <pod>
B.kubectl describe pod <pod>
C.kubectl get events
D.kubectl exec -it <pod> -- sh
AnswerB

kubectl describe pod <pod> is correct because it renders the pod's full status section, including each container's current state (Waiting, reason, message), and appends pod Events. For an ImagePullBackOff, the Events show the exact registry error (e.g., unauthorized, manifest unknown, network timeout), and the container status shows the backoff reason, giving you the diagnostic detail needed to fix the image pull failure.

Why this answer

The `kubectl describe pod <pod>` command provides detailed information about the pod, including the container state, the exact error message from the image pull (e.g., 'ImagePullBackOff'), and the underlying reason (e.g., 'Back-off pulling image', 'manifest for image not found', or 'unauthorized: authentication required'). This is the most direct way to see the specific error without needing to access the container or parse raw events.

Exam trap

CNCF often tests the misconception that `kubectl logs` can diagnose startup failures, but logs are only available after the container has started, making `kubectl describe pod` the correct tool for pre-start errors like ImagePullBackOff.

Why the other options are wrong

A

Container hasn't started; no logs.

C

Shows all events, not specific to pod.

D

Pod not running.

119
Multi-Selecteasy

Which TWO commands show cluster events that can help in troubleshooting?

Select 2 answers
A.kubectl cluster-info
B.kubectl logs <pod-name>
C.kubectl top pods
D.kubectl describe pod <pod-name>
E.kubectl get events
AnswersD, E

kubectl describe pod <pod-name> outputs a comprehensive summary of the pod's configuration and status, culminating in an Events section that chronologically lists recent events for that specific pod (e.g., successful pull, failed schedule, container started). These events often contain the direct cause of a pod problem, making this command one of the two candidates that actually surface events. However, it only shows events for the named pod, so its scope is limited to that Pod object.

Why this answer

Option E, `kubectl get events`, is correct because it directly lists the cluster's event stream (backed by the Events API), showing scheduling failures, image pull errors, probe failures, and other warnings tied to specific resources. Option D, `kubectl describe pod <pod-name>`, is correct because its output includes an Events section at the bottom that surfaces the same per-pod events with timestamps, reasons, and messages, which is essential for troubleshooting a specific pod. Option A, `kubectl cluster-info`, only prints the addresses of the control plane and core add-ons, providing no event data.

Option B, `kubectl logs <pod-name>`, shows application/container stdout and stderr, not Kubernetes cluster events. Option C, `kubectl top pods`, reports CPU and memory usage metrics from the metrics server, which is resource data rather than events.

Exam trap

The trap here is that candidates often confuse `kubectl logs` (which shows container output) with event viewing, or assume `kubectl cluster-info` provides troubleshooting events, when in fact only `kubectl describe` and `kubectl get events` surface the cluster's event history.

120
MCQhard

A pod is in CrashLoopBackOff. 'kubectl logs my-pod --previous' shows: 'Error: failed to start: exec: "/app/start.sh": stat /app/start.sh: no such file or directory'. What is the most likely cause?

A.The pod's service account lacks permissions.
B.The container image is missing the startup script.
C.The pod's liveness probe is misconfigured.
D.The pod has run out of memory.
AnswerB

The error 'no such file or directory' for /app/start.sh in the container logs directly indicates that the command or entrypoint defined in the pod spec references a script that does not exist within the container image. This typically happens when a Dockerfile fails to COPY the script into the image, or the image tag points to an older build without the file. Because the process cannot start, the container exits non-zero and Kubernetes enters CrashLoopBackOff.

Why this answer

The error indicates the specified entrypoint script is missing. This is often due to the container image not containing the script at the expected path, or the command/args in the pod spec referencing a non-existent file.

121
MCQeasy

You suspect the kubelet on a worker node is not functioning correctly. Which command should you use to check the kubelet service status?

A.kubectl get nodes
B.systemctl status kubelet
C.kubectl describe node <node-name>
D.journalctl -u kubelet
AnswerB

systemctl status kubelet is the definitive, node-local command for inspecting the kubelet's systemd service state. It directly queries systemd's unit properties, showing whether the unit is active (running), failed, activating, or inactive, along with the main process ID (MainPID) and recent journal entries. This is the first step when the kubelet is suspected of malfunctioning because it immediately confirms whether the process is alive and healthy from the init system's perspective, without relying on the Kubernetes API or network latency. It also reveals the exit code or last status change, which is essential for diagnosing crashes or manual stops on the node itself.

Why this answer

The kubelet is a systemd service on worker nodes, so `systemctl status kubelet` is the correct command to check its current state, whether it is active (running), inactive, or failed. This command directly queries the service manager for the kubelet's status, which is the first step in troubleshooting a suspected malfunction.

Exam trap

The trap here is that candidates confuse cluster-level commands like `kubectl get nodes` with node-level service management, assuming a node showing NotReady means the kubelet service is definitely down, when in fact the service could be running but failing to communicate with the API server.

How to eliminate wrong answers

Option A is wrong because `kubectl get nodes` only shows the cluster's view of node readiness, not the actual kubelet service status on the node; a node can appear NotReady even if the kubelet is running but misconfigured. Option C is wrong because `kubectl describe node <node-name>` provides detailed node conditions and events from the control plane's perspective, but it does not check the kubelet systemd service status directly. Option D is wrong because `journalctl -u kubelet` shows the kubelet's logs, which is useful for deeper investigation after confirming the service status, but it does not show whether the service is currently active or failed.

122
Multi-Selecthard

You have a pod that is in CrashLoopBackOff. Which two troubleshooting steps should you take first? (Choose two.)

Select 2 answers
A.kubectl describe pod pod-name
B.kubectl delete pod pod-name
C.kubectl logs pod-name --previous
D.kubectl exec -it pod-name -- sh
E.kubectl rollout restart deployment
AnswersA, C

kubectl describe pod pod-name is correct for CrashLoopBackOff because it displays the pod's full lifecycle events, container states, restart counts, and the last reason/exit code from the previous terminated container. Those events often reveal the root cause, such as image pull failures, failed readiness/liveness probes, or OOMKilled. It also shows the current backoff state and timestamps, making it the first diagnostic command to run.

Why this answer

`kubectl describe pod pod-name` provides detailed information about the pod's current state, including recent events, container restart counts, and the reason for the CrashLoopBackOff (e.g., exit code 137 from OOMKill or 1 from application error). This is the first step to understand the root cause of the crash loop.

Exam trap

The CKA exam often tests the misconception that `kubectl exec` can be used to debug a crashing pod, but in CrashLoopBackOff the container is not running, so exec fails; candidates must remember to use `kubectl logs --previous` to access logs from the terminated instance.

123
MCQmedium

A Deployment named 'web' is failing to schedule pods. You run 'kubectl describe pod web-xyz' and see the event: '0/3 nodes are available: 3 Insufficient cpu.' What is the most likely cause?

A.The CPU request in the pod spec is too high
B.The network plugin is misconfigured
C.The nodes have insufficient memory
D.The kubelet is not running on the nodes
AnswerA

The scheduler evaluates the pod's CPU request against the allocatable CPU on each node after reserving resources for existing pods and system components. If the requested CPU exceeds what any node can accommodate, the scheduler reports 'Insufficient cpu' and leaves the pod Pending. This is a resource-fit failure, not an API or runtime error, so checking the pod's resources.requests is the first diagnostic.

Why this answer

The error indicates that CPU requests are too high for the available node resources. Reducing CPU requests or adding more nodes can fix it.

124
Multi-Selectmedium

Which TWO of the following are common causes for a pod to be in the 'Pending' state?

Select 2 answers
A.The container image is not found
B.Insufficient cluster resources (CPU/memory) to schedule the pod
C.The pod's liveness probe is failing
D.A PersistentVolumeClaim (PVC) referenced by the pod is not bound
E.The node is unreachable due to network issues
AnswersB, D

If the cluster lacks nodes with enough allocatable CPU or memory to satisfy the pod's resource requests, the Kubernetes scheduler cannot assign the pod to any node and leaves it in Pending. The scheduler only considers unreserved resources calculated as capacity minus requests from existing pods; a node with insufficient free capacity or failing a node selector will keep the pod unschedulable. The scheduler emits events like '0/3 nodes are available: insufficient cpu' to point you to this condition.

Why this answer

Option B is correct because the Kubernetes scheduler places a pod into the Pending state when no node has enough allocatable CPU or memory to satisfy the pod's resource requests, so the pod remains unscheduled until resources free up or the cluster scales. Option D is correct because if a pod references a PersistentVolumeClaim that is not yet Bound (for example, waiting on a dynamic provisioner or a matching PersistentVolume), the scheduler cannot satisfy the volume's node affinity/topology constraints and the pod stays Pending. Option A is not a cause of Pending; an image that cannot be pulled produces an ImagePullBackOff/ErrImagePull status while the pod is already scheduled and typically Running or Waiting on a node.

Option C is not a cause of Pending; a failing liveness probe causes kubelet to restart the container, resulting in CrashLoopBackOff or repeated restarts, not a scheduling failure. Option E is not a cause of Pending; an unreachable node affects existing pods (marking them Unknown/NotReady) rather than preventing scheduling, since the scheduler simply avoids nodes it considers unhealthy.

Exam trap

The CKA exam often tests the distinction between pod lifecycle phases—candidates confuse 'Pending' (pre-scheduling or pre-startup) with post-startup failures like probe failures or image pull errors, which occur in later states.

125
MCQmedium

You see a pod in 'Pending' state. 'kubectl describe pod' shows '0/4 nodes are available: 1 node(s) had taint(s) that the pod didn't tolerate, 3 Insufficient cpu'. What should you do?

A.Delete the taint from the node
B.Scale down other deployments to free CPU
C.Increase the CPU limits only
D.Add the required toleration to the pod spec and increase CPU requests
AnswerB

Scaling down other Deployments reduces the total CPU requests reported against those nodes, which is exactly what the scheduler uses to compute resource availability. When the request sums drop below a node's allocatable capacity, the pending pod's request can fit, enabling successful placement. This resolves the root cause of the pending state without altering the pod's toleration settings or risking node stability, and it is a common operational remedy for clusters experiencing CPU exhaustion.

Why this answer

The pod is pending due to two issues: taint on one node and insufficient CPU on three nodes. Scaling down other deployments frees CPU resources on nodes, allowing the pod to be scheduled on nodes with sufficient CPU, potentially avoiding the tainted node. Option A only removes the taint but does not solve the CPU shortage.

Option C does not affect scheduling. Option D increases CPU requests, worsening the CPU shortage. Therefore, B is the correct action.

126
Multi-Selecthard

Which THREE of the following are common causes for a pod to remain in 'Pending' state?

Select 3 answers
A.Node taints that the pod does not tolerate
B.Image pull backoff
C.Container OOMKilled
D.Insufficient CPU or memory resources on any node
E.PersistentVolumeClaim is not bound
AnswersA, D, E

If one or more nodes have taints (e.g., node-role.kubernetes.io/control-plane:NoSchedule) and the pod spec lacks matching tolerations, the scheduler will exclude those nodes from feasible candidates. As a result, the pod remains in Pending because no node can satisfy both taint/toleration rules and other scheduling predicates. This is a common cause of Pending, especially in single-node clusters or clusters with specialized node pools.

Why this answer

A pod stays in Pending when the Kubernetes scheduler cannot place it on any node, and node taints that the pod does not tolerate (option A) cause the scheduler to filter out those nodes, leaving the pod unscheduled. Insufficient CPU or memory resources on any node (option D) is a classic scheduling failure: the scheduler's resource-fit predicate rejects every node, so the pod remains Pending. A PersistentVolumeClaim that is not bound (option E) also blocks scheduling, because a pod referencing an unbound PVC cannot be assigned to a node until the volume is provisioned and bound.

By contrast, Image pull backoff (option B) occurs after scheduling, when the kubelet cannot pull the image, so the pod is in Waiting/ContainerCreating rather than Pending. Container OOMKilled (option C) is a runtime termination state that happens after the container starts, so it does not cause a Pending pod.

Exam trap

The CKA exam often tests the distinction between pre-scheduling (Pending) and post-scheduling (Running, Waiting, Terminated) failures; candidates mistakenly associate image pull or OOM errors with Pending, but these occur only after the pod is bound to a node.

127
MCQmedium

A pod is in CrashLoopBackOff. You run 'kubectl logs mypod --previous' and see 'Error: unable to connect to database'. What is the MOST likely cause?

A.The database service endpoint is unreachable
B.The pod's readiness probe is failing
C.The pod's liveness probe is misconfigured
D.The pod is out of memory
AnswerA

`kubectl logs` captures the container's stdout/stderr, so a database connection failure (e.g., "connection refused" or "dial tcp: lookup db-service") is exactly what appears. This is a typical application-level startup error that causes the process to exit, triggering the kubelet's restart backoff. Probe failures or OOM would not produce this specific log content, making this the most plausible direct cause.

Why this answer

The error 'unable to connect to database' indicates a network connectivity issue between the pod and the database service. Since the error is from the application itself (not a Kubernetes probe), the most likely cause is that the database service endpoint is unreachable, either due to a misconfigured service, incorrect DNS resolution, or network policy blocking traffic. The `--previous` flag shows logs from the previous container instance, confirming the application consistently fails to reach the database.

Exam trap

The trap here is that candidates often confuse application-level errors with probe failures, but the CKA exam expects you to recognize that a specific database connection error points to a network or service endpoint issue, not a probe misconfiguration or resource problem.

How to eliminate wrong answers

Option B is wrong because a failing readiness probe would cause the pod to be removed from service endpoints, but it would not produce an application-level error like 'unable to connect to database'; readiness probe failures result in the pod being marked as not ready, not a CrashLoopBackOff from application crashes. Option C is wrong because a misconfigured liveness probe would cause the pod to be restarted by kubelet, but the application error message would not be 'unable to connect to database'; liveness probe failures typically result in container restarts without application-specific database connection errors. Option D is wrong because an out-of-memory (OOM) condition would cause the container to be killed by the kernel with an OOMKilled status, and logs would show no such error message; the error 'unable to connect to database' is a network-level error, not a resource exhaustion symptom.

128
MCQmedium

A pod is stuck in 'Pending' state. You run 'kubectl describe pod mypod' and see: '0/4 nodes are available: 4 node(s) didn't match pod anti-affinity constraints'. What does this mean?

A.The pod has a taint that no node tolerates.
B.The pod has a node selector that doesn't match any node.
C.The pod's anti-affinity rule conflicts with all existing pods on every node.
D.The cluster is out of resources.
AnswerC

Pod anti-affinity enforces that a pod cannot share a topology domain with other pods that match a specified label selector. If the pod's `requiredDuringSchedulingIgnoredDuringExecution` anti-affinity rule targets pods that appear on every node within the cluster, then each node is disqualified and the scheduler reports something like `0/3 nodes are available: 3 node(s) didn't match pod anti-affinity rules.` This exactly matches the described output, confirming the conflict with existing pods is the cause of the Pending state.

Why this answer

The error message '4 node(s) didn't match pod anti-affinity constraints' directly indicates that the pod's anti-affinity rule (defined in the pod spec under `affinity.podAntiAffinity`) prevents it from being scheduled on any node because the rule conflicts with the labels of existing pods on all nodes. Anti-affinity ensures the pod is not co-located with certain pods, and if every node already hosts a pod matching the anti-affinity selector, no node is eligible. This is distinct from taints, node selectors, or resource shortages.

Exam trap

The CKA exam often tests the distinction between anti-affinity, taints/tolerations, and node selectors; the trap here is that candidates misread 'anti-affinity constraints' as a generic scheduling failure and incorrectly choose resource exhaustion or taints.

How to eliminate wrong answers

Option A is wrong because taints and tolerations produce a message like 'node(s) had taints that the pod didn't tolerate', not an anti-affinity constraint error. Option B is wrong because a node selector mismatch yields 'node(s) didn't match node selector', not anti-affinity. Option D is wrong because resource exhaustion shows messages like 'Insufficient cpu' or 'Insufficient memory', not a constraint-based failure.

129
MCQmedium

You want to check the logs of a container that previously crashed. Which command should you use?

A.kubectl logs --previous <pod-name>
B.kubectl logs <pod-name>
C.kubectl exec <pod-name> -- cat /var/log/app.log
D.kubectl describe pod <pod-name>
AnswerA

The `--previous` flag instructs kubectl to retrieve the logs of the last terminated container instance within the pod. When a container has crashed and restarted, the current container's logs are empty or show only new output, while the terminated container's logs remain accessible via this flag. This is the correct way to diagnose why the previous container failed, as it directly fetches the stdout/stderr stream from that dead instance.

Why this answer

The `kubectl logs --previous` command retrieves logs from the previous instance of a container in a Pod that has crashed or been restarted. This is essential for debugging transient failures because the current container's logs may not contain the crash information. The `--previous` flag specifically accesses the terminated container's log stream, which is stored by the kubelet until the pod is deleted.

Exam trap

The trap here is that candidates often choose `kubectl logs <pod-name>` (option B) thinking it shows all logs, but they forget that a crashed container's logs are only accessible with the `--previous` flag.

How to eliminate wrong answers

Option B is wrong because `kubectl logs <pod-name>` only shows logs from the currently running container, not from a previously crashed instance. Option C is wrong because `kubectl exec` runs a command in a running container, which is impossible if the container has crashed and is not running. Option D is wrong because `kubectl describe pod` shows pod metadata, events, and status, but does not retrieve container logs, especially not from a previous crash.

130
Multi-Selecthard

A node is NotReady. Which THREE conditions could cause this?

Select 3 answers
A.Network plugin (e.g., Calico) is not functioning
B.A pod is in CrashLoopBackOff
C.kubelet service is stopped
D.Disk pressure on the node
E.The API server is overloaded
AnswersA, C, D

The kubelet determines node readiness by reporting conditions such as Ready, and one key condition the kubelet watches is the status of the container network interface (CNI). If the network plugin (e.g., Calico, Flannel, Cilium) fails to install routes, pods, or the CNI binaries, the kubelet cannot set up pod networking, causing the Ready condition to become False. The kubelet may also have its own network readiness checks (like the NodeStatus condition for network) that fail, and the node will stay NotReady until the CNI plugin is restored. This is why a healthy kubelet and Kubernetes control plane can still show a node as NotReady when the overlay network is broken.

Why this answer

Kubelet stopped, network plugin issues, and disk pressure can all cause a node to become NotReady.

131
Multi-Selecthard

Which THREE of the following are valid steps to troubleshoot a node that is in 'NotReady' state?

Select 3 answers
A.Check the kubelet status using 'systemctl status kubelet' on the node
B.View kubelet logs using 'journalctl -u kubelet'
C.Check node conditions with 'kubectl describe node <node-name>'
D.Restart the kubelet using 'systemctl restart kubelet'
E.Delete the node object and rejoin it to the cluster
AnswersA, B, C

Checking whether the kubelet is actually running on the node is the first diagnostic action: systemctl status kubelet reports whether the unit is active, the main PID, memory/CPU usage, and a short tail of recent log lines. If the service is inactive or failed, the exit status and timestamp help determine whether the node problem is a service crash, a stopped unit, or a configuration failure. This is a quick, non-destructive check that establishes the starting point before digging into logs.

Why this answer

Options A, B, and C are valid troubleshooting steps to investigate a NotReady node. Option A checks if kubelet is running, Option B examines kubelet logs for errors, and Option C shows node conditions. Option D (restarting kubelet) is a remediation action, not a troubleshooting step.

Option E (deleting and rejoining) is a recovery step.

132
Multi-Selecthard

A ClusterIP Service is not reachable from within the cluster. You verify that the Service has endpoints. Which of the following could be the cause? (Select two.)

Select 2 answers
A.kube-proxy is not running on the node.
B.The container is listening on a different port than the Service targetPort.
C.The pod's readiness probe is failing.
D.The Service name is too long.
AnswersA, B

kube-proxy programs the iptables or IPVS rules that implement ClusterIP load balancing on each node. If it is not running, packets to the virtual IP are never translated to a pod endpoint, so the Service fails despite having healthy endpoints.

Why this answer

Option A is correct because kube-proxy is the component that programs the iptables/IPVS rules on each node to implement ClusterIP Service virtual IP load balancing; if kube-proxy is not running on a node, pods on that node cannot reach the Service's ClusterIP even though endpoints exist. Option B is correct because the Service's targetPort must match the port the container actually listens on; if the container listens on a different port, traffic forwarded to the endpoint will be refused or dropped, making the Service unreachable despite having endpoints. Option C is not correct because a failing readiness probe would remove the pod from the Service's endpoints, but the scenario explicitly states that endpoints exist, so this cannot be the cause.

Option D is not correct because Kubernetes Service names are limited to 63 characters by DNS label rules, but an overly long name would be rejected at creation time rather than causing a running Service with endpoints to be unreachable.

Exam trap

Candidates may think a failing readiness probe can cause unreachability even when endpoints exist, but that is not possible because the pod would be removed from endpoints upon probe failure.

Why the other options are wrong

D

Service name length does not affect connectivity.

133
MCQmedium

You run 'kubectl get events --sort-by=.lastTimestamp' and see the following events for a pod: 'Warning FailedScheduling 0/3 nodes are available: 3 Insufficient cpu'. What is the most likely solution?

A.Reduce the CPU request for the pod or remove other workloads to free CPU
B.Change the scheduler to a different one
C.Increase the CPU limit for the pod
D.Add more nodes to the cluster
AnswerA

The Kubernetes scheduler uses a pod's CPU request to determine node feasibility during the filtering phase. Lowering this request value reduces the resource footprint required for scheduling, allowing the pod to fit onto existing nodes with limited allocatable CPU. Alternatively, evicting or deleting non-essential workloads frees up allocatable capacity on those nodes, resolving the scheduling bottleneck without requiring infrastructure changes.

Why this answer

The event '0/3 nodes are available: 3 Insufficient cpu' indicates that all three nodes in the cluster lack sufficient allocatable CPU to satisfy the pod's CPU request. The most direct solution is to either reduce the pod's CPU request (so it fits on an existing node) or remove other workloads to free up CPU capacity. This aligns with Kubernetes resource scheduling, where the scheduler only considers requests (not limits) when placing pods.

Exam trap

The trap here is that candidates often confuse CPU requests with CPU limits and mistakenly think increasing limits will help the pod get scheduled, but the scheduler only evaluates requests, not limits.

How to eliminate wrong answers

Option B is wrong because changing the scheduler does not address the root cause of insufficient CPU resources; the default scheduler already evaluates node capacity, and a different scheduler would face the same resource shortage. Option C is wrong because increasing the CPU limit does not affect scheduling decisions—limits are for resource enforcement at runtime, not for admission; the scheduler only considers CPU requests. Option D is wrong because adding more nodes is an over-engineered solution; the cluster already has three nodes, and the issue is that they are fully utilized, so reducing demand is more efficient and cost-effective than scaling out.

134
MCQmedium

You create a Deployment with the following YAML: apiVersion: apps/v1 kind: Deployment metadata: name: myapp spec: replicas: 3 selector: matchLabels: app: myapp template: metadata: labels: app: myapp spec: containers: - name: myapp image: myapp:latest resources: limits: cpu: "500m" memory: "256Mi" requests: cpu: "200m" memory: "128Mi" After applying it, the pods are in 'CrashLoopBackOff'. You check logs and see 'Error: container process exited with error'. What is the MOST likely cause?

A.The application is crashing at startup due to a missing dependency
B.The container is being OOMKilled because memory limit is too low
C.The readiness probe is failing
D.The container has insufficient CPU resources
AnswerA

When an application process exits immediately due to a missing dependency, runtime error, or misconfiguration, the container terminates with a non-zero exit code. Kubernetes detects this termination and automatically restarts the container, eventually entering a CrashLoopBackOff state as it backs off exponentially between restart attempts. This is the classic signature of an application-level startup failure.

Why this answer

The error message 'container process exited with error' indicates the application itself is failing to start, not that it is being killed by Kubernetes. Since the container exits immediately after starting, the most likely cause is a missing dependency or misconfiguration in the application code, such as a missing environment variable, database connection, or file. Resource limits and probes affect running containers, not initial startup failures.

Exam trap

The CKA exam often tests the distinction between container startup failures (application errors) and runtime resource issues (OOMKill, CPU throttling). Candidates may mistakenly attribute a generic exit error to resource limits or probes.

How to eliminate wrong answers

Option B is wrong because OOMKilled would show a specific 'OOMKilled' reason in `kubectl describe pod`, not a generic 'exited with error' message, and the memory limit of 256Mi is reasonable for many applications. Option C is wrong because readiness probes only affect traffic routing after the container is running; a failing readiness probe does not cause the container to exit with an error. Option D is wrong because insufficient CPU resources would cause the container to be throttled or pending, not to crash with an exit error; CPU limits are soft and do not kill containers.

135
MCQhard

A pod has status 'Init:Error'. What does this indicate?

A.The main container has crashed
B.An init container failed
C.The pod is being initialized
D.There is a network error during initialization
AnswerB

When an init container exits with a non-zero exit code, the pod status transitions to Init:Error (or Init:CrashLoopBackOff if it keeps failing). Kubernetes treats init containers as mandatory prerequisites: they run sequentially to completion before any regular containers start. The failing init container can be identified with kubectl describe pod, which shows the last exit code and reason, and its logs are available via kubectl logs <pod> -c <init-container-name>. This directly matches the init error status shown in the question stem.

Why this answer

The 'Init:Error' status indicates that a pod's init container has failed to complete successfully. Init containers run sequentially before any main containers start, and if one exits with a non-zero exit code, the pod enters this error state. This is distinct from a main container crash, which would show as 'CrashLoopBackOff' or 'Error' after the pod has started.

Exam trap

The trap here is that candidates confuse 'Init:Error' with a pod initialization phase or a main container error, when in fact it specifically indicates a failed init container that prevents the pod from reaching the running state.

Why the other options are wrong

A

Main container status would be CrashLoopBackOff or Error.

C

That would be Init:0/1 etc.

D

Network error would show as Init:NetworkNotReady or similar.

136
MCQeasy

You run 'kubectl get pods' and one pod shows 'ImagePullBackOff'. Which command would help you diagnose the issue?

A.kubectl logs <pod-name>
B.kubectl top pod <pod-name>
C.kubectl describe pod <pod-name>
D.kubectl exec -it <pod-name> -- sh
AnswerC

This command retrieves detailed configuration and lifecycle status of the pod, including its recent event log. The 'Events' section at the bottom of the output will explicitly detail the failure reason, such as a 'Failed to pull image' error due to an incorrect tag, network issue, or missing registry credentials.

Why this answer

The 'ImagePullBackOff' error indicates that Kubernetes is unable to pull the container image from the registry. 'kubectl describe pod <pod-name>' provides detailed pod events, including the exact error message from the kubelet (e.g., 'Failed to pull image', 'manifest not found', or 'unauthorized'), which directly reveals the root cause.

Exam trap

CNCF CKA often tests the misconception that 'kubectl logs' can diagnose pre-start failures, but logs only capture output from a running container, not from the image pull phase.

How to eliminate wrong answers

Option A is wrong because 'kubectl logs' retrieves container stdout/stderr, which is only available if the container has started; in ImagePullBackOff, the container never runs, so logs are empty. Option B is wrong because 'kubectl top pod' shows resource usage (CPU/memory) of running pods, which is irrelevant to image pull failures. Option D is wrong because 'kubectl exec' requires a running container to execute commands; in ImagePullBackOff, the container is not running, so exec fails.

137
Multi-Selecthard

Which THREE of the following are valid steps to troubleshoot a DNS issue within a Kubernetes cluster?

Select 3 answers
A.Verify that the kube-dns service has endpoints using 'kubectl get endpoints -n kube-system kube-dns'
B.Check the logs of the CoreDNS pods using 'kubectl logs -n kube-system -l k8s-app=kube-dns'
C.Check the /etc/resolv.conf on the node
D.Restart all nodes to reset DNS settings
E.Run 'kubectl exec -it busybox -- nslookup kubernetes.default'
AnswersA, B, E

This command is a first-line check to confirm the Service named kube-dns actually has backend Pods behind it. In Kubernetes, a Service is only able to route traffic to Pods that match its selector; if the endpoints list is empty, no Pod is available to receive DNS queries, meaning the Service is not routing to any CoreDNS Pod. You would then inspect the CoreDNS Deployment or DaemonSet to determine whether the Pods are unscheduled, crashing, or carry the wrong labels.

Why this answer

Option A is correct because verifying that the kube-dns service has endpoints with 'kubectl get endpoints -n kube-system kube-dns' confirms that the DNS service is properly backed by CoreDNS pods; if the endpoint list is empty, DNS resolution will fail cluster-wide. Option B is correct because inspecting CoreDNS pod logs via 'kubectl logs -n kube-system -l k8s-app=kube-dns' reveals errors such as plugin failures, upstream timeouts, or configuration problems that directly explain DNS resolution issues. Option E is correct because running 'kubectl exec -it busybox -- nslookup kubernetes.default' from inside a pod tests actual in-cluster DNS resolution against the kubernetes.default service, isolating whether the problem is DNS-specific or broader networking.

Option C is not a valid cluster DNS troubleshooting step because /etc/resolv.conf on the node governs the node's own resolver, not the DNS configuration injected into pods (which comes from the pod's dnsPolicy and kubelet settings). Option D is not valid because restarting all nodes is a disruptive, non-targeted action that does not address DNS misconfiguration and is not an accepted troubleshooting procedure.

Exam trap

CNCF often tests the misconception that node-level configuration files like `/etc/resolv.conf` are relevant for cluster-internal DNS troubleshooting, when in fact the issue is almost always within the CoreDNS pods or service endpoints.

138
MCQeasy

You have a pod that is in 'Pending' state. Which command would you run to get more information about why the pod cannot be scheduled?

A.kubectl logs <pod-name>
B.kubectl get events
C.kubectl get pod <pod-name> -o wide
D.kubectl describe pod <pod-name>
AnswerD

This command queries the Kubernetes API server for the complete state of the pod and appends a dedicated Events section at the bottom. This section explicitly details scheduler decisions, such as insufficient CPU/memory, taints and tolerations mismatches, or node selector conflicts that are keeping the pod in a pending state.

Why this answer

`kubectl describe pod <pod-name>` provides detailed information about the pod, including events, conditions, and scheduler-related messages. For a pod stuck in 'Pending', the 'Conditions' and 'Events' sections will reveal scheduling failures such as insufficient resources, node selector mismatches, or taint toleration issues, which are not visible in logs or basic status output.

Exam trap

The trap here is that candidates often confuse `kubectl logs` (which only works for running containers) with troubleshooting a pending pod, or they assume `kubectl get events` is the best tool, but the most targeted and efficient command for a single pod's scheduling issue is `kubectl describe pod`.

How to eliminate wrong answers

Option A is wrong because `kubectl logs` retrieves container logs from a running pod, but a pod in 'Pending' state has not started any containers, so there are no logs to fetch. Option B is wrong because `kubectl get events` shows cluster-wide events, which may include scheduling failures, but it does not filter specifically for the pod in question and can be noisy; it is less direct than describing the pod. Option C is wrong because `kubectl get pod <pod-name> -o wide` only adds node IP and host information, but does not expose the detailed scheduling conditions or error messages needed to diagnose why the pod is pending.

139
MCQhard

You have a pod that is stuck in 'Pending' state. Running 'kubectl describe pod' shows the event: '0/3 nodes are available: 1 node(s) had taint {node-role.kubernetes.io/master: }, that the pod didn't tolerate, 2 node(s) didn't match pod anti-affinity rules.' What is the MOST likely solution?

A.Cordon the master node to remove it from scheduling
B.Remove or modify the pod anti-affinity rules
C.Add a toleration for the master taint to the pod spec
D.Delete and recreate the pod
AnswerB

Pod anti-affinity rules prevent the scheduler from placing a pod on a node that already hosts a matching pod. By removing or relaxing these rules, such as changing a hard requiredDuringSchedulingIgnoredDuringExecution constraint to a soft preferredDuringSchedulingIgnoredDuringExecution rule, the scheduler can utilize the available worker nodes even if they already run similar pods.

Why this answer

The pod is unschedulable because two nodes are excluded by pod anti-affinity rules, not because of the master taint (only one node has that taint). The most direct solution is to remove or modify the anti-affinity rules so the pod can be scheduled on those two nodes. Adding a toleration would only address the single master node, leaving the anti-affinity issue unresolved.

Exam trap

The trap here is that candidates focus on the taint error (because it's a common CKA topic) and overlook the anti-affinity error, which is the actual majority blocker; the CKA exam often tests your ability to prioritize multiple scheduling failures.

How to eliminate wrong answers

Option A is wrong because cordoning the master node would remove it from scheduling entirely, making the situation worse (only 2 nodes would remain, still blocked by anti-affinity). Option C is wrong because the pod already has a toleration for the master taint? No—the event shows the pod did NOT tolerate the master taint, but that only affects one node; the primary blocker is the anti-affinity rules on two nodes. Option D is wrong because deleting and recreating the pod without changing the anti-affinity rules would result in the same 'Pending' state.

140
MCQmedium

You run 'kubectl logs my-pod -c my-container' and get no output, but you know the container produces logs. What should you do next to see previous container logs?

A.kubectl logs my-pod -c my-container -f
B.kubectl logs my-pod -c my-container --previous
C.kubectl logs my-pod -c my-container --tail=50
D.kubectl logs my-pod -c my-container --all-containers
AnswerB

`--previous` (or `-p`) is the correct flag because it instructs kubectl to retrieve the log file of the last terminated container instance for the named container in the same pod. When a container crashes and restarts, the kubelet retains the previous container's log file alongside the current one. By combining `-c my-container` with `--previous`, you specifically target the earlier incarnation that produced output, bypassing the currently empty log stream.

Why this answer

The --previous flag retrieves logs from the previous instance of a crashed container.

141
MCQeasy

You have a Pod that is stuck in Pending state. Which command should you use to get detailed information about why the Pod is not running?

A.kubectl logs <pod-name>
B.kubectl describe pod <pod-name>
C.kubectl exec <pod-name> -- /bin/sh
D.kubectl get pod <pod-name>
AnswerB

This command queries the Kubernetes API server for the complete state of the Pod resource, rendering its configuration, current conditions, and controller events. The "Events" section at the bottom of the output is critical for troubleshooting Pending Pods, as it reveals scheduler decisions, such as insufficient CPU/memory, node taints, or missing PersistentVolumeClaims. This makes it the primary tool for diagnosing scheduling failures.

Why this answer

`kubectl describe pod <pod-name>` provides detailed event logs, status conditions, and resource constraints (e.g., insufficient CPU/memory, persistent volume claims pending, node selector mismatches) that explain why the Pod is stuck in Pending state. The Pending state indicates the Pod has been accepted by the API server but not yet scheduled or started, and `describe` surfaces the exact scheduler or admission controller failures.

Exam trap

Candidates often mistakenly think that `kubectl logs` can diagnose startup failures, but logs only exist for running containers, whereas `kubectl describe` reveals pre-scheduling and admission issues that cause the Pending state.

How to eliminate wrong answers

Option A is wrong because `kubectl logs` retrieves container stdout/stderr, which is only available after the Pod has started running; a Pending Pod has no running containers to produce logs. Option C is wrong because `kubectl exec` requires a running container to execute commands, which does not exist in a Pending Pod. Option D is wrong because `kubectl get pod` only shows the current status (e.g., Pending) without the underlying reasons, such as failed volume mounts or taint toleration issues.

142
MCQhard

A Service of type LoadBalancer is created but the EXTERNAL-IP remains <pending>. The cluster is running on-premises without a cloud load balancer integration. Which of the following is the most likely reason?

A.The cluster has no default storage class.
B.The nodes are not reachable from the internet.
C.No load balancer controller (e.g., MetalLB) is installed.
D.The Service selector does not match any pods.
AnswerC

In on-premises clusters without a cloud provider integration, the Service controller has no built-in mechanism to allocate an external IP. A load balancer controller such as MetalLB is required to watch for Services of type LoadBalancer and update their status with an IP from its address pool. Until that controller runs, the external IP is stuck in <pending>, which is exactly the symptom shown.

Why this answer

A Service of type LoadBalancer in Kubernetes requires an external load balancer controller to provision an external IP address. In on-premises clusters without cloud integration, no such controller exists by default, so the EXTERNAL-IP remains <pending> until a bare-metal load balancer like MetalLB is installed and configured. Option C is correct because without a load balancer controller, Kubernetes cannot assign an external IP.

Exam trap

The trap here is that candidates may confuse the EXTERNAL-IP <pending> state with networking issues (Option B) or pod connectivity (Option D), when the root cause is the absence of a load balancer controller, a concept specific to on-premises Kubernetes deployments.

Why the other options are wrong

A

Storage class is unrelated to LoadBalancer IP assignment.

B

Even if nodes are unreachable, a cloud LB would still assign an IP; on-premises, the LB controller would assign an IP from a pool.

D

A mismatched selector would result in no endpoints, but the external IP could still be assigned (pending also occurs if there's no controller).

143
MCQmedium

You attempt to schedule a pod but it remains 'Pending'. 'kubectl describe pod' shows the event: '0/3 nodes are available: 3 node(s) didn't match node selector.' What is the MOST likely cause?

A.A PersistentVolumeClaim is not bound
B.All nodes have insufficient memory or CPU
C.The nodes have taints that the pod does not tolerate
D.The pod's nodeSelector does not match any node labels
AnswerD

A nodeSelector in the pod spec requires the selected node to have all the specified label key-value pairs. If no node in the cluster carries those labels, the scheduler marks those nodes as failing the `node selector` predicate and the pod remains Pending with an event like `0/N nodes are available: N node(s) didn't match node selector`. This is the canonical cause of a pending pod when there are no volume, resource, or taint issues, and it is confirmed by checking node labels with `kubectl get nodes --show-labels`.

Why this answer

The event '0/3 nodes are available: 3 node(s) didn't match node selector' explicitly indicates that the pod's nodeSelector field specifies labels that no node in the cluster possesses. The scheduler filters nodes based on the nodeSelector, and since none match, the pod remains Pending. The fix is to either label a node to match the selector or correct the pod's nodeSelector.

Exam trap

The trap is confusing node selector mismatches with taint/toleration or resource issues; candidates must read the exact scheduler event, because each failure mode produces a distinct message.

How to eliminate wrong answers

Option A is wrong because an unbound PersistentVolumeClaim would produce a different event such as 'pod has unbound immediate PersistentVolumeClaims' or 'waiting for first consumer to be created', not a node selector mismatch. Option B is wrong because insufficient memory or CPU would generate events like 'Insufficient cpu' or 'Insufficient memory' in the scheduler's failure message, not 'didn't match node selector'. Option C is wrong because taint/toleration issues produce events like 'node(s) had taint {key: value}, that the pod didn't tolerate', which is distinct from a node selector mismatch.

144
MCQhard

You are troubleshooting a connectivity issue between two pods in the same cluster. Pod A cannot reach Pod B's IP. Which command would you run from Pod A to test connectivity?

A.kubectl exec pod-a -- curl <pod-b-ip>
B.kubectl top pod pod-b
C.kubectl logs pod-b
D.kubectl describe pod-b
AnswerA

This command initiates an interactive execution context within the container of pod-a, allowing you to run curl directly from its network namespace. By targeting the IP address of pod-b, you can verify if the Kubernetes network plugin (CNI) is correctly routing traffic between these specific endpoints. It is the most direct way to validate data-plane connectivity from the source pod's perspective.

Why this answer

`kubectl exec pod-a -- curl <pod-b-ip>` runs the `curl` command inside Pod A, directly testing Layer 3 connectivity to Pod B's IP address. This bypasses any service or DNS abstraction and verifies whether the pod network (e.g., CNI plugin like Calico or Flannel) is functioning correctly between the two pods.

Exam trap

The trap here is that candidates may choose `kubectl describe pod-b` (option D) thinking it shows connectivity status, but it only displays static pod metadata and does not perform any active network test.

How to eliminate wrong answers

Option B is wrong because `kubectl top pod pod-b` shows resource usage (CPU/memory) of Pod B, not network connectivity. Option C is wrong because `kubectl logs pod-b` displays Pod B's container logs, which may show application errors but cannot test reachability from Pod A. Option D is wrong because `kubectl describe pod-b` provides metadata and status of Pod B, such as IP and labels, but does not execute any network probe to verify connectivity.

145
MCQhard

You are troubleshooting a DNS issue. From within a pod, you run 'nslookup kubernetes.default.svc.cluster.local' and get 'connection timed out; no servers could be reached'. What is the most likely cause?

A.The pod's /etc/resolv.conf has incorrect nameservers
B.The node's network plugin is misconfigured
C.The pod's DNS policy is set to 'None'
D.The kube-dns service is not running or is misconfigured
AnswerD

The `kube-dns` (or `CoreDNS`) service is the designated DNS resolver for pods within a Kubernetes cluster, with pods' `/etc/resolv.conf` typically pointing to its ClusterIP. If the underlying `kube-dns` or `CoreDNS` pods are not running, are crashing, or are misconfigured (e.g., resource starvation, incorrect upstream servers), the DNS service IP will be unresponsive to queries. This directly causes DNS resolution attempts from client pods to time out, as queries are sent to the correct IP but receive no response from the non-functional or overloaded DNS server.

Why this answer

The error 'connection timed out; no servers could be reached' from nslookup indicates that the DNS resolver (typically the kube-dns or CoreDNS service) is unreachable. Since the query targets the standard Kubernetes service name 'kubernetes.default.svc.cluster.local', the most likely cause is that the kube-dns service (or its backend pods) is not running or is misconfigured, preventing the pod from resolving cluster-internal DNS names.

Exam trap

The trap here is that candidates confuse DNS resolution failures with network plugin issues, but the specific 'connection timed out' error points to the DNS service itself being unreachable, not to a general network misconfiguration.

How to eliminate wrong answers

Option A is wrong because if the pod's /etc/resolv.conf had incorrect nameservers, the error would typically be 'server can't find ...' or 'no answer', not a connection timeout; a timeout suggests the DNS server IP is unreachable, not that it's misconfigured. Option B is wrong because a misconfigured node network plugin would cause broader connectivity issues (e.g., pod-to-pod or pod-to-service failures) rather than a DNS-specific timeout; DNS relies on the network plugin only for basic IP reachability, not for DNS resolution logic. Option C is wrong because setting the pod's DNS policy to 'None' would result in an empty /etc/resolv.conf, leading to an immediate 'no servers could be reached' or 'failure: no nameservers' error, not a timeout after attempting to reach servers.

146
MCQmedium

A pod is in the 'Pending' state for a long time. You run 'kubectl describe pod pending-pod' and see the event: '0/4 nodes are available: 1 node(s) had taint {node.kubernetes.io/not-ready: }, 3 node(s) had taint {node-role.kubernetes.io/control-plane: } that the pod didn't tolerate.' What is the MOST likely solution?

A.Remove the taint from the control-plane nodes
B.Delete the pod and recreate it
C.Increase the pod's resource requests
D.Add tolerations to the pod for the control-plane taint
AnswerD

Adding tolerations to the pod's manifest is the correct solution because taints repel pods unless those pods have a matching toleration. Control-plane nodes are typically tainted to prevent general workloads from running on them. By adding a toleration that matches the control-plane node's taint (e.g., `key: node-role.kubernetes.io/control-plane`, `operator: Exists`, `effect: NoSchedule`), the pod explicitly signals to the scheduler that it is permitted to be scheduled on such nodes, resolving the `Pending` state.

Why this answer

The pod is stuck in 'Pending' because it cannot be scheduled. The event shows that 3 control-plane nodes have the 'node-role.kubernetes.io/control-plane' taint, which by default prevents non-tolerant pods from scheduling on them. Adding the corresponding toleration to the pod's spec allows it to be scheduled on those nodes, resolving the pending state.

Exam trap

The trap here is that candidates often confuse taints/tolerations with node affinity or resource constraints, leading them to choose resource-related fixes or node modifications instead of adding the correct toleration to the pod spec.

How to eliminate wrong answers

Option A is wrong because removing the taint from control-plane nodes would allow all pods to schedule on them, which is not the intended solution for a specific pod and could compromise cluster security or workload isolation. Option B is wrong because deleting and recreating the pod without changing its configuration will result in the same scheduling failure, as the taint and lack of toleration remain unchanged. Option C is wrong because increasing resource requests does not address the taint-based scheduling restriction; resource constraints would produce a different event message (e.g., 'Insufficient cpu/memory').

147
MCQhard

A pod is in Pending state. You run 'kubectl describe pod pending-pod' and see an event: '0/3 nodes are available: 3 Insufficient memory'. However, you believe there is enough memory across the cluster. What could be the issue?

A.The pod's memory request is higher than any node's allocatable memory
B.The cluster is using a resource quota that is exhausted
C.The pod's memory limit is set too low
D.The nodes have taints that the pod does not tolerate
AnswerA

The pod's memory request is higher than any node's allocatable memory. The Kubernetes scheduler performs a feasibility check for each node, comparing the pod's sum of memory requests against the node's allocatable memory (which excludes reserved system resources). If no node can satisfy this request, the scheduler cannot bind the pod, leaving it in Pending state. The `kubectl describe` output would include events such as "0/3 nodes are available: insufficient memory" or "Fit failed" for all nodes, directly indicating that the request exceeds every node's capacity.

Why this answer

The '0/3 nodes are available: 3 Insufficient memory' event indicates that the scheduler could not place the pod because each node lacks enough allocatable memory to satisfy the pod's memory request. Even if the cluster has plenty of total memory, the scheduler evaluates each node individually against the pod's resource requests, not the cluster-wide sum. Therefore, if the pod's memory request exceeds the allocatable memory on every node, the pod will remain Pending.

Exam trap

The trap here is that candidates confuse cluster-wide total memory with per-node allocatable memory, assuming that if the sum of free memory across all nodes is sufficient, the pod should schedule — but the scheduler only considers individual node capacity, not aggregated cluster memory.

How to eliminate wrong answers

Option B is wrong because a resource quota limits total resource consumption within a namespace, but the scheduler error specifically says 'Insufficient memory' on nodes, not a quota violation (which would show a different event like 'exceeded quota'). Option C is wrong because a memory limit that is set too low does not prevent scheduling; limits are enforced at runtime by the kubelet, not by the scheduler, and a low limit would cause OOM kills, not a Pending state. Option D is wrong because taints and tolerations produce a different scheduler event: '0/3 nodes are available: 3 node(s) had taint {key: value} that the pod didn't tolerate', not an 'Insufficient memory' message.

148
Multi-Selectmedium

Which TWO of the following are correct methods to check the health of the kube-apiserver?

Select 2 answers
A.Run 'kubectl top nodes'
B.Run 'systemctl status kube-apiserver'
C.Run 'kubectl get pods -n kube-system -l component=kube-apiserver'
D.Run 'journalctl -u kubelet'
E.Run 'curl -k https://localhost:6443/healthz'
AnswersC, E

The kube-apiserver runs as a static pod in the kube-system namespace, labelled component=kube-apiserver on kubeadm clusters, so listing pods with that label selector reveals its phase and restart count. This checks control-plane health through the API itself.

Why this answer

Option C is correct because kube-apiserver typically runs as a static pod in the kube-system namespace, so 'kubectl get pods -n kube-system -l component=kube-apiserver' lets you verify the pod's status, readiness, and restart count. Option E is correct because the kube-apiserver exposes a /healthz endpoint on its secure port (6443 by default), and 'curl -k https://localhost:6443/healthz' returns 'ok' when the API server is healthy. Option A is wrong because 'kubectl top nodes' reports CPU and memory usage from metrics-server, not API server health.

Option B is wrong because kube-apiserver is usually a static pod managed by the kubelet, not a systemd unit, so 'systemctl status kube-apiserver' would typically fail. Option D is wrong because 'journalctl -u kubelet' shows kubelet logs, which may reveal pod issues but does not directly check the kube-apiserver's health.

Exam trap

On the CKA exam, remember that control plane components (kube-apiserver, kube-controller-manager, kube-scheduler, etcd) in a kubeadm cluster run as static pods. Do not attempt to manage or check them using `systemctl`. Only the `kubelet` and container runtime (e.g., `containerd`) run as systemd services on the nodes.

149
Multi-Selectmedium

Which TWO of the following are valid commands to view cluster events sorted by timestamp?

Select 2 answers
A.kubectl get events
B.kubectl get events --sort-by=.metadata.creationTimestamp
C.kubectl get events -w
D.kubectl get events --sort-by=.metadata.name
E.kubectl get events --all-namespaces
AnswersA, B

kubectl get events is correct because the default output of the events command is already sorted by lastTimestamp, the moment each event was last observed, from most to least recent. This gives an effective chronological view without needing extra flags, satisfying the requirement to view events sorted by a time field.

Why this answer

Options A and B are correct. 'kubectl get events' shows events sorted by last timestamp by default, which satisfies the requirement. 'kubectl get events --sort-by=.metadata.creationTimestamp' explicitly sorts by creation timestamp, also valid. Option C uses -w to watch, not sort. Option D sorts by name, not timestamp.

Option E shows events from all namespaces but does not sort by timestamp.

150
MCQmedium

A node in your cluster is in the 'NotReady' state. You SSH into the node and run 'systemctl status kubelet' which shows the kubelet is active but not functioning correctly. Which command should you use to get detailed logs to troubleshoot the kubelet?

A.kubectl describe node <node-name>
B.journalctl -u kubelet
C.kubectl logs kubelet -n kube-system
D.systemctl restart kubelet
AnswerB

Because the kubelet typically runs as a systemd service on the host operating system rather than as a containerized pod, its logs are managed by systemd-journald. Running `journalctl -u kubelet` allows you to view the service's standard output and error streams directly on the node. This is the standard method for diagnosing startup failures, certificate issues, or connection timeouts to the control plane.

Why this answer

`journalctl -u kubelet` retrieves the systemd journal logs specifically for the kubelet service, which is the standard way to access detailed, timestamped logs when the kubelet is running but malfunctioning. Since the kubelet is active (not stopped), its logs are captured by systemd and can be inspected without restarting the service, preserving the current state for troubleshooting.

Exam trap

The trap here is that candidates confuse the kubelet (a systemd service) with a Kubernetes pod and incorrectly choose `kubectl logs`, not realizing that the kubelet is not managed by the Kubernetes API server and its logs must be accessed via the node's system journal.

How to eliminate wrong answers

Option A is wrong because `kubectl describe node` shows cluster-level node status and conditions from the control plane's perspective, not the kubelet's detailed logs; it relies on the node being reachable and the kubelet reporting, which may be stale or incomplete when the node is NotReady. Option C is wrong because `kubectl logs` retrieves container logs from a pod, not the kubelet binary's logs; the kubelet runs as a systemd service, not as a pod in the kube-system namespace, so this command would fail or return unrelated output. Option D is wrong because `systemctl restart kubelet` restarts the service, which may temporarily fix the issue but destroys the current log buffer and prevents diagnosis of the root cause; it is a recovery action, not a diagnostic command.

← PreviousPage 2 of 3 · 207 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Troubleshooting questions.