Courseiva

CCNA Troubleshooting Questions

75 of 207 questions · Page 1/3 · Troubleshooting · Answers revealed

1
MCQeasy

Which command should you use to view the logs of a container that has previously crashed in a Pod?

A.kubectl logs <pod-name>
B.kubectl describe pod <pod-name>
C.kubectl logs <pod-name> -c <container-name>
D.kubectl logs <pod-name> --previous
AnswerD

`--previous` retrieves logs from the container's prior terminated instance, not the current one. This directly satisfies the stem's constraint of a container that has already crashed and restarted, since the default `kubectl logs` reads only the running instance's output.

Why this answer

`kubectl logs <pod-name> --previous` retrieves the logs from the previous instance of a container in a Pod, which is exactly what you need when a container has crashed and restarted. The `--previous` flag accesses the logs of the terminated (crashed) container, not the current running one, allowing you to see the error that caused the crash.

Exam trap

The trap here is that candidates often assume `kubectl logs <pod-name>` alone will show crash logs, but it only shows the current container's logs, so they miss the `--previous` flag required for accessing logs from a terminated container.

How to eliminate wrong answers

Option A is wrong because `kubectl logs <pod-name>` only shows logs from the currently running container; if the container has crashed and restarted, the logs from the crash are lost from the current instance. Option B is wrong because `kubectl describe pod <pod-name>` shows the Pod's metadata, status, and events (including crash loop backoff details), but it does not display the container's log output. Option C is wrong because `kubectl logs <pod-name> -c <container-name>` is used to specify a container name when a Pod has multiple containers, but it still only shows logs from the current (running) container, not the previous crashed one.

2
MCQhard

A user reports that their application cannot resolve DNS names for services in the cluster. The application runs in a pod with dnsPolicy: ClusterFirst. What is the most likely cause?

A.The CoreDNS deployment has 0 ready replicas.
B.The pod's dnsPolicy is set to Default instead of ClusterFirst.
C.The node's network plugin is misconfigured, blocking UDP port 53.
D.The pod's /etc/resolv.conf contains incorrect nameserver entries.
AnswerA

CoreDNS is the default cluster DNS provider in Kubernetes, responsible for resolving internal service names and external domains. If the CoreDNS deployment has zero ready replicas, there are no active pods to handle DNS queries sent to the kube-dns service IP. Consequently, any pod attempting to resolve a DNS name will experience a timeout or resolution failure.

Why this answer

When dnsPolicy is ClusterFirst, the pod's DNS queries are forwarded to the cluster's DNS service (CoreDNS by default). If the CoreDNS deployment has 0 ready replicas, the DNS service has no backend endpoints to handle queries, causing all DNS resolutions to fail. This is the most direct and common cause of complete DNS failure in a cluster.

Exam trap

The trap here is that candidates may overthink network-level issues (like UDP port blocking) or misread the dnsPolicy, when the simplest and most common cause is that the DNS service itself (CoreDNS) is not running.

How to eliminate wrong answers

Option B is wrong because the pod's dnsPolicy is already set to ClusterFirst (as stated in the question), so suggesting it is set to Default is factually incorrect and would not explain the failure. Option C is wrong because while a misconfigured network plugin blocking UDP port 53 could cause DNS issues, it is less likely than CoreDNS being down, and the question asks for the 'most likely' cause; also, CoreDNS itself listens on port 53, so if it has 0 replicas, the port is irrelevant. Option D is wrong because with dnsPolicy: ClusterFirst, the pod's /etc/resolv.conf is automatically generated by kubelet to point to the cluster DNS service IP (e.g., 10.96.0.10), and incorrect entries would only occur if the policy were Default or if the kubelet configuration is broken, which is less common than CoreDNS being unavailable.

3
MCQeasy

You need to see the startup logs of the kubelet service. Which command should you use?

A.kubectl get events --all-namespaces
B.systemctl status kubelet
C.journalctl -u kubelet
D.kubectl logs kubelet -n kube-system
AnswerC

Because the kubelet runs as a native systemd service on the control plane and worker nodes, its standard output and error streams are captured by systemd-journald. Using this command queries the journal daemon specifically for the `kubelet` unit, outputting its complete, chronological startup and runtime log history.

Why this answer

`journalctl -u kubelet` retrieves the systemd journal logs for the kubelet service unit. Since kubelet runs as a systemd service on the node, its startup logs (including failures, configuration errors, or TLS bootstrap issues) are captured by journald and can be viewed with this command. This is the standard way to inspect kubelet's early boot-time behavior, which is not accessible via kubectl or systemctl status alone.

Exam trap

The trap here is that candidates confuse the kubelet with a Kubernetes pod and try to use `kubectl logs`, forgetting that kubelet is a node-level systemd service, not a container managed by the API server.

How to eliminate wrong answers

Option A is wrong because `kubectl get events --all-namespaces` shows Kubernetes API server events (e.g., pod scheduling, node conditions), not the kubelet service's own startup logs. Option B is wrong because `systemctl status kubelet` shows only the current status and the last few log lines of the kubelet service, not the full startup log history. Option D is wrong because `kubectl logs kubelet -n kube-system` attempts to fetch container logs from a pod named 'kubelet' in kube-system, but kubelet is not a Kubernetes pod; it runs as a systemd service on the node, so this command will fail with an error like 'Error from server: pods "kubelet" not found'.

4
MCQmedium

A node in the cluster is reporting NotReady. You run 'journalctl -u kubelet' and see repeated errors: 'failed to get system container stats for "/system.slice/kubelet.service": failed to get cgroup stats'. What is the most likely cause?

A.Disk pressure on the node
B.The node is out of memory
C.Cgroup driver mismatch between kubelet and container runtime
D.The kubelet service is not running
AnswerC

This error occurs when the kubelet is configured to use one cgroup driver (such as systemd) while the container runtime (like containerd) is configured to use another (such as cgroupfs). Because both systems must agree on how to manage cgroup hierarchies, this mismatch prevents the kubelet from registering the node and managing container resources properly, leading to a NotReady status. Aligning both configurations to use the same driver, preferably systemd on modern Linux distributions, resolves the issue.

Why this answer

The error 'failed to get cgroup stats' indicates that the kubelet cannot read container cgroup statistics, which typically occurs when the cgroup driver used by kubelet (e.g., cgroupfs) does not match the driver used by the container runtime (e.g., systemd). This mismatch prevents the kubelet from properly monitoring container resource usage, causing the node to report NotReady.

Exam trap

The trap here is that candidates may misinterpret the cgroup stats error as a resource pressure issue (disk or memory) because 'stats' sounds like resource monitoring, but the actual root cause is a configuration mismatch between kubelet and the container runtime.

How to eliminate wrong answers

Option A is wrong because disk pressure would manifest as 'NodeHasDiskPressure' condition or 'eviction manager' warnings in kubelet logs, not a cgroup stats failure. Option B is wrong because out-of-memory conditions typically cause OOM kills or 'NodeHasMemoryPressure' condition, not a cgroup stats retrieval error. Option D is wrong because if the kubelet service were not running, 'journalctl -u kubelet' would show no logs or a 'unit not found' error, not repeated cgroup stats failures.

5
Multi-Selecthard

You suspect a DNS issue within the cluster. Which TWO commands can you run from within a pod to test DNS resolution?

Select 2 answers
A.ping kubernetes.default.svc.cluster.local
B.kubectl exec -it pod-name -- /bin/sh
C.dig kubernetes.default.svc.cluster.local
D.curl http://kubernetes.default.svc.cluster.local
E.nslookup kubernetes.default.svc.cluster.local
AnswersC, E

dig is a dedicated DNS lookup utility that sends explicit DNS queries to the configured resolver and prints the full DNS response, including the resolved A/AAAA record and the answer section. When invoked from inside a pod, it directly interrogates the cluster DNS service (usually CoreDNS) for kubernetes.default.svc.cluster.local, so a successful answer proves that DNS resolution works end-to-end. It also distinguishes error types such as NXDOMAIN versus SERVFAIL, which makes it highly effective for isolating DNS-specific misconfigurations.

Why this answer

Options C and E are correct because both dig and nslookup are dedicated DNS query tools that directly resolve the name kubernetes.default.svc.cluster.local against the cluster DNS service (CoreDNS/kube-dns), letting you verify the returned A/AAAA records and diagnose resolution failures from inside the pod. dig queries the configured resolver and shows the ANSWER SECTION, while nslookup performs an equivalent name-resolution lookup, so either one specifically tests DNS rather than general connectivity. Option A (ping) is not a DNS test per se; it only triggers resolution as a side effect and may fail due to ICMP being blocked even when DNS works. Option B (kubectl exec -it pod-name -- /bin/sh) merely opens a shell in the pod and does not itself test DNS resolution.

Option D (curl http://kubernetes.default.svc.cluster.local) tests HTTP connectivity to the API server, not DNS, and can fail for reasons unrelated to name resolution.

6
MCQmedium

You need to test DNS resolution from within a pod. Which command should you run?

A.kubectl describe svc <service-name>
B.kubectl logs <pod> | grep dns
C.kubectl exec <pod> -- nslookup <service-name>
D.kubectl run nslookup --image=busybox -- nslookup <service-name>
AnswerC

This command uses kubectl exec to run the nslookup utility directly inside the container of an existing pod, utilizing its specific network namespace and /etc/resolv.conf configuration. By querying the target service name from within the pod, you can accurately verify if the pod's configured DNS resolver (typically CoreDNS) is functioning and resolving names correctly. This is the standard method for diagnosing pod-level DNS resolution issues in a Kubernetes cluster.

Why this answer

`kubectl exec <pod> -- nslookup <service-name>` runs the `nslookup` command directly inside an existing pod, which tests DNS resolution from within the pod's network namespace. This is the most direct way to verify that the pod's DNS resolver (typically CoreDNS or kube-dns) can resolve a Kubernetes service name to a cluster IP, as it uses the pod's `/etc/resolv.conf` and the cluster DNS service.

Exam trap

The trap here is that candidates may choose Option D because it seems like a quick way to run a DNS test without needing an existing pod, but they overlook that creating a new pod can introduce different DNS settings (e.g., default `dnsPolicy: ClusterFirst` vs. a pod with `dnsPolicy: Default`) and that the pod must be cleaned up manually, whereas Option C directly tests the actual pod's DNS resolution without side effects.

How to eliminate wrong answers

Option A is wrong because `kubectl describe svc <service-name>` shows the service's metadata and endpoints, but does not perform any DNS resolution or test connectivity from within a pod. Option B is wrong because `kubectl logs <pod> | grep dns` only searches the pod's container logs for the string 'dns', which is not a standard DNS test and does not trigger or verify DNS resolution. Option D is wrong because `kubectl run nslookup --image=busybox -- nslookup <service-name>` creates a new pod to run nslookup, but it does not specify the `--rm` flag (to clean up) and, more critically, it runs in a separate pod that may have different DNS configuration or network policies, making it less reliable than using an existing pod that is already part of the application's network context.

7
Multi-Selectmedium

A pod is stuck in 'Pending' state. Which TWO of the following are common causes?

Select 2 answers
A.The service account does not exist
B.A taint on the node that the pod does not tolerate
C.The pod's liveness probe is failing
D.Insufficient CPU or memory resources on any node
E.The pod's container image does not exist
AnswersB, D

Taints that are not tolerated prevent scheduling, causing Pending.

Why this answer

A pod enters 'Pending' state when it cannot be scheduled onto a node. Taints on a node with `NoSchedule` or `NoExecute` effects prevent pods that do not have matching tolerations from being scheduled there. This is a common cause because the scheduler skips tainted nodes unless the pod explicitly tolerates the taint, leaving the pod unscheduled and stuck in Pending.

Exam trap

The CKA exam often tests the distinction between pre-scheduling failures (Pending) and post-scheduling failures (CrashLoopBackOff, ImagePullBackOff), so candidates mistakenly select image or probe issues that occur after the pod is running.

8
MCQmedium

You are troubleshooting a pod that is failing to start due to a volume mount error. The pod spec references a PersistentVolumeClaim (PVC) named 'data-pvc'. You run 'kubectl get pvc data-pvc -n default' and see the status is 'Pending'. Which of the following is the MOST likely cause?

A.The PVC is not in the same namespace as the pod that is trying to use it.
B.The PVC's storage class does not exist, so the PVC cannot be provisioned.
C.There is no available PersistentVolume that matches the PVC's storage class, access mode, and capacity requirements.
D.The PVC is waiting for a pod to be scheduled before it can bind to a PersistentVolume.
AnswerC

A PVC remains Pending when the control plane cannot find a PV that satisfies its request. This could be because no PV exists, or existing PVs do not match the storage class, access modes, or capacity. If the PVC uses a storage class with a dynamic provisioner, the provisioner might be failing to create a volume. Checking the PVC events and the storage class configuration is necessary.

Why this answer

A PersistentVolumeClaim remains in Pending state when the Kubernetes control plane cannot find or provision a matching PersistentVolume. This typically happens when no PV meets the storage class, access mode, and capacity requirements, or when the dynamic provisioner fails. Investigating the PVC's events with 'kubectl describe pvc' will show the reason, such as 'no persistent volumes available for this claim' or provisioner errors.

Ensuring a suitable PV or a working storage class is the resolution.

Exam trap

The trap here is assuming the PVC is waiting for a pod, when in fact PVC binding is independent of pod scheduling and depends on PV availability or provisioning.

9
Multi-Selecteasy

A pod is in 'ImagePullBackOff' state. Which TWO are valid first troubleshooting steps?

Select 2 answers
A.Check for network policies blocking egress to the registry
B.Check node CPU/memory resources
C.Verify the image exists in the configured registry
D.Check the image name spelling in the pod spec
E.Restart the kubelet on the node
AnswersC, D

ImagePullBackOff almost always follows a kubelet event that includes the exact registry response, such as 'manifest unknown', 'not found', or 'unauthorized'. Verifying the image exists in the configured registry, with the exact tag or digest specified in the pod spec, is the most direct confirmation of the root cause. Use `docker manifest inspect` or `skopeo inspect` against that registry, and remember that images can be deleted or made private after a pod was previously running. This check targets the registry artifact itself.

Why this answer

Option C is correct because ImagePullBackOff commonly results from the referenced image not existing in the registry (wrong tag, deleted image, or private repo requiring credentials), so verifying the image exists in the configured registry is a valid first step. Option D is correct because a simple typo in the image name or tag in the pod spec will cause the kubelet to fail pulling the image, and checking the spelling is a fast, direct troubleshooting action. Option A is not a typical first step since network policies rarely block egress to registries by default and the error would more likely be a timeout or connection refused rather than ImagePullBackOff.

Option B is unrelated because insufficient node CPU/memory produces Pending or Evicted states, not image pull failures. Option E is not a first step because restarting the kubelet is disruptive and does not address the root cause of an image pull error.

10
MCQmedium

A pod is in ImagePullBackOff state. You run 'kubectl describe pod mypod' and see 'Failed to pull image "myapp:latest": rpc error: code = Unknown desc = Error response from daemon: manifest for myapp:latest not found: manifest unknown'. What is the most likely cause?

A.The image tag does not exist in the registry
B.The node has insufficient disk space
C.The image is too large and exceeds the node's disk quota
D.The container registry requires authentication
AnswerA

The 'manifest unknown' error in the kubectl describe pod event is the definitive signature: the registry does not have a manifest for the referenced tag, so either the tag was never pushed, was deleted, or contains a typo. Unlike authentication or storage failures, this is a content-addressing issue: the digest cannot be resolved because that tag's manifest is absent. Confirm by inspecting the image name and tag exactly as written and, if needed, querying the registry API directly.

Why this answer

The error indicates the image tag 'latest' does not exist in the registry. The tag may have been deleted or never pushed.

11
MCQhard

A Pod is running but cannot connect to a Service. You have verified that the Service endpoints are correct. Which of the following is the most likely cause if the Pod is using hostNetwork: true?

A.The kube-proxy is not running on the node
B.The Service is not defined correctly
C.The container image is missing networking tools
D.The Pod uses hostNetwork and cannot resolve the ClusterIP due to DNS configuration
AnswerD

When a Pod is configured with `hostNetwork: true`, it directly uses the node's network namespace, bypassing the Kubernetes CNI network. Consequently, such a pod inherits the node's `/etc/resolv.conf` for DNS resolution instead of the cluster's internal DNS (CoreDNS/kube-dns). The node's DNS resolver typically cannot resolve Kubernetes Service ClusterIPs, which are internal to the cluster's DNS domain, leading to name resolution failures for services. This specific DNS misconfiguration prevents the `hostNetwork` pod from finding the Service's IP address.

Why this answer

When a Pod uses `hostNetwork: true`, it shares the node's network namespace and directly uses the host's network stack. ClusterIP Services are virtual IPs managed by iptables or IPVS rules on the node, but these rules are typically applied only to the host's network namespace. However, the most common issue is that the Pod's DNS resolver (e.g., `/etc/resolv.conf`) is configured to resolve the Service name via the cluster's DNS (CoreDNS/kube-dns), which returns a ClusterIP.

Since the Pod is on the host network, it may not have the necessary iptables rules to route traffic to the ClusterIP, or the DNS configuration may point to a DNS server that is not reachable from the host network (e.g., the cluster DNS service IP itself). Option D correctly identifies that the Pod cannot resolve the ClusterIP due to DNS configuration, as the Pod's DNS settings are inherited from the node but may not include the cluster DNS server, or the cluster DNS is not accessible from the host network.

Exam trap

The trap here is that candidates assume `hostNetwork: true` gives the Pod full access to all cluster services, but they overlook that DNS resolution for ClusterIP Services depends on the cluster DNS being reachable and properly configured in the Pod's resolv.conf, which is not automatically set when using hostNetwork.

How to eliminate wrong answers

Option A is wrong because kube-proxy runs on every node and is responsible for implementing Service rules (e.g., iptables/IPVS); if it were not running, no Pod (hostNetwork or not) would reach any Service, but the question states endpoints are correct, implying kube-proxy is functional. Option B is wrong because the question explicitly states that the Service endpoints are correct, meaning the Service definition itself is valid and has healthy endpoints. Option C is wrong because missing networking tools (e.g., curl, ping) would prevent the user from testing connectivity, but the Pod's inability to connect is a network-layer issue, not a tool availability issue; the Pod could still connect via raw sockets or other means if the network path worked.

12
MCQmedium

A pod is in 'ImagePullBackOff' state. You run 'kubectl describe pod' and see the event: 'Failed to pull image "nginx:latest": rpc error: code = Unknown desc = Error response from daemon: Get https://registry-1.docker.io/v2/: dial tcp: lookup registry-1.docker.io on 8.8.8.8:53: no such host'. What is the MOST likely cause?

A.The container runtime is not installed
B.The image tag is incorrect
C.The registry requires authentication
D.DNS resolution is failing on the node
AnswerD

The presence of a 'no such host' or 'lookup failed' error in the pod's event log indicates that the node's local resolver cannot translate the registry's domain name into an IP address. Because the Kubelet cannot resolve the external registry's FQDN, it fails to initiate the TCP handshake, ultimately causing the container pull to fail and triggering the 'ImagePullBackOff' state.

Why this answer

The error message 'no such host' when trying to resolve 'registry-1.docker.io' on DNS server 8.8.8.8:53 indicates that the node's DNS resolution is failing. Kubernetes nodes rely on DNS to resolve container registry hostnames; if the node cannot resolve the registry's domain, the container runtime cannot pull the image, resulting in ImagePullBackOff. This is a node-level DNS issue, not a container runtime or image tag problem.

Exam trap

The trap here is that candidates may confuse DNS resolution failures with authentication or image tag issues, but the specific 'no such host' error in the event message directly points to a DNS problem on the node.

How to eliminate wrong answers

Option A is wrong because the container runtime is clearly working (it returns an error response from the daemon, proving the runtime is installed and running). Option B is wrong because the error is about DNS resolution, not an incorrect image tag; 'nginx:latest' is a valid tag and would not cause a 'no such host' error. Option C is wrong because authentication failures produce different errors (e.g., 'denied: requested access to the resource is denied' or 'unauthorized: authentication required'), not a DNS lookup failure.

13
MCQeasy

You are tasked with troubleshooting a production Kubernetes cluster. A user reports that they cannot access a web application running in the cluster. The application is deployed as a Deployment named 'frontend' with 2 replicas, exposed via a Service of type LoadBalancer. You have kubectl access to the cluster. You run 'kubectl get pods -l app=frontend' and see both pods are Running and Ready. You run 'kubectl get svc frontend' and see the Service has an external IP of 192.168.1.100. However, when you curl http://192.168.1.100 from a machine outside the cluster, you get a connection timeout. You are able to curl the pod IPs directly from within the cluster and get a response. Which of the following is the most likely cause of the issue?

A.The Service selector does not match the pod labels.
B.The cloud provider's load balancer is not properly configured or the security group/firewall is blocking traffic to the node ports.
C.The NodePort service type is not enabled in the cluster.
D.The Ingress resource is missing or misconfigured.
AnswerB

When internal cluster communication works but external traffic times out, the issue lies in the external network path. Misconfigured cloud load balancers or restrictive security groups/firewalls blocking the NodePort range (typically 30000-32767) prevent external packets from reaching the cluster nodes.

Why this answer

The pods are running and ready, and the service has an external IP, but external access fails with a connection timeout while internal access to pod IPs works. This indicates the cloud provider's load balancer is not properly routing traffic to the node ports, or a security group/firewall is blocking inbound traffic on the node port range (30000-32767). The load balancer must forward traffic to the node ports, and the nodes must allow that traffic.

Exam trap

The trap here is that candidates assume a LoadBalancer Service automatically works end-to-end, but the CKA exam tests the understanding that cloud provider integration (security groups, load balancer health checks) is a separate layer that can fail even when Kubernetes components are healthy.

How to eliminate wrong answers

Option A is wrong because if the Service selector did not match the pod labels, the endpoints would be empty and curl to pod IPs would fail, but the user reports internal curl to pod IPs works. Option C is wrong because NodePort is not a service type that needs to be 'enabled'; it is automatically assigned when a Service of type LoadBalancer is created, and the cluster always supports NodePort. Option D is wrong because an Ingress resource is not required for a LoadBalancer Service; the Service itself provides external access via the load balancer, and the issue is a connection timeout, not a routing or hostname mismatch.

14
MCQmedium

An application team reports that a Deployment's Pods in namespace 'web' become unreachable after a rolling update, even though 'kubectl get pods' shows them as Running and Ready. The Service 'web-svc' of type ClusterIP exists and has endpoints listed. You exec into a client Pod in the same namespace and run 'curl http://web-svc:8080'; the connection times out. You then run 'kubectl get networkpolicy -n web' and see a policy named 'deny-all-ingress' with podSelector matching the web Pods and policyTypes: Ingress, but no ingress rules. Which of the following is the MOST likely cause of the timeout?

A.The Service 'web-svc' is missing a selector that matches the Pod labels, so no endpoints are created and traffic cannot be routed.
B.The Pods are not actually Ready because their readiness probe is failing, so the Service removes them from endpoints.
C.The NetworkPolicy 'deny-all-ingress' is selecting the web Pods and, because it has no ingress rules, it denies all ingress traffic to those Pods.
D.The kube-proxy component on the client's node is not programmed with the Service's iptables rules, so the ClusterIP is not reachable.
AnswerC

A NetworkPolicy with policyTypes: Ingress and an empty ingress rule list isolates the selected Pods, denying all inbound traffic except from the Pod's own node. Since the policy selects the web Pods, the client's connection is dropped, causing the timeout. This directly explains the symptom and the presence of the policy.

Why this answer

The presence of a NetworkPolicy named 'deny-all-ingress' that selects the web Pods and has no ingress rules is the definitive cause. Such a policy isolates the Pods, dropping all inbound connections. The Service and endpoints are healthy, so the timeout is due to packet filtering, not routing or readiness.

Removing or modifying the policy restores connectivity.

Exam trap

The trap here is assuming that a NetworkPolicy with no rules is a no-op, when in fact an empty ingress rule list combined with policyTypes: Ingress denies all ingress traffic to the selected Pods.

15
MCQeasy

You run 'kubectl get pods' and see a pod in 'ImagePullBackOff' state. What is the most likely cause?

A.The node is out of disk space
B.The pod exceeded its memory limit
C.The container command is invalid
D.The container image name is misspelled
AnswerD

A misspelled image name or tag prevents the container runtime from locating the manifest in the specified registry. The kubelet first transitions to ErrImagePull, and after repeated failed attempts to download the non-existent image, it backs off to avoid overloading the registry, resulting in the ImagePullBackOff status. This indicates a configuration error in the pod specification.

Why this answer

The 'ImagePullBackOff' state indicates that the kubelet on the node is unable to pull the container image from the registry. The most common cause is an incorrect image name or tag, which results in a manifest not found error. This triggers exponential backoff retries by the kubelet, leading to the 'ImagePullBackOff' status.

Exam trap

The trap here is that candidates confuse 'ImagePullBackOff' with 'CrashLoopBackOff', assuming the container started but failed, when in fact the image was never successfully pulled.

How to eliminate wrong answers

Option A is wrong because a node being out of disk space would typically cause an 'Evicted' or 'NodeDiskPressure' condition, not an 'ImagePullBackOff' error, which is specific to image pull failures. Option B is wrong because exceeding the pod's memory limit results in an 'OOMKilled' state (container terminated with Out Of Memory), not an image pull issue. Option C is wrong because an invalid container command causes a 'CrashLoopBackOff' state after the container starts, not 'ImagePullBackOff', which occurs before the container can run.

16
MCQmedium

You run 'kubectl top nodes' but get 'error: metrics not available yet'. What is the most likely cause?

A.kubectl top requires root access
B.Nodes have insufficient CPU
C.The Metrics Server is not deployed
D.Kubelet is not running on nodes
AnswerC

The Metrics Server is an optional cluster component that aggregates resource usage from kubelets and exposes it through the metrics.k8s.io API, which kubectl top and the Horizontal Pod Autoscaler consume. In a standard kubeadm or vanilla cluster it is not installed by default, so kubectl top nodes fails with 'metrics not available' until the Metrics Server manifests are applied. Deploying the Metrics Server and allowing it time to collect its first metrics resolves the error.

Why this answer

Metrics server is not deployed or not ready, so resource metrics are unavailable.

17
MCQeasy

You are troubleshooting a node that is in 'NotReady' state. Which command should you use to check the kubelet logs for errors?

A.journalctl -u kubelet
B.journalctl -u docker
C.journalctl -u kube-apiserver
D.journalctl -u kube-controller-manager
AnswerA

When a node is in a `NotReady` state, it signifies that the Kubelet, the agent running on the node, is unable to register itself with the API server or report its health status. The Kubelet is responsible for managing pods, reporting node status, and executing container operations. Therefore, examining the Kubelet's logs via `journalctl -u kubelet` provides direct insight into why it might be failing to communicate, experiencing resource issues, or encountering problems with the container runtime, making it the most critical first step for diagnosis.

Why this answer

The kubelet is the primary node agent that communicates with the control plane and manages pods. When a node is in 'NotReady' state, checking the kubelet logs is the first step to diagnose issues such as certificate errors, resource pressure, or network problems. The command `journalctl -u kubelet` retrieves logs from the systemd unit for the kubelet service, which is the correct way to view its output on systems using systemd.

Exam trap

The trap here is that candidates may confuse control plane components (kube-apiserver, kube-controller-manager) with node-level components, or mistakenly think the container runtime (Docker) is the primary source of node readiness logs, when in fact the kubelet is the authoritative agent for node status.

How to eliminate wrong answers

Option B is wrong because `journalctl -u docker` shows logs for the Docker daemon, not the kubelet; while container runtime issues can affect node readiness, the primary component to check for node status is the kubelet. Option C is wrong because `journalctl -u kube-apiserver` retrieves logs from the API server, which runs on the control plane, not on the worker node; this would not help diagnose a node-level issue. Option D is wrong because `journalctl -u kube-controller-manager` shows logs from the controller manager, another control plane component; it does not run on worker nodes and is irrelevant for troubleshooting a node's kubelet.

18
Multi-Selectmedium

Which TWO of the following are valid methods to diagnose why a node is in 'NotReady' state?

Select 2 answers
A.Restart all containers on the node
B.Check the kubelet service logs using 'journalctl -u kubelet'
C.Check the kube-apiserver logs on the control plane
D.Check kube-proxy configuration
E.Verify the network plugin (e.g., Calico, Flannel) pods are running
AnswersB, E

The kubelet is the component that reports Node status and manages pod lifecycle on each node; its logs often contain explicit error messages such as CNI failures, timeouts contacting the API server, or runtime issues. 'journalctl -u kubelet' allows you to view the kubelet's systemd journal, which is the authoritative source for why the kubelet has marked the node as NotReady. This is a direct, non-invasive diagnostic step that should be performed first when investigating node readiness problems.

Why this answer

Option B is correct because the kubelet is the node agent responsible for reporting node status to the API server; if it crashes, fails to register, or cannot reach the API server, the node transitions to NotReady, so inspecting its logs with 'journalctl -u kubelet' directly reveals the root cause. Option E is correct because a NotReady condition is frequently caused by the Container Network Interface (CNI) plugin (Calico, Flannel, etc.) not being ready, which prevents pod networking and triggers the 'NetworkPluginNotReady' condition on the node, so verifying those pods are running is a valid diagnostic step. Option A is not a diagnostic method but a remediation action, and restarting containers indiscriminately does not identify the cause.

Option C is incorrect for this scenario because kube-apiserver logs on the control plane relate to API server issues, not the node's own readiness reporting. Option D is incorrect because kube-proxy handles Service load balancing via iptables/IPVS rules and its misconfiguration does not typically cause a node to report NotReady.

Exam trap

CKA often tests the difference between node-level diagnostics (kubelet logs, CNI pods) and control-plane diagnostics (API server logs), causing candidates to check the wrong component.

19
MCQeasy

You see events like 'Failed to pull image' when running 'kubectl get events'. Which command can you use to get more detailed information about a specific pod's events?

A.kubectl get pod <pod-name> -o yaml
B.kubectl top pod <pod-name>
C.kubectl describe pod <pod-name>
D.kubectl logs <pod-name>
AnswerC

This command aggregates detailed information about the Pod from the Kubernetes API, including its configuration, controller references, and container statuses. Crucially, it appends a dedicated "Events" section at the end of the output, which queries the cluster's event log for warnings and errors directly related to that Pod. This is where image pull failures, back-offs, and scheduling issues are explicitly logged.

Why this answer

C is correct because 'kubectl describe pod <pod-name>' provides a detailed summary of the pod's current state, including a chronological list of events related to that specific pod (e.g., image pull failures, container restarts, scheduling issues). This command surfaces the exact error messages and timestamps for each event, making it the most direct way to investigate a 'Failed to pull image' event.

Exam trap

The trap here is that candidates often confuse 'kubectl logs' (which shows container output) with event inspection, not realizing that a pod that fails to start due to image pull issues never generates container logs, so the error is only visible in the pod's events via 'kubectl describe'.

How to eliminate wrong answers

Option A is wrong because 'kubectl get pod <pod-name> -o yaml' outputs the pod's full YAML manifest and current status, but it does not include the event history; events are stored separately in the cluster and are not part of the pod object's spec or status. Option B is wrong because 'kubectl top pod <pod-name' shows real-time CPU and memory usage metrics, which are irrelevant to image pull failures. Option D is wrong because 'kubectl logs <pod-name>' retrieves the container's stdout/stderr logs, but if the container never started due to a failed image pull, there are no logs to display; the error is captured in the pod's events, not in the container's log stream.

20
MCQmedium

A pod is stuck in Pending state. 'kubectl describe pod' shows '0/1 nodes are available: 1 node(s) had taint {key=value: NoSchedule}, 1 node(s) had taint {node.kubernetes.io/unreachable: }. What does this indicate?

A.The nodes are all in NotReady state
B.The pod does not tolerate the taints on the available nodes
C.The pod requires at least 2 nodes
D.The pod has insufficient resources
AnswerB

When a node has a taint, the scheduler ignores that node for any pod that does not have a matching toleration in its pod spec. The describe output would include an event such as '0/1 nodes are available: 1 node(s) had untolerated taint', which precisely identifies this scheduling failure. To make the pod run, you must either remove the taint from the node or add a toleration that matches the taint key, value, and effect; because neither was done, the pod remains Pending.

Why this answer

The pod cannot tolerate the taints. One node has a custom taint, another is unreachable.

21
MCQmedium

You run 'kubectl get nodes' and one node shows 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. The output shows 'active (running)'. What should you check NEXT?

A.Check if the kubelet is registered with the API server using 'kubectl get cs'
B.Check the kubelet logs with 'journalctl -u kubelet' for errors
C.Restart the kubelet service
D.Check if the node has sufficient disk space
AnswerB

The kubelet is the agent that reports the node's status, so its logs are the authoritative source for why it stopped reporting Ready. 'journalctl -u kubelet' shows systemd-managed kubelet output, including fatal errors such as failure to connect to the container runtime, CNI plugin failures, or authentication problems with the API server. Examining these logs is the first diagnostic step before any corrective action is taken.

Why this answer

Since the kubelet is running but the node is NotReady, the issue is likely that the kubelet cannot communicate with the API server or has encountered an error during its operation. Checking the kubelet logs with 'journalctl -u kubelet' is the next logical step to identify specific errors, such as certificate issues, network connectivity problems, or pod lifecycle failures, which would not be visible from the service status alone.

Exam trap

The trap here is that candidates assume a running service means the kubelet is healthy, but the kubelet can be 'active (running)' while failing to register or sync with the API server, making log inspection essential before any restart or resource check.

How to eliminate wrong answers

Option A is wrong because 'kubectl get cs' (componentstatuses) is deprecated and does not show kubelet registration; it checks the health of control plane components like etcd and scheduler, not node-level kubelet status. Option C is wrong because restarting the kubelet without first diagnosing the root cause is a blind action that could mask transient errors or fail to resolve underlying issues like misconfiguration or resource exhaustion. Option D is wrong because while disk pressure can cause node issues, it typically manifests as a specific condition (e.g., DiskPressure) rather than a generic NotReady state, and the kubelet logs would reveal such conditions more directly.

22
MCQeasy

A pod is in CrashLoopBackOff state. Which command shows the last termination reason?

A.kubectl logs <pod>
B.kubectl describe pod <pod>
C.kubectl get pod <pod> -o yaml
D.kubectl top pod <pod>
AnswerB

kubectl describe pod <pod> aggregates the full object state with a timeline of Events, including the reason, message, and count for failed probes, image pulls, and container starts. It surfaces the container's current waiting reason (e.g., CrashLoopBackOff), last terminated exit code, and the actual error message from the runtime, which together pinpoint why the pod is repeatedly crashing. This is the standard diagnostic first step for a crash loop.

Why this answer

The `kubectl describe pod <pod>` command displays detailed information about the pod, including the container state transitions and the last termination reason in the `Last State` field under the container status. This is the correct way to see why the container previously exited, which is essential for diagnosing a CrashLoopBackOff.

Exam trap

The trap here is that candidates often reach for `kubectl logs` first, not realizing that in a CrashLoopBackOff the container may have already restarted, so the logs from the last crash are not visible without the `--previous` flag, whereas `kubectl describe pod` directly shows the last termination reason.

Why the other options are wrong

A

Shows current logs, not termination reason.

C

Shows state but not human-readable termination reason.

D

Shows resource usage, not status.

23
Multi-Selecteasy

Which TWO of the following are valid commands to check resource usage of pods?

Select 2 answers
A.kubectl top nodes
B.kubectl describe pods
C.kubectl top pod <pod-name>
D.kubectl get pods --watch
E.kubectl top pods
AnswersC, E

kubectl top pod <pod-name> is a valid command that queries the Metrics API (backed by metrics-server) to display real-time CPU and memory usage for the specific named pod. It reports usage in cores and MiB (or GiB) for each container, making it the precise tool for troubleshooting a single pod's resource consumption against its requests/limits.

Why this answer

Option C, `kubectl top pod <pod-name>`, is correct because the `kubectl top` command retrieves live CPU and memory usage metrics for a specific pod from the Metrics Server, which is exactly what checking resource usage of a pod requires. Option E, `kubectl top pods`, is also correct because it lists current CPU and memory consumption for all pods in the namespace, directly satisfying the resource-usage check. Option A, `kubectl top nodes`, is not correct here because it reports resource usage of cluster nodes, not pods.

Option B, `kubectl describe pods`, shows configuration, events, and status details but not live CPU/memory usage metrics. Option D, `kubectl get pods --watch`, only streams pod status changes and does not report resource consumption.

Exam trap

The trap here is that candidates confuse `kubectl describe pods` (which shows resource requests/limits) with actual resource usage, or they mistakenly think `kubectl get pods --watch` provides real-time performance metrics instead of status changes.

24
Multi-Selectmedium

Which of the following are valid methods to troubleshoot a Node that is 'NotReady'? (Select all that apply)

Select 3 answers
A.Check the kubelet logs on the node via journalctl
B.Verify that the kubelet certificate is valid and not expired
C.Restart the node's container runtime (e.g., containerd)
D.Delete the node object and re-create it
E.Reinstall the entire Kubernetes cluster
AnswersA, B, C

Checking kubelet logs via journalctl is the primary, non-destructive diagnostic step: the kubelet is the agent that registers the node and reports its status, so any failure in its startup or operation (e.g., CNI misconfiguration, resource pressure, or runtime communication errors) will be logged there. Running `journalctl -u kubelet` (or `journalctl -u kubelet --since today`) reveals the exact error before deciding on any remediation. This step should always precede restarting services or modifying cluster objects, because it pinpoints the underlying cause without causing downtime.

Why this answer

The kubelet is the primary node agent that communicates with the control plane. When a node is 'NotReady', checking the kubelet logs via `journalctl -u kubelet` can reveal errors such as network connectivity issues, certificate problems, or resource exhaustion. Option B is correct because an expired or invalid kubelet certificate will cause TLS authentication failures with the API server, leading to the node being marked 'NotReady'.

Option C is correct because the container runtime (e.g., containerd) is responsible for managing containers; if it is down or misconfigured, the kubelet cannot start pods, and restarting it can resolve transient failures.

Exam trap

CNCF often tests the misconception that deleting and re-creating a Node object is a valid troubleshooting step, when in reality it only removes the API server's representation and does not fix the underlying cause of the 'NotReady' state.

Why the other options are wrong

D

Deleting the node does not fix the node condition; it only removes it from the cluster.

E

Overkill; a single node issue doesn't require cluster reinstall.

25
MCQhard

A node in your cluster is reporting 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. The output shows 'Active: inactive (dead)'. Which command should you run FIRST to attempt to resolve this?

A.journalctl -u kubelet
B.systemctl status kubelet
C.reboot
D.systemctl start kubelet
AnswerD

Since the node is NotReady due to an inactive kubelet daemon, executing this command directly instructs systemd to launch the kubelet process. Once running, the kubelet will resume posting node status updates to the Kubernetes API server, restoring the node to a Ready state.

Why this answer

The kubelet service is the primary Kubernetes agent on each node, and when it is 'inactive (dead)', the node cannot report its status to the control plane. The first step to restore the node is to start the kubelet service using 'systemctl start kubelet', which will attempt to bring the kubelet process back online and allow it to reconnect to the API server.

Exam trap

The trap here is that candidates may jump to checking logs (Option A) or rebooting (Option C) without first attempting the simplest corrective action—starting the service—which is the direct fix for an inactive kubelet.

How to eliminate wrong answers

Option A is wrong because 'journalctl -u kubelet' is used to view the kubelet logs for debugging, but it does not resolve the inactive state of the service. Option B is wrong because 'systemctl status kubelet' only displays the current status and does not change the service state; the output already showed 'inactive (dead)', so this command provides no resolution. Option C is wrong because rebooting the node is an unnecessary and heavy-handed step that may cause additional downtime; the correct approach is to start the kubelet service first, and only if that fails should further troubleshooting or a reboot be considered.

26
MCQmedium

You are unable to resolve a Service DNS name from within a pod. To test DNS resolution, which command should you run inside the pod?

A.kubectl describe svc <service-name>
B.kubectl logs <pod>
C.kubectl attach <pod>
D.kubectl exec <pod> -- nslookup <service-name>
AnswerD

kubectl exec <pod> -- nslookup <service-name> is correct because it starts the nslookup process directly inside the pod's container, using the same /etc/resolv.conf, search domains, and network stack that the application uses. This actually queries the cluster's DNS service and returns the ClusterIP (or pod IPs for headless services) associated with that Service in the same namespace. If the Service is in a different namespace, you can test the fully qualified name <service-name>.<namespace>.svc.cluster.local. It is the standard, non-invasive way to validate service discovery from a pod's perspective.

Why this answer

The correct command is 'kubectl exec <pod> -- nslookup <service-name>'. nslookup or dig are common DNS troubleshooting tools.

27
MCQeasy

Which kubectl command is used to view the logs of a container that has previously crashed in a pod?

A.kubectl logs pod-name -c container-name --tail=100
B.kubectl logs pod-name --all-containers
C.kubectl logs pod-name --previous
D.kubectl logs pod-name
AnswerC

The --previous flag retrieves logs from the container's prior terminated instance, which is exactly what a crashed container requires. Without it, kubectl logs returns only the current container's output, which may not exist after a crash.

Why this answer

The `--previous` flag in `kubectl logs` retrieves logs from the previous instance of a container that has crashed and been restarted. This is essential for debugging crash loops, as the current container may have no logs or only startup logs, while the crashed container's logs contain the error.

Exam trap

The trap here is that candidates assume `kubectl logs` without flags or with `--tail` will show crash logs, but they only see the current container's logs, missing the critical error from the previous crashed instance.

How to eliminate wrong answers

Option A is wrong because `--tail=100` limits the log output to the last 100 lines of the current container's logs, but does not access logs from a previously crashed instance; it is useful for reducing output, not for crash debugging. Option B is wrong because `--all-containers` streams logs from all containers in the pod simultaneously, but it still only shows current container logs, not the logs of a container that has crashed and restarted. Option D is wrong because `kubectl logs pod-name` without flags shows only the current container's logs; if the container has crashed and been replaced, the current container's logs may be empty or irrelevant, missing the crash context.

28
Multi-Selectmedium

A pod is in 'Pending' state. 'kubectl describe pod' shows: '0/3 nodes are available: 1 Insufficient memory, 2 node(s) had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate.' Which THREE actions would resolve the issue? (Choose three)

Select 3 answers
A.Remove the taint from the control-plane nodes.
B.Decrease the memory request of the pod.
C.Increase the CPU request of the pod.
D.Add a toleration to the pod for the control-plane taint.
E.Add a node selector to the pod that matches the control-plane nodes.
AnswersA, B, D

Removing the control-plane taint lets the scheduler place the pod on those nodes, directly resolving the 'taint the pod didn't tolerate' rejection. It addresses one of the two stated scheduling failures, freeing capacity that the Insufficient memory node cannot provide.

Why this answer

The scheduler message shows two distinct problems: one node has Insufficient memory, and two nodes carry the taint node-role.kubernetes.io/control-plane that the pod does not tolerate. Option A is correct because removing that taint from the control-plane nodes makes them eligible for scheduling, directly eliminating the 'didn't tolerate' rejection. Option B is correct because lowering the pod's memory request lets it fit on the node that currently reports Insufficient memory, resolving that filter failure.

Option D is correct because adding a matching toleration (e.g., key node-role.kubernetes.io/control-plane with the appropriate effect) allows the pod to be scheduled onto the tainted control-plane nodes without changing the cluster. Option C is wrong because increasing the CPU request only makes scheduling harder and does not address either reported failure. Option E is wrong because a node selector matching control-plane nodes does not bypass the taint; without a toleration the pod would still be rejected by those nodes.

Exam trap

CKA often tests whether candidates confuse 'node selector' with 'toleration' — a selector narrows eligible nodes but never overrides a taint, so it cannot fix a taint-based Pending condition.

29
MCQhard

You need to check the logs of a container that previously crashed. The pod is currently running, but the previous instance of the container exited with an error. Which command will show you the logs from the crashed container?

A.kubectl logs pod-name --previous
B.kubectl exec pod-name -- cat /var/log/crash.log
C.kubectl attach pod-name
D.kubectl logs pod-name -c container-name
AnswerA

The `--previous` (or `-p`) flag instructs the Kubernetes API server to retrieve the log output from the most recently terminated or crashed instance of a container within the specified Pod. This is crucial for troubleshooting post-mortem events like CrashLoopBackOff, where the active container has restarted and its current log buffer is empty or reset.

Why this answer

The `kubectl logs` command with the `--previous` flag retrieves the logs from the previous instance of a container in a pod, which is exactly what you need when the current container is running but the previous one crashed. This flag works by accessing the terminated container's logs stored by the kubelet, allowing you to debug the crash without needing to access the node directly.

Exam trap

The trap here is that candidates often forget the `--previous` flag exists and instead try to use `kubectl exec` or `kubectl logs` without it, assuming the crashed container's logs are still accessible via standard commands or file paths.

How to eliminate wrong answers

Option B is wrong because `kubectl exec` runs a command in the currently running container, not the crashed one, and the path `/var/log/crash.log` is not a standard location for Kubernetes container logs; container logs are typically streamed to stdout/stderr and managed by the container runtime. Option C is wrong because `kubectl attach` attaches to the standard input/output of a currently running container, which is useless for a crashed container that is no longer active. Option D is wrong because `kubectl logs pod-name -c container-name` shows logs from the current instance of the specified container, not the previous crashed instance, unless combined with `--previous`.

30
MCQeasy

You run 'kubectl top nodes' and it returns an error: 'error: metrics not available yet'. What does this indicate?

A.The kubelet is not running on the nodes.
B.The cluster is using an older version of Kubernetes.
C.The nodes are under heavy load.
D.The Metrics Server is not installed or not functioning.
AnswerD

The Metrics Server is a cluster add-on that collects resource metrics from each kubelet and exposes them through the aggregated `metrics.k8s.io` API; `kubectl top nodes` queries this API to display usage figures. If the Metrics Server is not installed, or if its pods are failing to start (e.g., due to image pull issues or invalid configuration), the API is unavailable and the command returns an error such as 'metrics.k8s.io is not available' or 'the server is currently unable to handle the request.' Even when installed, it may need about 15 seconds to collect its first batch of metrics; if queried immediately after cluster startup, you might see an empty or error response. Checking for metrics-server pods in the `kube-system` namespace and verifying that the aggregated API is registered will confirm this root cause.

Why this answer

The 'top' command relies on the metrics server to provide resource usage data. The error means the metrics server is not deployed or not ready.

31
MCQhard

You have a multi-node cluster. One node shows 'NotReady'. You run 'journalctl -u kubelet' on that node and see 'network plugin is not ready'. What is the most likely cause?

A.The CNI plugin pod (e.g., Calico) is not running on that node
B.The container runtime (e.g., containerd) is down
C.The kubelet service is not running
D.The node's IP address has changed
AnswerA

The kubelet agent on each node continuously monitors the health and operational status of the CNI plugin. If the CNI plugin's pod (e.g., Calico, Flannel) crashes, fails its readiness probes, or is not running on a specific node, the kubelet cannot properly configure network interfaces for pods or establish essential network routes. Consequently, the kubelet reports the node's status as `NotReady` with the specific `NetworkPluginNotReady` condition, indicating a fundamental failure in the node's networking capabilities.

Why this answer

The 'network plugin is not ready' error from kubelet indicates that the kubelet is waiting for a CNI (Container Network Interface) plugin to configure the pod network. If the CNI plugin pod (e.g., Calico, Flannel, Weave) is not running on that specific node, the kubelet cannot set up the network for pods, causing the node to remain in 'NotReady' state. This is the most direct cause because the kubelet relies on the CNI plugin to report readiness before marking the node as Ready.

Exam trap

The trap here is that candidates often confuse 'network plugin is not ready' with a general network connectivity issue or a kubelet failure, but the error specifically points to the CNI plugin not being operational on that node, not the kubelet itself or the container runtime.

How to eliminate wrong answers

Option B is wrong because if the container runtime (e.g., containerd) were down, the kubelet would report a different error, such as 'container runtime is down' or 'failed to connect to containerd', not a network plugin error. Option C is wrong because if the kubelet service were not running, you would not be able to run 'journalctl -u kubelet' successfully, and the node would not show 'NotReady' but would be unreachable entirely. Option D is wrong because a changed node IP address would cause kubelet to fail to register with the control plane, typically resulting in 'Node not found' or authentication errors, not a 'network plugin is not ready' message.

32
Multi-Selectmedium

Which TWO of the following are common causes of CrashLoopBackOff? (Choose two)

Select 2 answers
A.Node resource pressure (CPU/memory)
B.Misconfigured startup probe
C.NetworkPolicy blocking egress
D.Application initialization failure
E.PersistentVolume not mounted
AnswersB, D

A misconfigured startup probe can force the kubelet to kill a container that has actually started successfully but fails the probe due to wrong parameters, such as an incorrect port, path, or timeout. With the startup probe, the kubelet specially guards the container until it reports readiness; if the probe thresholds are too restrictive, the container is restarted before the application can finish initializing. This creates an endless cycle of restart and re-probe, producing CrashLoopBackOff. It is a direct and common cause of crash loops because the container is terminated by the kubelet based on probe failure, not by an actual app error.

Why this answer

CrashLoopBackOff occurs when a container repeatedly starts and then exits, and Kubernetes backs off restarting it. Option B is correct because a misconfigured startup probe (wrong path, port, or timing) causes kubelet to kill and restart the container before it can finish initializing, producing repeated crashes. Option D is correct because an application initialization failure (e.g., missing config, bad DB connection, unhandled exception at boot) makes the process exit immediately each time, which is the classic trigger for CrashLoopBackOff.

Option A is not the direct cause: node resource pressure typically leads to eviction or Pending/OOMKilled states rather than the restart-backoff loop itself. Option C is not correct because a NetworkPolicy blocking egress usually causes runtime errors or timeouts, not necessarily a container exit loop. Option E is not correct because an unmounted PersistentVolume generally leaves the pod stuck in Pending or ContainerCreating, not in CrashLoopBackOff.

Exam trap

The CKA exam often tests the distinction between CrashLoopBackOff (caused by container crashes or probe failures) and other Pod failure states like ImagePullBackOff, ErrImagePull, or CreateContainerConfigError, leading candidates to confuse resource pressure or network issues with actual container exit failures.

33
MCQmedium

A pod is in 'Pending' state. After running 'kubectl describe pod', you see the event: '0/3 nodes are available: 3 PersistentVolumeClaim is not bound'. What is the most likely cause?

A.The PersistentVolumeClaim referenced by the pod does not exist or is not bound
B.The storage class used by the PVC does not match any PV
C.The scheduler is not configured to handle persistent volumes
D.The nodes do not have enough CPU or memory
AnswerA

When a Pod references a PersistentVolumeClaim (PVC) in its volume definition, the Kubernetes scheduler will keep the Pod in a Pending state if that PVC is not bound to a PersistentVolume (PV) or if the PVC does not exist in the namespace. Running kubectl describe pod will explicitly show a warning event such as FailedScheduling with a message indicating the PVC is not bound, preventing the Pod from being assigned to a node.

Why this answer

The event '0/3 nodes are available: 3 PersistentVolumeClaim is not bound' indicates that the pod's PersistentVolumeClaim (PVC) is not in a 'Bound' state. This means the PVC either does not exist in the namespace or has not been successfully bound to a PersistentVolume (PV). The scheduler cannot schedule the pod because the PVC requirement is unsatisfied, which is a prerequisite for pod startup.

Exam trap

Kubernetes often tests the distinction between 'PVC not bound' and 'insufficient resources' — candidates mistakenly think a PVC binding error implies a storage class mismatch (Option B), but the event message directly points to the PVC's binding status, not the storage class configuration.

How to eliminate wrong answers

Option B is wrong because a mismatch between the storage class used by the PVC and any PV would result in the PVC remaining 'Pending' (not bound), but the event message specifically states 'PersistentVolumeClaim is not bound', which is a general condition that can occur even without storage class issues. Option C is wrong because the scheduler does not require special configuration to handle persistent volumes; it natively checks PVC binding status via the PersistentVolumeClaim predicate. Option D is wrong because insufficient CPU or memory would produce a different event, such as '0/3 nodes are available: 3 Insufficient cpu' or 'Insufficient memory', not a PVC binding error.

34
MCQmedium

You need to investigate why a service is not reachable from within the cluster. Which of the following is the first step?

A.Check kube-proxy logs on the nodes
B.Check DNS resolution
C.Check if the service has endpoints
D.Restart the service
AnswerC

The first and most decisive troubleshooting step is to run `kubectl get endpoints <service>` (or `kubectl describe service`) and confirm the endpoint list is non-empty. If the selector doesn't match any pod, or the matched pods are NotReady, the EndpointSlice controller generates zero addresses, leaving kube-proxy with no valid destination for the service IP. This explains why the Service is unreachable even though the Service object itself appears healthy. Checking endpoints isolates the problem to label selectors or pod readiness immediately.

Why this answer

The most fundamental check when a service is unreachable from within the cluster is to verify whether the service has any endpoints. A service without endpoints means no pods are matching its selector, so traffic cannot be forwarded regardless of DNS or kube-proxy status. The `kubectl get endpoints <service>` command directly reveals this, making it the logical first step before deeper diagnostics.

Exam trap

The trap here is that candidates often jump to DNS or kube-proxy issues because those are common networking topics, but the CKA exam expects you to follow a logical troubleshooting hierarchy, starting with the simplest check—whether the service has any backing pods.

Why the other options are wrong

A

More advanced step; start with endpoints.

B

DNS is for service discovery, not connectivity.

D

Restarting is not troubleshooting.

35
MCQeasy

You need to check the CPU and memory usage of all pods in the 'production' namespace. Which command should you use?

A.kubectl top pods -n production
B.kubectl get pods -n production --show-labels
C.kubectl get resourceusage pods -n production
D.kubectl top nodes -n production
AnswerA

kubectl top pods -n production is correct because it queries the metrics API (metrics.k8s.io) via the metrics-server and returns a live CPU and memory consumption report for every pod in the production namespace. The output shows CPU usage in cores (or millicores) and memory in bytes or MiB, which is exactly the resource usage data the task asks for. This command requires the metrics-server to be deployed and configured, otherwise it will return an error from the metrics API.

Why this answer

The `kubectl top pods` command retrieves real-time CPU and memory metrics for pods from the metrics server, which is required for the CKA exam. The `-n production` flag restricts the output to the 'production' namespace, making option A the correct choice for checking resource usage of all pods in that namespace.

Exam trap

The trap here is that candidates confuse `kubectl top` with `kubectl get` or invent non-existent commands like `kubectl get resourceusage`, assuming there is a generic resource usage command, when only `kubectl top` provides live metrics.

How to eliminate wrong answers

Option B is wrong because `kubectl get pods --show-labels` displays pod metadata and labels, not CPU or memory usage metrics. Option C is wrong because `kubectl get resourceusage` is not a valid kubectl command; the correct command for resource usage is `kubectl top`. Option D is wrong because `kubectl top nodes` shows node-level resource usage, not pod-level usage, and would not filter by namespace even if it did.

36
MCQmedium

A pod is failing to start with the error 'CrashLoopBackOff'. You check the logs with 'kubectl logs pod' and see nothing. What is the most likely reason?

A.The pod's log path is misconfigured
B.The container crashed before generating any log output
C.The kubelet has logging disabled
D.The pod is using a sidecar container for logging
AnswerB

If an application terminates immediately upon execution—such as due to a missing shared library, a syntax error in an entrypoint script, or a rapid kernel panic—it may exit before writing any bytes to stdout or stderr. In this scenario, the container runtime creates an empty log file, resulting in no output when running kubectl logs.

Why this answer

The 'CrashLoopBackOff' error indicates the container in the pod starts, crashes, and is repeatedly restarted by the kubelet. If 'kubectl logs pod' returns nothing, it means the container exited before writing any output to stdout/stderr, which is the default logging target for Kubernetes. Option B is correct because the container likely failed during initialization or startup, before any application code produced log entries.

Exam trap

The trap here is that candidates assume 'kubectl logs' always shows something, even for a crashing pod, but if the container fails before any output is written to stdout/stderr, the logs will be empty, leading to the correct conclusion that the crash occurred during early startup.

How to eliminate wrong answers

Option A is wrong because the pod's log path is not a configurable concept in Kubernetes; containers log to stdout/stderr by default, and 'kubectl logs' reads from the container runtime's log driver, not a file path. Option C is wrong because the kubelet does not have a 'logging disabled' feature; it always forwards container stdout/stderr to the configured log driver (e.g., journald, JSON file). Option D is wrong because a sidecar container for logging would not prevent the main container's logs from being visible via 'kubectl logs'; the sidecar would typically consume logs from a shared volume or stream, but the main container's stdout/stderr would still be captured.

37
MCQhard

You have a Deployment with the following resource limits for containers: memory: 256Mi. The pod is repeatedly killed with OOMKilled. You need to change the limit to 512Mi. Which field should you modify in the Deployment YAML?

A.spec.template.spec.containers[].resources.limits.memory
B.spec.template.spec.containers[].resources.requests.memory
C.spec.template.spec.containers[].resources.requests.cpu
D.spec.template.spec.containers[].resources.limits.cpu
AnswerA

The container's memory limit is implemented as a cgroup v2 (or v1) limit on the memory cgroup to which that container belongs. When the container's resident set size plus page cache exceeds this limit, the kernel's out-of-memory (OOM) killer selects a process in that cgroup — usually the container's main PID — and kills it, resulting in OOMKilled. Increasing spec.template.spec.containers[].resources.limits.memory raises this cgroup ceiling, giving the container more usable memory headroom before the kernel decides to kill its process, which is exactly why this field addresses the reported OOMKilled status.

Why this answer

The OOMKilled error occurs when a container exceeds its memory limit. To resolve this, you must increase the memory limit in the Deployment's pod template. Option A correctly identifies the field `spec.template.spec.containers[].resources.limits.memory`, which directly controls the maximum memory the container can use before being killed by the OOM killer.

Exam trap

The trap here is that candidates confuse `requests` (which only affects scheduling and QoS class) with `limits` (which enforces hard resource caps), leading them to mistakenly modify `requests.memory` instead of `limits.memory` to fix an OOMKilled issue.

How to eliminate wrong answers

Option B is wrong because `resources.requests.memory` is the minimum memory guaranteed to the container, not the limit that triggers OOMKilled; changing requests does not prevent the OOM killer from terminating the container if it exceeds the limit. Option C is wrong because `resources.requests.cpu` sets the minimum CPU allocation, which has no effect on memory-related OOM kills. Option D is wrong because `resources.limits.cpu` caps CPU usage, not memory; exceeding the CPU limit causes throttling, not OOMKilled.

38
MCQeasy

You want to see the last 50 lines of logs from a pod named 'api-pod' for the container 'api-container'. Which command accomplishes this?

A.kubectl logs -l app=api --tail=50
B.kubectl logs api-pod --tail=50
C.kubectl logs api-pod -c api-container --tail=50
D.kubectl logs api-container -p api-pod --tail 50
AnswerC

This is correct because kubectl logs expects the Pod name as the first positional argument, then -c names a specific container within that Pod, and --tail=50 limits the output to the last 50 lines. The flag syntax --tail=N is valid; it tells the kubelet to read only the requested number of lines from the container's log file, which is far more efficient than streaming the entire log for a long-running Pod. This exactly satisfies the requirement to see the last 50 lines from the named Pod's api-container.

Why this answer

'kubectl logs api-pod -c api-container --tail=50' shows the last 50 lines for the specified container.

39
Multi-Selectmedium

Which three are possible reasons for a pod being in Pending state? (Choose three.)

Select 3 answers
A.PersistentVolumeClaim not bound
B.Liveness probe failing
C.Container image does not exist
D.Insufficient CPU or memory resources on any node
E.Node taints that the pod does not tolerate
AnswersA, D, E

A pod that references a PersistentVolumeClaim (PVC) remains in Pending when that claim is unbound because the scheduler cannot determine which node has the required storage volume unless the PVC is already in the Bound phase. The kubelet will not start containers until all referenced volumes are successfully attached and mounted, and if the PVC is lost or awaiting dynamic provisioning, the pod cannot progress. For example, a manually pre-created PV with a mismatched access mode or storage class can leave the PVC unbound, indefinitely keeping the pod in Pending.

Why this answer

Pending means the pod hasn't been scheduled. Reasons include insufficient resources, persistent volume claim not bound, and taints/tolerations mismatch. Image pull issues cause ImagePullBackOff, not Pending.

40
MCQmedium

You run 'kubectl get pods' and see a pod in 'ImagePullBackOff' state. Which command would help you determine the exact reason for the image pull failure?

A.kubectl describe pod <pod-name>
B.kubectl top pod <pod-name>
C.kubectl logs <pod-name>
D.kubectl get events
AnswerA

"kubectl describe pod <pod-name>" provides a comprehensive summary of a pod's current state, including its status, events, and container specifications. When a pod is in `ImagePullBackOff`, this command will display the specific events related to the image pull attempt, such as `Failed` or `ErrImagePull`, along with the exact error message from the container runtime. This detailed output is crucial for diagnosing the root cause, such as an incorrect image name, a private registry authentication failure, or network issues, making it the most direct and effective diagnostic tool.

Why this answer

A is correct because 'kubectl describe pod <pod-name>' provides detailed information about the pod, including the container status, events, and the exact error message from the image pull attempt. This output includes the reason for the ImagePullBackOff, such as a missing image, incorrect tag, authentication failure, or network issue, which is essential for troubleshooting.

Exam trap

The trap here is that candidates often choose 'kubectl logs' thinking it will show the error, but logs only exist if the container started; for ImagePullBackOff, the container never runs, so logs are empty and the describe command is the correct tool.

How to eliminate wrong answers

Option B is wrong because 'kubectl top pod' shows resource usage (CPU/memory) and does not provide any information about image pull failures. Option C is wrong because 'kubectl logs' retrieves container logs, but if the container never started due to ImagePullBackOff, there are no logs to fetch; the error is in the pod status, not in stdout/stderr. Option D is wrong because 'kubectl get events' shows cluster-wide events, but it may not include the specific image pull error for the pod, and it is less detailed than the pod description; the describe command is the standard tool for this scenario.

41
MCQhard

You have a Pod that is stuck in Pending state. Running 'kubectl describe pod' shows events: '0/4 nodes are available: 1 node(s) had taint {node-role.kubernetes.io/control-plane: }, 3 node(s) had taint {key: value}, that the pod didn't tolerate.' How can you resolve this issue?

A.Increase the Pod's resource requests
B.Remove the taints from all nodes using 'kubectl taint nodes --all key:value-'
C.Delete the Pod and recreate it with a different name
D.Add appropriate tolerations to the Pod's spec
AnswerD

Adding matching tolerations to the Pod's specification directly instructs the kube-scheduler that this Pod is allowed to run on nodes with corresponding taints. This is the standard, least-privilege method to resolve scheduling issues on tainted nodes without compromising the cluster's overall node isolation strategy. Once the toleration is applied, the scheduler can successfully bind the Pod to the tainted node.

Why this answer

The Pod is stuck in Pending because none of the nodes can schedule it due to taints that the Pod does not tolerate. Option D is correct because adding the appropriate tolerations to the Pod's spec tells the scheduler that the Pod can tolerate those taints, allowing it to be scheduled on the tainted nodes. This directly addresses the mismatch between node taints and Pod tolerations.

A common misconception in the CKA exam is that removing taints from nodes is the only fix, but the correct Kubernetes approach is to add tolerations to the Pod spec, preserving node isolation for other workloads.

Exam trap

A common misconception in the CKA exam is that removing taints from nodes is the only fix, but the correct Kubernetes approach is to add tolerations to the Pod spec, preserving node isolation for other workloads.

How to eliminate wrong answers

Option A is wrong because increasing resource requests would only worsen scheduling constraints, not resolve taint/toleration mismatches. Option B is wrong because removing taints from all nodes is an overly broad and potentially disruptive action; the correct approach is to add tolerations to the Pod, not modify cluster-wide node settings. Option C is wrong because simply deleting and recreating the Pod with a different name does not change its toleration configuration, so it would still be stuck in Pending.

42
MCQmedium

A node in the cluster is showing NotReady status. Which steps should you take to diagnose the issue? (Select the BEST initial step.)

A.Test DNS resolution from a pod on the node
B.Run 'kubectl describe pod' for pods on that node
C.Check the kube-apiserver logs on the master node
D.Run 'journalctl -u kubelet' on the node to check kubelet logs
AnswerD

The kubelet is the primary node agent responsible for posting node status and heartbeats back to the control plane via the Node Lifecycle Controller. When a node transitions to NotReady, checking the kubelet's systemd logs using 'journalctl -u kubelet' is the most direct way to diagnose the issue. These logs will reveal critical failures such as container runtime connection errors, out-of-memory (OOM) events, misconfigured certificates, or disk pressure conditions that caused the agent to fail.

Why this answer

When a node is NotReady, the kubelet is the primary agent responsible for reporting node status. The kubelet communicates node conditions to the control plane via periodic heartbeats. Checking the kubelet logs with 'journalctl -u kubelet' on the node itself is the most direct initial step to identify why the kubelet is failing to report readiness, such as network issues, resource exhaustion, or certificate problems.

Exam trap

The trap here is that candidates often jump to checking control-plane components (like kube-apiserver) or pod-level issues, when the kubelet is the direct source of node status and its logs are the first place to look for local node problems.

How to eliminate wrong answers

Option A is wrong because DNS resolution from a pod tests cluster DNS functionality, not the node's kubelet health or status reporting; a node can be NotReady even if DNS works. Option B is wrong because 'kubectl describe pod' shows pod status and events, but does not reveal why the node itself is NotReady; the node condition is independent of individual pod states. Option C is wrong because the kube-apiserver logs on the master node may show that the node is NotReady, but they do not provide the root cause on the node; the kubelet logs are the authoritative source for local failures.

43
MCQhard

A pod runs but you cannot connect to its container port from another pod in the same namespace. 'kubectl exec' into the pod and 'curl localhost:8080' works. What is the MOST likely cause?

A.There is a NetworkPolicy blocking ingress
B.The container's port is not exposed in the pod spec
C.The Service selector does not match the pod labels
D.The pod is bound to localhost only, not 0.0.0.0
AnswerD

Binding to localhost (127.0.0.1) restricts connections to only the same network namespace, which is isolated to the container itself. Kubernetes networking relies on the pod's IP address and virtual Ethernet interfaces, so a loopback bind cannot receive traffic from outside the pod. Listening on 0.0.0.0 makes the port available on all interfaces, including the pod's external-facing one.

Why this answer

When `curl localhost:8080` works inside the pod but connections from other pods fail, the most likely cause is that the application is listening only on the loopback interface (127.0.0.1) instead of 0.0.0.0 (all interfaces). This means the container process binds to localhost, which is only reachable from within the same network namespace (the pod itself), not from external sources like other pods. Kubernetes networking relies on the container listening on 0.0.0.0 so that traffic arriving via the pod's eth0 interface (which has its own IP) can be accepted.

Exam trap

The trap here is that candidates often assume a Service issue (Option C) is the cause, but the question specifies direct pod-to-pod connectivity (not via a Service), so the real problem is the application binding to localhost, which is a classic application-layer misconfiguration rather than a Kubernetes networking misconfiguration.

How to eliminate wrong answers

Option A is wrong because a NetworkPolicy blocking ingress would prevent connections from other pods even if the application is correctly listening on 0.0.0.0, but the fact that `curl localhost:8080` works inside the pod does not rule out a NetworkPolicy; however, the symptom of localhost working while external connections fail points directly to a binding issue, not a policy. Option B is wrong because the container's port not being exposed in the pod spec (i.e., missing `containerPort`) does not affect whether the application listens on the correct interface; it only affects service discovery and documentation, not actual network connectivity. Option C is wrong because the Service selector not matching pod labels would prevent traffic from reaching the pod via the Service, but the question states the connection attempt is 'from another pod in the same namespace' — this could be direct pod-to-pod via IP, which bypasses the Service entirely, so a mismatched selector would not explain the failure.

44
Multi-Selecthard

Which THREE of the following are valid steps to troubleshoot a Node in NotReady state? (Choose three)

Select 3 answers
A.Run 'systemctl status kubelet' on the node
B.Restart the kube-apiserver on the control plane
C.Verify that the container runtime (e.g., containerd) is running
D.Run 'kubectl delete node <node-name>' to re-register
E.Check the kubelet logs with 'journalctl -u kubelet'
AnswersA, C, E

Running 'systemctl status kubelet' on the node checks whether the kubelet systemd service is active or has failed, showing the current unit state, exit code, and recent process info. This is a logical first step because a Node stuck NotReady often means the kubelet process itself is not running or repeatedly crashing. It directly inspects the node agent, unlike control-plane diagnostics, and can quickly reveal if the service is masked, stopped, or in a failed state.

Why this answer

Option A is correct because a Node in NotReady state is most often caused by the kubelet service being stopped or failed, and 'systemctl status kubelet' on the node quickly reveals whether the kubelet is active, inactive, or in a failed state. Option C is correct because the kubelet requires a working CRI container runtime such as containerd (or CRI-O) to report node health; if containerd is down, the kubelet cannot manage pods and the node transitions to NotReady, so verifying the runtime is running is a valid troubleshooting step. Option E is correct because 'journalctl -u kubelet' surfaces the kubelet's own error messages (e.g., failed to get node status, CRI connection errors, certificate or cgroup driver problems) that explain why the node is NotReady.

Option B is not a valid step because restarting the kube-apiserver on the control plane does not fix a node-local kubelet or runtime failure and would disrupt the whole cluster. Option D is not valid because 'kubectl delete node <node-name>' removes the Node object from the API server rather than re-registering it; the kubelet re-registers only when it restarts and reconnects, so deleting the node is not a troubleshooting step for NotReady.

Exam trap

The trap here is that candidates may confuse control plane components with node-level agents, incorrectly assuming that restarting the kube-apiserver or deleting the node object will fix a node-level NotReady state, when the real issue lies with the kubelet or container runtime on the node itself.

45
MCQhard

You are troubleshooting a network connectivity issue between two pods in different namespaces. The pods have the following labels: pod-a in namespace 'foo' with labels {app: web}, pod-b in namespace 'bar' with labels {app: db}. You verify that both pods have IP addresses and can ping the Kubernetes service IP. However, pod-a cannot connect to pod-b on port 5432. What should you check first?

A.Check if the kube-proxy is running on the node hosting pod-b
B.Check if a NetworkPolicy exists that denies ingress traffic to pod-b from namespace 'foo'
C.Check if the container runtime is Docker
D.Check if the DNS resolution for pod-b's service is correct
AnswerB

A NetworkPolicy is a namespaced Kubernetes resource that acts as a pod-level firewall, and if cluster networking is configured with a CNI that enforces it, any ingress rule can explicitly deny traffic from pods in other namespaces. The default behavior is allow-all only when no NetworkPolicy selects the pod; once one exists, the default becomes deny for anything not matched by its rules. If a policy selects pod-b and its ingress list does not include namespace 'foo' or a matching podSelector, it will silently drop the TCP SYN packets from pod-a, making the port 5432 connection time out or refuse while the service IP ping still succeeds.

Why this answer

When pods can reach each other's IPs and services but a specific port connection fails, the most likely cause in a multi-namespace Kubernetes cluster is a NetworkPolicy restricting ingress to pod-b. NetworkPolicies are namespace-scoped and can deny cross-namespace traffic on specific ports, which perfectly matches the symptom of ICMP/service reachability succeeding while TCP/5432 fails.

Exam trap

CKA often tests the assumption that if basic connectivity (ping, service IP) works, the network is fine — candidates overlook that NetworkPolicies selectively block specific ports and namespaces while leaving other traffic untouched.

How to eliminate wrong answers

Option A is wrong because if kube-proxy were down on pod-b's node, service IP connectivity would fail broadly, not just the specific pod-to-pod port 5432 connection — and the scenario states service IPs are reachable. Option C is wrong because the container runtime (Docker vs. containerd) is irrelevant to pod-to-pod TCP connectivity; both runtimes use the same CNI networking. Option D is wrong because DNS resolution for pod-b's service is not the issue — the analyst is connecting directly to pod-b's IP on port 5432, and DNS failures would manifest as name-resolution errors, not port-specific connection refusals.

46
Multi-Selecthard

You run 'kubectl get pods' and see that a pod is in CrashLoopBackOff. Which THREE of the following are valid next steps? (Select 3)

Select 3 answers
A.Run 'kubectl describe pod pod-name' to check events
B.Run 'kubectl logs pod-name'
C.Run 'kubectl get pod pod-name -o jsonpath={.status.containerStatuses[0].state.waiting.reason}'
D.Run 'kubectl top pod pod-name'
E.Run 'kubectl delete pod pod-name'
AnswersA, B, C

kubectl describe pod pod-name aggregates the pod's current state, recent events, container statuses (waiting/running/terminated), restart counts, and last termination reason into a human-readable summary. This surfaces CrashLoopBackOff events with timestamps, showing whether the container exited due to an error, OOMKill, or probe failure. It is the first step for diagnosing why Kubernetes is restarting the container.

Why this answer

Option A is correct because 'kubectl describe pod pod-name' surfaces the pod's Events section, which shows why the kubelet restarted the container (e.g., OOMKilled, failed liveness probe, image pull issues) and the backoff timing. Option B is correct because 'kubectl logs pod-name' retrieves the container's stdout/stderr, which typically contains the application error that caused the process to exit and enter CrashLoopBackOff; adding --previous shows logs from the prior crashed instance. Option C is correct because the jsonpath query reads .status.containerStatuses[0].state.waiting.reason, which directly returns the waiting reason such as CrashLoopBackOff along with the last termination state, confirming the container's current status.

Option D is not appropriate because 'kubectl top pod' only reports live CPU/memory usage from the metrics server and does not explain why the container is crash-looping. Option E is not a valid diagnostic step because deleting the pod merely recreates it (or removes it if managed by a controller) without revealing the root cause, so it does not help troubleshoot the failure.

Exam trap

The CKA exam often tests the misconception that 'kubectl top' or deleting the pod is a valid troubleshooting step for CrashLoopBackOff, when in fact these actions either provide irrelevant metrics or mask the issue without diagnosis.

47
MCQhard

A pod is stuck in Pending state. You describe the pod and see: '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate.' What is the most likely reason?

A.The pod needs a toleration for node.kubernetes.io/not-ready
B.All nodes in the cluster are NotReady
C.The pod has a resource request that cannot be met
D.The scheduler is not running
AnswerB

This is the correct answer because the scheduler cannot place the pod on any node that has a taint not tolerated by the pod. When every node has the node.kubernetes.io/not-ready taint (a direct result of the node controller marking them NotReady), there is no feasible node in the cluster. The pod stays in Pending until at least one node returns to Ready status, which clears the taint and allows scheduling.

Why this answer

The error message '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate' indicates that every node in the cluster is tainted with node.kubernetes.io/not-ready, which is automatically applied by the node controller when a node becomes unreachable or fails its health checks. Since no node is Ready, the pod cannot be scheduled, and the only way to schedule it would be to add a toleration for this taint, but that would not fix the underlying node issue. Therefore, the most likely reason is that all nodes are in the NotReady state.

Exam trap

The trap here is that candidates often confuse 'tolerating a taint' with 'fixing the node condition', and assume that adding a toleration is the correct solution, when the error message explicitly states that all nodes have the taint, meaning the nodes themselves are NotReady.

Why the other options are wrong

A

Tolerating this taint would schedule pods on unhealthy nodes.

C

Would show insufficient resources, not taint.

D

Would show 0 nodes available due to other reasons.

48
MCQmedium

A worker node in your cluster has been marked NotReady for several minutes. You SSH into the node and find that the kubelet process is not running. You start the kubelet service, but it immediately exits. Running 'journalctl -u kubelet -n 50' shows repeated errors about being unable to load the kubeconfig file at /etc/kubernetes/kubelet.conf. Which of the following is the MOST likely cause?

A.The container runtime socket is not available, so the kubelet cannot manage pods.
B.The kubelet.conf file has incorrect file permissions or is missing, preventing the kubelet from authenticating to the API server.
C.The kubelet's systemd unit file has an incorrect path to the kubelet binary.
D.The node's network plugin is misconfigured, preventing the kubelet from reaching the API server.
AnswerB

The kubelet requires a valid kubeconfig to authenticate to the API server. If the file is missing or has permissions that prevent the kubelet user from reading it, the kubelet will fail to start with an error referencing the kubeconfig path. This aligns with the journalctl output. Restoring or regenerating the file fixes the issue.

Why this answer

The kubelet relies on a kubeconfig file to authenticate to the API server. When that file is missing or unreadable, the kubelet cannot establish its connection and exits with a clear error in its logs. The correct resolution is to restore the file from backup or regenerate it using kubeadm, ensuring proper permissions.

Other issues like container runtime or network plugin would produce different error messages.

Exam trap

The trap here is assuming that any kubelet failure is due to the container runtime or network, when the specific error message clearly points to a kubeconfig problem.

49
Multi-Selectmedium

Which THREE of the following are common causes for a Pod to remain in Pending state? (Select THREE.)

Select 3 answers
A.Insufficient CPU or memory resources in the cluster
B.Taints on nodes that are not tolerated by the Pod
C.PersistentVolumeClaim is not bound to a PersistentVolume
D.Container exits with OOMKilled
E.Image pull error due to incorrect image name
AnswersA, B, C

When no node has enough allocatable CPU or memory to satisfy the Pod's resource requests, the kube-scheduler's filter phase rejects every candidate node, leaving the Pod unscheduled in Pending. This is the classic capacity-driven cause.

Why this answer

Option A is correct because when no node has enough allocatable CPU or memory to satisfy the Pod's resource requests, the scheduler cannot place the Pod and it stays in Pending. Option B is correct because taints on nodes that the Pod does not tolerate cause the scheduler to reject those nodes, leaving the Pod unscheduled and Pending. Option C is correct because a PersistentVolumeClaim that is not yet bound to a PersistentVolume blocks scheduling of Pods that reference it, keeping them in Pending.

Option D does not belong because OOMKilled is a container termination reason that occurs after the Pod has been scheduled and started, producing a CrashLoopBackOff or Error state rather than Pending. Option E does not belong because an image pull error happens during kubelet container creation after scheduling, resulting in ImagePullBackOff or ErrImagePull, not Pending.

Exam trap

The CKA exam frequently tests your ability to troubleshoot Pod states. Remember that 'Pending' is a scheduling-phase state. If a Pod is Pending, the issue is almost always related to scheduling (resources, taints/tolerations, node selectors, affinity) or volume binding.

Once a Pod is assigned to a node, failures will manifest as 'ImagePullBackOff', 'CrashLoopBackOff', or 'Failed', not 'Pending'.

50
MCQhard

You run 'kubectl get nodes' and one node shows 'NotReady'. You SSH into the node and run 'systemctl status kubelet'. Kubelet is active but 'journalctl -u kubelet -n 50' shows 'network plugin is not ready: cni config uninitialized'. What is the most likely cause?

A.The container runtime is not installed
B.The CNI configuration file is missing or incorrect
C.Kubelet is not running
D.The node's IP address has changed
AnswerB

The kubelet depends on a CNI configuration file in /etc/cni/net.d/ to set up pod networking. When this file is absent or malformed, the kubelet cannot initialize the network plugin and marks the node as NotReady. The kubelet logs will show a 'CNI configuration uninitialized' error, which directly matches the symptom described in the question.

Why this answer

The CNI plugin configuration is missing, causing the network plugin to be unready.

51
MCQeasy

Which command shows CPU and memory usage of nodes in the cluster?

A.kubectl describe nodes
B.kubectl get nodes
C.kubectl logs nodes
D.kubectl top nodes
AnswerD

kubectl top nodes is the correct command because it queries the Metrics API, typically served by metrics-server, and aggregates resource usage from each node's kubelet and cAdvisor. It outputs a table showing CPU usage (e.g., 348m) and memory usage (e.g., 1200Mi) alongside the node's total capacity, making it the standard kubectl command for viewing current CPU and memory consumption across nodes.

Why this answer

`kubectl top nodes` is the command specifically designed to display real-time CPU and memory usage metrics for all nodes in a Kubernetes cluster. It relies on the metrics server, which collects resource usage data from kubelets via the Summary API, and presents it in a concise table format.

Exam trap

The trap here is that candidates often confuse `kubectl describe nodes` (which shows allocatable resources but not real-time usage) with `kubectl top nodes` (which shows actual current usage), leading them to pick option A incorrectly.

How to eliminate wrong answers

Option A is wrong because `kubectl describe nodes` shows detailed node information including capacity, allocatable resources, conditions, and taints, but it does not show real-time CPU and memory usage metrics. Option B is wrong because `kubectl get nodes` only lists nodes with their status, roles, age, and version, without any resource usage data. Option C is wrong because `kubectl logs nodes` is not a valid kubectl command; `kubectl logs` is used to fetch logs from pods, not nodes.

52
Multi-Selectmedium

A pod is in ImagePullBackOff. Which TWO of the following are possible causes? (Select 2)

Select 2 answers
A.The node has insufficient memory
B.The image is in a private registry and no imagePullSecrets are defined
C.The pod has a resource limit that is too low
D.The image tag is misspelled
E.The kubelet is not running
AnswersB, D

When an image resides in a private registry, the kubelet must authenticate itself using image pull secrets specified in the pod's spec via `imagePullSecrets`. If those secrets are missing or not attached to the pod, the registry responds with an authorization error (e.g., "unauthorized: authentication required"), and the kubelet reports ErrImagePull before entering ImagePullBackOff. This is a common production misconfiguration because merely having the secret in the namespace does not grant the pod access.

Why this answer

Common causes: invalid image tag (typo) and authentication failure when the image is in a private registry.

53
MCQeasy

You have a pod that is in 'CrashLoopBackOff' state. Which command should you use to view the logs from the previous instance of the container?

A.kubectl logs pod-name --previous
B.kubectl get events
C.kubectl describe pod pod-name
D.kubectl logs pod-name
AnswerA

`kubectl logs pod-name --previous` correctly retrieves the logs from the last terminated container instance in the pod. Because CrashLoopBackOff means the container keeps crashing and restarting, the logs from the most recent crashed instance contain the panic, exception, or startup failure that caused the crash. The `--previous` flag explicitly tells the kubelet to read the terminated container's log file, which is exactly what you need to inspect when the current container has already restarted.

Why this answer

The correct option is A, `kubectl logs pod-name --previous`, because when a container is in CrashLoopBackOff the current container instance may not have started or may have already been replaced, so the `--previous` (or `-p`) flag retrieves the logs from the last terminated instance of that container in the pod. This is the standard way to inspect why the container crashed on its prior run. Option D, `kubectl logs pod-name`, only shows logs from the currently running container instance, which may be empty or unavailable during a crash loop.

Option B, `kubectl get events`, shows cluster-level event messages but not the container's stdout/stderr logs, and option C, `kubectl describe pod pod-name`, shows pod status, conditions, and recent events but not the previous container's application logs.

54
Multi-Selecthard

You are troubleshooting a node that is 'NotReady'. Which THREE of the following are possible causes? (Choose three.)

Select 3 answers
A.The kubelet cannot contact the API server
B.The kubelet service is stopped
C.The node has disk pressure
D.A pod on the node is consuming excessive memory
E.The network plugin (e.g., Calico, Flannel) is not running
AnswersA, B, E

The kubelet is responsible for registering the node with the API server and continuously reporting its health and status. If the kubelet loses its ability to communicate with the API server, it cannot send its periodic heartbeats or update the node's conditions. After a default timeout period, the control plane will mark the node as NotReady because it has stopped receiving updates, indicating a potential issue with the node's availability or connectivity.

Why this answer

The kubelet is the primary node agent that communicates with the API server to report node status, heartbeats, and pod lifecycle events. If the kubelet cannot reach the API server (e.g., due to network partition, TLS certificate issues, or API server downtime), it cannot send the periodic NodeStatus updates, and the control plane marks the node as 'NotReady' after the `node-monitor-grace-period` (default 40 seconds) expires.

Exam trap

The trap here is that candidates confuse node conditions like 'DiskPressure' or 'MemoryPressure' with the 'NotReady' status, but these conditions do not change the 'Ready' status unless the kubelet itself fails to report.

55
MCQhard

You have a pod that is CrashLoopBackOff. The logs show 'error: dial tcp: lookup service.default.svc.cluster.local: no such host'. What is the most likely cause?

A.The CoreDNS pod is down
B.The service 'service' does not exist in the 'default' namespace
C.The pod's DNS policy is set to 'None'
D.A network policy is blocking UDP port 53
AnswerB

The application's connection string or manifest refers to a Kubernetes Service by the DNS name "service", but no Service object with that name exists in the "default" namespace. When the pod attempts to resolve it, CoreDNS returns NXDOMAIN because no A/AAAA record has been created for a non-existent Service, so the application fails to connect to its required dependency and exits, causing Kubernetes to restart the pod in a CrashLoopBackOff state. This is a configuration error in the workload definition, not an infrastructure or DNS-server failure.

Why this answer

The error message 'no such host' indicates that the DNS lookup for 'service.default.svc.cluster.local' failed because the hostname does not exist. In Kubernetes, this FQDN resolves only if a Service named 'service' exists in the 'default' namespace. Since the lookup fails with 'no such host', the most likely cause is that the Service does not exist, not a DNS infrastructure issue.

Exam trap

The trap here is that candidates often assume any DNS error means CoreDNS is down, but the specific 'no such host' message points to a missing DNS record, not a DNS service failure.

Why the other options are wrong

A

While CoreDNS being down could cause this, the error is about a specific service name not found, not a generic DNS failure. More likely the service doesn't exist.

C

If DNS policy was None, the pod would not even attempt cluster DNS; the error shows it tried but failed.

D

Network policies block traffic; but DNS would likely timeout or connection refused, not 'no such host'.

56
MCQhard

You are troubleshooting a pod that cannot start. Running 'kubectl describe pod' shows the event: 'Failed to pull image "myregistry.io/myapp:1.0": rpc error: code = Unknown desc = Error response from daemon: manifest for myregistry.io/myapp:1.0 not found'. What is the MOST likely cause?

A.The registry is unreachable due to network issues
B.The image tag '1.0' does not exist in the registry
C.The image registry requires authentication and the imagePullSecret is missing
D.The image has been deleted from the registry
AnswerB

The 'manifest not found' error explicitly indicates that the container runtime successfully contacted the image registry but could not locate the specific image manifest associated with the requested tag '1.0'. This means the registry confirmed its existence but reported that no image with that precise tag is available. This is the most direct and accurate interpretation of the given error message, signifying the tag itself is absent.

Why this answer

The error message 'manifest for myregistry.io/myapp:1.0 not found' indicates that the registry successfully received the pull request but could not locate the specific image tag '1.0'. This is a manifest lookup failure, not a connectivity or authentication issue. The most likely cause is that the tag '1.0' does not exist in the repository, either because it was never pushed or was removed.

Exam trap

The trap here is that candidates confuse 'manifest not found' with network or authentication errors, but the specific wording of the error message directly points to a missing tag in the registry, not connectivity or credentials.

How to eliminate wrong answers

Option A is wrong because network issues would produce a different error, such as 'dial tcp: lookup myregistry.io: no such host' or 'connection refused', not a manifest-not-found error. Option C is wrong because missing authentication would result in a 'denied: requested access to the resource is denied' or 'unauthorized: authentication required' error, not a manifest-not-found error. Option D is wrong because if the image had been deleted from the registry, the registry would typically still have the manifest metadata and would return a 'not found' for the blob, but the error specifically says 'manifest not found', which means the tag itself is missing—this is functionally the same as the tag never existing, but the phrasing 'deleted' implies the tag existed before, which is less likely given the exact error message; however, the most precise cause is that the tag does not exist in the registry's index.

57
Multi-Selectmedium

A node is 'NotReady'. Which THREE steps should you take to troubleshoot?

Select 3 answers
A.Check the kube-apiserver logs on the control plane
B.Reboot the node immediately
C.Check kubelet logs with 'journalctl -u kubelet'
D.SSH to the node and run 'systemctl status kubelet'
E.Run 'kubectl describe node <node-name>' to see conditions
AnswersC, D, E

The kubelet is the node agent that registers the node, runs pods, and continuously posts its status and heartbeat to the control plane, so its logs are the definitive source for why it stopped functioning. Running 'journalctl -u kubelet' reveals systemd unit output including fatal errors like failure to connect to the container runtime via the CRI socket, expired certificates, or resource exhaustion. You can also use '--since' or '-f' to focus on the time the node became NotReady and to watch for live errors, making this a primary diagnostic step.

Why this answer

Option C is correct because when a node reports NotReady, the kubelet is the primary agent responsible for reporting node status, so inspecting its logs with 'journalctl -u kubelet' reveals errors such as failed certificate rotation, container runtime connectivity issues, or PLEG problems. Option D is correct because 'systemctl status kubelet' on the node quickly shows whether the kubelet service is active, failed, or crash-looping, which is a fundamental first check for a NotReady node. Option E is correct because 'kubectl describe node <node-name>' displays the node's Conditions (Ready, MemoryPressure, DiskPressure, PIDPressure) and events, giving the exact reason and timestamp for the NotReady state.

Option A is not appropriate because kube-apiserver logs on the control plane do not diagnose why a specific node's kubelet stopped posting status; the API server merely reflects the missing heartbeats. Option B is not appropriate because rebooting the node immediately is a disruptive action that destroys diagnostic state and should only be done after evidence is gathered, not as a troubleshooting step.

58
MCQmedium

A pod is in ImagePullBackOff state. Which command is MOST useful to diagnose the issue?

A.kubectl logs <pod-name>
B.kubectl exec <pod-name> -- ls
C.kubectl get events --field-selector type=Warning
D.kubectl describe pod <pod-name>
AnswerD

kubectl describe pod is the correct and most useful command because it displays the pod's complete status, conditions, container states, and recent events in one place, including the exact kubelet-generated error from the failed image pull. This output reveals crucial details such as the image name/tag, whether the registry returned a 'manifest unknown' error, an authentication failure, or a network issue, and shows the exponential backoff timing, enabling a precise root-cause diagnosis.

Why this answer

`kubectl describe pod <pod-name>` provides detailed pod status, including the exact reason for ImagePullBackOff (e.g., invalid image name, registry authentication failure, or network issues). It surfaces the underlying error message from the kubelet, such as 'Failed to pull image' or 'manifest not found', which directly points to the root cause.

Exam trap

The trap here is that candidates often assume `kubectl logs` is the universal diagnostic tool, but it fails for pre-start states like ImagePullBackOff, where the container never runs to produce logs.

How to eliminate wrong answers

Option A is wrong because `kubectl logs` retrieves container logs, but ImagePullBackOff occurs before the container starts, so there are no logs to fetch. Option B is wrong because `kubectl exec` requires a running container to execute commands, but a pod in ImagePullBackOff never reaches the running state. Option C is wrong because `kubectl get events --field-selector type=Warning` may show related warnings, but it is less specific to the pod and may miss the exact error message; `kubectl describe pod` directly includes the relevant event in its output.

59
MCQhard

You run 'kubectl get pods' and see a pod with status 'ImagePullBackOff'. Which of the following is a possible cause?

A.The node has a disk pressure condition
B.The pod's resource limits are too low
C.The container image name is misspelled
D.The pod's liveness probe is failing
AnswerC

The ImagePullBackOff status indicates that the kubelet tried to pull the image from the repository but the attempt failed and it is backing off with exponential retry delay. A misspelled image name means the registry cannot find a matching repository or tag, often returning a 404 or 'manifest unknown' error, which leaves the kubelet in this backoff loop. For example, specifying 'ngin:x' instead of 'nginx:latest' will surface as ImagePullBackOff; fixing the typo and re-pulling resolves it.

Why this answer

The correct answer is C: the container image name is misspelled. ImagePullBackOff means the kubelet failed to pull the container image and is backing off before retrying, which commonly happens when the image reference is invalid, such as a typo in the image name or tag, or when the image does not exist in the registry. Option A is incorrect because disk pressure on a node typically causes Evicted pods or a DiskPressure node condition, not ImagePullBackOff.

Option B is incorrect because insufficient resource limits usually lead to Pending scheduling or OOMKilled states, not image pull failures. Option D is incorrect because a failing liveness probe causes container restarts and a CrashLoopBackOff-like state, not an image pull error.

60
MCQmedium

You try to run 'kubectl logs mypod' and get the error: 'Error from server (BadRequest): container "myapp" in pod "mypod" is waiting to start: PodInitializing'. What does this mean?

A.The pod has crashed and is restarting.
B.The kubelet is not running on the node.
C.The container has a different name than specified.
D.The pod is still being initialized and the container has not started yet.
AnswerD

This is correct because during the PodInitializing phase, the pod's init containers are still executing, meaning the primary application containers have not yet been created or started by the container runtime. Because the target application container does not exist in a running or terminated state yet, the kubelet cannot stream any log output, resulting in a waiting to start error.

Why this answer

The error message 'container "myapp" in pod "mypod" is waiting to start: PodInitializing' indicates that the pod's init containers (if any) are still running or the container runtime is pulling the image and setting up the container. The container has not yet entered the 'Running' state, so logs cannot be retrieved until it starts. This is a standard Kubernetes lifecycle phase where the pod is in 'PodInitializing' status, meaning the main container is not ready to serve logs.

Exam trap

The trap here is that candidates confuse 'PodInitializing' with a crash or restart loop, but the key distinction is that 'PodInitializing' is a transient startup phase, not a failure state, and logs are unavailable until the container actually starts.

How to eliminate wrong answers

Option A is wrong because 'PodInitializing' does not indicate a crash loop; a crashed container would show 'CrashLoopBackOff' or 'Error' status, not 'PodInitializing'. Option B is wrong because if the kubelet were not running, the pod would not be scheduled or would show 'NodeLost' or 'Unknown' status, not a specific container-level initialization error. Option C is wrong because the error explicitly names the container 'myapp', confirming the container name matches the one in the pod spec; a name mismatch would produce a different error like 'container "myapp" is not valid' or 'not found'.

61
MCQmedium

A pod in the 'production' namespace is in CrashLoopBackOff state. Running 'kubectl describe pod web-app -n production' shows the event 'OOMKilled'. What is the most appropriate action to resolve this issue?

A.Increase the CPU request for the container
B.Delete the namespace and redeploy
C.Increase the memory limit in the container spec
D.Delete and recreate the pod
AnswerC

The container was OOMKilled because its memory footprint hit the hard memory limit set in the container spec, and the Linux OOM killer terminated it to protect the node. Increasing the memory limit (and, if applicable, the memory request) gives the container more headroom under its cgroup, so it can allocate the additional memory it needs without being killed. This is the direct fix: update the Deployment or Pod template with a higher memory limit, which triggers a rolling update to apply the new limit.

Why this answer

OOMKilled means the container exceeded its memory limit. Increasing the memory limit is the correct fix.

62
MCQmedium

You need to check the memory usage of all pods in the 'production' namespace. Which command fulfills this requirement?

A.kubectl get pod --namespace=production -o wide
B.kubectl describe pod --namespace=production
C.kubectl top node
D.kubectl top pod --namespace=production
AnswerD

kubectl top pod --namespace=production is the correct kubectl command because it directly queries the metrics-server, which receives its data from kubelet cAdvisor, and returns the current CPU and memory usage for every pod in the production namespace. This is the only option among the four that actually measures live runtime memory consumption per pod, exactly satisfying the requirement. It produces a table with columns for NAME, CPU(cores), and MEMORY(bytes); note that without an explicit --namespace flag, kubectl top pod would default to the current namespace, so specifying production is essential.

Why this answer

`kubectl top pod --namespace=production` directly queries the metrics-server to retrieve real-time CPU and memory usage for each pod in the specified namespace. This command leverages the Kubernetes Metrics API to display resource consumption, making it the precise tool for checking memory usage of pods.

Exam trap

The trap here is that candidates confuse `kubectl describe` (which shows resource requests/limits but not actual usage) with `kubectl top` (which shows live consumption), or they mistakenly think `kubectl get pod -o wide` includes resource metrics.

How to eliminate wrong answers

Option A is wrong because `kubectl get pod -o wide` only shows pod metadata and node assignment, not memory usage metrics. Option B is wrong because `kubectl describe pod` provides detailed pod configuration and status, but does not include real-time memory usage data from the metrics-server. Option C is wrong because `kubectl top node` displays resource usage at the node level, not per-pod, and does not filter by namespace.

63
Multi-Selecteasy

Which TWO of the following kubectl commands can be used to view the logs of a container in a pod? (Choose two.)

Select 2 answers
A.kubectl describe pod pod-name
B.kubectl logs pod-name
C.kubectl exec pod-name -- logs
D.kubectl get pod pod-name -o yaml
E.kubectl logs pod-name --previous
AnswersB, E

This is the standard command used to retrieve the stdout and stderr streams from the primary container running inside a pod. If the pod contains multiple containers, you must specify the target container using the `-c` or `--container` flag, otherwise it defaults to the first container defined in the spec.

Why this answer

Option B, `kubectl logs pod-name`, is correct because the `kubectl logs` command is the dedicated subcommand for retrieving the stdout/stderr output of a container in a pod, and with a single container it defaults to that container. Option E, `kubectl logs pod-name --previous`, is also correct because the `--previous` (or `-p`) flag retrieves the logs from the previous, terminated instance of the container in the pod, which is still a valid way to view container logs. Option A, `kubectl describe pod pod-name`, is not correct because describe shows pod metadata, events, and status, not the container's log stream.

Option C, `kubectl exec pod-name -- logs`, is not correct because it attempts to run a `logs` binary inside the container rather than using the Kubernetes logs API. Option D, `kubectl get pod pod-name -o yaml`, is not correct because it only outputs the pod's YAML manifest/status, not its container logs.

Exam trap

The trap here is that candidates may confuse `kubectl exec` with `kubectl logs`, thinking they can run a 'logs' command inside the container, or they may mistakenly believe `kubectl describe` includes log output, when in fact it only shows events and container state.

64
MCQhard

A Pod is in 'CrashLoopBackOff' state. You run 'kubectl logs <pod> --previous' and see an error about a missing environment variable. The Pod spec defines the environment variable in a ConfigMap. What is the best next step to diagnose the issue?

A.Use 'kubectl exec -it <pod> -- env' to list environment variables
B.Run 'kubectl get configmap <configmap-name>' to verify the ConfigMap exists and contains the expected key
C.Increase the pod's memory limit to prevent OOM
D.Check the node's kubelet logs for errors
AnswerB

If a container crashes due to a missing environment variable sourced from a ConfigMap, verifying the existence of the target ConfigMap and its keys using kubectl get configmap is the correct troubleshooting step. This confirms whether the reference in the Pod specification is valid and populated.

Why this answer

The first step in diagnosing a missing environment variable that should come from a ConfigMap is to verify that the ConfigMap itself exists and contains the expected key. If the ConfigMap is missing or the key is absent, the Pod will fail to start or crash, leading to CrashLoopBackOff. Running 'kubectl get configmap' directly confirms whether the resource is present and correctly populated, which is the most efficient next step before deeper investigation.

Exam trap

The trap here is that candidates may jump to 'kubectl exec' to inspect environment variables, forgetting that a crashed Pod cannot be exec'd into, and instead should first verify the ConfigMap resource that the Pod depends on.

How to eliminate wrong answers

Option A is wrong because 'kubectl exec -it <pod> -- env' requires the Pod to be running, but a Pod in CrashLoopBackOff is not in a running state, so exec will fail. Option C is wrong because increasing memory limits addresses OOMKilled scenarios, not missing environment variables; the error message explicitly points to a missing variable, not resource exhaustion. Option D is wrong because checking node kubelet logs is a low-level diagnostic step for node-level issues (e.g., CNI, kubelet failures), not for Pod-level configuration errors like a missing ConfigMap key.

65
MCQeasy

A node in your cluster is reporting 'NotReady' status. You log into the node and run 'systemctl status kubelet'. The kubelet service is not running. Which command should you use to start the kubelet and enable it to start on boot?

A.systemctl start --enable kubelet
B.systemctl enable kubelet
C.systemctl enable --now kubelet
D.systemctl start kubelet
AnswerC

This is the correct command to resolve the `NotReady` status and ensure future stability. The `systemctl enable --now kubelet` command not only configures the `kubelet` service to start automatically during subsequent system boots but also immediately starts the service in the current session. This dual action ensures the `kubelet` is running right away, allowing the node to quickly transition to a `Ready` state without requiring a manual reboot.

Why this answer

`systemctl enable --now kubelet` both starts the kubelet service immediately and creates the necessary symlinks to enable it to start automatically on boot. This is the most efficient way to handle a stopped service that needs to be persistent across reboots, which is critical for a Kubernetes node to rejoin the cluster after a reboot.

Exam trap

The trap here is that candidates often confuse `systemctl start` with `systemctl enable`, or assume that `systemctl start` alone is sufficient, overlooking the requirement to persist the service across reboots, which is a common cause of nodes failing to rejoin after a reboot in production.

How to eliminate wrong answers

Option A is wrong because `systemctl start --enable` is not a valid systemctl syntax; the correct flag for simultaneous start and enable is `--now`. Option B is wrong because `systemctl enable kubelet` only creates the boot-time symlinks but does not start the service immediately, leaving the node in a NotReady state until a manual start or reboot. Option D is wrong because `systemctl start kubelet` starts the service only for the current session; after a reboot, the kubelet will not start automatically, and the node will again report NotReady.

66
Multi-Selecthard

Which THREE of the following are valid commands to troubleshoot network connectivity between pods? (Select 3)

Select 3 answers
A.kubectl exec pod-a -- nslookup service-name
B.kubectl describe node | grep Network
C.kubectl logs pod-a | grep network
D.kubectl exec pod-a -- curl http://pod-b
E.kubectl exec pod-a -- ping pod-b-ip
AnswersA, D, E

This command launches nslookup inside pod-a's network namespace, querying the cluster's CoreDNS/kube-dns service to resolve the name "service-name". Successful resolution proves that DNS pod(s), kube-dns service, and the pod's resolv.conf are functioning, and it returns the ClusterIP or headless endpoints needed for service discovery. It is a valid connectivity test because without DNS, application-level communication via service names fails regardless of reachable pods.

Why this answer

kubectl exec can run networking tools inside a pod. curl is a common tool. nslookup tests DNS. ping tests basic connectivity.

67
MCQmedium

After deploying a new Deployment, you run 'kubectl get events' and see 'FailedScheduling' events. What is a possible cause?

A.The container port is already in use on the node
B.The pod has a node selector that matches no nodes
C.The node has a taint that tolerates the pod
D.The pod's image pull secret is missing
AnswerB

When a pod defines a nodeSelector that does not match the labels of any active node in the cluster, the default-scheduler cannot find a valid placement. Consequently, the pod remains in a Pending state, and the scheduler emits a FailedScheduling warning event indicating that zero nodes match the selector.

Why this answer

A FailedScheduling event indicates that the Kubernetes scheduler could not find a suitable node to place the pod. Option B is correct because if a pod has a node selector that does not match any node's labels, the scheduler will fail to schedule it, resulting in a FailedScheduling event. The scheduler evaluates node selectors against node labels, and if no node satisfies the selector, the pod remains unscheduled.

Exam trap

The trap here is confusing scheduling failures with runtime failures, as candidates often associate port conflicts or image pull issues with scheduling, when in fact those errors occur after the pod is placed on a node.

How to eliminate wrong answers

Option A is wrong because a container port already in use on a node would cause a port conflict at runtime, not a scheduling failure; the scheduler does not check port availability on nodes. Option C is wrong because a node with a taint that tolerates the pod would actually allow scheduling, not prevent it; the issue is when a taint is not tolerated by the pod. Option D is wrong because a missing image pull secret would cause an ImagePullBackOff or ErrImagePull event after scheduling, not a FailedScheduling event.

68
MCQhard

A Pod is stuck in Pending state. 'kubectl describe pod' shows the event: '0/4 nodes are available: 1 node had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate, 3 Insufficient cpu.' Which of the following is the most likely combination of issues?

A.Three nodes have insufficient CPU for the pod's request, and one node has a taint not tolerated by the pod
B.The pod has a resource request that exceeds available CPU on all nodes
C.The pod does not tolerate any taints, and all nodes have taints
D.The cluster has only one node with sufficient CPU, but it is cordoned
AnswerA

The kubectl describe pod output lists Events from the scheduler. In this case, the events directly report two distinct issues: three nodes have insufficient CPU to satisfy the pod's resource request, and one node has a taint for which the pod has no matching toleration. Since these are the exact messages shown, this option correctly captures the full diagnosis.

Why this answer

The event message explicitly states that 1 node has a taint (node-role.kubernetes.io/control-plane) that the pod does not tolerate, and 3 nodes have insufficient CPU. This means the pod's CPU request cannot be satisfied on three nodes, and the remaining node is tainted, leaving no schedulable node. Option A correctly identifies this combination of issues.

Exam trap

The trap here is that candidates may misinterpret '0/4 nodes are available' as all nodes having the same issue, but the event message lists distinct reasons per node, requiring careful reading to identify the combination of taint and resource insufficiency.

How to eliminate wrong answers

Option B is wrong because the event shows only 3 nodes have insufficient CPU, not all 4; one node has a taint issue, not a CPU shortage. Option C is wrong because the event indicates only 1 node has a taint, not all nodes; the other 3 nodes have insufficient CPU, not taints. Option D is wrong because the event does not mention any node being cordoned; it specifically cites taint and insufficient CPU as the reasons.

69
MCQeasy

A pod named 'app' is not starting. You run 'kubectl describe pod app' and see the event: 'MountVolume.SetUp failed for volume "pvc-volume" : rpc error: code = NotFound desc = volume not found'. What is the most likely issue?

A.The container image is not found in the registry
B.The node running the pod is out of disk space
C.The pod has insufficient CPU resources
D.The PersistentVolumeClaim (PVC) referenced by the pod does not exist or is not bound
AnswerD

When a pod definition references a PersistentVolumeClaim that is either missing from the namespace or stuck in a 'Pending' state (unbound), the kubelet cannot mount the volume. Consequently, the pod remains blocked in the 'ContainerCreating' or 'Pending' phase, and 'kubectl describe' will display a warning event such as 'FailedMount' or 'FailedScheduling' due to the missing volume.

Why this answer

The error 'volume not found' indicates that the PersistentVolumeClaim (PVC) named 'pvc-volume' does not exist or is not bound to a PersistentVolume (PV). Kubernetes requires a PVC to be in 'Bound' state before a pod can mount it; if the PVC is missing or unbound, the volume mount fails at pod startup.

Exam trap

The trap here is that candidates may confuse volume mount errors with image or resource issues, but the specific 'volume not found' RPC error directly points to a missing or unbound PVC, not to node-level or image-related problems.

How to eliminate wrong answers

Option A is wrong because a missing container image would produce an 'ErrImagePull' or 'ImagePullBackOff' event, not a volume mount error. Option B is wrong because node disk space issues typically cause 'Evicted' pods or 'NodeHasDiskPressure' conditions, not a 'volume not found' RPC error from the CSI driver. Option C is wrong because insufficient CPU resources would result in a 'FailedScheduling' event or 'OOMKill' if the pod runs, not a volume mount failure.

70
MCQmedium

You run 'kubectl get pods' and see a pod named 'db' in CrashLoopBackOff. 'kubectl logs db' shows nothing. 'kubectl logs db --previous' shows 'Error: database connection failed'. What is the most likely cause?

A.The node is out of memory
B.The environment variable for database host is incorrect
C.The pod's liveness probe is misconfigured
D.The container image is missing
AnswerB

The application inside the container is crashing because it cannot establish a network connection to its dependency, which is a classic symptom of a misconfigured database host environment variable. When the application attempts to resolve or connect to an invalid hostname or IP address specified in its configuration, it throws an unhandled connection exception and terminates, causing Kubernetes to place the pod in a CrashLoopBackOff state.

Why this answer

The 'kubectl logs db --previous' output shows 'Error: database connection failed', which indicates the application inside the container is failing to connect to a database. Since the current logs are empty (the container restarted), the previous logs reveal the root cause: a configuration issue, most likely an incorrect database host environment variable. This is a classic application-level startup failure, not a resource or probe issue.

Exam trap

A common trap is that candidates may assume CrashLoopBackOff always indicates a probe or resource problem, ignoring the diagnostic value of 'kubectl logs --previous' which reveals application-level errors.

How to eliminate wrong answers

Option A is wrong because a node out-of-memory condition would cause the pod to be evicted or show OOMKilled status, not CrashLoopBackOff with a specific database connection error. Option C is wrong because a misconfigured liveness probe would cause the pod to be restarted by kubelet, but the logs would not show a database connection error; the probe failure would be logged by kubelet, not the application. Option D is wrong because a missing container image would result in ErrImagePull or ImagePullBackOff, not CrashLoopBackOff with application logs.

71
MCQhard

A Service of type ClusterIP is not reachable from within the cluster. Pods backing the Service are running and healthy. What is the most likely cause?

A.DNS resolution failure
B.kube-proxy not running or misconfigured
C.Ingress controller not set up
D.Service type should be NodePort
AnswerB

kube-proxy is the component responsible for implementing the Service abstraction by installing iptables, IPVS, or eBPF rules on every node. If it is not running, misconfigured, or its rules are out of sync, packets destined for the ClusterIP have no forwarding rules and are silently dropped or rejected. This directly matches the symptom of an unreachable ClusterIP while the backing pods and their endpoints may still be healthy.

Why this answer

When Pods are healthy but a ClusterIP Service is unreachable from within the cluster, the most common cause is that kube-proxy is not running or is misconfigured. kube-proxy is responsible for implementing the ClusterIP virtual IP by programming iptables (or IPVS) rules on each node to forward traffic to the selected Pods. Without these rules, packets destined for the ClusterIP are dropped or rejected, even though the Service and Endpoints objects exist.

Exam trap

The trap here is that candidates often assume a healthy Pod and Service object guarantee connectivity, overlooking that kube-proxy must actively program the underlying network rules to make the ClusterIP routable.

Why the other options are wrong

A

DNS resolves Service name to ClusterIP; connectivity fails after.

C

Ingress is for external; ClusterIP is internal.

D

ClusterIP works internally.

72
MCQmedium

You need to check the resource usage of nodes in your cluster. Which command should you run?

A.kubectl top nodes
B.kubectl get nodes -o wide
C.kubectl logs --all-containers
D.kubectl describe nodes
AnswerA

The "kubectl top nodes" command queries the Metrics Server API to retrieve and display the current, real-time CPU and memory utilization of all nodes in the cluster. This is the standard, built-in command used by administrators to quickly identify resource-constrained nodes. It requires the Metrics Server to be properly installed and running in the cluster to function.

Why this answer

`kubectl top nodes` retrieves and displays real-time CPU and memory usage metrics for all nodes in the cluster. This command relies on the metrics server being deployed and functioning, which aggregates resource usage data from kubelet’s cAdvisor endpoint. It is the standard Kubernetes command for checking node-level resource consumption.

Exam trap

The trap here is that candidates often confuse `kubectl describe nodes` (which shows static capacity and allocatable resources) with `kubectl top nodes` (which shows dynamic, real-time usage), leading them to choose option D when they need actual consumption data.

How to eliminate wrong answers

Option B is wrong because `kubectl get nodes -o wide` shows additional node information such as internal IP, external IP, and OS image, but does not display resource usage metrics. Option C is wrong because `kubectl logs --all-containers` retrieves container logs from a pod, not node-level resource usage. Option D is wrong because `kubectl describe nodes` provides detailed node status, conditions, and capacity/allocatable resources, but does not show current real-time resource consumption like `kubectl top nodes` does.

73
Multi-Selectmedium

Which of the following are valid methods to debug a failing CoreDNS pod? (Select TWO)

Select 2 answers
A.kubectl logs -n kube-system -l k8s-app=kube-dns
B.kubectl delete pod -n kube-system -l k8s-app=kube-dns
C.kubectl get endpoints -n kube-system kube-dns
D.kubectl scale deployment coredns --replicas=2
E.kubectl run -it --rm test --image=busybox -- nslookup kubernetes.default
AnswersA, C

This command retrieves the current logs from every CoreDNS pod by targeting the standard label selector `k8s-app=kube-dns` in the `kube-system` namespace. Since CoreDNS is normally deployed as a Deployment with multiple replicas, the `-l` flag aggregates output from all pods, allowing you to spot query processing errors, panics, or plugin misconfigurations (e.g., kubernetes plugin failures, upstream resolution timeouts) that would not be visible if you inspected only a single pod. Logs directly reveal whether DNS requests are reaching CoreDNS and how it handles them, making this the first-line diagnostic method.

Why this answer

`kubectl logs -n kube-system -l k8s-app=kube-dns` retrieves the logs from all CoreDNS pods (which are labeled `k8s-app=kube-dns` in the `kube-system` namespace). This is a direct method to inspect errors, crashes, or misconfigurations in the DNS service. Option C is correct because `kubectl get endpoints -n kube-system kube-dns` shows the IP addresses of the pods backing the `kube-dns` service; if the endpoint list is empty, it indicates that no CoreDNS pods are ready to serve DNS requests, which is a common failure point.

Exam trap

The trap here is that candidates confuse debugging actions (like checking logs or endpoints) with recovery actions (like deleting or scaling pods), and they may also mistake client-side DNS tests (like `nslookup`) for pod-level debugging, which only confirms the symptom, not the cause.

Why the other options are wrong

B

Deleting a pod is a fix, not a debug method.

D

Scaling is a fix, not debugging.

E

This tests DNS resolution, not debugging CoreDNS pod itself.

74
Multi-Selecthard

You have a Pod that is in CrashLoopBackOff. Which TWO of the following commands would be most helpful in diagnosing the issue?

Select 2 answers
A.kubectl get events --all-namespaces
B.kubectl logs <pod-name> --previous
C.kubectl describe pod <pod-name>
D.kubectl logs <pod-name>
E.kubectl top pod <pod-name>
AnswersB, C

When a container crashes and restarts, standard log commands only show the output of the currently running instance. Appending the `--previous` flag retrieves the stdout and stderr logs from the prior, terminated instance of the container, which typically contains the stack trace or error that caused the crash.

Why this answer

Option B, `kubectl logs <pod-name> --previous`, is correct because CrashLoopBackOff means the container has already restarted, so the `--previous` flag retrieves the logs from the last terminated container instance, which typically contains the actual crash/error output that caused the failure. Option C, `kubectl describe pod <pod-name>`, is correct because it shows the Pod's status conditions, container states (Waiting/Terminated with reason and exit code), restart count, and recent events such as image pull errors, OOMKilled, or failed probes, which are essential for pinpointing why the container keeps crashing. Option A, `kubectl get events --all-namespaces`, is not the best choice because it returns a broad cluster-wide event stream that is noisy and not scoped to the specific Pod, making it less efficient than describing the Pod itself.

Option D, `kubectl logs <pod-name>`, is not the most helpful here because without `--previous` it targets the current container instance, which may not have produced logs yet or may not exist if the container is stuck restarting. Option E, `kubectl top pod <pod-name>`, is incorrect because it only reports CPU and memory usage metrics and provides no diagnostic information about the crash cause.

Exam trap

The trap here is that candidates often pick `kubectl logs <pod-name>` (option D) without `--previous`, not realizing that in a CrashLoopBackOff the current container may have no useful logs, while the previous terminated container holds the error details.

75
MCQhard

A pod is stuck in Pending with event '0/4 nodes are available: 1 node(s) had taint "node.kubernetes.io/disk-pressure", and 3 node(s) had taint "node.kubernetes.io/memory-pressure", that the pod didn't tolerate'. What is the best approach to schedule the pod?

A.Use 'kubectl taint nodes --all node.kubernetes.io/disk-pressure- node.kubernetes.io/memory-pressure-' to remove the taints
B.Add tolerations for both taints to the pod spec
C.Free up disk space and reduce memory usage on the affected nodes
D.Delete the pod and recreate it with higher resource requests
AnswerC

The correct action is to free disk space and reduce memory consumption on the affected nodes. When the underlying pressure conditions clear, the kubelet automatically removes the corresponding taints, allowing the pending pod to schedule normally. This directly addresses the root cause of the unschedulable state and aligns with Kubernetes' intended behavior of reflecting node health through taints and conditions.

Why this answer

The cluster has nodes with disk-pressure and memory-pressure taints. These are typically added by the node controller when resources are low. The pod does not tolerate these taints.

The best approach is to free resources on the nodes or add tolerations; however, tolerating pressure taints is not recommended as it can cause further instability. Usually, you should resolve the underlying resource issues.

Page 1 of 3 · 207 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Troubleshooting questions.