Courseiva
Troubleshooting →mediumMultiple Select

CKA Troubleshoot NotReady node Practice Question

Which of the following are valid methods to troubleshoot a Node that is 'NotReady'? (Select all that apply)

⚠ Common exam trap

CNCF often tests the misconception that deleting and re-creating a Node object is a valid troubleshooting step, when in reality it only removes the API server's representation and does not fix the underlying cause of the 'NotReady' state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check the kubelet logs on the node via journalctl

The kubelet is the primary node agent that communicates with the control plane. When a node is 'NotReady', checking the kubelet logs via `journalctl -u kubelet` can reveal errors such as network connectivity issues, certificate problems, or resource exhaustion. Option B is correct because an expired or invalid kubelet certificate will cause TLS authentication failures with the API server, leading to the node being marked 'NotReady'. Option C is correct because the container runtime (e.g., containerd) is responsible for managing containers; if it is down or misconfigured, the kubelet cannot start pods, and restarting it can resolve transient failures.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Check the kubelet logs on the node via journalctl

    Why this is correct

    Checking kubelet logs via journalctl is the primary, non-destructive diagnostic step: the kubelet is the agent that registers the node and reports its status, so any failure in its startup or operation (e.g., CNI misconfiguration, resource pressure, or runtime communication errors) will be logged there. Running `journalctl -u kubelet` (or `journalctl -u kubelet --since today`) reveals the exact error before deciding on any remediation. This step should always precede restarting services or modifying cluster objects, because it pinpoints the underlying cause without causing downtime.

  • ✓

    Verify that the kubelet certificate is valid and not expired

    Why this is correct

    Verifying the kubelet certificate is valid and not expired addresses a common root cause: the kubelet authenticates to the API server using a client certificate, typically located in /etc/kubernetes/kubelet.conf or /var/lib/kubelet/pki/kubelet-client-current.pem. If the certificate has expired or failed to renew (for example, when automatic rotation is disabled or the cluster has been running past its certificate lifetime), the kubelet cannot communicate with the control plane, and the node will enter NotReady. Use `openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -text -noout` to check the validity period, and inspect the kubelet log for authentication errors like `Unauthorized` or `certificate has expired`.

  • ✓

    Restart the node's container runtime (e.g., containerd)

    Why this is correct

    Restarting the node's container runtime (e.g., containerd) is a valid recovery step when the runtime has encountered a transient failure, such as a stuck container state, a dead CRI plugin, or a resource leak that prevents the kubelet from creating/stopping pods. Because the kubelet depends on the CRI to execute container operations, a malfunctioning runtime can cause the node to stop reporting or to fail readiness checks. However, a restart should be performed only after checking logs and understanding the root cause; blindly restarting the runtime without diagnosis may mask a persistent issue like a broken CNI config or a corrupt overlay filesystem.

  • ✗

    Delete the node object and re-create it

    Why it's wrong here

    Deleting the node object and re-creating it is an incorrect action because the Node object in Kubernetes is only a representation of the kubelet's registration and status; the kubelet will automatically re-create the object when it next heartbeats to the API server. If the underlying problem—such as a dead kubelet, expired certificate, or broken runtime—persists, the re-created node will immediately become NotReady again. Furthermore, deleting the node object forces the control plane to mark all its pods as lost and may evict workloads, causing unnecessary disruption. The proper approach is to fix the actual fault on the node, not to manipulate the cluster-side representation.

  • ✗

    Reinstall the entire Kubernetes cluster

    Why it's wrong here

    Reinstalling the entire Kubernetes cluster is a drastic and inappropriate remedy for a single NotReady node because it involves tearing down all control-plane and worker nodes, affecting every workload and requiring full re-provisioning of infrastructure. The NotReady condition is typically node-local—caused by kubelet, runtime, networking, or resource issues—so reinstalling the whole cluster would not address the specific root cause and would introduce massive downtime. Even if the node has an unrecoverable OS or filesystem problem, the correct procedure is to replace or repair that single node, not reinstall the cluster. Troubleshooting should be scoped to the failing node and its components, using logs and targeted restarts first.

Option-by-option analysis

Why each answer is right or wrong

Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The CKA exam frequently reuses these exact scenarios with slightly different constraints.

✓Check the kubelet logs on the node via journalctlCorrect answer▾

Why this is correct

Checking kubelet logs via journalctl is the primary, non-destructive diagnostic step: the kubelet is the agent that registers the node and reports its status, so any failure in its startup or operation (e.g., CNI misconfiguration, resource pressure, or runtime communication errors) will be logged there. Running `journalctl -u kubelet` (or `journalctl -u kubelet --since today`) reveals the exact error before deciding on any remediation. This step should always precede restarting services or modifying cluster objects, because it pinpoints the underlying cause without causing downtime.

✗Delete the node object and re-create itWrong answer — click to see why▾

Why this is wrong here

Deleting the node does not fix the node condition; it only removes it from the cluster.

✗Reinstall the entire Kubernetes clusterWrong answer — click to see why▾

Why this is wrong here

Overkill; a single node issue doesn't require cluster reinstall.

Analysis generated from the official CKAblueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

One of 726 original CKA practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This CKA practice question is part of Courseiva's free CNCF certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CKA exam.