Courseiva

CCNA Workloads Scheduling Questions

20 questions · Workloads Scheduling topic · All types, answers revealed

1
MCQeasy

You are managing a Kubernetes cluster with three worker nodes. A deployment named 'frontend' is configured with 3 replicas. After a node failure, you notice that only 2 pods are running, and the third pod is stuck in 'Pending' state. The remaining nodes have sufficient CPU and memory. You check the deployment events and find no errors. You also verify that the PersistentVolumeClaims (PVCs) used by the deployment are bound. What is the most likely reason the third pod is not scheduled?

A.The ReplicaSet controller is not creating a new pod because the deployment's progressDeadlineSeconds has expired.
B.The PersistentVolumeClaims are using 'WaitForFirstConsumer' binding mode and the pod is pending because the volume is not yet bound.
C.The kube-scheduler is down or misconfigured.
D.The pod has a nodeSelector that only matches the failed node.
AnswerD

When a pod specifies a nodeSelector that targets a unique label only present on the failed worker node, the kube-scheduler cannot find any other eligible nodes that satisfy this constraint. Consequently, the pod remains unschedulable and stuck in a Pending state, even though other healthy worker nodes are available in the cluster.

Why this answer

A nodeSelector that exclusively matches the failed node would prevent the scheduler from placing the pod on any other node, even if those nodes have sufficient resources. Since the failed node is unavailable, the pod remains in 'Pending' state indefinitely, as no other node satisfies the constraint.

Exam trap

The trap here is that candidates assume resource constraints or scheduler failures are the default cause for pending pods, overlooking that a nodeSelector or affinity rule can silently prevent scheduling even when resources are abundant.

How to eliminate wrong answers

Option A is wrong because progressDeadlineSeconds only triggers a deployment rollout failure (marking the deployment as progressing false) but does not prevent the ReplicaSet from creating or scheduling pods; it is a rollout health check, not a scheduling blocker. Option B is wrong because the PVCs are already bound (as stated), so 'WaitForFirstConsumer' would cause the volume to bind only after scheduling, but the pod would still be pending due to volume binding, not because the volume is unbound—the scenario says PVCs are bound, eliminating this. Option C is wrong because if the kube-scheduler were down or misconfigured, all pods would be stuck in 'Pending', not just one; the fact that two pods are running indicates the scheduler is functional.

2
MCQeasy

A cluster administrator wants to ensure that no pods are scheduled on the master node(s). Which approach is the best practice?

A.Add a taint to the master node
B.Delete the master node from the cluster
C.Use a resource quota on the master namespace
D.Set nodeSelector on the master node
AnswerA

Applying a NoSchedule or NoExecute taint to the master (control plane) node ensures that the Kubernetes scheduler will not place any pods on it unless they have a matching toleration. While modern Kubernetes clusters apply this taint by default to protect control plane resources, manually adding or verifying this taint is the standard declarative method to enforce this scheduling restriction.

Why this answer

Adding a taint to the master node(s) with the `node-role.kubernetes.io/master:NoSchedule` effect is the best practice because it prevents the Kubernetes scheduler from placing any pods on that node unless a pod explicitly tolerates the taint. This ensures that only critical system pods (which include the toleration) can run on the master, keeping it dedicated to cluster control plane operations.

Exam trap

The trap here is that candidates often confuse `nodeSelector` (a pod scheduling constraint) with node-level restrictions, or think that deleting a node or using resource quotas can control scheduling to a specific node, when only taints (or node affinity with requiredDuringSchedulingIgnoredDuringExecution) provide that node-level control.

How to eliminate wrong answers

Option B is wrong because deleting the master node from the cluster would remove the control plane, making the cluster non-functional; the goal is to prevent pod scheduling, not to remove the node entirely. Option C is wrong because a resource quota on a namespace (e.g., 'master' namespace) limits resource consumption but does not prevent pods from being scheduled onto a specific node; pods could still be placed on the master node from any namespace. Option D is wrong because `nodeSelector` is used to constrain which nodes a pod can be scheduled on, but it is a pod-level attribute, not a node-level restriction; setting it on the master node itself is not a valid operation and does not prevent other pods from being scheduled there.

3
MCQmedium

An administrator wants to ensure that a critical Pod always runs on a node that has an SSD. Which approach should be used?

A.Use podAffinity to prefer nodes running other Pods.
B.Set spec.nodeName to an SSD node's name.
C.Add a taint to nodes without SSD and add a toleration to the Pod.
D.Label SSD nodes with 'disk=ssd' and use nodeSelector in the Pod spec.
AnswerD

Applying a custom label like 'disk=ssd' to the target nodes and referencing it via nodeSelector in the Pod specification is the standard, declarative way to constrain scheduling. The Kubernetes scheduler will strictly filter out any nodes lacking this label, ensuring the Pod is only scheduled on SSD-equipped hardware.

Why this answer

Using a nodeSelector with a label like 'disk=ssd' ensures that the Pod is scheduled only on nodes that have that specific label, which directly matches the requirement to run the Pod on a node with an SSD. This is the simplest and most reliable method for node-level selection in Kubernetes, as it leverages the scheduler's built-in filtering mechanism.

Exam trap

The trap here is that candidates often confuse taints and tolerations with node selection, thinking they can force a Pod onto a specific node type, when in fact taints only repel Pods from nodes and require tolerations to allow scheduling, not to guarantee placement on a desired node.

How to eliminate wrong answers

Option A is wrong because podAffinity is used to schedule Pods relative to other Pods (e.g., co-location), not to select nodes based on hardware characteristics like SSD presence. Option B is wrong because setting spec.nodeName bypasses the scheduler entirely and directly assigns the Pod to a specific node by name, which is inflexible and does not scale; it also requires manual knowledge of the node name and does not use labels or taints for dynamic selection. Option C is wrong because adding a taint to nodes without SSD would repel Pods that do not tolerate that taint, but the Pod would need a toleration to run on those nodes; this approach would prevent the Pod from running on non-SSD nodes but does not actively ensure it runs on an SSD node—it could still be scheduled on any node without the taint, including those without SSD if they are not tainted.

4
MCQhard

A cluster administrator wants to ensure that a set of batch processing Pods are preemptible and should not cause disruption to other critical workloads. Which combination of scheduling features should be used?

A.Taint all worker nodes with a custom taint and add tolerations to batch Pods.
B.Assign a low priority class to the batch Pods and set a PodDisruptionBudget for critical workloads.
C.Use nodeAffinity to schedule batch Pods on dedicated nodes.
D.Use resource quotas to limit the batch Pods' resource consumption.
AnswerB

Defining a low PriorityClass for batch pods ensures that the kube-scheduler will preempt them if higher-priority critical pods require those node resources. Simultaneously, configuring a PodDisruptionBudget (PDB) for the critical workloads guarantees that voluntary disruptions do not drop their replica counts below a safe threshold, balancing resource availability with application resilience.

Why this answer

Assigning a low priority class to batch Pods ensures they are preempted by higher-priority critical workloads when resources are scarce, while a PodDisruptionBudget (PDB) for critical workloads guarantees that a minimum number of those Pods remain available during voluntary disruptions (e.g., node drains). This combination allows batch Pods to be preemptible without causing disruption to critical workloads, aligning with the requirement.

Exam trap

The trap here is that candidates often confuse taints/tolerations or node affinity with preemption and disruption protection, failing to realize that only priority classes enable preemption and PDBs control voluntary disruptions.

How to eliminate wrong answers

Option A is wrong because tainting worker nodes with a custom taint and adding tolerations to batch Pods only prevents other Pods from scheduling on those nodes unless they have the toleration; it does not make batch Pods preemptible or protect critical workloads from disruption. Option C is wrong because using nodeAffinity to schedule batch Pods on dedicated nodes isolates them but does not provide preemption capability or disruption protection for critical workloads; dedicated nodes can still be disrupted by node failures or maintenance. Option D is wrong because resource quotas limit the total resource consumption of batch Pods but do not make them preemptible or prevent disruption to critical workloads; quotas only enforce resource caps, not scheduling priority or availability guarantees.

5
MCQeasy

A Kubernetes cluster has a deployment with 3 replicas. After a node failure, you notice that only 2 pods are running, and the deployment has not rescheduled the missing pod. What is the most likely cause?

A.The deployment has a resource quota that prevents new pods
B.The pod's terminationGracePeriodSeconds is set to 0
C.The node controller has not yet evicted the pod
D.The deployment's replicas field is set to 2
AnswerC

When a node becomes unreachable, the node controller waits for a default grace period of five minutes (configured via the --pod-eviction-timeout flag) before marking the pods for eviction. During this window, the control plane keeps the pods in a Terminating or Unknown state on the failed node and does not reschedule them. Only after this timeout expires will the deployment controller spin up a replacement pod on a healthy node.

Why this answer

When a node fails, the node controller marks the node as `NodeReady=False` and waits for a configurable timeout (`pod-eviction-timeout`, default 5 minutes) before evicting pods. Until eviction, the deployment's ReplicaSet sees the pod as still existing (though on an unreachable node) and does not create a replacement. Option C correctly identifies that the node controller has not yet evicted the pod, which is the default behavior.

Exam trap

The trap here is that candidates assume a deployment immediately reschedules pods after a node failure, but the CKA tests knowledge of the node controller's eviction timeout and the fact that the ReplicaSet controller waits for pod eviction before creating replacements.

How to eliminate wrong answers

Option A is wrong because a resource quota would cause a failure to create a new pod (with an error event), not prevent the deployment from attempting rescheduling; the deployment would still try and fail, but the question states it has not rescheduled the missing pod. Option B is wrong because `terminationGracePeriodSeconds` affects how long a pod is given to shut down gracefully after a deletion request, not the node controller's eviction timeout or the deployment's rescheduling logic. Option D is wrong because if the deployment's `replicas` field were set to 2, the deployment would be satisfied with 2 running pods and would not attempt to reschedule; however, the question states the deployment originally had 3 replicas, so this would be a configuration change, not a transient failure scenario.

6
MCQeasy

A Pod with a restartPolicy of 'OnFailure' exits with code 0. What will happen?

A.The container will restart immediately.
B.The Pod will be terminated.
C.The Pod will remain in Running state.
D.The container will not restart, and the Pod will be in Succeeded phase.
AnswerD

With restartPolicy: OnFailure, an exit code of 0 means the container completed its workload successfully, so the kubelet deliberately avoids restarting it. Because all containers have exited with code 0 and no restart is warranted, the Pod's phase is set to Succeeded, which is the terminal state for successful Pod completion. This matches the expected behavior for batch or job-style workloads that should run to completion exactly once.

Why this answer

When a Pod has a restartPolicy of 'OnFailure' and its container exits with code 0 (indicating successful completion), the container will not be restarted. Instead, the Pod transitions to the Succeeded phase, as defined by Kubernetes Pod lifecycle semantics. This is because 'OnFailure' only triggers a restart on a non-zero exit code, which signifies a failure.

Exam trap

The trap here is that candidates often confuse 'OnFailure' with 'Always', assuming any exit triggers a restart, or they mistakenly think a Pod is 'terminated' (deleted) when it actually enters a terminal phase like Succeeded.

How to eliminate wrong answers

Option A is wrong because the container will restart only if the exit code is non-zero; exit code 0 indicates success, so no restart occurs. Option B is wrong because the Pod is not terminated; it enters the Succeeded phase, which is a terminal phase but not a termination of the Pod object itself. Option C is wrong because the Pod cannot remain in the Running state after the container exits; it must transition to a terminal phase (Succeeded or Failed) based on the exit code and restartPolicy.

7
MCQmedium

You are a platform engineer managing a Kubernetes cluster with 5 worker nodes (node1-node5). The cluster runs a mix of stateless web services and stateful databases. Users report that a critical database Pod (part of a StatefulSet) is frequently evicted during node maintenance. The StatefulSet has a single replica. You need to improve the availability of this database Pod. The current configuration: the Pod has resource requests (2 CPU, 4Gi memory) and limits (4 CPU, 8Gi memory). The cluster uses the default scheduler with no custom policies. Nodes have varying capacities: node1 and node2 have 8 CPU/32Gi memory, node3-node5 have 4 CPU/16Gi memory. During rolling node reboots, the database Pod gets evicted and takes a long time to reschedule because no node has enough resources. What should you do to minimize downtime and ensure the Pod is rescheduled promptly after eviction?

A.Add nodeAffinity to prefer node1 and node2.
B.Create a PodDisruptionBudget with minAvailable: 1.
C.Assign a high priority class to the database Pod.
D.Increase the resource requests to match the limits.
AnswerC

Assigning a high priority class to the database Pod is the most effective solution for ensuring critical workloads are scheduled promptly. When the scheduler attempts to place a high-priority Pod and cannot find a node with sufficient resources, it will actively preempt (evict) lower-priority Pods from existing nodes to free up the necessary capacity. This mechanism directly addresses resource contention, significantly reducing the scheduling delay for essential applications like a database.

Why this answer

Assigning a high priority class (Option C) ensures that when the database Pod is evicted during node maintenance, the scheduler treats it as a higher-priority workload than other Pods. This allows it to preempt lower-priority Pods on nodes with sufficient capacity (e.g., node1 or node2), even if those nodes appear fully allocated, thereby minimizing downtime and ensuring prompt rescheduling.

Exam trap

CNCF often tests the distinction between disruption budgets (which prevent eviction) and priority/preemption (which ensure rescheduling after eviction), leading candidates to mistakenly choose PDB when the real issue is resource contention after eviction.

How to eliminate wrong answers

Option A is wrong because nodeAffinity with a 'prefer' rule is a soft scheduling preference, not a guarantee; during eviction, the scheduler may still place the Pod on a smaller node if node1/node2 are full, leading to scheduling failures. Option B is wrong because a PodDisruptionBudget (PDB) with minAvailable: 1 only protects against voluntary disruptions (e.g., node drains) by preventing eviction if it would violate the budget, but it does not help with resource availability after eviction; the Pod still cannot be scheduled if no node has enough free resources. Option D is wrong because increasing resource requests to match limits (4 CPU, 8Gi memory) would make the Pod even harder to schedule, as it would require more resources than the smaller nodes (node3-node5) can provide, worsening the problem.

8
MCQmedium

Your team is deploying a new application that consists of a web frontend and a backend API. The frontend must be accessible from outside the cluster, and the backend should only be accessible from within the cluster. The cluster has multiple namespaces: 'frontend' and 'backend'. You have been asked to design the deployment. The frontend Deployment should have 5 replicas, and the backend Deployment should have 3 replicas. Additionally, you need to ensure that the frontend pods can communicate with the backend pods using a stable DNS name. You also want to isolate the backend from other namespaces. Which set of resources should you create?

A.Frontend: Deployment, Service (NodePort); Backend: Deployment, Service (ClusterIP); no NetworkPolicy
B.Frontend: Deployment, Service (ClusterIP); Backend: Deployment, Service (ClusterIP); NetworkPolicy to allow ingress from frontend namespace
C.Frontend: Deployment, Service (LoadBalancer); Backend: Deployment, Service (LoadBalancer); NetworkPolicy to allow only frontend to backend
D.Frontend: Deployment, Service (LoadBalancer); Backend: Deployment, Service (ClusterIP); NetworkPolicy to allow ingress from frontend namespace and deny others
AnswerD

This architecture correctly leverages a LoadBalancer Service to route external public traffic to the frontend deployment, while keeping the backend isolated using an internal-only ClusterIP Service. Additionally, the NetworkPolicy enforces strict zero-trust security by explicitly allowing ingress traffic only from the frontend namespace while dropping all other non-compliant cluster traffic.

Why this answer

It uses a LoadBalancer Service for the frontend to provide external access, a ClusterIP Service for the backend to restrict access to within the cluster, and a NetworkPolicy that allows ingress traffic from the frontend namespace to the backend while denying all other ingress, thus isolating the backend. This ensures the frontend pods can reach the backend via a stable DNS name (the ClusterIP Service's DNS name) and meets the requirement of backend isolation from other namespaces.

Exam trap

The trap here is that candidates often forget that a ClusterIP Service cannot be accessed from outside the cluster, and they may incorrectly choose a LoadBalancer or NodePort for the backend, or omit the NetworkPolicy needed to enforce isolation.

How to eliminate wrong answers

Option A is wrong because it uses a NodePort Service for the frontend, which exposes the frontend on a high port on every node but does not provide a stable external endpoint like a LoadBalancer, and it lacks a NetworkPolicy to isolate the backend. Option B is wrong because it uses a ClusterIP Service for the frontend, which does not make the frontend accessible from outside the cluster, violating the requirement. Option C is wrong because it uses a LoadBalancer Service for the backend, which exposes the backend externally, contradicting the requirement that the backend should only be accessible from within the cluster.

9
MCQeasy

A Pod is in Pending state for a long time. 'kubectl describe pod' shows the event: '0/3 nodes are available: 3 node(s) had taints that the pod didn't tolerate'. What is the most likely issue?

A.The nodes do not have enough CPU or memory to run the Pod.
B.The Pod does not have tolerations for the taints on the nodes.
C.The nodes are not tainted, but the Pod has tolerations.
D.The scheduler is not running.
AnswerB

When a node is tainted with a NoSchedule or NoExecute effect, the kube-scheduler will reject any pod that lacks matching tolerations. The kubectl describe pod output confirms this scenario by displaying a scheduling event indicating that the pod had 0/N nodes available due to untolerated taints. Adding the appropriate tolerations to the pod specification is required to allow scheduling on these nodes.

Why this answer

The event '0/3 nodes are available: 3 node(s) had taints that the pod didn't tolerate' directly indicates that all nodes in the cluster have taints applied, and the Pod's spec does not include corresponding tolerations. Taints and tolerations work together to ensure that Pods are only scheduled onto nodes whose taints they tolerate; without matching tolerations, the scheduler will skip those nodes entirely, leaving the Pod in a Pending state.

Exam trap

CNCF often tests the distinction between taints/tolerations and resource constraints, so candidates may incorrectly assume the error is about resource exhaustion when the event message explicitly mentions taints.

How to eliminate wrong answers

Option A is wrong because insufficient CPU or memory would produce a different event, such as 'Insufficient cpu' or 'Insufficient memory', not a taint-related message. Option C is wrong because if nodes are not tainted, the Pod would be scheduled normally regardless of tolerations; the event explicitly states nodes had taints, so this scenario contradicts the observed error. Option D is wrong because if the scheduler were not running, the Pod would remain in Pending state without any scheduling events at all, and 'kubectl describe pod' would not show taint-related messages.

10
MCQhard

You are managing a Kubernetes cluster that hosts a microservices application. One of the services, 'payment-processor', is critical and must always be available. It has a Deployment with 3 replicas, each requesting 1 CPU and 2Gi memory. Recently, the team added a new service 'data-analyzer' that runs as a DaemonSet on all nodes, consuming significant CPU and memory. After the addition, you notice that 'payment-processor' pods are occasionally being evicted, and new pods are slow to be scheduled. You check node resource usage and find that some nodes are overcommitted. You want to ensure that 'payment-processor' pods are never evicted and are scheduled before less critical workloads. Which action should you take?

A.Add a taint to nodes that have low resources and add tolerations only to 'payment-processor' pods
B.Increase the resource requests for 'payment-processor' pods to guarantee resources
C.Create a PriorityClass with a high value and assign it to the 'payment-processor' Deployment
D.Use node affinity to ensure 'payment-processor' pods run on dedicated nodes
AnswerC

Creating a PriorityClass with a high integer value and assigning it to the 'payment-processor' Deployment is the most effective solution. Pods with higher priority are preferentially scheduled by the kube-scheduler. Crucially, if a high-priority 'payment-processor' pod cannot be scheduled due to insufficient resources on any node, the scheduler will attempt to preempt (evict) lower-priority pods on suitable nodes to free up the necessary resources, thereby ensuring the critical workload runs.

Why this answer

PriorityClass with a high value ensures that 'payment-processor' pods are considered higher priority than other pods during scheduling and eviction. When nodes are overcommitted, the Kubernetes scheduler will preempt lower-priority pods to make room for higher-priority pods, and the kubelet will evict lower-priority pods first when resources are scarce. This directly addresses the requirement that 'payment-processor' pods are never evicted and are scheduled before less critical workloads.

Exam trap

The trap here is that candidates often confuse taints/tolerations or node affinity with priority and preemption, but those features only affect scheduling placement, not eviction ordering or preemption behavior.

How to eliminate wrong answers

Option A is wrong because taints and tolerations control which pods can be scheduled on a node, but they do not provide a mechanism for eviction priority or guarantee that 'payment-processor' pods will be scheduled before other pods on the same node; taints only repel pods without tolerations, and adding tolerations to 'payment-processor' would allow them to schedule on tainted nodes but not prevent eviction. Option B is wrong because increasing resource requests for 'payment-processor' pods would require more resources to schedule them, potentially making scheduling harder, and it does not affect eviction ordering; requests only affect scheduling decisions, not eviction priority. Option D is wrong because node affinity only influences scheduling placement, not eviction behavior; it can ensure pods run on specific nodes but does not prevent eviction when those nodes are overcommitted, nor does it prioritize scheduling over other workloads.

11
Drag & Dropmedium

Drag and drop the steps to troubleshoot a Node that is in NotReady state into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Start with kubectl to identify the node, then SSH, check kubelet and runtime, review logs, then restart.

12
MCQhard

An operations team must guarantee that exactly one copy of a log-shipping Pod runs on every worker node, including nodes added to the cluster later, and that each node gets at most one such Pod. Some nodes carry the taint 'dedicated=batch:NoSchedule'. Which resource and configuration should the team use?

A.A DaemonSet whose Pod template includes a toleration for the 'dedicated=batch:NoSchedule' taint.
B.A StatefulSet with a headless Service and podAntiAffinity across hostnames.
C.A Job with completions set to the number of nodes and parallelism equal to one.
D.A Deployment with replicas equal to the number of nodes, plus a podAntiAffinity rule.
AnswerA

A DaemonSet creates exactly one Pod on each eligible node and automatically adds Pods when new nodes join the cluster, which matches the every-node requirement. Because tainted nodes would otherwise be skipped, the Pod template must include a matching toleration so the DaemonSet controller schedules onto those nodes as well. This combination satisfies both the coverage and the one-per-node constraints.

Why this answer

A DaemonSet is the only workload controller designed to place exactly one Pod on every eligible node and to extend that coverage automatically as nodes join. Since the dedicated batch taint would otherwise repel the Pods, the Pod template must carry a matching toleration so those nodes are included in the rollout.

Exam trap

The trap here is believing that a toleration alone attracts Pods to tainted nodes, when a DaemonSet is what provides the one-per-node coverage and the toleration merely removes the repulsion.

13
MCQmedium

A company runs a batch job that processes a queue. The job should run to completion exactly once. Which resource should be used?

A.Job
B.DaemonSet
C.CronJob
D.Deployment
AnswerA

A Kubernetes Job is specifically designed to run a finite batch workload to completion. It creates one or more Pods and ensures that a specified number of them successfully terminate, making it the ideal resource for processing a finite work queue. Once the tasks are complete, the Job stops creating Pods, conserving cluster resources.

Why this answer

A Kubernetes Job is designed to run a specified number of pods to successful completion. When the pod exits with a zero exit code, the Job is marked as complete and will not be restarted, making it the correct choice for a batch job that must run exactly once.

Exam trap

The trap here is that candidates often confuse a CronJob with a Job, thinking that a CronJob can also run a one-time task, but CronJob is specifically for scheduled, recurring execution, not a single run-to-completion workload.

How to eliminate wrong answers

Option B (DaemonSet) is wrong because it ensures that a copy of a pod runs on every node (or a subset of nodes) in the cluster, continuously, not for a one-time batch job. Option C (CronJob) is wrong because it creates Jobs on a recurring schedule; while it could run a batch job, it is intended for periodic execution, not a single run-to-completion task. Option D (Deployment) is wrong because it manages a set of identical pods with a desired replica count, ensuring they are always running and self-healing, which is the opposite of a run-to-completion workload.

14
MCQeasy

A developer wants to deploy a pod that will run only once to initialize a database schema. Which Kubernetes resource should they use?

A.DaemonSet
B.Job
C.Deployment
D.CronJob
AnswerB

A Job controller creates one or more pods and tracks them until a specified number successfully terminate. For a one-time task, a simple Job with default completions=1 runs once to completion, and the workload is not recreated if it exits with code 0. It is the native Kubernetes API for exactly-once batch processing.

Why this answer

A Job is the correct Kubernetes resource for a one-time task that runs to completion, such as initializing a database schema. Unlike controllers that maintain a desired number of replicas, a Job creates one or more Pods and ensures they successfully terminate, making it ideal for batch or initialization workloads.

Exam trap

The trap here is that candidates often confuse a Job with a CronJob, thinking they need scheduling, or with a Deployment, assuming all workloads must be continuously running, when the key differentiator is the 'run to completion' lifecycle.

How to eliminate wrong answers

Option A is wrong because a DaemonSet ensures that a copy of a Pod runs on every node (or a subset of nodes) in the cluster, which is designed for continuous daemon processes like logging or monitoring, not a one-time initialization task. Option C is wrong because a Deployment manages a set of Pods to maintain a desired state with rolling updates and self-healing, intended for long-running stateless applications, not a single run-to-completion job. Option D is wrong because a CronJob is used for scheduling Jobs to run at specific times or intervals, which is overkill and incorrect for a task that should run only once immediately.

15
MCQmedium

A Pod in a Deployment is CrashLoopBackOff. 'kubectl logs' shows the application exits with code 1 after printing one line. The Pod has a liveness probe that checks an HTTP endpoint. What should be checked first?

A.Check the application logs and configuration that causes the exit code 1.
B.Check if the liveness probe's initialDelaySeconds is too short.
C.Check if the liveness probe endpoint is correctly configured.
D.Check if the container has insufficient memory limits.
AnswerA

When a container terminates with exit code 1, it indicates an application-level failure or configuration error rather than an external infrastructure issue. Inspecting the container logs using `kubectl logs` is the primary troubleshooting step to identify runtime exceptions, missing environment variables, or syntax errors causing the process to crash immediately upon startup.

Why this answer

The application exits with code 1 after printing one line, which indicates a runtime error in the application itself. The liveness probe only restarts the container after the application has already failed; it does not prevent the initial crash. Therefore, the first step is to examine the application logs and configuration to understand why the process exits with code 1, as this is the root cause of the CrashLoopBackOff.

Exam trap

The trap here is that candidates often focus on the liveness probe configuration (options B and C) because the question mentions it, but the immediate cause is the application crash itself, which must be diagnosed first via logs.

How to eliminate wrong answers

Option B is wrong because adjusting initialDelaySeconds would only delay the start of liveness checks; it does not address the application exiting with code 1. Option C is wrong because the liveness probe endpoint configuration is irrelevant when the application crashes before the probe can even be evaluated. Option D is wrong because insufficient memory limits typically cause OOMKilled (exit code 137) or resource pressure, not a clean exit code 1 after printing one line.

16
MCQeasy

A user creates a Deployment with replicas=3. Two Pods are running, but the third is stuck in ContainerCreating. 'kubectl describe pod' shows 'Failed to create pod sandbox: rpc error: code = Unknown desc = failed to create containerd task: OCI runtime create failed: container_linux.go:349: starting container process caused: exec: "/app": stat /app: no such file or directory'. What is the most likely cause?

A.The container image is not pulled due to authentication failure.
B.The container exceeds its memory limit and is killed.
C.The node is not schedulable due to taints.
D.The container image does not have the /app executable at the specified path.
AnswerD

When a container starts, the container runtime attempts to execute the command specified in the container's configuration. If the /app binary is missing from the image or the path is misconfigured, the runtime fails to spawn the process, resulting in a container creation or start failure with a 'no such file or directory' error.

Why this answer

The error message 'exec: "/app": stat /app: no such file or directory' indicates that the container runtime (containerd) successfully created the sandbox and started the container process, but the command specified in the container's entrypoint or command tried to execute '/app', which does not exist inside the container image. This is a classic misconfiguration where the image lacks the expected binary or script at the given path, causing the container to fail immediately after creation.

Exam trap

The trap here is that candidates may confuse a 'ContainerCreating' status with image pull issues or resource limits, but the specific OCI runtime exec error points directly to a missing executable in the container image, not to infrastructure or scheduling problems.

How to eliminate wrong answers

Option A is wrong because an authentication failure would produce an 'ImagePullBackOff' or 'ErrImagePull' status, not a 'ContainerCreating' state with an OCI runtime exec error. Option B is wrong because exceeding the memory limit results in an OOMKilled container (status 'OOMKilled'), not a failure to start the container process with a 'no such file or directory' error. Option C is wrong because taints cause the pod to remain in 'Pending' state (not 'ContainerCreating'), and the error message is about container runtime execution, not scheduling.

17
Multi-Selecthard

Which THREE of the following are valid considerations when using resource requests and limits? (Select 3)

Select 3 answers
A.Limits must be equal to requests for a Pod to be scheduled.
B.Requests are used by the scheduler to decide which node can accommodate the Pod.
C.CPU limits guarantee the Pod will get that amount of CPU.
D.The QoS class is determined based on requests and limits.
E.Memory limits can cause the Pod to be OOMKilled if exceeded.
AnswersB, D, E

During scheduling, kube-scheduler evaluates each candidate node by subtracting the sum of container requests for CPU and memory from the node's allocatable capacity. A node is deemed feasible only if it can satisfy all requested quantities, because requests represent the minimum resource reservation needed to run the Pod. Limits are deliberately ignored in this admission calculation, making requests the primary input for node fit decisions.

Why this answer

The Kubernetes scheduler uses resource requests (CPU and memory) to determine node suitability for a Pod. The scheduler checks if the sum of requests for all Pods on a node, plus the new Pod's requests, is less than or equal to the node's allocatable capacity. Limits are not used for scheduling decisions.

Exam trap

The trap here is that candidates often confuse CPU limits as a guarantee of CPU allocation, when in fact CPU is compressible and limits only throttle usage, while memory limits are hard and can cause OOM kills.

18
MCQhard

You have a StatefulSet with 5 pods, each requiring a unique stable network identity. The StatefulSet is scaled down from 5 to 3. Which pods will be terminated?

A.Random pods
B.Pods with the highest ordinals (4 and 3)
C.Pods with the lowest ordinals (0 and 1)
D.Pods with the highest resource usage
AnswerB

Correct. The StatefulSet controller honors the ordinal ordering by deleting pods with the highest index numbers first. For a 5-pod StatefulSet (ordinals 0 through 4) scaling from 5 to 3 replicas, pods 4 and 3 are removed because they are the highest ordinals. This preserves the sequential identity of the remaining pods (0, 1, 2) and their associated storage, which is the core behavior of StatefulSet lifecycle management.

Why this answer

When a StatefulSet is scaled down, Kubernetes terminates pods in reverse order of their ordinal indices, starting from the highest. For a StatefulSet with 5 pods (ordinals 0-4) scaled to 3, pods with ordinals 4 and 3 are terminated first, ensuring that the remaining pods (0, 1, 2) maintain their stable network identities and storage.

Exam trap

The trap here is that candidates may assume pods are terminated randomly or based on resource usage, but the CKA exam tests the specific deterministic behavior of StatefulSet scaling, which always removes pods with the highest ordinals first.

How to eliminate wrong answers

Option A is wrong because StatefulSet does not terminate pods randomly; it follows a deterministic ordinal-based termination order. Option C is wrong because pods with the lowest ordinals (0 and 1) are the last to be terminated, not the first, as the scaling-down process removes pods from the highest ordinal downward. Option D is wrong because StatefulSet termination is based on ordinal index, not resource usage; resource-aware termination is not a feature of StatefulSet scaling.

19
MCQhard

A Kubernetes cluster has a node pool with GPU nodes labeled 'accelerator=nvidia-tesla'. A Pod requires a GPU. Which configuration is necessary?

A.Use nodeAffinity with requiredDuringSchedulingIgnoredDuringExecution for the GPU label.
B.Set resources.limits for 'nvidia.com/gpu' only.
C.Set nodeSelector to 'accelerator=nvidia-tesla' and request 'nvidia.com/gpu' in resources.
D.Add a toleration for GPU node taints.
AnswerC

This is the correct and comprehensive approach because it addresses both the placement of the Pod and the allocation of the specialized hardware resource. The `nodeSelector` ensures the Pod is scheduled exclusively onto nodes labeled `accelerator=nvidia-tesla`, which are the GPU-equipped nodes in this scenario. Simultaneously, requesting `nvidia.com/gpu` in the Pod's `resources` section (either `requests` or `limits`) informs the Kubernetes device plugin for NVIDIA GPUs to allocate a specific GPU device to the container, making it available inside the Pod. Both mechanisms are crucial for successful GPU workload deployment.

Why this answer

A Pod that requires a GPU must both be scheduled onto a node with the appropriate GPU label and explicitly request the GPU resource. The `nodeSelector` ensures the Pod lands on a node labeled `accelerator=nvidia-tesla`, and requesting `nvidia.com/gpu` in `resources.requests` or `resources.limits` (typically limits) tells the kubelet to allocate a GPU device to the container. Without the resource request, the scheduler has no way to account for GPU capacity, and without the nodeSelector, the Pod might be scheduled on a non-GPU node.

Exam trap

The trap here is that candidates often think either node selection (nodeSelector/affinity) or resource requests alone is sufficient, but the CKA exam requires both to be present for a GPU workload to function correctly.

How to eliminate wrong answers

Option A is wrong because `nodeAffinity` with `requiredDuringSchedulingIgnoredDuringExecution` is a valid way to select GPU nodes, but it is not sufficient on its own — the Pod must also request the `nvidia.com/gpu` resource to actually get a GPU assigned. Option B is wrong because setting `resources.limits` for `nvidia.com/gpu` alone does not guarantee the Pod lands on a GPU node; without a nodeSelector or affinity, the scheduler may place the Pod on a non-GPU node where the resource is unavailable. Option D is wrong because GPU nodes do not inherently have taints; while administrators may add taints to GPU nodes, a toleration is only needed if a taint is present, and it is not a required configuration for GPU access.

20
MCQeasy

A developer deployed a Pod that is stuck in Pending state. The cluster has one worker node with taint 'node.kubernetes.io/disk-pressure:NoSchedule'. The Pod does not specify any tolerations. What is the most likely cause?

A.The Pod was evicted due to resource pressure.
B.The Pod requests more CPU than available on the node.
C.The scheduler failed to communicate with the API server.
D.The node has a taint that the Pod does not tolerate.
AnswerD

The node is marked with a 'node.kubernetes.io/disk-pressure' taint with a 'NoSchedule' effect. Because the Pod's manifest does not define a matching toleration for this specific taint, the kube-scheduler is forced to bypass this node, leaving the Pod in a Pending state due to unschedulability.

Why this answer

A Pod stuck in Pending state indicates the scheduler cannot find a suitable node. The cluster has a worker node with the taint 'node.kubernetes.io/disk-pressure:NoSchedule', and the Pod has no tolerations. Since taints with effect NoSchedule prevent scheduling of Pods that do not tolerate them, the Pod cannot be placed on that node, leaving it in Pending.

Exam trap

CNCF often tests the distinction between taints/tolerations and resource constraints, so the trap here is that candidates may confuse a taint-based scheduling block with a resource shortage, especially when the taint name includes 'disk-pressure' which sounds like a resource issue.

How to eliminate wrong answers

Option A is wrong because eviction occurs when a Pod is already running and the node experiences resource pressure, not when a Pod is stuck in Pending; eviction would move the Pod to a different state (e.g., Failed or Evicted). Option B is wrong because insufficient CPU would cause the scheduler to report a '0/1 nodes are available: insufficient cpu' event, but the question explicitly states the node has a disk-pressure taint, and the Pod has no tolerations, making the taint the primary blocking factor. Option C is wrong because scheduler-to-API-server communication failures typically result in scheduler errors or 'no persistent volumes available' messages, not a simple Pending state; the scheduler communicates via the API server to list nodes and bind Pods, and a failure would produce different symptoms.

Ready to test yourself?

Try a timed practice session using only Workloads Scheduling questions.