CKAD · domain
Application Observability and Maintenance
This CKAD domain covers keeping workloads observable and debuggable: configuring liveness, readiness, and startup probes, reading container logs, inspecting resource usage, and troubleshooting running containers. Questions are task-oriented, asking you to pick the correct kubectl command or probe configuration under time pressure, so you must know exact flags and probe handler types rather than concepts alone.
Focused practice
Practice Application Observability and Maintenance questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Application Observability and Maintenance
Be able to write correct liveness, readiness, and startup probes and run the right kubectl command first try. The single most important thing: know kubectl logs --previous for crashed containers and kubectl exec -it for shells, since these appear as direct task questions.
Probe handler types: exec, httpGet, tcpSocket, and gRPC, plus liveness, readiness, and startup probe semantics.
Log retrieval with kubectl logs, including the --previous flag for restarted or crashed containers.
Resource usage inspection via kubectl top pod and kubectl top pod --containers for per-container metrics.
Interactive debugging using kubectl exec -it with a shell or command inside a running container.
Watch out for
Common Application Observability and Maintenance exam traps
- ▸Forgetting --previous when a container restarted, so kubectl logs shows only the current instance instead of the crashed one.
- ▸Confusing probe types or fields, such as using tcpSocket where httpGet is required, or omitting initialDelaySeconds and periodSeconds.
- ▸Assuming kubectl top works without metrics-server installed, or using it on pods where metrics are unavailable.
Question index
All Application Observability and Maintenance questions (146)
Click any question to see the full explanation, or start a practice session above.
Refer to the exhibit. A pod named 'app-backend-6b4c9d8f7-2x4z5' is in CrashLoopBackOff state. What is the MOST likely cause of this issue?
Easy2A pod is stuck in 'Pending' state. You run 'kubectl describe pod my-pod' and see the event: '0/3 nodes are available: 3 Insufficient cpu.' Which action should you take?
Medium3A pod is running but not responding to traffic. You suspect the application inside the container is unhealthy but the pod is still marked as 'Running'. Which probe should be configured to remove the pod from the service's endpoints automatically?
Easy4A DevOps engineer needs to set up resource monitoring for pods in a namespace. Which built-in Kubernetes resource provides CPU and memory metrics out-of-the-box?
Easy5Which THREE of the following are true about the 'kubectl describe pod' output? (Select 3)
Hard6What is the purpose of a readiness probe?
Easy7You want to debug a pod that is failing to start. The pod does not have a shell installed. Which command can you use to attach an ephemeral debug container to the running (or failed) pod?
Hard8Which TWO are valid ways to debug a pod using ephemeral containers?
Easy9Which THREE of the following are correct statements about the terminationGracePeriodSeconds field? (Select 3)
Medium10You deploy a pod with the following specification: apiVersion: v1 kind: Pod metadata: name: probe-pod spec: containers: - name: app image: nginx livenessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 5 periodSeconds: 10 failureThreshold: 3 The application listens on port 80, not 8080. What will happen to the pod?
Hard11You have a pod 'app-pod' that keeps restarting. You want to see the logs from the previous (crashed) container instance. Which command should you use?
Easy12A pod named 'web-app' is experiencing high CPU usage. You want to investigate which process inside the container is consuming the most CPU. Which command should you run?
Medium13Which TWO of the following are valid approaches to debug a pod that is in a CrashLoopBackOff state? (Select 2)
Hard14A deployment 'web-app' is running with 3 replicas. Users report that the application is slow. The team suspects a memory leak. Which command provides the most immediate insight into the memory usage of the pods?
Easy15Which THREE of the following are valid parameters for configuring probes? (Select 3)
Medium16Which THREE of the following are best practices for configuring readiness probes?
Hard17Which probe type is used to indicate that a container is ready to serve traffic and should be added to Service endpoints?
Easy18A developer needs to view the logs of a pod named 'web-app-84b7f6f5b6-abcde' that crashed and has been restarted. Which kubectl command should they use to see the logs from the previous (crashed) instance?
Easy19Which command outputs the pod's full YAML specification including status?
Medium20You are asked to ensure that a pod with a slow-starting container (requires 60 seconds to initialize) is not prematurely restarted by the liveness probe. The liveness probe should start only after the container is fully initialized. Which probe type should you add to the pod spec?
Hard21You want to ensure your containerized application handles SIGTERM gracefully and shuts down within 30 seconds. Which field should you set in the Pod spec?
Medium22You have a Deployment running a web application that takes 60 seconds to start up. You need to configure probes so that Kubernetes waits for the application to fully start before checking its health and directing traffic to it. Which combination of probes should you use?
Medium23Which THREE of the following are valid parameters for a probe?
Medium24Which TWO of the following are valid types of probes in Kubernetes? (Select 2)
Medium25Which TWO statements about readiness probes are correct? (Choose two.)
Medium26Which field in the container spec controls the time Kubernetes waits after sending SIGTERM before sending SIGKILL during pod shutdown?
Medium27A CKAD candidate must make a container's exit code observable after a crash so that a monitoring agent reading pod status can distinguish an application error from an out-of-memory kill. The container has terminated and restarted several times. Which command shows the exit code and termination reason of the most recent terminated instance of the container?
Medium28A pod has a readiness probe using tcpSocket on port 3306. The application listens on port 3306 but returns errors on database queries. What is the effect of the readiness probe?
Hard29Which TWO of the following are valid reasons for using a startup probe?
Hard30A pod is in CrashLoopBackOff state. You need to view the last few lines of its logs to understand why it is crashing. Which command is most appropriate?
Easy31You are a platform engineer managing a production Kubernetes cluster. A team deploys a stateful application called 'inventory-service' with 3 replicas using a StatefulSet. Each pod writes logs to a persistent volume via a PersistentVolumeClaim. Recently, the team reports that the application becomes unresponsive after running for a few hours. You notice that the pods are still running (READY 1/1) but the application does not respond to HTTP requests. You exec into one pod and find that the disk is 100% full. The PVC is backed by a cloud disk (e.g., AWS EBS). You check the pod's resource limits and see that memory and CPU are not exhausted. The container logs are not rotated. Which course of action should you take to resolve the immediate issue and prevent recurrence?
Hard32A pod is running but the application inside is not serving traffic. The team runs 'kubectl exec -it <pod> -- curl localhost:8080' and gets 'Connection refused'. What is the most likely cause?
Hard33You run 'kubectl get pods -o wide' and see that a pod's NODE column shows 'node-1'. You want to see more details about that node's resource usage. Which command should you use?
Medium34Which THREE of the following are valid fields for configuring a probe in Kubernetes? (Select 3)
Hard35A pod with a startup probe configured is taking longer than usual to start. The startup probe has 'failureThreshold: 10' and 'periodSeconds: 5'. What is the maximum time the pod has to start before it is restarted?
Medium36Which THREE of the following are valid reasons to use a startup probe? (Select THREE.)
Hard37You run 'kubectl get events' and see an event 'FailedScheduling' for a pod. What is the most common cause?
Medium38Which THREE commands can be used to get detailed information about a pod named 'my-pod'? (Choose three.)
Hard39You need to collect metrics from an application running in a pod. The application exposes metrics on port 8080 at /metrics in Prometheus format. Which resource should you configure to allow Prometheus to scrape these metrics?
Medium40You have a Deployment that uses an httpGet liveness probe on port 8080 with path /healthz. The probe fails after the container starts, but you can successfully curl http://localhost:8080/healthz from within the container. What is the most likely cause?
Hard41A pod named 'app' is stuck in 'Pending' state. You run 'kubectl describe pod app' and see the event: '0/3 nodes are available: 3 Insufficient cpu'. What is the most likely cause?
Medium42You want to see the IP address and the node on which each pod in the 'default' namespace is running. Which command provides this information?
Medium43A pod is running but not serving traffic. You suspect the readiness probe is failing. Which THREE commands or actions would help you diagnose the readiness probe issue?
Hard44A team ships a container that writes structured JSON diagnostics to stdout and nothing to a file. A support engineer needs to filter those lines for entries with level 'error' across all containers of a deployment named 'checkout', without installing additional tooling on the nodes. Which approach accomplishes this using only kubectl?
Medium45You need to configure a container to shut down gracefully when it receives a SIGTERM signal, with a timeout of 30 seconds before force kill. Which field in the Pod spec should you set?
Medium46Which command shows resource usage (CPU/memory) of pods in the default namespace?
Easy47A deployment 'api-deploy' has resource limits set but is frequently being OOMKilled. The team suspects the memory limit is too low. Which approach should be taken to confirm this without causing downtime?
Hard48Sequence the steps to troubleshoot a Pod stuck in CrashLoopBackOff state.
Medium49Which command shows resource usage (CPU and memory) for pods in a namespace?
Easy50You need to view the logs of a container that previously crashed and has been restarted. Which flag do you use with 'kubectl logs'?
Medium51A developer deployed a pod named 'web-pod' with a single container. Users report intermittent 502 errors. The developer wants to see CPU and memory consumption of that specific container to determine whether it is being throttled. Which command should be used?
Easy52You have a Deployment running a web server that takes 30 seconds to initialize. You want to ensure that the load balancer does not send traffic to the pod until it is ready. Which probe should you configure?
Easy53A production pod named 'api-pod' runs a single container whose liveness probe has begun failing after a configuration change, causing repeated restarts. An engineer must determine whether the failures are caused by the probe configuration itself or by the application inside the container. Which TWO actions best provide that evidence? (Choose two.)
Hard54You want to run a command inside a running container to check environment variables. Which command should you use?
Medium55You want to configure a pod so that it receives a SIGTERM signal and has 60 seconds to shut down gracefully before being forcefully killed. Which field should you set?
Hard56Your team manages a microservices application on a Kubernetes cluster. A critical service 'order-service' is deployed with 3 replicas. Lately, customers have reported occasional timeouts when placing orders. You suspect that the service is overloaded during peak hours. You have configured a HorizontalPodAutoscaler (HPA) based on CPU utilization, but the autoscaler does not appear to be scaling up quickly enough. Upon inspection, you notice that the HPA is configured with a target CPU utilization of 80%, and the current CPU usage of the pods is around 70%. However, the pods' memory usage is high and growing. The application is also logging slow database queries. Which action is most likely to improve the responsiveness of the service during peak load?
Medium57You have a pod that is in a 'Pending' state. You run 'kubectl describe pod' and see the event: '0/3 nodes are available: 3 Insufficient cpu'. What is the most likely cause?
Medium58Which TWO commands can be used to view the logs of a pod that has crashed?
Medium59What is the purpose of the 'kubectl describe pod' command?
Easy60Which TWO of the following are valid reasons to use a readiness probe? (Select 2)
Easy61Which command shows the CPU and memory usage of all pods in the current namespace?
Easy62You have a Deployment with a liveness probe using an exec command. The probe currently runs 'cat /tmp/healthy' and fails after 3 failures. You notice the pod is being restarted even though the application is healthy. What is the most likely cause?
Hard63Which TWO statements are true about readiness probes? (Select two.)
Medium64You run 'kubectl get pods' and see a pod in 'CrashLoopBackOff' state. Which TWO conditions could cause this state?
Hard65A pod is stuck in 'Pending' state. You run 'kubectl describe pod' and see '0/1 nodes are available: 1 Insufficient cpu, 1 Insufficient memory.' What does this mean?
Medium66A container in a pod is expected to shut down gracefully within 30 seconds when it receives SIGTERM. How should you configure the pod to ensure the container is given enough time to shut down before being forcefully killed?
Medium67A pod spec has terminationGracePeriodSeconds: 30. The main process ignores SIGTERM. After 30 seconds, what happens?
Hard68You want to run a command inside an existing container 'app-container' in pod 'my-pod'. The pod has only one container. Which command enters an interactive shell?
Medium69You need to debug a Pod that is in CrashLoopBackOff. Which command should you run first?
Medium70You need to view events for a specific pod named 'web-1' in the 'default' namespace. Which command shows events related to this pod?
Medium71A container in your pod takes a long time to start (up to 5 minutes). You want to avoid the container being restarted by the liveness probe during this period. What should you configure?
Easy72You need to view the logs of the previous (terminated) instance of a container in a pod. Which command should you use?
Easy73You need to debug a pod that has no running containers because it is in a CrashLoopBackOff state. You want to start an ephemeral container with debugging tools in the same namespace. Which command accomplishes this?
Hard74Which TWO actions can you perform using 'kubectl describe'? (Select two.)
Medium75You need to configure a liveness probe that checks if the container port 8080 is open. Which probe type should you use?
Hard76Which THREE are valid types of probes in Kubernetes?
Hard77A pod has a readiness probe configured with httpGet on port 8080. The probe is failing, but the pod is running. What is the immediate effect on the pod?
Medium78Which kubectl command can you use to stream the logs from a pod in real-time?
Easy79A pod has a startup probe with failureThreshold: 30 and periodSeconds: 10. The application takes up to 5 minutes to start. What should be changed to ensure the startup probe does not kill the container prematurely?
Hard80A pod is running but not receiving traffic from a Service. The readiness probe is failing. What is the likely effect on the pod?
Medium81Which TWO of the following are valid parameters for configuring probes in Kubernetes? (Select TWO.)
Medium82You have a Deployment that uses a ConfigMap. You update the ConfigMap, but the pods are not picking up the changes. What is the MOST efficient way to force the pods to use the new ConfigMap values without downtime?
Hard83You have a Deployment called 'web-deploy' with 3 replicas. One of the pods is not receiving traffic, but it shows 'Running' and passes its liveness probe. The readiness probe is configured as a TCP socket check on port 8080. You verify that the application is listening on port 8080. What is a likely reason the pod is not receiving traffic?
Hard84You are a platform engineer for a large e-commerce site. The application is deployed on a Kubernetes cluster with 10 worker nodes. Recently, a new deployment 'checkout' was rolled out, and soon after, the cluster experienced network latency and intermittent connectivity issues between services. The 'checkout' pods are configured with a liveness probe that makes an HTTP request to an internal health endpoint. Upon investigation, you find that the kubelet on several nodes is consuming high CPU, and the number of iptables rules has increased significantly. You suspect that the 'checkout' deployment's configuration is causing excessive churn in the network rules. Which aspect of the deployment configuration is most likely the root cause?
Hard85You need to configure a liveness probe for a container that listens on TCP port 8080. The probe should wait 5 seconds before starting, check every 10 seconds, and timeout after 2 seconds. Which YAML snippet correctly configures this?
Hard86Which THREE of the following are valid ways to debug a pod that is not responding? (Select three.)
Hard87A developer reports that a pod named 'api-pod' is restarted repeatedly. You run 'kubectl get events --field-selector involvedObject.name=api-pod' and see multiple events with reason 'BackOff' and message 'Back-off restarting failed container'. Which probe failure is MOST likely causing this?
Medium88A pod is stuck in the Pending state. You run 'kubectl describe pod pod-name' and see the event: '0/4 nodes are available: 1 node had taint {node-role.kubernetes.io/control-plane: }, that the pod didn't tolerate, 3 nodes had taint {node.kubernetes.io/disk-pressure: }, that the pod didn't tolerate.' What is the MOST likely cause?
Hard89What is the purpose of a readiness probe in Kubernetes?
Easy90Which THREE of the following are valid types of probes in Kubernetes? (Select 3)
Medium91You are writing a Deployment manifest for a container that takes up to 60 seconds to become ready. You want to ensure the liveness probe does not interfere during this startup period. What should you configure?
Medium92You need to configure a liveness probe for a container that starts a web server on port 8080. The probe should check the '/healthz' endpoint. Which YAML snippet correctly defines this probe?
Medium93You need to see the YAML definition of a running pod named 'app' including the current status. Which command should you use?
Medium94You need to run a command inside a running pod's container. The container has a shell available. Which command allows you to execute 'ls -la /data' inside the container?
Medium95A pod named 'db-pod' is running but not responding as expected. You want to check its logs from the previous instantiation (after a crash). Which command should you use?
Medium96A cluster has a node that is NotReady. The kubelet on that node is not responding. Which command should be used to investigate the kubelet logs on the node?
Hard97A pod has been scheduled on a node but is stuck in 'ContainerCreating' state. The team suspects a missing storage class. Which command would best confirm this?
Medium98You need to check the CPU and memory usage of a pod named 'monitored-pod' in the 'default' namespace. Which command should you run?
Medium99A deployment is configured with a liveness probe that checks an HTTP endpoint. The probe fails intermittently, causing pod restarts. What is the best first step to diagnose the issue?
Easy100A pod with a liveness probe using 'httpGet' is restarting repeatedly. The probe checks '/healthz' on port 8080. The application is healthy and responds with HTTP 200. What is the most likely cause?
Medium101A pod named 'web' is in a CrashLoopBackOff state. You suspect the application is failing due to a configuration error. You want to see the logs from the previous instance of the container. Which command should you use?
Medium102You want to see a list of all events in the 'default' namespace, sorted by timestamp. Which command should you use?
Easy103You are debugging a pod that is running but not responding to network requests on port 8080. You suspect the application inside the container is faulty. You need to run an interactive shell inside the container to inspect the process. Which command should you use?
Medium104Which probe type is used to determine if a container is ready to serve traffic?
Easy105You have a Deployment with 3 replicas. The pods have a readiness probe that checks an HTTP endpoint /ready. One pod's readiness probe is failing. What will happen?
Hard106A pod uses a liveness probe with exec command 'cat /tmp/healthy'. The file /tmp/healthy exists initially but is deleted by the application after 60 seconds. Which behavior will occur?
Medium107You need to get a list of all events in the cluster sorted by timestamp. Which command should you use?
Easy108A Pod is running but not responding to requests. The liveness probe is a TCP check on port 8080. What is the most likely issue?
Hard109A pod is stuck in 'Pending' state. You run 'kubectl describe pod pending-pod' and see the following condition: 'Status: False, Type: PodScheduled, Reason: Unschedulable, Message: 0/4 nodes are available: 1 node(s) had taint {node.kubernetes.io/disk-pressure: }, 1 node(s) had taint {node.kubernetes.io/memory-pressure: }, 2 node(s) didn't match Pod's node affinity/selector.' What is the most specific reason the pod cannot be scheduled?
Medium110You want to ensure your application shuts down gracefully when a pod is terminated. The application needs 30 seconds to clean up. Which field should you set in the pod spec?
Medium111Based on the exhibit, what should you do to determine why the container is failing?
Medium112You are debugging a network issue: a pod 'frontend' cannot reach a service 'backend' in the same namespace. The service endpoints are empty. What is the most likely cause?
Hard113Which TWO parameters can be configured for a probe to control its behavior? (Choose two.)
Medium114You want to view the events related to a specific pod named 'my-pod' in the 'default' namespace. Which command filters events to show only those pertaining to this pod?
Medium115You are troubleshooting a service connectivity issue. You have a pod named 'client' and a service named 'server'. You want to check if the service's endpoints are populated. Which command should you run?
Medium116Which TWO of the following commands can be used to view events related to a specific pod? (Select TWO.)
Medium117A developer reports that a container in a pod is not responding correctly. You need to get an interactive shell in the container to investigate. Which command should you run?
Hard118You run 'kubectl get pods -o wide' and see that a pod 'worker-pod' is in 'Completed' state with 'Restart Count: 5'. The pod's restart policy is 'OnFailure'. What is the most likely reason the pod has restarted 5 times?
Hard119You want to see detailed information about a node's resource usage. Which command should you run?
Easy120Which command shows resource usage (CPU and memory) for pods?
Easy121You need to view metrics from a Pod running a web server. Which approach follows Kubernetes best practices?
Easy122You are tasked with improving the observability of a microservices application running in Kubernetes. The application is deployed with multiple replicas and experiences occasional high latency. Which TWO actions should you take to gain better insight into the application's performance?
Hard123You run 'kubectl get pods' and see that a pod is in 'Pending' state. What is the most likely cause?
Medium124A Pod is stuck in CrashLoopBackOff. You run 'kubectl logs mypod' and get no output. What is the most likely cause?
Medium125Which kubectl command shows detailed information about a pod, including events, labels, and container state?
Easy126You have a Deployment that must run a legacy application that takes up to 5 minutes to start. You need to ensure the liveness probe does not kill the container prematurely. Which probe configuration should you use?
Hard127You have a pod that takes 2 minutes to start its application. You want to avoid the liveness probe from killing the pod during startup, but still have it active afterward. Which probe should you add to the pod spec?
Hard128A developer configures a liveness probe for a container that takes a long time to start (about 120 seconds). The probe uses httpGet on port 8080 with a path '/healthz'. The probe is configured with initialDelaySeconds=10, periodSeconds=10, failureThreshold=3. The pod enters CrashLoopBackOff. What is the MOST likely cause?
Hard129You have a Deployment named 'api' with 3 replicas. You need to ensure that new pods are not added to the Service's endpoints until the application is ready to serve traffic. Which probe configuration should you add to the pod spec?
Hard130A developer wants to view the resource usage of all containers in a specific pod. Which command should they use?
Easy131Which kubectl command streams logs from a pod named 'web-pod' in real-time?
Easy132You need to check the resource usage of nodes and pods in your cluster. Which TWO commands should you use? (Choose two)
Medium133You want to debug a pod that has no shell or debugging tools installed. Which feature allows you to temporarily add a sidecar container with debugging tools to a running pod?
Medium134The Pod 'myapp' is in CrashLoopBackOff. Based on the exhibit, what is the most likely cause?
Hard135A pod named 'db-backup' is in CrashLoopBackOff. The team needs to understand why it keeps crashing. Which approach should be taken first to diagnose the issue?
Medium136You want to see all events in the default namespace sorted by timestamp. Which command should you use?
Easy137A container in a pod is crashing repeatedly. You want to see the logs from the previous (crashed) instance of the container. Which command should you use?
Medium138You need to view resource usage (CPU and memory) for all pods in the 'default' namespace. Which command should you use?
Easy139Based on the exhibit, why is the container being killed and restarted?
Hard140You are a platform engineer at a company that runs a microservices architecture on Kubernetes. The application consists of a frontend service (Node.js), a backend API (Go), and a PostgreSQL database. All components are deployed in the same namespace 'production'. Recently, the backend API has been experiencing intermittent 503 errors from the frontend. The backend API Pods have CPU limits set to 500m and memory limits to 256Mi. The backend API exposes metrics at /metrics and has a liveness probe (HTTP GET /healthz) and a readiness probe (HTTP GET /ready). You notice that during traffic spikes, the backend API Pods are restarted frequently. You examine the metrics and see that memory usage spikes to 250Mi during high load. What is the most likely cause of the restarts and 503 errors?
Hard141Which FOUR of the following are valid probe handlers in Kubernetes?
Medium142You have a Deployment with a liveness probe that fails intermittently, causing the pod to restart. You want to reduce the sensitivity of the probe so that it only restarts after 3 consecutive failures. Which probe parameter should you adjust?
Hard143Which command streams logs from a pod in real-time?
Easy144You have a pod that is stuck in 'Pending' state. Which command would you run first to diagnose the issue?
Medium145Order the steps to update a Kubernetes Secret and ensure a Pod uses the new secret.
Medium146You need to view the logs of a container named 'sidecar' inside a pod named 'app'. Which command should you use?
EasyOther domains
All CKAD exam domains
Frequently asked questions
- What does the Application Observability and Maintenance domain cover on the CKAD exam?
- Be able to write correct liveness, readiness, and startup probes and run the right kubectl command first try. The single most important thing: know kubectl logs --previous for crashed containers and kubectl exec -it for shells, since these appear as direct task questions.
- How many questions are in this domain?
- This page lists all 146 Application Observability and Maintenance questions in the CKAD question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Application Observability and Maintenance questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.