Courseiva

CKAD Application Observability and Maintenance Practice Question

Your team manages a microservices application on a Kubernetes cluster. A critical service 'order-service' is deployed with 3 replicas. Lately, customers have reported occasional timeouts when placing orders. You suspect that the service is overloaded during peak hours. You have configured a HorizontalPodAutoscaler (HPA) based on CPU utilization, but the autoscaler does not appear to be scaling up quickly enough. Upon inspection, you notice that the HPA is configured with a target CPU utilization of 80%, and the current CPU usage of the pods is around 70%. However, the pods' memory usage is high and growing. The application is also logging slow database queries. Which action is most likely to improve the responsiveness of the service during peak load?

⚠ Common exam trap

Watch out — candidates often assume CPU is the only metric for HPA scaling, but the CKAD exam tests the ability to identify when other metrics (like memory) are more relevant based on application behavior and symptoms.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Modify the HPA to also scale based on memory utilization, setting a target average value for memory.

The HPA is configured only for CPU, but the symptom (high memory usage, slow DB queries) indicates memory pressure is the bottleneck. Adding a memory-based metric to the HPA allows it to scale when memory exceeds the target, addressing the root cause of overload during peak hours. The current CPU at 70% is below the 80% threshold, so the HPA won't trigger on CPU alone, but memory scaling can react to the actual resource constraint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Modify the HPA to also scale based on memory utilization, setting a target average value for memory.

    Why this is correct

    Relying solely on CPU utilization can miss memory-bound workloads, where pods degrade due to memory pressure long before CPU approaches its target. Adding a memory metric to the HPA—using targetAverageValue or targetAverageUtilization—makes the autoscaler react to actual resource exhaustion. Since memory is a non-compressible resource, Kubernetes cannot throttle it as it does with CPU, so the only effective mitigation is to add replicas when average memory consumption exceeds the configured threshold.

  • ✗

    Reduce the number of replicas to 2 to force the HPA to react more aggressively.

    Why it's wrong here

    Manually reducing the replica count to 2 will spike CPU utilization per pod, but the HPA's scaling decision still compares current CPU usage to the existing target, such as 70%. If per-pod CPU remains below that target, the HPA will not trigger a scale-out; it could even compute a desired replica count of 1, shrinking the deployment further. This approach does not alter the scaling metric or its threshold, and it deliberately removes capacity, increasing the risk of overload and service disruption.

  • ✗

    Increase the CPU target utilization to 90% so the HPA triggers earlier.

    Why it's wrong here

    Raising the CPU target utilization to 90% actually raises the bar that must be crossed before the HPA adds replicas. For example, if the current average CPU usage is 80% and the target is 70%, the HPA scales out; with a 90% target, that same 80% usage results in no scaling action. Since the problem is that scaling is not happening soon enough, increasing the threshold makes the autoscaler even less sensitive and delays the very reaction you need, worsening response times under load.

  • ✗

    Optimize the database queries to reduce response time.

    Why it's wrong here

    Optimizing database queries can lower per-request CPU and latency, but it does not change the HPA's configuration or its decision-making process. If the true issue is that the HPA ignores memory pressure and therefore keeps too few replicas, query optimization may temporarily mask the symptom but leaves the scaling mechanism blind to the real bottleneck. Furthermore, such optimizations are application-level fixes, whereas the HPA operates at the infrastructure layer; they address a different cause and fail to ensure automatic scaling under future memory spikes.

About these practice questions

This CKAD question is part of Courseiva's 826-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This CKAD practice question is part of Courseiva's free CNCF certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CKAD exam.