Courseiva

DOP-C02 Resilient Cloud Solutions Practice Question

A company runs a containerized application on Amazon EKS. The application must be highly available across multiple Availability Zones and must automatically recover from node failures. Which THREE steps should be taken?

⚠ Common exam trap

Many exam-takers think deploying in a single AZ or using a single instance type simplifies management and reduces costs, but the DOP-C02 exam specifically tests the principle of designing for failure across multiple AZs and instance diversity to achieve true high availability.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Pod Disruption Budgets to ensure a minimum number of pods are available during voluntary disruptions.

Pod Disruption Budgets (PDBs) are correct because they allow you to specify the minimum number of pods that must remain available during voluntary disruptions, such as node drains or cluster upgrades. This ensures that the application maintains high availability even when Kubernetes performs planned maintenance, preventing all replicas from being taken down simultaneously.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Pod Disruption Budgets to ensure a minimum number of pods are available during voluntary disruptions.

    Why this is correct

    Pod Disruption Budgets (PDBs) constrain voluntary disruptions such as node drains during cluster upgrades, node-group updates, or Cluster Autoscaler scale-in by specifying minAvailable or maxUnavailable for a pod selector. During a voluntary disruption, the Kubernetes eviction API rejects requests that would cause the number of available pods to fall below the budget. PDBs do not protect against involuntary failures like an EC2 instance crash, so they are a complement, not a substitute, for multi-AZ node deployment.

  • ✓

    Configure the Cluster Autoscaler to add nodes when pods are unschedulable.

    Why this is correct

    Configuring Cluster Autoscaler to add nodes when pods are unschedulable enables the cluster to respond to capacity shortages by scaling out node groups. In the event of a node failure, the pods that were running on it become unschedulable after rescheduling attempts, causing Cluster Autoscaler to provision replacement capacity from the launch template. However, Cluster Autoscaler does not automatically remove or replace unhealthy nodes; that responsibility falls to node health checks or managed node group end-to-end remediation, so it must be paired with node-level recovery mechanisms.

  • ✓

    Deploy worker nodes across multiple Availability Zones.

    Why this is correct

    Deploying worker nodes across multiple Availability Zones isolates the cluster from a single-zone failure; if one AZ becomes unavailable, pods can be rescheduled onto nodes in the remaining AZs. EKS automatically runs the control plane across three AZs in a region, so a multi-AZ worker fleet matches that distribution. For resilience, you should also use topology spread constraints (e.g., topologySpreadConstraints or podAntiAffinity) to ensure pods are not packed into just one AZ despite the node spread.

  • ✗

    Deploy worker nodes in a single Availability Zone to reduce cross-AZ data transfer costs.

    Why it's wrong here

    Placing worker nodes in a single Availability Zone creates a single point of failure: if that AZ experiences an outage, all worker nodes become unavailable and every workload running on them is lost. While cross-AZ data transfer charges apply, the cost difference is usually minor compared to the business impact of an AZ failure. EKS control plane is multi-AZ, but the worker nodes are the execution layer, so concentrating them in one AZ undermines the entire high-availability design.

  • ✗

    Use a single large instance type for all worker nodes to simplify management.

    Why it's wrong here

    Using one large instance type for all nodes does nothing to increase resilience and actually increases the blast radius of a single node failure; one impacted large node disrupts many replicas. It also makes the cluster vulnerable to instance-type capacity constraints or price spikes, which can prevent replacement capacity from launching. A better practice is to use a diverse set of instance types (and optionally Spot with capacity-optimized allocation) so that the cluster has multiple independent supply chains for compute capacity.

About these practice questions

This DOP-C02 question is part of Courseiva's 1,298-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on DOP-C02

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company runs a containerized microservices application on Amazon EKS. The application includes a critical service that processes real-time financial transactions. This service must be highly available and resilient to node failures. The current setup uses a Deployment with 3 replicas and a ClusterIP service. During a recent node failure, the application experienced a brief period of unavailability. Which action should the DevOps engineer take to improve resilience without changing the underlying infrastructure?

hard
  • A.Change the service type from ClusterIP to NodePort and configure an external load balancer.
  • B.Increase the number of replicas to 10 and use a node selector to schedule all pods on the largest instance type.
  • ✓ C.Configure a PodDisruptionBudget with a maxUnavailable of 1, and add pod anti-affinity rules to spread pods across different nodes.
  • D.Enable HorizontalPodAutoscaler with a target CPU utilization of 50% to automatically scale the Deployment.

Why C: A PodDisruptionBudget with maxUnavailable=1 ensures that at most one pod is unavailable during voluntary disruptions, while pod anti-affinity rules force the scheduler to distribute pods across different nodes. This combination prevents a single node failure from taking down all replicas, maintaining service availability without altering the underlying infrastructure.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.