Courseiva

CCNA Resilient Cloud Questions

75 of 184 questions · Page 1/3 · Resilient Cloud topic · Answers revealed

1
MCQhard

A company runs a critical application on AWS Lambda that processes messages from an Amazon SQS queue. The application must be resilient to downstream service failures. The team notices that when the downstream service is unhealthy, messages are repeatedly retried and eventually sent to the dead-letter queue (DLQ) before the service recovers. What design change would improve resilience by allowing automatic retries after the downstream service recovers?

A.Configure the SQS queue with a large visibility timeout (e.g., 6 hours) and use a redrive policy only after a high number of receives. Keep the messages in the queue and retry when the downstream service becomes healthy.
B.Reduce the maxReceiveCount to 1 so that messages are sent to DLQ immediately, then reprocess them from DLQ later.
C.Increase the message retention period to 14 days and use a DLQ with high retention.
D.Use Amazon SNS to fan out messages to multiple SQS queues, each with different retry policies.
AnswerA

This approach leverages the SQS visibility timeout as a built-in retry buffer: setting it to 6 hours prevents messages from being redelivered to consumers while the downstream service is unhealthy, effectively holding them in the queue. Combined with a high maxReceiveCount (e.g., 100) in the redrive policy, messages are not prematurely moved to the DLQ, allowing the Lambda consumer to keep retrying for hours until the service recovers. This ensures no data loss and avoids manual intervention, making it the correct pattern for intermittent downstream outages.

Why this answer

Increasing the visibility timeout to a long duration (e.g., 6 hours) prevents messages from being repeatedly retried and sent to the DLQ while the downstream service is unhealthy. Instead, messages remain in the SQS queue and become visible again only after the visibility timeout expires, allowing automatic retries once the downstream service recovers. This approach avoids premature DLQ delivery and leverages SQS's built-in redrive policy based on maxReceiveCount.

Exam trap

The trap here is that candidates often think increasing the DLQ retention or reducing retries (maxReceiveCount) is the solution, but the real key is controlling the retry timing via the visibility timeout to allow the downstream service to recover before messages are exhausted.

How to eliminate wrong answers

Option B is wrong because reducing maxReceiveCount to 1 sends messages to the DLQ immediately after the first failure, which defeats resilience by not allowing any retries and requiring manual reprocessing from the DLQ. Option C is wrong because increasing the message retention period and using a DLQ with high retention does not prevent messages from being sent to the DLQ prematurely; it only keeps them in the DLQ longer, but the downstream service may recover before the messages are consumed from the DLQ. Option D is wrong because using SNS to fan out to multiple SQS queues with different retry policies adds complexity and does not address the core issue of preventing premature DLQ delivery; it still relies on the same visibility timeout and retry mechanism.

2
MCQmedium

A company is deploying a critical microservice on Amazon ECS with Fargate. They need to ensure that the service can tolerate an Availability Zone failure. What is the BEST approach?

A.Use a cluster placement constraint to spread tasks across instances
B.Use EC2 launch type and spread tasks across instance types
C.Define the service to spread tasks across multiple Availability Zones
D.Configure service auto scaling to add tasks when CPU is high
AnswerC

Defining the service to spread tasks across multiple Availability Zones ensures that the microservice remains available even if a single AZ fails. In Amazon ECS, you can use a placement strategy with 'spread' across 'attribute:ecs.availability-zone' for EC2, or for Fargate you can use the `AvailabilityZoneSpread` placement strategy (or specify AZs in the service definition) to distribute tasks evenly across AZs. This directly mitigates the risk of an AZ outage by maintaining running tasks in other zones.

Why this answer

Amazon ECS with Fargate allows you to define a service with a 'spread across Availability Zones' strategy. By setting the service's placement strategy to spread tasks across multiple Availability Zones, the service automatically distributes tasks evenly across the specified AZs. If one AZ fails, the remaining tasks in other AZs continue to serve traffic, ensuring high availability.

This is the most direct and effective method for tolerating an AZ failure in a Fargate-based deployment.

Exam trap

The trap here is that candidates often confuse 'spreading tasks across instances' (Option A) or 'across instance types' (Option B) with the correct concept of spreading across Availability Zones, or they mistakenly think auto scaling (Option D) alone provides AZ failure tolerance, when in fact it only scales based on load and does not guarantee multi-AZ distribution.

How to eliminate wrong answers

Option A is wrong because cluster placement constraints are used with the EC2 launch type to control task placement on specific instances, not to spread tasks across Availability Zones; Fargate manages the underlying infrastructure, so placement constraints are not applicable. Option B is wrong because the EC2 launch type and spreading tasks across instance types addresses instance-level diversity, not AZ-level resilience; it does not protect against an entire AZ failure. Option D is wrong because service auto scaling based on CPU handles performance scaling, not fault tolerance; it does not distribute tasks across AZs and cannot recover from an AZ failure without pre-existing multi-AZ distribution.

3
Multi-Selecteasy

A company wants to ensure that its Amazon S3 bucket is resilient to accidental deletion of objects. Which TWO actions should be taken?

Select 2 answers
A.Enable MFA Delete on the bucket.
B.Enable S3 Object Lock.
C.Enable S3 Versioning.
D.Enable S3 Transfer Acceleration.
E.Configure a lifecycle policy to expire objects after 30 days.
AnswersA, C

MFA Delete requires an authenticated AWS user to provide both a valid AWS credential and a one-time code from an MFA device before permanently deleting an object version or toggling the versioning state. This effectively blocks accidental or malicious permanent deletion because an attacker who compromises credentials would still need physical access to the MFA token. It is a strong, targeted defense for critical S3 data, especially when combined with versioning.

Why this answer

Enabling MFA Delete on an S3 bucket requires multi-factor authentication for any delete operations, including object version deletion and bucket deletion. This adds a critical layer of protection against accidental or unauthorized deletions, as the user must present both their AWS credentials and a valid MFA code to perform these destructive actions.

Exam trap

The trap here is that candidates often confuse S3 Object Lock (which prevents overwrites and deletes during a retention period) with MFA Delete (which requires additional authentication for delete operations), but Object Lock does not protect against accidental bucket deletion or version deletion without MFA, and it is not a direct resilience mechanism for accidental deletion scenarios.

4
MCQhard

A DevOps engineer is responsible for a critical application that uses an Amazon Aurora MySQL cluster. The application experiences sudden spikes in read traffic. The engineer needs to ensure that read replicas automatically scale to handle the load and that the application can tolerate the failure of an Availability Zone without manual intervention. Which solution should the engineer implement?

A.Use Amazon RDS for MySQL with Multi-AZ and create read replicas in multiple Availability Zones, then configure an Auto Scaling group for the replicas.
B.Create an Aurora Auto Scaling policy that adds Aurora Replicas based on CPU utilization, and place the replicas in multiple Availability Zones.
C.Enable Aurora Global Database and promote a secondary Region during read spikes to distribute the load.
D.Configure an Application Load Balancer to distribute read traffic across multiple Aurora Replicas, and use AWS Lambda to add replicas when CPU exceeds a threshold.
AnswerB

Aurora Auto Scaling automatically adjusts the number of Aurora Replicas based on metrics like CPU utilization. Placing replicas in multiple Availability Zones ensures that if one AZ fails, replicas in other AZs continue to serve read traffic. This solution provides automatic scaling and AZ resilience with minimal manual intervention, meeting both requirements.

Why this answer

Aurora Auto Scaling automatically adds or removes Aurora Replicas based on performance metrics, such as CPU utilization, to handle read traffic spikes. Deploying replicas across multiple Availability Zones ensures that the cluster remains available if an AZ fails. This managed feature provides automatic scaling and high availability with minimal operational effort, directly satisfying the requirements.

Exam trap

The trap here is confusing Aurora Replicas with RDS read replicas or assuming that manual scaling via Lambda is needed, when Aurora Auto Scaling handles it natively.

5
MCQmedium

A company experiences intermittent high latency for a web application running on EC2 behind an ALB. They want to monitor and automatically replace instances that have high CPU. Which solution meets this requirement?

A.Create a CloudWatch alarm on CPU utilization that triggers an Auto Scaling policy to replace the instance
B.Use Auto Scaling scheduled scaling actions to replace instances at peak times
C.Use AWS Lambda to periodically check CPU and terminate high-CPU instances
D.Configure the ALB health check to mark instances unhealthy when CPU is high
AnswerA

A CloudWatch alarm on CPUUtilization monitors the metric in near real-time and, when breached, triggers an Auto Scaling policy that terminates the affected instance, after which the Auto Scaling group automatically launches a fresh replacement. This directly addresses the intermittent high latency by reacting to the actual symptom (high CPU) as it occurs, rather than relying on predetermined schedules, and the managed replacement ensures capacity is maintained without manual intervention.

Why this answer

You can configure a CloudWatch alarm on the EC2 instance's CPU utilization metric, and then use that alarm to trigger an Auto Scaling lifecycle hook or a scaling policy that terminates the unhealthy instance and launches a replacement. This directly ties performance monitoring to automated instance replacement, meeting the requirement to replace instances with high CPU.

Exam trap

The trap here is that candidates often confuse ALB health checks with instance health monitoring, assuming ALB can react to CPU metrics, when in fact ALB health checks only verify application-level responsiveness (e.g., HTTP status codes) and cannot directly measure CPU utilization.

How to eliminate wrong answers

Option B is wrong because scheduled scaling actions replace instances at fixed times, not in response to real-time high CPU utilization, so they cannot address intermittent latency. Option C is wrong because while Lambda could terminate instances, it adds unnecessary complexity and latency, and Auto Scaling already provides native health-check-based replacement without custom code. Option D is wrong because ALB health checks are designed to detect application or network failures (e.g., HTTP 5xx, connection timeouts), not CPU utilization; they cannot be configured to mark instances unhealthy based on CPU metrics.

6
MCQmedium

A DevOps team is designing a disaster recovery solution for an Amazon RDS for MySQL database. The primary database is in us-east-1, and the recovery point objective (RPO) is 5 minutes, recovery time objective (RTO) is 1 hour. Which solution meets these requirements?

A.Enable Multi-AZ deployment for high availability.
B.Create a cross-Region read replica in the secondary Region.
C.Take manual snapshots and copy them to the secondary Region daily.
D.Configure automated backups with a retention period of 35 days.
AnswerB

A cross-Region read replica uses asynchronous replication to continuously copy changes from the source DB instance to a replica in the secondary Region, keeping data loss typically within 5 minutes. In a disaster, you can promote this replica to a standalone primary instance, which is a fast, reversible operation that meets the 1-hour RTO. This is the only option that both maintains an up-to-date copy in another Region and provides a ready-to-activate target for write traffic.

Why this answer

A cross-Region read replica in the secondary Region meets the RPO of 5 minutes because replication from the primary RDS instance to the read replica is asynchronous but typically completes within seconds to a few minutes, well under the 5-minute threshold. In a disaster, promoting the read replica to a standalone instance can be done manually or automated, and the RTO of 1 hour is achievable because promotion takes only a few minutes, leaving ample time for DNS and application failover. This solution provides a continuous replication stream without manual intervention, unlike snapshot-based approaches.

Exam trap

The trap here is that candidates confuse Multi-AZ (high availability within a Region) with cross-Region disaster recovery, assuming Multi-AZ protects against Regional failures, but it only protects against Availability Zone failures within the same Region.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability within a single Region (us-east-1) by synchronously replicating to a standby in a different Availability Zone, but it does not protect against a Regional disaster, so it cannot meet the cross-Region recovery requirement. Option C is wrong because taking manual snapshots daily and copying them to the secondary Region results in an RPO of up to 24 hours, far exceeding the required 5 minutes, and the copy operation adds additional latency. Option D is wrong because automated backups with a retention period of 35 days are stored within the same Region and cannot be used for cross-Region recovery; they also do not provide a mechanism to restore in a secondary Region within the required RPO/RTO.

7
MCQmedium

A company runs a stateless web application on Amazon ECS with Fargate. The application must be highly available across multiple Availability Zones. What is the BEST way to achieve this?

A.Create an ECS service with tasks in multiple AZs and place an ALB in front.
B.Use an Auto Scaling group of EC2 instances in a single AZ and run ECS tasks on them.
C.Deploy a CloudFront distribution with multiple origins in different AZs.
D.Deploy a single ECS service with tasks in one AZ and use an ALB.
AnswerA

Running ECS tasks in multiple Availability Zones ensures that an AZ failure does not cause total loss of capacity. The ALB performs health checks and routes traffic only to healthy targets, so if tasks in one AZ become unhealthy, traffic automatically shifts to tasks in the other AZ. Additionally, ECS service scheduler maintains desired task count across AZs, and with service auto scaling you can scale based on request count or CPU. This architecture provides both resiliency and elasticity.

Why this answer

Running an ECS service with tasks in multiple Availability Zones (AZs) and placing an Application Load Balancer (ALB) in front ensures that if one AZ becomes unavailable, the ALB can route traffic to healthy tasks in the remaining AZs. This architecture provides both high availability and fault tolerance for the stateless web application, as the ALB performs health checks and distributes requests across tasks in different AZs.

Exam trap

The trap here is that candidates may think CloudFront (Option C) provides high availability for dynamic web applications, but it is a CDN for static content caching and does not replace the need for an ALB with multi-AZ task placement for active traffic routing.

How to eliminate wrong answers

Option B is wrong because using an Auto Scaling group of EC2 instances in a single AZ creates a single point of failure; if that AZ goes down, all instances and tasks become unavailable, violating the high availability requirement. Option C is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations; it does not provide active-active load balancing across AZs for dynamic web application traffic, and its origins are not designed to replace an ALB for distributing requests to ECS tasks. Option D is wrong because deploying a single ECS service with tasks in one AZ means all tasks are in a single failure domain; even with an ALB, if that AZ fails, the application becomes unavailable.

8
MCQeasy

A company runs a stateless web application on EC2 instances in an Auto Scaling group across three Availability Zones. The application uses an Application Load Balancer. The operations team needs to ensure that the application remains available if one AZ fails. Which solution is MOST resilient?

A.Configure the Auto Scaling group to launch instances in a single Availability Zone with a desired capacity of 6.
B.Configure the Auto Scaling group to launch instances in two Availability Zones with a desired capacity of 4.
C.Configure the Auto Scaling group to launch instances in three Availability Zones with a desired capacity of 3.
D.Configure the Auto Scaling group to launch instances in two Availability Zones with a desired capacity of 6, all in one AZ.
AnswerC

Configuring three Availability Zones with a desired capacity of three places one instance in each AZ, so if any single AZ fails, the remaining two instances continue serving traffic, preserving 66% of capacity. The Auto Scaling group will automatically detect the unhealthy instances and launch replacements in other healthy AZs, gradually restoring capacity to the desired level. Because the application is stateless, a load balancer can distribute traffic across the surviving instances, providing high availability with minimal disruption.

Why this answer

Distributing instances across three Availability Zones (AZs) with a desired capacity of 3 ensures that even if one AZ fails, the remaining two AZs still have at least 2 instances running, maintaining service capacity. The Application Load Balancer (ALB) automatically routes traffic away from the failed AZ, and the Auto Scaling group will replace lost instances in the healthy AZs, providing the highest resilience against a single-AZ failure.

Exam trap

The trap here is that candidates often think using two AZs is sufficient for high availability, but the question specifically asks for the 'MOST resilient' solution, and three AZs provide better fault isolation and recovery capacity than two, especially when the desired capacity is low.

How to eliminate wrong answers

Option A is wrong because launching all instances in a single AZ creates a single point of failure; if that AZ fails, all instances are lost and the application becomes unavailable. Option B is wrong because distributing instances across only two AZs with a desired capacity of 4 means that if one AZ fails, the remaining AZ may have only 2 instances (if evenly split), but the total capacity drops by 50%, and the Auto Scaling group cannot launch instances in the failed AZ, potentially leading to insufficient capacity. Option D is wrong because it configures instances in two AZs but places all 6 instances in one AZ, which is functionally identical to a single-AZ deployment and provides no resilience against an AZ failure.

9
Multi-Selectmedium

A company is deploying a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application must be resilient to regional outages. Which THREE steps should the company take to achieve multi-Region resilience? (Choose THREE.)

Select 3 answers
A.Use Amazon CloudFront with multiple origins pointing to each Region's API Gateway.
B.Configure Route 53 with a failover routing policy to direct traffic to the secondary Region if the primary fails.
C.Use DynamoDB global tables to replicate data across Regions.
D.Deploy Lambda@Edge functions to handle requests at edge locations.
E.Deploy a second API Gateway and Lambda function in another Region.
AnswersB, C, E

Route 53 failover routing enables traffic redirection.

Why this answer

Amazon Route 53 with a failover routing policy allows the company to route traffic to a secondary Region when health checks detect a failure in the primary Region. This provides DNS-level failover, which is a fundamental component of multi-Region resilience for HTTP-based applications.

Exam trap

The trap here is that candidates often confuse CloudFront's origin failover capability (which requires manual configuration of origin groups) with automatic multi-Region failover, or they mistakenly believe Lambda@Edge can serve as a full application backend across Regions, when in fact it is limited to edge processing and cannot replace regional Lambda deployments.

10
MCQhard

A company runs a stateless web application on AWS Lambda behind an Application Load Balancer (ALB). During a deployment, the team updates the Lambda function to a new version. Some users report seeing the old version of the application for several minutes after the deployment. What is the MOST likely cause?

A.The Lambda function versions are not immutable, causing a gradual rollout.
B.Lambda@Edge is overriding the function version at the edge locations.
C.Amazon CloudFront is caching the old response and has not been invalidated.
D.The ALB target group is still pointing to the old Lambda function version due to connection draining.
AnswerD

ALB invokes Lambda functions via a target group that references a specific function version or alias. When you publish a new version and update the target group, ALB's connection draining process allows existing in-flight connections to complete on the old version before deregistering it. As a result, the old Lambda version can continue serving requests for a short period, causing a temporary gradual rollout until draining finishes.

Why this answer

When an ALB is used with Lambda, the ALB invokes a specific Lambda function version or alias. If the deployment updates the Lambda function but the ALB target group alias is not updated atomically, or if connection draining keeps old connections active, some requests may still be routed to the old version. This can cause users to see the old application for several minutes.

Option A is wrong because Lambda versions are immutable, so gradual rollout is not related. Option B is wrong because Lambda@Edge is not used in this setup (the application runs behind an ALB, not CloudFront). Option C is wrong because CloudFront is not mentioned in the architecture—the traffic goes directly from ALB to Lambda.

11
MCQmedium

A company runs a microservices architecture on Amazon ECS with Fargate. Services communicate via an internal Application Load Balancer. Recently, one service became unavailable due to a memory leak, causing cascading failures in downstream services. What design change would MOST effectively improve resilience and limit the blast radius?

A.Increase the memory limit for each ECS task to accommodate memory leaks.
B.Implement circuit breaker patterns in the service discovery and client libraries to stop calling unhealthy services.
C.Enable connection draining on the ALB to allow in-flight requests to complete.
D.Implement automatic scaling policies for ECS services based on memory utilization.
AnswerB

Circuit breakers are a client-side resilience pattern that monitor outgoing requests to a dependency and, after exceeding a failure threshold, automatically fail fast without attempting the network call. In an ECS microservices environment with service discovery, this stops unhealthy services from being flooded with retries, preventing latency spikes and thread exhaustion in healthy services. By quarantining the failing service from call traffic, circuit breakers effectively stop cascading failures and allow the unhealthy service time to recover.

Why this answer

Implementing a circuit breaker pattern in service discovery and client libraries stops requests to unhealthy services, preventing cascading failures and limiting blast radius. Option A is wrong because increasing memory limits only delays the inevitable failure and does not prevent downstream services from being affected. Option C (connection draining) only affects in-flight requests during deregistration, not active health issues.

Option D (auto scaling) helps but does not stop requests from being sent to a failing service; scaling cannot fix a memory leak.

12
MCQhard

A company runs an application on EC2 with a shared Elastic IP. The instance fails and an engineer manually attaches the Elastic IP to a standby instance. To automate this failover, which service should be used?

A.Use an Auto Scaling group with a lifecycle hook
B.AWS Elastic Beanstalk
C.CloudWatch Events with a Lambda target
D.Configure a second Elastic IP
AnswerC

CloudWatch Events (now EventBridge) can capture Simple Notification Service messages or CloudWatch alarm state changes triggered by a Route 53 health check or EC2 status check, and then invoke an AWS Lambda function. That function can programmatically call ec2-associate-address to detach the Elastic IP from the failed instance and attach it to a healthy one, fully automating active/passive failover. This approach is the standard serverless pattern for EIP failover because it reacts to health signals and directly manages the association.

Why this answer

C is correct because CloudWatch Events (now Amazon EventBridge) can detect the EC2 instance state change (e.g., 'stopped' or 'failed') and trigger a Lambda function. The Lambda function can then programmatically disassociate the Elastic IP from the failed instance and reassociate it to a standby instance using the AWS SDK (e.g., ec2.disassociate_address and ec2.associate_address). This provides a fully automated, event-driven failover without manual intervention.

Exam trap

The trap here is that candidates often confuse event-driven automation (CloudWatch Events + Lambda) with scaling or deployment services (Auto Scaling, Elastic Beanstalk), or they assume adding a second Elastic IP solves failover without considering the need for automated reassignment.

How to eliminate wrong answers

Option A is wrong because an Auto Scaling group with a lifecycle hook is designed to manage instance launch/termination events (e.g., running custom scripts during scale-out/scale-in), not to reassociate an Elastic IP to a standby instance; it does not handle Elastic IP failover. Option B is wrong because AWS Elastic Beanstalk is a PaaS service for deploying and scaling web applications, not a mechanism for automating Elastic IP reassignment; it manages environments but does not provide direct control over Elastic IP failover. Option D is wrong because configuring a second Elastic IP does not automate failover; it simply adds another static IP, but the engineer would still need to manually update DNS or routing to switch traffic, which does not solve the automation requirement.

13
Multi-Selecthard

A company runs a critical application on AWS that uses an Auto Scaling group of EC2 instances. The application must remain available even if an entire Availability Zone fails. Which THREE actions should the company take?

Select 3 answers
A.Configure an ALB health check to automatically replace unhealthy instances.
B.Use a single instance in each Availability Zone to minimize cost.
C.Use multiple subnets in each Availability Zone for the instances.
D.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
E.Use an Elastic Load Balancer (ELB) to distribute traffic across the instances in different AZs.
AnswersA, D, E

An ALB's health check marks an instance unhealthy when it repeatedly fails the configured protocol and path check, and the ALB then stops sending traffic to that instance. When the Auto Scaling group is configured to use ELB health checks, the 'ReplaceUnhealthy' process automatically terminates the unhealthy instance and launches a new one to maintain desired capacity. Thus, the health check is the trigger that enables the self-healing replacement of failed instances, ensuring application availability.

Why this answer

Configuring an ALB health check allows the Auto Scaling group to automatically detect and replace unhealthy instances. The ALB health check pings the instances at a specified interval (e.g., every 30 seconds) and marks them as unhealthy if they fail to respond. This triggers the Auto Scaling group to terminate the unhealthy instance and launch a new one, maintaining application availability even if an instance fails within a single AZ.

Exam trap

The trap here is that candidates might think using multiple subnets within a single AZ (Option C) provides redundancy, but it does not protect against an AZ failure, which requires instances to be spread across at least two distinct AZs.

14
Multi-Selecteasy

A startup runs a stateless web application on AWS Elastic Beanstalk with a single environment. The application uses an Amazon RDS for MySQL database instance. The startup is preparing for a marketing campaign that is expected to increase traffic by 10x. The CTO is concerned about the application's ability to handle the load and wants to ensure high availability and resilience. The current architecture has a single RDS instance (db.t3.medium) and a single Elastic Beanstalk environment with one EC2 instance (t3.medium). The startup has a limited budget but wants to improve resilience without over-provisioning. Which combination of actions should the DevOps engineer recommend? (Choose THREE.)

Select 3 answers
A.Add an Amazon ElastiCache cluster to cache frequent database queries.
B.Use dedicated instances for the EC2 instances to ensure consistent performance.
C.Switch the Elastic Beanstalk environment to a load-balanced, auto-scaled environment with a minimum of 2 instances across 2 Availability Zones.
D.Enable Multi-AZ deployment for the RDS instance to provide a standby in another AZ.
E.Add Amazon RDS Proxy in front of the RDS instance to handle connection pooling.
AnswersC, D, E

Deploying the Elastic Beanstalk environment as a load-balanced, auto-scaling configuration with a minimum of two instances in separate Availability Zones eliminates the web tier as a single point of failure. The load balancer distributes traffic across instances and health-checks them, while Auto Scaling replaces unhealthy instances and can scale out during load spikes. This is the foundational action for high availability of stateless applications because it provides both redundancy and elasticity.

Why this answer

Option C is correct because converting the single-instance Elastic Beanstalk environment to a load-balanced, auto-scaled environment with a minimum of two instances across two Availability Zones removes the single point of failure at the web tier and lets the environment scale horizontally to absorb the 10x traffic spike. Option D is correct because enabling Multi-AZ on the RDS for MySQL instance creates a synchronous standby replica in a second AZ with automatic failover, improving database resilience without requiring application changes. Option E is correct because RDS Proxy pools and shares database connections, which prevents the connection exhaustion and overhead that occur when many new EC2 instances and users open connections directly to MySQL during a traffic surge.

Option A is not among the marked answers, and while caching could reduce read load, it is not required to achieve the stated high availability and resilience goals. Option B is not marked correct because dedicated instances raise cost and do not provide the multi-AZ redundancy or elasticity that the scenario demands.

Exam trap

DOP-C02 often tests the misconception that adding caching or dedicated instances is necessary for resilience, but the core actions are auto-scaling, Multi-AZ, and connection pooling; candidates may overlook RDS Proxy or choose cost-ineffective options.

15
MCQmedium

A company is building a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application must be resilient to sudden spikes in traffic without manual intervention. Which combination of services should be used?

A.API Gateway with throttling, Lambda with reserved concurrency, and DynamoDB auto scaling.
B.API Gateway with usage plans, Lambda with provisioned concurrency, and DynamoDB on-demand.
C.API Gateway with WAF, Lambda with function URLs, and DynamoDB Accelerator (DAX).
D.API Gateway with caching, Lambda with no concurrency limits, and DynamoDB global tables.
AnswerA

This combination directly addresses a demand spike: API Gateway throttling imposes a maximum request rate per client or across the API, preventing a flood from reaching Lambda. Reserved concurrency allocates a guaranteed number of Lambda executions, insulating the function from account-level throttling and ensuring the spikes are absorbed within the reserved ceiling. DynamoDB auto scaling adjusts provisioned read/write capacity based on live utilization, so the database keeps pace without manual intervention.

Why this answer

It combines API Gateway throttling to absorb traffic spikes by queuing or rejecting excess requests, Lambda reserved concurrency to guarantee execution capacity for the function, and DynamoDB auto scaling to adjust read/write capacity based on demand. This triad ensures the application remains available and responsive under sudden load without manual intervention.

Exam trap

The trap here is that candidates confuse 'provisioned concurrency' (Option B) with 'reserved concurrency' (Option A), mistakenly believing pre-warming instances handles spikes, when in fact reserved concurrency guarantees capacity but does not reduce cold starts, and provisioned concurrency is for latency, not burst resilience.

How to eliminate wrong answers

Option B is wrong because Lambda provisioned concurrency is designed for low-latency cold starts, not for handling traffic spikes; it pre-warms a fixed number of instances, which can be overwhelmed if spikes exceed that count, and DynamoDB on-demand is suitable for unpredictable workloads but can incur higher costs and does not prevent throttling at the API or Lambda layer. Option C is wrong because AWS WAF provides web application firewall protection against exploits, not traffic spike resilience; Lambda function URLs are a direct invocation method without built-in throttling, and DynamoDB Accelerator (DAX) is an in-memory cache for read-heavy workloads, not a scaling mechanism for write spikes. Option D is wrong because API Gateway caching reduces backend load but does not throttle incoming requests; Lambda with no concurrency limits can lead to uncontrolled scaling and potential account-level throttling; DynamoDB global tables provide multi-region replication for disaster recovery, not automatic scaling under traffic spikes.

16
MCQmedium

A company runs a critical web application on AWS using an Application Load Balancer (ALB) in front of an Auto Scaling group of EC2 instances. The application experiences periodic traffic spikes. To handle these spikes, the company wants to use a combination of proactive scaling based on a predictable schedule and reactive scaling based on CPU utilization. What is the MOST resilient scaling strategy?

A.Use a scheduled scaling policy for the predictable spikes and a step scaling policy for CPU utilization.
B.Use predictive scaling based on historical traffic patterns.
C.Use manual scaling by increasing the desired capacity before expected spikes.
D.Use a target tracking scaling policy based on average CPU utilization.
AnswerA

Scheduled scaling pre-provisions instances at fixed times, aligning capacity with known upcoming demand, while a step scaling policy reacts to actual CPU utilization through CloudWatch alarms with defined adjustment steps (e.g., add 2 instances when CPU exceeds 80%, add 1 when it exceeds 70%). This combination ensures both proactive readiness for predictable traffic spikes and rapid reactive response to any deviation, giving the highest resilience for mixed traffic patterns.

Why this answer

It combines scheduled scaling for predictable traffic spikes with step scaling for reactive adjustments based on CPU utilization, providing both proactive and reactive resilience. Scheduled scaling adjusts capacity in advance of known events, while step scaling allows for larger, more aggressive adjustments when CPU utilization exceeds thresholds, avoiding the slower, linear response of target tracking. This dual approach ensures the application can handle spikes without over-provisioning or under-provisioning.

Exam trap

The trap here is that candidates often assume predictive scaling (Option B) is the best for all predictable patterns, but it fails for non-recurring or sudden spikes, and they overlook that target tracking (Option D) cannot proactively add capacity before a spike begins.

How to eliminate wrong answers

Option B is wrong because predictive scaling relies on historical traffic patterns and may not accurately predict sudden, non-recurring spikes, leading to under-provisioning during critical events. Option C is wrong because manual scaling requires human intervention, which is not resilient for periodic spikes that occur outside business hours or without warning, and it lacks automation for reactive scaling. Option D is wrong because target tracking scaling only adjusts capacity to maintain a specific CPU utilization target, which can be too slow to respond to rapid spikes and does not allow for proactive scaling based on a schedule.

17
MCQhard

A company is building a global application that requires low-latency access to static content across multiple AWS Regions. The content changes infrequently. Which solution is MOST resilient and cost-effective?

A.Use Amazon S3 Transfer Acceleration
B.Set up a VPN to a single Region
C.Deploy EC2 instances in each Region with a global load balancer
D.Use Amazon CloudFront with an S3 bucket as origin
AnswerD

Amazon CloudFront with an S3 bucket as origin is the intended AWS solution for low-latency global delivery of static content. CloudFront caches objects at hundreds of edge locations, so users receive data from the nearest edge POP rather than the origin, drastically reducing latency. The S3 bucket provides durable, cost-effective storage, and CloudFront minimizes origin requests via cache hits, reducing cost. This combination also improves resilience by absorbing traffic spikes at the edge and supports features like SSL, geo-restriction, and signed URLs.

Why this answer

Amazon CloudFront with an S3 bucket as origin provides low-latency global content delivery by caching static content at edge locations worldwide. It is highly resilient because CloudFront automatically routes around failures, and cost-effective since it reduces load on origin and data transfer costs.

Exam trap

DOP-C02 often tests the difference between CDN (CloudFront) and other acceleration services like S3 Transfer Acceleration, causing candidates to choose the latter for global content delivery.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration speeds up uploads to S3, not global content delivery to users. Option B is wrong because a VPN to a single Region does not provide low-latency access across multiple Regions and adds complexity. Option C is wrong because deploying EC2 instances in each Region with a global load balancer is more expensive and complex than using a CDN, and does not inherently cache content.

18
MCQhard

Refer to the exhibit. An IAM policy is attached to a user. A developer tries to upload an object to s3://my-bucket/confidential/report.pdf without specifying server-side encryption. What will happen?

A.The upload succeeds because the Allow statement grants PutObject.
B.The upload succeeds because the Deny condition uses a wrong condition key.
C.The upload fails because the Deny statement requires SSE-KMS.
D.The upload fails only if the object name matches the prefix.
AnswerC

An explicit Deny in the IAM policy overrides any Allow, and it is conditioned on the absence of SSE-KMS encryption headers. Because the request specifies no server-side encryption, the Deny matches and the upload is rejected.

Why this answer

The upload fails because the Deny statement in the IAM policy explicitly denies the s3:PutObject action unless the request includes the s3:x-amz-server-side-encryption header with a value of 'aws:kms'. When the developer does not specify server-side encryption, the condition is not met, so the Deny applies, causing the upload to fail. Option A is incorrect because the Allow statement does not override an explicit Deny.

Option B is incorrect because the condition key is correct; the failure is due to the missing encryption header. Option D is incorrect because the Deny is not based on the object name prefix; it applies to all objects in the bucket.

19
MCQhard

An AWS account owner (Account A) owns an S3 bucket named my-bucket. The bucket policy shown in the exhibit is attached to the bucket. A user from Account B attempts to upload an object to the bucket without specifying the x-amz-acl header. What will happen?

A.The upload fails because the bucket policy requires the object ACL to be set, but the default ACL allows the upload anyway.
B.The upload succeeds because the bucket policy does not explicitly deny the request.
C.The upload succeeds because the bucket policy allows s3:PutObject for any principal.
D.The upload fails because the bucket policy requires the x-amz-acl header to be set to bucket-owner-full-control.
AnswerD

The bucket policy uses a condition key such as s3:x-amz-acl with a value of bucket-owner-full-control, making that header a mandatory requirement for any successful s3:PutObject call. When the requester omits the x-amz-acl header, the condition evaluates to false, so the allow statement cannot grant the action. With no other applicable allow, the request is implicitly denied and the upload fails. This design ensures the bucket owner can later manage or delete the object by forcing ownership transfer via the canned ACL.

Why this answer

The condition requires the x-amz-acl header to be set to bucket-owner-full-control. If the header is not specified, the condition fails, and the request is denied. Option A is wrong because the condition is not met.

Option B is wrong because the policy does not grant permission without the header. Option C is wrong because the bucket policy evaluates before the object ACL.

20
MCQhard

A company runs a stateful application on EC2 instances with instance store volumes. The application requires low-latency access to data. The operations team needs to ensure that instance failure does not result in data loss. Which solution is MOST resilient?

A.Use instance store volumes with RAID 1 across multiple instances.
B.Replicate data in real time to an EBS volume and take periodic snapshots.
C.Use larger instance types with more instance store capacity.
D.Create an AMI of the instance periodically to capture the data.
AnswerB

Replicating application data in real time to an Amazon EBS volume gives you a persistent, network-attached copy that survives the EC2 instance's lifecycle. EBS volumes are independently replicated within an Availability Zone, so they remain available if the instance is stopped or terminated. Taking periodic EBS snapshots copies the volume to Amazon S3, providing point-in-time recovery points that can be restored in the same or a different Availability Zone, which protects against both instance failure and EBS volume loss.

Why this answer

It combines the low-latency performance of instance store volumes with the durability of EBS snapshots. By replicating data in real time to an EBS volume, the application benefits from the instance store's speed while the EBS volume provides a persistent copy that survives instance failure. Periodic snapshots of the EBS volume add further resilience by enabling point-in-time recovery, ensuring data is not lost even if the instance or its instance store fails.

Exam trap

The trap here is that candidates assume instance store volumes are inherently durable because they are fast, overlooking that they are ephemeral and tied to the instance lifecycle, while the correct solution uses a hybrid approach to combine performance with persistence.

How to eliminate wrong answers

Option A is wrong because RAID 1 across multiple instances requires network-based replication, which introduces latency and complexity, and does not guarantee data durability if all instances fail simultaneously or if the instance store volumes themselves are ephemeral and tied to instance lifecycle. Option C is wrong because using larger instance types with more instance store capacity only increases storage size, not durability; instance store data is still lost on instance failure, reboot, or termination. Option D is wrong because creating an AMI periodically captures the entire instance state, but AMIs are not designed for real-time data replication; they are point-in-time snapshots that can lead to significant data loss between creation intervals and do not provide low-latency access to the latest data.

21
MCQeasy

A company is using Amazon RDS for MySQL with Multi-AZ deployment. During a recent failover, the application experienced a brief downtime because the DNS cache on the application servers still pointed to the old primary. How can a DevOps engineer minimize this downtime?

A.Use an RDS Proxy to manage connections and reduce DNS dependency.
B.Configure the application to use the Multi-AZ endpoint instead of the primary endpoint.
C.Configure application servers to use a hardcoded IP address instead of the RDS endpoint.
D.Increase the TTL on the RDS DNS record.
AnswerA

RDS Proxy sits between your application and the database, exposing a fixed writer endpoint that masks the underlying instance DNS changes. When a Multi-AZ failover occurs, RDS Proxy automatically redirects existing connections to the new primary, which eliminates the delay caused by clients waiting for DNS TTL to expire. It also pools and reuses database connections, reducing connection-related errors during failover and lowering CPU/memory pressure on the database. This is exactly why it reduces DNS dependency and delivers failover times typically under one second.

Why this answer

RDS Proxy acts as a connection broker that maintains persistent connections to the database and abstracts the underlying DNS changes during failover. When a failover occurs, RDS Proxy automatically reconnects to the new primary without requiring the application to resolve a new DNS record, thereby eliminating the downtime caused by stale DNS caches. This reduces the application's dependency on DNS resolution and provides faster failover recovery.

Exam trap

The trap here is that candidates often think increasing TTL or using a different endpoint will help, but the real issue is DNS cache staleness, which RDS Proxy bypasses entirely by managing connections at the proxy layer.

How to eliminate wrong answers

Option B is wrong because there is no such thing as a 'Multi-AZ endpoint' in RDS; the Multi-AZ feature uses a single DNS endpoint (the primary) that is automatically updated after failover, so using a different endpoint does not solve the DNS caching issue. Option C is wrong because hardcoding an IP address is highly discouraged — RDS instances can change IP addresses after failover or maintenance, leading to permanent connectivity loss. Option D is wrong because increasing the TTL on the RDS DNS record would actually make the DNS cache stale for longer, increasing downtime instead of minimizing it.

22
MCQmedium

A company uses Amazon RDS Multi-AZ for disaster recovery. The primary DB instance in us-east-1a fails. What happens next?

A.The standby DB instance in us-east-1b is promoted automatically and the CNAME record is updated
B.The administrator must manually promote the standby instance
C.The primary instance is automatically rebuilt in the same AZ
D.A read replica in us-east-1b is automatically promoted to primary
AnswerA

RDS Multi-AZ maintains a synchronous standby in a different Availability Zone, so when the primary in us-east-1a fails, the standby in us-east-1b is promoted automatically and the DB instance's CNAME endpoint is repointed to it. This satisfies the stem's automatic failover requirement without manual intervention.

Why this answer

RDS Multi-AZ maintains a synchronous standby replica in a different AZ and performs automatic failover by promoting the standby and updating the DB instance's DNS CNAME to point to the new primary. The failover is triggered by the primary's failure and typically completes in 60–120 seconds, with no manual intervention required. Applications using the endpoint hostname reconnect transparently once DNS TTL expires.

Exam trap

DOP-C02 often tests whether candidates confuse Multi-AZ (synchronous standby, automatic failover, HA) with read replicas (asynchronous, manual promotion, read scaling), causing them to pick the read-replica promotion answer.

How to eliminate wrong answers

Option B is wrong because Multi-AZ failover is automatic — manual promotion is the behavior of a read replica promotion (which is a separate feature), not Multi-AZ. Option C is wrong because RDS does not rebuild the primary in the same AZ during failover; it promotes the standby in the other AZ to preserve availability. Option D is wrong because read replicas are not part of the Multi-AZ failover mechanism — they are asynchronous, separately managed, and require manual promotion; Multi-AZ uses a synchronous standby, not a read replica.

23
MCQeasy

A company runs a stateless web application on EC2 instances behind an Application Load Balancer. To improve resilience, which configuration should be used for the EC2 instances?

A.Use one EC2 instance with a larger instance type
B.Use a single, large EC2 instance in one Availability Zone
C.Use multiple EC2 instances in one Availability Zone with health checks disabled
D.Use multiple EC2 instances across two or more Availability Zones
AnswerD

Spreading instances across two or more Availability Zones ensures the application survives an AZ outage, since the Application Load Balancer routes to healthy targets in remaining zones. This directly satisfies the resilience requirement for the stateless workload.

Why this answer

D is correct because deploying multiple EC2 instances across two or more Availability Zones (AZs) ensures high availability and fault tolerance. If one AZ fails, the Application Load Balancer (ALB) automatically routes traffic to healthy instances in other AZs, maintaining service continuity. This aligns with the AWS Well-Architected Framework's resilience best practices for stateless applications.

Exam trap

The trap here is that candidates may think scaling vertically (larger instance) or using multiple instances in a single AZ is sufficient, but the DOP-C02 exam specifically tests the requirement for multi-AZ deployment to achieve resilience against AZ failures.

How to eliminate wrong answers

Option A is wrong because using a single, larger EC2 instance creates a single point of failure; if that instance fails, the entire application goes down. Option B is wrong because placing a single large instance in one AZ does not protect against AZ-level failures, such as power outages or network disruptions. Option C is wrong because using multiple instances in one AZ with health checks disabled means the ALB cannot detect and route away from failed instances, and a single AZ failure still takes down all instances.

24
MCQeasy

A company wants to ensure that its Amazon S3 bucket can withstand the loss of an entire AWS Availability Zone. Which configuration meets this requirement?

A.Use the S3 Standard storage class.
B.Configure cross-Region replication to another bucket.
C.Enable S3 Versioning on the bucket.
D.Use the S3 One Zone-IA storage class.
AnswerA

S3 Standard is the default storage class and is engineered to deliver 99.999999999% object durability and 99.99% availability by synchronously storing each object across a minimum of three Availability Zones (AZs) within the same AWS Region. When an AZ becomes unavailable, S3 automatically serves requests from the remaining copies, so the bucket stays accessible without manual intervention. Because the data exists in at least three independent AZs, a single AZ failure does not cause data loss or downtime, making it the appropriate choice for resilience against an AZ disruption.

Why this answer

S3 Standard storage class automatically replicates data across at least three Availability Zones within an AWS Region, ensuring resilience against the loss of an entire AZ. Option B is incorrect because cross-Region replication replicates data to a different AWS Region, which provides geographic resilience but not specifically AZ resilience within the same Region. Option C is incorrect because S3 Versioning helps protect against accidental deletion or overwrite by preserving previous versions, but it does not provide data replication across AZs.

Option D is incorrect because S3 One Zone-IA stores data in a single AZ, which would not withstand the loss of that AZ.

25
Multi-Selecthard

A company runs a critical application on AWS using Amazon EC2 instances in an Auto Scaling group, an Application Load Balancer (ALB), and an Amazon RDS for PostgreSQL Multi-AZ DB cluster. The application must maintain an RTO of 5 minutes and an RPO of 1 second for database transactions. The current setup meets these requirements, but the DevOps team wants to improve the resilience of the application tier to withstand a regional failure. Which THREE actions should be taken? (Choose three.)

Select 3 answers
A.Replace the RDS Multi-AZ cluster with Amazon Aurora Global Database to replicate data across regions.
B.Use an active-passive architecture with a second Auto Scaling group and ALB in another region.
C.Use Amazon EFS Replication to replicate application data across regions with a recovery point objective (RPO) of 1 second.
D.Extend the existing Auto Scaling group to launch instances in two regions by specifying a second region in the launch template.
E.Set up Amazon Route 53 with health checks and failover routing policy to direct traffic to the secondary region if the primary fails.
AnswersA, B, E

RDS Multi-AZ only synchronously replicates to a standby in the same Region, so it cannot protect against a regional outage. Amazon Aurora Global Database replicates data to up to five secondary Regions using storage-level replication, typically with an RPO of 1 second or less, while still providing a familiar MySQL/PostgreSQL-compatible endpoint. Promoting a secondary Region during failover is a fast, deliberate action that preserves the database's durability and availability in a disaster recovery scenario.

Why this answer

Amazon Aurora Global Database is the correct choice because it provides cross-region replication with a typical RPO of 1 second and RTO of 1 minute, meeting the stated requirements. Unlike standard RDS Multi-AZ, which is limited to a single region, Aurora Global Database replicates data asynchronously across multiple regions with minimal lag, ensuring the database tier can survive a regional failure while maintaining the required RPO of 1 second.

Exam trap

The trap here is that candidates often confuse Multi-AZ with cross-region disaster recovery, assuming Multi-AZ alone provides regional failover, when in fact it only protects against Availability Zone failures within a single region.

26
Multi-Selectmedium

A company is designing a disaster recovery (DR) strategy for a critical application that runs on EC2 instances with an RDS database. The DR site must be in a different AWS Region. The Recovery Point Objective (RPO) is 15 minutes, and Recovery Time Objective (RTO) is 1 hour. Which TWO actions should the company take to meet these objectives? (Choose TWO.)

Select 2 answers
A.Use AWS Backup to copy EC2 AMIs and RDS snapshots to the DR region every 15 minutes.
B.Use AWS CloudFormation to pre-provision resources in the DR region manually.
C.Configure Amazon Route 53 with health checks and failover routing to the DR region.
D.Create an RDS cross-Region read replica in the DR region.
E.Configure S3 cross-Region replication for application data stored in S3.
AnswersC, D

Amazon Route 53 health checks monitor the primary endpoint (e.g., an Application Load Balancer or an IP address) and automatically detect when it becomes unhealthy. With failover routing, Route 53 stops returning the primary resource's DNS answer and instead returns the DR region's endpoint, effectively steering user traffic within minutes. This directly supports the RTO by eliminating the need for manual DNS changes, and it works with any application architecture as long as the endpoint can be health-checked. However, Route 53 only handles traffic redirection—it does not replicate data or pre-warm compute, so it must be paired with a data replication strategy to also meet the RPO.

Why this answer

Options C and D are correct. D: Creating an RDS cross-Region read replica in the DR region allows the replica to be promoted to the primary database with minimal data loss, meeting the 15-minute RPO. C: Configuring Amazon Route 53 with health checks and failover routing enables automatic traffic redirection to the DR region within the 1-hour RTO.

Option A is wrong because copying AMIs and RDS snapshots every 15 minutes would require launching EC2 instances and restoring the database from snapshots, which can exceed the 1-hour RTO. Option B is wrong because manually pre-provisioning resources with CloudFormation does not provide the automated failover needed to meet the RTO. Option E is wrong because S3 cross-Region replication does not address the EC2 and RDS components of the application.

27
Multi-Selecteasy

A company is designing a disaster recovery strategy for its application. The application runs on EC2 instances and uses an RDS MySQL database. The RTO is 1 hour, and the RPO is 15 minutes. Which TWO approaches meet these requirements?

Select 2 answers
A.Use a warm standby strategy: run a scaled-down version of the application in the DR region with RDS Multi-AZ across regions.
B.Use a pilot light strategy: replicate data using RDS cross-region automated backups and have a small environment running in the DR region.
C.Use a read replica in the DR region and promote it on failover.
D.Use a Multi-Zone deployment with RDS in the same region.
E.Use a backup and restore strategy: take snapshots every hour and restore in the DR region on failover.
AnswersA, B

The warm standby pattern keeps a fully functional, scaled-down copy of the application running in the DR region, so the entire stack is ready for production traffic almost immediately after failover. RDS data is continuously replicated to the DR region (via cross-region read replicas or Aurora global replication), keeping the RPO near zero—well within the 15-minute target. Because compute, storage, and networking are already provisioned, RTO is also met by simply scaling up and shifting traffic rather than building infrastructure from scratch.

Why this answer

Options A and B are correct. A warm standby with RDS Multi-AZ across regions ensures a standby database is ready and can be promoted quickly, meeting the 1-hour RTO. A pilot light with RDS cross-region automated backups provides replication with a 15-minute RPO; a small environment is running, allowing faster failover than a full pilot light.

Option C is wrong because RDS read replicas do not support automatic failover; manual promotion can take longer than 1 hour. Option D is wrong because Multi-AZ in the same region does not protect against region failure. Option E is wrong because hourly snapshots meet RPO but restoring from snapshots typically exceeds the 1-hour RTO.

28
MCQhard

A company runs a microservices architecture on Amazon ECS with Fargate. Each service is deployed in its own ECS service. The company wants to ensure that if one Availability Zone (AZ) fails, the services can continue to operate with minimal impact. What is the MOST resilient task placement strategy?

A.Use a task placement constraint to run tasks on distinct instances.
B.Use a task placement strategy that uses the random algorithm.
C.Use a task placement strategy that uses the binpack algorithm to maximize resource utilization.
D.Use a task placement strategy that spreads tasks across Availability Zones.
AnswerD

A spread placement strategy with the field attribute:ecs.availability-zone explicitly distributes tasks evenly across the Availability Zones used by the cluster or service. For Fargate, this strategy works in conjunction with your VPC subnets, ensuring tasks are placed in each configured AZ before any are duplicated. This directly satisfies the requirement that a single AZ failure does not take down all instances of the microservice.

Why this answer

The 'spread across Availability Zones' strategy explicitly distributes ECS tasks across multiple AZs, ensuring that if one AZ fails, the remaining AZs continue to run the service. This is the most resilient approach for Fargate tasks, as it leverages the AZ isolation provided by AWS to minimize the blast radius of a single-AZ failure.

Exam trap

The trap here is that candidates often confuse 'high availability' with 'resource efficiency' and choose binpack (Option C) because it reduces cost, but the question explicitly asks for resilience, not cost optimization.

How to eliminate wrong answers

Option A is wrong because 'distinct instances' is a constraint for EC2 launch type, not Fargate; Fargate tasks run on AWS-managed infrastructure, so this constraint is irrelevant and does not provide AZ-level resilience. Option B is wrong because the 'random' algorithm distributes tasks without any awareness of AZ boundaries, potentially placing all tasks in a single AZ and leaving the service vulnerable to that AZ's failure. Option C is wrong because 'binpack' maximizes resource utilization by packing tasks onto the fewest underlying resources, which often concentrates tasks in one AZ, reducing resilience and increasing the impact of an AZ failure.

29
MCQeasy

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group. To ensure high availability, the instances are deployed across three Availability Zones. Which additional step should the company take to protect against a regional failure?

A.Place all instances in a single Availability Zone to simplify management.
B.Use EC2 Dedicated Hosts to ensure capacity.
C.Increase the minimum size of the Auto Scaling group to 10 instances.
D.Deploy the application in a second AWS Region and use Route 53 with failover routing.
AnswerD

Deploying the application in a second AWS Region and using Route 53 failover routing gives you active-passive or active-active DNS-level failover: Route 53 health checks continuously monitor the primary endpoint, and when it is unhealthy, DNS queries are answered with the secondary Region's IP addresses. This directly addresses an entire Region becoming unavailable, as long as the secondary Region has the resources and the data needed to serve traffic. For a critical application, combine this with RDS Cross-Region Read Replicas or Aurora Global Database for proper data durability.

Why this answer

Deploying the application in a second AWS Region and using Route 53 failover routing provides protection against a regional failure, because the application remains available even if an entire AWS Region becomes unavailable. Route 53 health checks detect the failure and automatically redirect traffic to the standby Region. This is the only option that addresses regional-level disasters.

Exam trap

DOP-C02 often tests the misconception that multi-AZ deployment alone protects against regional failure, when true regional resilience requires a multi-Region architecture with Route 53 failover.

How to eliminate wrong answers

Option A is wrong because placing all instances in a single Availability Zone reduces availability and does not protect against regional failure; it actually increases risk. Option B is wrong because EC2 Dedicated Hosts provide physical server isolation for licensing or compliance, not regional redundancy. Option C is wrong because increasing the Auto Scaling group size within the same Region does not protect against a regional outage; all instances would still be in the affected Region.

30
Multi-Selectmedium

A company is designing a highly available architecture for a web application using AWS services. The application must be resilient to the failure of an entire AWS Region. Which TWO strategies should the company implement? (Choose TWO.)

Select 2 answers
A.Deploy the application in multiple AWS Regions and use Route 53 with failover routing policy.
B.Use Amazon CloudFront with multiple origins in the same region.
C.Enable S3 cross-Region replication for static assets.
D.Configure Amazon RDS for Multi-AZ and enable cross-Region read replicas.
E.Use Auto Scaling groups in a single region with multiple Availability Zones.
AnswersA, D

This configuration provides global DNS-level failover by associating Route 53 health checks with endpoints in each region. If the primary region's health check fails, Route 53 automatically rewrites DNS responses to direct traffic to the standby region, with failover typically occurring within a few minutes depending on TTL and health-check intervals. However, this requires the application to be designed for multi-region operation, with data replication between regions and compute capacity pre-provisioned in the secondary region to actually serve traffic. This is the core pattern for regional disaster recovery and meets the requirement for a highly available architecture across regions.

Why this answer

Deploying to multiple regions with Route 53 failover provides cross-region disaster recovery. Option D is correct because using Amazon RDS Multi-AZ with cross-Region read replicas or Aurora Global Database ensures database resilience across regions. Option B is wrong because CloudFront alone does not provide compute failover.

Option C is wrong because S3 cross-Region replication is for data, not compute. Option E is wrong because single-region Auto Scaling does not protect against region failure.

31
Multi-Selecteasy

A company is deploying a web application on Amazon ECS with Fargate. The application consists of a frontend service and a backend service. The DevOps team needs to ensure that the frontend service can communicate with the backend service securely without exposing the backend to the internet. Which THREE steps should the team take? (Choose THREE.)

Select 3 answers
A.Deploy the backend service in a private subnet with no internet access.
B.Use AWS Cloud Map service discovery for the backend service.
C.Configure a security group for the backend service that allows inbound traffic only from the frontend service's security group.
D.Deploy the backend service in a public subnet with an internet-facing Application Load Balancer.
E.Use an internet-facing Network Load Balancer for the backend service.
AnswersA, B, C

Placing the backend ECS service in a private subnet with no route to an internet gateway ensures it has no public IP and cannot be reached from the internet. This is correct for internal-only workloads because the backend does not need outbound internet access, and any required image pulls can be handled via VPC endpoints or pre-pulled images. This isolation reduces the attack surface and aligns with the principle of least privilege.

Why this answer

Deploying the backend service in a private subnet with no internet access ensures that the backend is not reachable from the internet, which is a fundamental security requirement. In Amazon ECS with Fargate, tasks in a private subnet use an elastic network interface (ENI) with no public IP address, and outbound traffic can be routed through a NAT gateway if needed, but inbound traffic from the internet is blocked. This isolates the backend from direct external exposure while still allowing communication from the frontend service within the same VPC.

Exam trap

The trap here is that candidates might think a load balancer is required for service-to-service communication in ECS, but AWS Cloud Map service discovery combined with security group rules can achieve secure, direct communication without exposing the backend to the internet.

32
MCQhard

A company runs a critical application on EC2 instances in an Auto Scaling group behind an ALB. They want to ensure that if an instance fails, the application remains available with minimal disruption. Which combination of services provides the best resilience?

A.Auto Scaling group with minimum 2 in a single AZ.
B.EC2 instance recovery with CloudWatch alarms.
C.Auto Scaling group with desired capacity of 2 and a lifecycle hook.
D.Auto Scaling group with ELB health checks and multiple AZs.
AnswerD

An Auto Scaling group spanning multiple Availability Zones and using Elastic Load Balancing health checks meets the resilience requirement: the ELB sends HTTP/HTTPS health checks to each instance, and when an instance is marked unhealthy, the ASG terminates it and launches a new one to maintain desired capacity. Multi-AZ distribution ensures that even if one AZ fails, the remaining instances continue serving traffic, and the ASG can launch replacements in other AZs. This combination provides both automated fault replacement and cross-AZ high availability.

Why this answer

Deploying an Auto Scaling group across multiple Availability Zones (AZs) with Elastic Load Balancer (ALB) health checks ensures that if an EC2 instance fails in one AZ, the ALB automatically routes traffic to healthy instances in other AZs, and Auto Scaling replaces the failed instance. This combination provides both fault isolation and automated recovery, minimizing disruption to the application.

Exam trap

The trap here is that candidates often think a single AZ with multiple instances (Option A) or instance recovery (Option B) provides sufficient resilience, but they overlook the need for AZ-level fault isolation and integrated health-check-driven replacement that only multi-AZ Auto Scaling with ELB health checks provides.

How to eliminate wrong answers

Option A is wrong because a single AZ is a single point of failure; if that AZ experiences an outage, all instances become unavailable regardless of the minimum count. Option B is wrong because EC2 instance recovery with CloudWatch alarms only recovers the same instance (e.g., after a hardware failure) but does not handle AZ-level failures or provide load balancing; it also does not automatically replace instances that fail health checks. Option C is wrong because a lifecycle hook is used for custom actions during instance launch or termination (e.g., draining connections), not for resilience; it does not distribute instances across AZs or provide health-check-based replacement.

33
MCQmedium

A DevOps engineer runs the above command and sees that instance i-0abcd1234efgh5678 is unhealthy with reason 'Target.Timeout'. The instance is running and the application on port 80 responds to curl from the instance itself. What is the MOST likely cause?

A.The ALB health check interval is set too high.
B.The web server process is not running on the instance.
C.The health check path returns a 404 status code.
D.The security group for the instance does not allow inbound traffic from the ALB on port 80.
AnswerD

A timeout typically indicates a network connectivity issue between ALB and instance.

Why this answer

The 'Target.Timeout' health check reason indicates that the load balancer's health check request timed out, meaning it could not establish a connection or receive a response within the timeout period. Since the instance responds to curl locally, the application is running, but the ALB cannot reach it. The most likely cause is that the instance's security group does not allow inbound traffic from the ALB on port 80, blocking the health check requests.

Exam trap

DOP-C02 often tests troubleshooting of load balancer health checks, and candidates might overlook security group rules, assuming the application is at fault because local curl works.

How to eliminate wrong answers

Option A is wrong because a high health check interval would delay checks but not cause timeouts; the interval is the time between checks, not the timeout duration. Option B is wrong because the application responds to curl, so the web server process is running. Option C is wrong because a 404 status code would result in a 'Target.ResponseCodeMismatch' reason, not a timeout.

34
MCQmedium

An IAM policy is attached to an S3 bucket to allow access from a specific VPC CIDR range. However, users from the VPC are receiving 'Access Denied' errors when trying to access objects in the bucket. What is the MOST likely reason?

A.The users are assuming an IAM role that does not have permission to access S3
B.The condition key 'aws:SourceIp' evaluates the public IP address, but the VPC uses private IP addresses
C.The policy should use 'aws:sourceVpce' instead of 'aws:SourceIp' to restrict access to a VPC endpoint
D.The bucket policy requires HTTPS and the requests are using HTTP
AnswerB

The aws:SourceIp condition key in a bucket policy compares the source IP address of the requester as seen by S3, which is the public IP address after any network address translation (NAT) or internet gateway processing. In a VPC, instances typically use private RFC 1918 IP addresses, which are not exposed to S3 when traffic traverses a NAT gateway or the internet. Thus, a policy specifying private IP ranges will never match, causing access to be denied for legitimate VPC clients.

Why this answer

The 'aws:SourceIp' condition key evaluates the source IP address of the request as seen by AWS, which is always a public IP address. When requests originate from within a VPC (e.g., from an EC2 instance), the source IP seen by S3 is the public IP of the instance or its NAT device, not the VPC's private CIDR range. Since the policy specifies a private VPC CIDR range using 'aws:SourceIp', it never matches the public source IP, resulting in 'Access Denied' errors.

Exam trap

The trap here is that candidates often confuse 'aws:SourceIp' with being able to match private IP ranges, not realizing that AWS services always see the public IP address of the request, even if the request originates from within a VPC.

How to eliminate wrong answers

Option A is wrong because the question states that the IAM policy is attached to the S3 bucket, and the error occurs despite that policy; if the users were assuming an IAM role without S3 permissions, the error would be consistent regardless of the bucket policy, but the issue here is specifically tied to the VPC CIDR condition. Option C is wrong because 'aws:sourceVpce' is used to restrict access to a specific VPC endpoint (interface or gateway), not to a VPC CIDR range; using it would require the request to originate from a VPC endpoint, which is not the scenario described. Option D is wrong because the error message 'Access Denied' is distinct from a '403 Forbidden' or '400 Bad Request' that would occur if HTTPS were required and HTTP was used; S3 bucket policies with 'aws:SecureTransport' condition would explicitly deny HTTP, but the question does not mention HTTPS enforcement.

35
MCQhard

A company is designing a disaster recovery (DR) strategy for a stateless web application deployed on Amazon ECS with Fargate. The application is fronted by an Application Load Balancer (ALB) and uses Amazon ElastiCache for Redis for session state. The primary region is us-east-1. The DR plan requires a Recovery Point Objective (RPO) of 15 minutes and a Recovery Time Objective (RTO) of 30 minutes. Which solution meets these requirements with the LEAST operational overhead?

A.Deploy an ALB with a warm standby ECS service in us-west-2. Use Route 53 health checks to route traffic to the secondary region if primary fails. Use ElastiCache Global Datastore for Redis to replicate data across regions.
B.Deploy an Active-Active configuration across two AWS regions using Route 53 latency routing. Use ElastiCache for Redis Global Datastore with multi-region writes.
C.Deploy a Pilot Light environment in us-west-2 with a scaled-down ECS service and Redis cluster. Use Route 53 DNS failover. On disaster, scale up the ECS service and promote the Redis cluster.
D.Use Amazon ECS with Fargate in us-east-1 only, and schedule daily snapshots of ElastiCache for Redis. In case of disaster, restore the snapshot in a new region and update DNS.
AnswerA

A warm standby architecture places a fully functional but possibly smaller ECS service behind an ALB in us-west-2, and Route 53 health checks automatically shift user traffic when us-east-1's ALB or service is unhealthy. ElastiCache Global Datastore for Redis maintains a cross-region replica with sub-second replication lag, so the secondary can serve reads and writes after promote without a point-in-time restore. This combination keeps both RPO and RTO under 30 minutes because failover is automated and data is already in the secondary region.

Why this answer

It uses ElastiCache Global Datastore for Redis, which provides cross-region replication with an RPO of seconds (well within 15 minutes) and automatic failover, minimizing operational overhead. The warm standby ECS service in us-west-2 with Route 53 health checks allows traffic to be redirected within the 30-minute RTO without manual intervention, as the ALB and ECS service are pre-provisioned.

Exam trap

The trap here is that candidates may confuse Pilot Light (Option C) as lower overhead, but it requires manual scaling and promotion steps, whereas a warm standby with Global Datastore automates failover, making it the least operational overhead for the given RPO/RTO.

How to eliminate wrong answers

Option B is wrong because an Active-Active configuration with multi-region writes for ElastiCache Global Datastore is not supported; Global Datastore only supports active-passive (one primary, one replica) to avoid write conflicts. Option C is wrong because a Pilot Light approach requires manual scaling of the ECS service and promoting the Redis cluster on disaster, which adds operational overhead and risks exceeding the 30-minute RTO due to provisioning delays. Option D is wrong because daily snapshots of ElastiCache cannot achieve a 15-minute RPO (snapshots are at most daily), and restoring a snapshot in a new region plus updating DNS would likely exceed the 30-minute RTO due to manual steps and data transfer time.

36
MCQeasy

A company uses Amazon CloudFront to distribute content from an S3 bucket origin. Some users report intermittent access errors. The DevOps team suspects the origin is overwhelmed. What is the MOST effective way to improve resilience?

A.Set up an origin failover with two S3 buckets behind an Application Load Balancer (ALB).
B.Reduce the CloudFront cache TTL to serve fresher content.
C.Increase the CloudFront cache TTL to reduce requests to the origin.
D.Configure CloudFront to perform health checks on the origin.
AnswerD

An origin group with health checks lets CloudFront monitor the primary origin using HTTP or HTTPS requests; when the origin returns errors or times out, CloudFront automatically fails over to a designated secondary origin. This gives near-instant recovery without manual intervention, directly addressing the overwhelmed origin scenario described in the question. Health checks are essential for the failover mechanism to know when to stop sending traffic to the failing origin.

Why this answer

Configuring CloudFront to perform health checks on the origin allows it to detect origin issues and trigger failover to a secondary origin, improving resilience. CloudFront origin groups support failover based on error rates, effectively acting as health checks. Option A is incorrect because an ALB is not required for origin failover; CloudFront can directly use multiple S3 buckets as origins.

Option B (reducing TTL) increases requests to the origin, worsening the problem. Option C (increasing TTL) reduces load but does not address origin failures or provide failover.

Exam trap

Candidates may assume an ALB is necessary for origin failover, but CloudFront can directly failover between S3 buckets using origin groups without additional load balancers.

37
MCQhard

A company's application runs on Amazon EC2 instances in an Auto Scaling group. The application writes logs to local instance storage. The operations team needs to ensure logs are not lost during instance termination or scaling events. What should be done?

A.Increase the size of the instance store volumes.
B.Use an Amazon EFS file system and mount it to each instance for log storage.
C.Configure the Auto Scaling group to terminate instances after logs are copied to S3.
D.Install the CloudWatch Logs agent on each instance and stream logs to CloudWatch Logs.
AnswerD

Installing the CloudWatch Logs agent (or the newer unified CloudWatch agent) on each EC2 instance enables real-time delivery of log data to the CloudWatch Logs service, which stores it durably and independently of the instance lifecycle. Because logs are streamed continuously as they are generated, they survive instance termination, crashes, or scale-in events without requiring manual offload steps. This is the AWS-recommended managed solution for centralized log collection and retention.

Why this answer

The CloudWatch Logs agent streams log data in near real-time to Amazon CloudWatch Logs, ensuring logs are persisted independently of the EC2 instance lifecycle. This decouples log storage from ephemeral instance store volumes, so logs are not lost during Auto Scaling termination or scaling events. The agent handles log rotation, compression, and encryption in transit, providing a durable and centralized log solution.

Exam trap

The trap here is that candidates often assume instance store volumes are persistent or that increasing their size provides durability, when in fact instance store volumes are ephemeral and tied to the instance lifecycle, making them unsuitable for critical log data that must survive termination.

How to eliminate wrong answers

Option A is wrong because increasing instance store volume size does not address the fundamental issue that instance store volumes are ephemeral and data is lost when the instance is stopped, terminated, or replaced during scaling events. Option B is wrong because while Amazon EFS provides persistent shared storage, it introduces network latency and additional cost, and the application would need to be reconfigured to write logs to the EFS mount point; more critically, the question does not indicate that the application can handle network file system writes, and EFS is not the simplest or most AWS-native solution for log streaming. Option C is wrong because configuring the Auto Scaling group to delay termination until logs are copied to S3 is unreliable and complex; there is no native Auto Scaling lifecycle hook that guarantees logs are fully copied before termination, and this approach can cause scaling delays, race conditions, and increased operational overhead.

38
MCQmedium

A company runs a stateless web application on a fleet of EC2 instances in an Auto Scaling group. The application stores session state in a shared ElastiCache Redis cluster. During traffic spikes, the application becomes slow. Monitoring shows that the Redis cluster has high CPU utilization. Which solution is MOST cost-effective and scalable?

A.Upgrade the Redis instance to a larger node type to handle more operations
B.Enable cluster mode on the ElastiCache Redis cluster and add more shards
C.Add read replicas to offload read traffic from the primary node
D.Migrate session state to DynamoDB with DAX for caching
AnswerB

Enabling cluster mode on ElastiCache for Redis partitions the keyspace across multiple shards, with each shard having its own primary node to handle writes independently. Adding shards increases aggregate write throughput and lets you scale beyond the limits of a single node, while also supporting larger datasets by distributing memory. This is the correct answer because it directly addresses a write-heavy workload by horizontally scaling the data plane and maintains Redis' low-latency session state access.

Why this answer

Read replicas offload GET/SMEMBERS-type traffic from the primary, but every write must still be processed by the primary node, so if the workload has a significant write component, the primary's CPU and write path remain unchanged. Note: ElastiCache Redis read replicas do NOT require cluster mode — a cluster-mode-disabled replication group already supports up to 5 read replicas on a single shard. Cluster mode is only needed to add additional shards for horizontal write scaling, which is why option B (enabling cluster mode) is the more complete, scalable fix when the bottleneck may include writes.

Exam trap

The trap here is that candidates often confuse read replicas with horizontal scaling for write-heavy workloads, not realizing that replicas only help with read scaling and cannot reduce CPU from write operations, while cluster mode directly addresses both read and write scaling by splitting the data set.

How to eliminate wrong answers

Option A is wrong because upgrading to a larger node type (vertical scaling) is less cost-effective and has an upper limit; it does not provide the linear scalability of horizontal sharding and can lead to over-provisioning during low traffic. Option C is wrong because adding read replicas offloads only read traffic, but session state in Redis involves both reads and writes, and the high CPU is likely from write-heavy operations (e.g., SET/GET) that replicas cannot offload; replicas also introduce eventual consistency issues for session data. Option D is wrong because migrating to DynamoDB with DAX introduces unnecessary complexity and cost for session state that is already well-served by Redis; DAX is a separate caching layer that adds latency and cost, and DynamoDB's throughput pricing can be less predictable than ElastiCache for bursty traffic.

39
MCQmedium

A company hosts a static website on Amazon S3 with a CloudFront distribution. The website is critical for business operations and must be available even if the primary AWS Region fails. Currently, the S3 bucket is in us-east-1, and CloudFront uses that bucket as the origin. The company has a secondary bucket in us-west-2 with a replica of the data. The company wants to use CloudFront to automatically fail over to the secondary bucket if the primary becomes unavailable. The DevOps engineer needs to implement a solution that requires minimal operational overhead. What should the engineer do?

A.Use an Application Load Balancer in front of both S3 buckets and point CloudFront to the ALB.
B.Create a second CloudFront distribution pointing to the secondary bucket and use Route 53 failover routing between the two distributions.
C.Modify the application to switch the CloudFront origin URL using Lambda@Edge when health checks fail.
D.Configure CloudFront Origin Failover by adding both buckets as origins, with the primary in us-east-1 and secondary in us-west-2.
AnswerD

CloudFront Origin Failover is the native, minimal-configuration solution: you create an origin group containing two S3 buckets as the primary and secondary origins, and attach that group to your cache behavior. When the primary origin returns a configurable HTTP error code (commonly 5xx) or a connection timeout, CloudFront automatically retries the request against the secondary bucket in us-west-2, all within the same edge location and without involving DNS or custom code. This requires no additional compute, no Route 53 policies, and no application modifications—only enabling Cross-Region Replication between the two buckets so content stays consistent. It is the only option that leverages a built-in CloudFront feature designed specifically for this use case, meeting both the fault-tolerance and low-operational-overhead requirements.

Why this answer

CloudFront Origin Failover is a native feature that allows you to designate a primary and secondary origin within a single distribution; CloudFront automatically routes requests to the secondary origin when the primary returns specific error codes (e.g., 500, 502, 503, 504) or fails health checks. This requires no additional infrastructure, no DNS changes, and no custom code, making it the lowest-operational-overhead solution. It directly meets the requirement of automatic failover to the us-west-2 bucket.

Exam trap

The trap here is that candidates often reach for Route 53 failover or Lambda@Edge because they are familiar multi-region patterns, missing that CloudFront Origin Failover is a built-in, zero-overhead feature designed exactly for this scenario.

How to eliminate wrong answers

Option A is wrong because Application Load Balancers cannot target S3 buckets as origins — ALBs route to EC2 instances, Lambda functions, or IP addresses, not S3 static website endpoints. Option B is wrong because creating a second CloudFront distribution and using Route 53 failover routing adds significant operational overhead (managing two distributions, DNS health checks, and failover policies) and is unnecessary when Origin Failover exists. Option C is wrong because Lambda@Edge cannot dynamically switch origin URLs based on health checks in a supported way — origin selection is determined at distribution configuration time, and Lambda@Edge runs at edge locations for request/response manipulation, not origin failover.

40
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer. The application stores session state in an in-memory cache on each instance. During deployment of a new version, users experience session timeouts and errors. Which design change will MOST effectively improve resilience and avoid session loss during deployments?

A.Enable sticky sessions (session affinity) on the ALB.
B.Migrate session state to ElastiCache for Redis.
C.Increase the ALB idle timeout to 600 seconds.
D.Increase the EC2 instance size to handle higher memory.
AnswerB

ElastiCache for Redis externalises session state from individual instance memory, so any instance behind the Application Load Balancer can serve any request. This removes the constraint that deployments destroy in-memory sessions, preventing timeouts and errors as instances are replaced.

Why this answer

Migrating session state from in-memory EC2 instance storage to ElastiCache for Redis decouples session data from individual instances. This ensures that when a new deployment replaces instances, sessions persist independently, preventing timeouts and errors. ElastiCache provides a centralized, highly available session store that survives instance termination and scaling events.

Exam trap

The trap here is that candidates often confuse sticky sessions (which only route traffic consistently) with session persistence (which requires external storage), leading them to choose option A despite it not preserving session data across instance replacements.

How to eliminate wrong answers

Option A is wrong because enabling sticky sessions (session affinity) on the ALB would lock users to a specific instance, but during deployment that instance is terminated and replaced, causing session loss regardless of stickiness. Option C is wrong because increasing the ALB idle timeout to 600 seconds only extends how long the ALB keeps a connection open without data transfer; it does not preserve session state stored in the instance's memory when the instance is replaced. Option D is wrong because increasing the EC2 instance size to handle higher memory does not solve the fundamental problem of session state being ephemeral and lost during instance replacement in a deployment.

41
MCQmedium

A company is designing a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application must tolerate a Regional failure. Which design provides the most resilience?

A.Use Lambda@Edge to run functions at AWS edge locations
B.Use DynamoDB auto-scaling and run Lambda in a single Region
C.Use DynamoDB global tables with Lambda functions deployed in multiple Regions and Route 53 multi-Region routing
D.Use DynamoDB Accelerator (DAX) to cache data across Regions
AnswerC

DynamoDB global tables asynchronously replicate writes to multiple Regions, providing active-active tables with automatic conflict resolution. Deploying Lambda in each Region makes stateless compute locally available, while Route 53 multi-Region routing (for example, failover or latency routing) sends traffic to a healthy Region. Together, these components form a resilient multi-Region serverless architecture with no single Region dependency.

Why this answer

DynamoDB global tables provide multi-Region, fully replicated tables with automatic conflict resolution, ensuring data availability during a Regional outage. Deploying Lambda functions in multiple Regions with Route 53 multi-Region routing (using health checks and latency-based or weighted routing) allows traffic to fail over to a healthy Region, making the entire serverless stack resilient to a Regional failure.

Exam trap

The trap here is that candidates often confuse caching (DAX) or edge computing (Lambda@Edge) with true multi-Region replication and failover, assuming they provide Regional resilience when they do not.

How to eliminate wrong answers

Option A is wrong because Lambda@Edge runs at CloudFront edge locations, not in multiple AWS Regions, and is designed for lightweight request/response modification, not for hosting a full serverless application backend; it does not provide Regional failover for DynamoDB or API Gateway. Option B is wrong because DynamoDB auto-scaling only adjusts throughput within a single Region and does not replicate data across Regions, so a Regional failure would still cause complete data unavailability; running Lambda in a single Region creates a single point of failure for compute. Option D is wrong because DAX is a caching layer that operates within a single Region and does not replicate data across Regions; it cannot provide data durability or availability during a Regional outage, and it is not designed for cross-Region failover.

42
MCQhard

A company uses AWS Lambda functions to process events from Amazon SQS. The Lambda function sometimes fails due to timeouts. The team wants to preserve the event for reprocessing. How should they configure the integration?

A.Set up a DLQ on the SQS queue that receives the events
B.Use Lambda reserved concurrency
C.Enable Lambda function DLQ with SNS topic
D.Increase Lambda timeout to maximum
AnswerA

Configuring a dead-letter queue on the SQS queue ensures that any messages that are not successfully processed by the Lambda function are redirected to the DLQ for later analysis and reprocessing. This directly addresses the need to preserve failed events.

Why this answer

By configuring a dead-letter queue (DLQ) on the SQS queue, failed messages are preserved for later reprocessing.

43
MCQhard

A company runs a stateful web application on EC2 instances in an Auto Scaling group. The application uses an Application Load Balancer (ALB) and an Amazon ElastiCache Redis cluster. Users report that after a scaling event, they are logged out and lose session data. What is the most likely cause?

A.The ALB health check interval is too short, causing healthy instances to be marked unhealthy
B.The ElastiCache cluster is configured with in-transit encryption, causing session tokens to be invalidated
C.The Auto Scaling group is using a termination policy that terminates the oldest instance first, which holds active sessions
D.The ElastiCache cluster is not configured for Multi-AZ and a node failure caused all sessions to be lost
AnswerD

An ElastiCache cluster without Multi-AZ deployment means there is only a single primary node with no replica in another Availability Zone. If that node fails, any in-memory session data is lost because there is no replication or automatic failover to a secondary node. Even if ElastiCache replaces the node, the replacement starts empty unless persistence/backups were configured. This exactly matches the symptom of all users being logged out simultaneously.

Why this answer

The scenario describes a stateful web application that relies on ElastiCache Redis for session storage. If the ElastiCache cluster is not configured for Multi-AZ, a node failure can cause all cached session data to be lost, logging users out. This is the most likely cause of session loss after a scaling event, as scaling events do not directly affect ElastiCache data persistence.

Exam trap

The trap here is that candidates may focus on the Auto Scaling group's termination policy or ALB health checks, overlooking that the session data is stored externally in ElastiCache and that its lack of high availability is the root cause of session loss.

How to eliminate wrong answers

Option A is wrong because a short health check interval would cause instances to be marked unhealthy and replaced, but it would not directly cause session data loss; sessions are stored in ElastiCache, not on the instances. Option B is wrong because in-transit encryption on ElastiCache protects data during transmission and does not invalidate session tokens; it is unrelated to session persistence. Option C is wrong because terminating the oldest instance first is a common termination policy that does not cause session loss if sessions are stored externally in ElastiCache; the issue is with the session store itself, not the instance termination order.

44
MCQhard

A company runs a stateless web application on Amazon ECS with Fargate launch type. The application experiences intermittent traffic spikes. The company wants to ensure that the application can scale automatically and remain resilient to underlying infrastructure failures. Which combination of actions should the DevOps engineer take?

A.Configure a scheduled scaling policy for the Amazon ECS service to add tasks during known peak hours.
B.Launch tasks in a single Availability Zone and use an Application Auto Scaling target tracking policy based on CPU utilization.
C.Configure a step scaling policy for the Amazon ECS service and increase the task memory size.
D.Configure an Application Auto Scaling target tracking policy based on memory utilization and enable Amazon ECS service auto-recovery.
AnswerD

An Application Auto Scaling target tracking policy with memory utilization as the metric dynamically scales out tasks when memory pressure increases and scales in when it subsides, providing immediate response to unpredictable workload spikes. Amazon ECS service auto-recovery, enabled through service health checks and automatic task replacement, ensures that any tasks that fail or become unhealthy are automatically restarted, maintaining desired availability. Together, these mechanisms deliver both elasticity—scaling with real-time demand—and resilience—self-healing from task or infrastructure failures—for the stateless web application.

Why this answer

It combines Application Auto Scaling target tracking based on memory utilization, which is a relevant metric for a stateless web application to handle traffic spikes, with Amazon ECS service auto-recovery, which automatically replaces unhealthy tasks to ensure resilience against underlying infrastructure failures. This approach provides both automatic scaling and fault tolerance without manual intervention.

Exam trap

The trap here is that candidates often assume CPU utilization is the only valid scaling metric for web applications, but memory utilization can be more appropriate for stateless workloads, and they may overlook the critical need for service auto-recovery to handle infrastructure failures in Fargate.

How to eliminate wrong answers

Option A is wrong because scheduled scaling is reactive to known peak hours but cannot handle intermittent, unpredictable traffic spikes, and it does not address resilience to infrastructure failures. Option B is wrong because launching tasks in a single Availability Zone creates a single point of failure, violating resilience best practices, and while target tracking based on CPU utilization can scale, it does not provide auto-recovery for failed tasks. Option C is wrong because step scaling policies can be effective, but increasing task memory size does not directly improve scaling or resilience; it may reduce the need for scaling but does not automate recovery from failures.

45
Multi-Selecthard

A company runs a containerized application on Amazon ECS with Fargate. The application needs to be resilient to Availability Zone failures. Which THREE actions should the company take? (Choose THREE.)

Select 3 answers
A.Configure the ECS service to spread tasks across multiple Availability Zones.
B.Disable managed service scaling to avoid resource contention.
C.Use a multi-AZ Amazon RDS or DynamoDB for persistent data.
D.Deploy an Application Load Balancer (ALB) with targets in multiple Availability Zones.
E.Use a single service discovery namespace for all tasks.
AnswersA, C, D

Configure the ECS service with a spread placement strategy across Availability Zones (AZs) using either the ability to spread evenly or by AZ as a custom attribute. This guarantees that tasks are distributed redundantly, so if one AZ becomes unavailable, the remaining tasks in other AZs continue to serve traffic and provide capacity for service scaling to replace unhealthy tasks. Without this spread, Amazon ECS could place all tasks in a single AZ, creating a single point of failure that contradicts the goal of high availability.

Why this answer

Spreading tasks across multiple Availability Zones ensures that an AZ failure does not impact all tasks, increasing resilience. Option C is correct because using a multi-AZ Amazon RDS or DynamoDB provides persistent data storage that survives AZ failures. Option D is correct because an Application Load Balancer with targets in multiple AZs distributes traffic and can route requests to healthy targets in other AZs if one fails.

Option B is wrong because disabling managed service scaling reduces the application's ability to handle load changes and may impact availability. Option E is wrong because a single service discovery namespace does not provide AZ resilience; it only provides service discovery without redundancy across AZs.

46
MCQhard

A company is designing a multi-Region disaster recovery strategy for a stateless web application. The application runs on EC2 instances in an Auto Scaling group behind an ALB in us-east-1. The recovery point objective (RPO) is 15 minutes and recovery time objective (RTO) is 30 minutes. The application data is stored in Amazon RDS for PostgreSQL. Which combination of actions should the company take to meet the RPO and RTO?

A.Use RDS cross-Region replication to a standby DB instance in another Region. Maintain a warm standby environment (Auto Scaling group, ALB) in the disaster Region. Configure Route 53 health checks to fail over automatically.
B.Use RDS cross-Region snapshot copy every 15 minutes. In the disaster Region, manually launch a new environment and restore the latest snapshot.
C.Use RDS Multi-AZ in us-east-1. In the disaster Region, keep a standby Auto Scaling group and ALB. On failure, promote the Multi-AZ standby to primary and update DNS.
D.Use RDS read replicas in another Region. On failure, promote the read replica to a standalone instance and update the application.
AnswerA

Cross-Region replication maintains a continuously updated standby database in the DR Region, so the RPO is reduced to the replication lag, which is typically seconds for RDS MySQL/PostgreSQL. The pre-provisioned Auto Scaling group and Application Load Balancer constitute a warm standby that can receive traffic immediately after Route 53 health checks detect a regional impairment and automatically update DNS records to point to the DR endpoint. This combination of async data replication and warm infrastructure is what enables a low RPO and an RTO well under the 30-minute requirement.

Why this answer

It meets both the 15-minute RPO and 30-minute RTO. RDS cross-Region replication provides continuous asynchronous replication with minimal lag, typically well under 15 minutes, ensuring data is nearly up-to-date. The warm standby environment (pre-provisioned Auto Scaling group and ALB) in the disaster Region allows automatic failover via Route 53 health checks, enabling recovery within the 30-minute RTO without manual intervention.

Exam trap

The trap here is confusing Multi-AZ (single-Region HA) with cross-Region DR, leading candidates to choose Option C, which fails to protect against a Regional outage.

How to eliminate wrong answers

Option B is wrong because manual snapshot copies every 15 minutes cannot guarantee a 15-minute RPO due to snapshot creation and transfer delays, and manually launching a new environment and restoring the latest snapshot far exceeds the 30-minute RTO. Option C is wrong because RDS Multi-AZ in us-east-1 provides high availability within a single Region only; it does not replicate data to another Region, so a Regional failure would result in complete data loss and no DR capability. Option D is wrong because promoting a cross-Region read replica to a standalone instance can take several minutes and requires manual DNS updates, which together exceed the 30-minute RTO; also, read replicas may have replication lag that could violate the 15-minute RPO.

47
Multi-Selecthard

A company uses Amazon ECS with Fargate for containerized applications. They need to ensure that if a task fails, it is automatically restarted and the application remains available. Which THREE actions should they take? (Choose THREE.)

Select 3 answers
A.Configure the ECS service to automatically restart failed tasks.
B.Place tasks across multiple Availability Zones.
C.Use an Application Load Balancer with health checks.
D.Set up a CloudWatch alarm to trigger AWS Lambda to restart tasks.
E.Configure an EC2 Auto Scaling group for the ECS cluster.
AnswersA, B, C

The ECS service scheduler continuously monitors running tasks and compares the current count against the desired count. When a task fails due to container exit, OOM, or health check failure, the service automatically starts a replacement task using the same task definition. This is a built-in control loop that requires no custom scripts or external triggers, and it works identically for Fargate and EC2 launch types.

Why this answer

Amazon ECS services can be configured with a desired task count and a task placement strategy that automatically replaces any failed or stopped tasks. When a task exits unexpectedly, the ECS service scheduler detects the discrepancy between the desired and running count and launches a new task to maintain availability. This is the native mechanism for self-healing in ECS, without requiring external triggers.

Exam trap

The trap here is that candidates often over-engineer the solution by adding CloudWatch and Lambda (Option D) or confuse EC2 Auto Scaling (Option E) with task-level recovery, when the native ECS service configuration already handles automatic restarts.

48
Multi-Selectmedium

A company runs a stateful web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application stores session data in local instance memory. To improve resiliency, the company wants to make the application stateless and distribute the load across multiple Availability Zones. Which THREE actions should the company take? (Choose three.)

Select 3 answers
A.Enable sticky sessions (session affinity) on the ALB.
B.Replace the ALB with a Network Load Balancer (NLB) to improve performance.
C.Configure the ALB to distribute traffic across EC2 instances in multiple Availability Zones.
D.Implement an Amazon ElastiCache cluster to store session data externally.
E.Use Amazon DynamoDB to store session state.
AnswersC, D, E

Distributing traffic across multiple Availability Zones directly satisfies the resiliency and load-spreading requirement. The ALB already operates at the regional level, so enabling cross-AZ distribution removes the single-AZ failure domain, ensuring instances in surviving zones continue serving requests when one zone becomes unavailable.

Why this answer

Option C is correct because an ALB is a regional load balancer that can register targets in multiple Availability Zones, so configuring it to distribute traffic across EC2 instances in several AZs removes the single-AZ failure point and improves resiliency. Option D is correct because moving session data out of local instance memory into an Amazon ElastiCache cluster (e.g., Redis or Memcached) externalizes session state, allowing any instance in any AZ to serve any request and making the application stateless. Option E is correct because Amazon DynamoDB is a fully managed, multi-AZ, highly durable key-value store that can also hold session state externally, achieving the same stateless design without relying on instance memory.

Option A is not appropriate because sticky sessions keep a user bound to a single instance and actually perpetuate statefulness rather than eliminate it. Option B is not appropriate because replacing the ALB with an NLB does not address session state at all and an NLB lacks the HTTP/HTTPS layer-7 routing features typically used by web applications.

Exam trap

DOP-C02 often tests the misconception that enabling sticky sessions improves resiliency — in fact, session affinity is the anti-pattern that statelessness is designed to eliminate.

49
MCQhard

A company has a stateless web application on EC2 instances behind an ALB. They want to ensure that if an entire Availability Zone fails, the application remains available with minimal impact. Which architecture best meets this requirement?

A.Use a global secondary index in DynamoDB to replicate data across regions
B.Deploy EC2 instances in two AZs but use a single-AZ ALB
C.Deploy the application in two AZs with an Auto Scaling group and an ALB that is enabled for multiple AZs
D.Deploy the application in a single AZ with an Auto Scaling group that launches instances in the same AZ
AnswerC

Spreading instances across two Availability Zones with an Auto Scaling group lets it replace capacity in the surviving AZ, while the multi-AZ ALB routes only to healthy targets. This satisfies the AZ-failure constraint by removing the single-AZ dependency and enabling automatic recovery.

Why this answer

The correct architecture deploys the stateless application across two Availability Zones using an Auto Scaling group that spans both AZs, and an Application Load Balancer (ALB) configured with subnets in both AZs. If one AZ fails, the ALB automatically routes traffic to healthy instances in the surviving AZ, and the Auto Scaling group maintains capacity. This provides high availability with minimal impact, as required.

Exam trap

DOP-C02 often tests the misconception that a single-AZ ALB can provide high availability if EC2 instances are in multiple AZs, but the ALB itself is a single point of failure; candidates must remember that the ALB must also be enabled for multiple AZs.

How to eliminate wrong answers

Option A is wrong because a DynamoDB global secondary index (GSI) is a within-region indexing feature, not a cross-region replication mechanism; cross-region replication requires DynamoDB global tables, and the question is about AZ failure, not region failure. Option B is wrong because a single-AZ ALB creates a single point of failure: if that AZ fails, the ALB itself becomes unavailable, so deploying EC2 instances in two AZs does not help. Option D is wrong because deploying everything in a single AZ means an AZ failure takes down the entire application, regardless of Auto Scaling.

50
Multi-Selectmedium

A company is designing a multi-region disaster recovery strategy for a stateless web application. They want to minimize RTO and RPO. Which TWO of the following should they implement? (Choose TWO.)

Select 2 answers
A.Use cross-region replication for data stores.
B.Use a passive standby in a single Availability Zone.
C.Perform periodic backups and restore in the DR region.
D.Configure cross-region read replicas for the database.
E.Deploy an active-active workload using Route 53 weighted routing.
AnswersA, E

Cross-region replication for data stores, such as Amazon S3 CRR, DynamoDB global tables, or Aurora Global Database, synchronizes data between regions automatically, typically with sub-minute RPO. This ensures that when a regional failure occurs, the application can fail over to the DR region with minimal data loss, making it a strong foundation for a multi-region DR strategy. Because replication is continuous rather than point-in-time, it provides a much lower RPO than backup and restore, and it often pairs with per-region compute stacks to achieve low RTO.

Why this answer

Cross-region replication for data stores ensures that data is continuously synchronized to a secondary AWS region, minimizing Recovery Point Objective (RPO) to near-zero and reducing Recovery Time Objective (RTO) as the data is already available in the DR region. This approach avoids the need to restore from backups, which would increase RTO and potentially lose recent transactions.

Exam trap

The trap here is that candidates often confuse cross-region read replicas (which are read-only and not suitable for active-active writes) with true multi-region replication solutions, leading them to select Option D instead of Option A or E.

51
MCQeasy

A company wants to design a resilient architecture for a web application using AWS services. Which of the following is a best practice for improving resilience?

A.Deploy EC2 instances in multiple Availability Zones.
B.Use an Auto Scaling group in a single AZ.
C.Use a single AZ with RDS Multi-AZ.
D.Use one large EC2 instance to handle all traffic.
AnswerA

Placing EC2 instances in multiple Availability Zones (AZs) and fronting them with an Elastic Load Balancer and an Auto Scaling group ensures that if an entire AZ becomes unavailable, the load balancer can route traffic only to healthy instances in the remaining AZs. This pattern provides fault tolerance at the AZ granularity, which is the foundation of a resilient web architecture. It also allows the application to absorb a single-AZ failure without requiring any manual intervention.

Why this answer

Deploying EC2 instances across multiple Availability Zones (AZs) is a fundamental best practice for resilience because it eliminates a single point of failure at the data center level. If one AZ experiences an outage, traffic can be automatically routed to healthy instances in other AZs via an Elastic Load Balancer (ELB), ensuring application availability. This approach aligns with the AWS Well-Architected Framework's Reliability Pillar, which mandates distributing workloads across multiple AZs to achieve high availability.

Exam trap

The trap here is that candidates often confuse database-level high availability (RDS Multi-AZ) with full application resilience, mistakenly thinking that a single-AZ compute layer is acceptable as long as the database is redundant.

How to eliminate wrong answers

Option B is wrong because using an Auto Scaling group in a single AZ creates a single point of failure; if that AZ becomes unavailable, all instances are lost, and the application goes down. Option C is wrong because RDS Multi-AZ provides high availability for the database layer, but the compute layer (EC2) remains in a single AZ, meaning an AZ failure still takes down the web application. Option D is wrong because relying on one large EC2 instance violates the principle of horizontal scaling and introduces a single point of failure; if the instance fails or the AZ fails, the entire application becomes unavailable.

52
MCQmedium

A company runs a microservices application on Amazon ECS with Fargate launch type. The application experiences intermittent failures when calling an external API. The errors are transient and usually resolve within a few seconds. How should the company improve resilience?

A.Increase the timeout of the external API call to 60 seconds.
B.Implement retry logic with exponential backoff in the application code.
C.Increase the number of tasks in the ECS service to handle failures.
D.Use an Amazon SQS queue to decouple the API call from the application.
AnswerB

Transient external API failures resolve within seconds, so retrying with exponential backoff lets the application recover without overwhelming the dependency. This directly addresses the intermittent, short-lived errors described, unlike circuit breakers or timeouts that would abort calls rather than succeed on retry.

Why this answer

Transient errors from an external API are best handled with retry logic using exponential backoff and jitter. This allows the application to recover from temporary failures without overwhelming the downstream API, improving resilience without architectural changes.

Exam trap

DOP-C02 often tests the difference between scaling (more tasks) and resilience patterns (retries, backoff) — candidates pick 'increase tasks' because it sounds like high availability, but it does nothing for transient downstream errors.

How to eliminate wrong answers

Option A is wrong because increasing the timeout to 60 seconds does not address transient failures — it just makes the caller wait longer and can exhaust resources. Option C is wrong because adding more ECS tasks increases capacity but does not fix the underlying call failures; each task would still fail. Option D is wrong because SQS decouples asynchronous processing but the application needs a synchronous response from the external API; queuing does not solve transient call failures and adds complexity.

53
MCQmedium

A company uses AWS Lambda to process messages from an Amazon SQS queue. The Lambda function occasionally times out after 15 seconds. To improve resilience, the team wants to ensure messages are not lost and are retried. Which configuration is MOST appropriate?

A.Reduce the Lambda timeout to 5 seconds to fail fast and retry quickly.
B.Set the SQS queue visibility timeout to less than the Lambda timeout.
C.Increase the batch size and remove the DLQ to speed up processing.
D.Increase the Lambda timeout to 30 seconds and configure a dead-letter queue (DLQ) for the SQS queue.
AnswerD

Increasing the Lambda function timeout to 30 seconds gives the SQS-triggered processor enough wall-clock time to complete network calls, database writes, or third-party integrations that were failing under a shorter timeout, while still remaining within a reasonable operational bound. Configuring a dead-letter queue (DLQ) on the SQS queue ensures that messages that still repeatedly fail after retries are moved to a separate queue for later inspection and redrive, preventing poison messages from consuming the main queue indefinitely. Together, these changes address the immediate failure mode and give you a durable, observable mechanism for handling the few messages that cannot be processed successfully.

Why this answer

Increasing the Lambda timeout to 30 seconds accommodates the occasional processing delays that cause the current 15-second timeout, preventing premature failures. Configuring a dead-letter queue (DLQ) for the SQS queue ensures that messages that repeatedly fail after all retries are exhausted are preserved for analysis and manual reprocessing, rather than being lost. This combination directly addresses the requirement to not lose messages and to allow retries, as Lambda will automatically retry failed invocations up to the function's configured retry count (default 2) before sending the message to the DLQ.

Exam trap

The trap here is that candidates may think reducing the timeout or adjusting the visibility timeout alone improves resilience, but they overlook the critical need for a DLQ to prevent message loss and the necessity of matching the visibility timeout to the function's execution window to avoid duplicate processing.

How to eliminate wrong answers

Option A is wrong because reducing the Lambda timeout to 5 seconds would cause even more frequent timeouts, increasing failures without solving the underlying processing issue, and does not preserve messages for retry. Option B is wrong because setting the SQS visibility timeout to less than the Lambda timeout would cause messages to become visible again in the queue while the Lambda function is still processing them, leading to duplicate processing and potential data inconsistency. Option C is wrong because increasing the batch size would increase the processing load per invocation, likely worsening timeouts, and removing the DLQ would cause messages that exceed the maximum retries to be silently discarded, violating the requirement to not lose messages.

54
MCQhard

A company runs a critical batch processing workload on Amazon EMR that must complete within a 2-hour window each night. The workload is fault-tolerant but must be resilient to instance failures. Currently, the EMR cluster uses instance fleets with Spot Instances. Recently, Spot Instance interruptions caused the cluster to take over 3 hours to complete. Which change will MOST effectively ensure the workload completes within the 2-hour window despite Spot interruptions?

A.Increase the number of core nodes to 20 to improve parallelism.
B.Switch to using On-Demand instances for all nodes.
C.Use a mixed instances policy that includes multiple instance types across different Availability Zones.
D.Configure the cluster to terminate idle nodes after 5 minutes to reduce costs.
AnswerC

Using a mixed instances policy with multiple instance types across different Availability Zones is the recommended way to reduce Spot interruption risk in EMR. This approach makes the cluster's Spot capacity pool more diverse, so a capacity reclamation event in one pool is unlikely to affect all nodes simultaneously. EMR's instance fleets can automatically provision from the specified pools, improving both initial capacity acquisition and fault tolerance during interruptions.

Why this answer

A mixed instances policy across multiple Availability Zones increases the diversity of Spot capacity pools. When one instance type or zone experiences interruptions, the cluster can fall back to other pools, reducing the likelihood of prolonged delays. This approach directly addresses Spot interruption risk without sacrificing cost efficiency, as On-Demand instances would.

Exam trap

The trap here is that candidates may assume increasing parallelism (Option A) or using On-Demand instances (Option B) are the only ways to handle Spot interruptions, overlooking the cost-effective and resilient design of mixed instances across zones.

How to eliminate wrong answers

Option A is wrong because simply increasing core nodes to 20 does not mitigate Spot interruptions; it only adds parallelism, which may not help if all nodes are interrupted simultaneously. Option B is wrong because switching entirely to On-Demand instances eliminates Spot interruption risk but significantly increases cost, which is not the most effective solution given the fault-tolerant nature of the workload. Option D is wrong because terminating idle nodes after 5 minutes reduces cost but does not address the root cause of Spot interruptions causing delays; it may even worsen performance by removing nodes that could be reused.

55
MCQmedium

A company uses an Application Load Balancer (ALB) to distribute traffic to EC2 instances. The ALB is in us-east-1a and us-east-1b. They want to ensure that if one AZ fails, traffic is routed only to healthy instances in the other AZ. What configuration is necessary?

A.Enable sticky sessions (session affinity)
B.Configure health checks on the target group
C.Add more subnets in additional AZs
D.Enable cross-zone load balancing on the ALB
AnswerD

Cross-zone load balancing on the ALB enables each load balancer node to distribute incoming traffic evenly across all healthy targets in all enabled AZs, rather than limiting each node to its own AZ. With this enabled, if an AZ becomes unhealthy or unreachable, the remaining healthy ALB nodes can seamlessly route traffic to healthy targets in other AZs, ensuring continuous availability. This directly addresses the root cause of an ALB failing to route to instances in other AZs during an outage.

Why this answer

Cross-zone load balancing must be enabled on the ALB so that traffic can be distributed across instances in all AZs. By default, an ALB routes requests only to targets in the same Availability Zone as the requesting client. Enabling cross-zone load balancing allows the ALB to distribute traffic evenly across all registered targets in all enabled AZs, ensuring that if one AZ fails, traffic can be routed to healthy instances in other AZs.

Option B is incorrect because health checks are already enabled by default and do not affect cross-AZ routing. Option C is incorrect because adding more AZs does not change the default AZ-affinity behavior; cross-zone load balancing must be explicitly enabled to utilize multiple AZs for failover.

56
Drag & Dropmedium

Drag and drop the steps to troubleshoot a failed deployment in AWS CodeDeploy into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Troubleshooting starts with console, then agent logs, then AppSpec, then instance configuration, then redeploy.

57
MCQeasy

A company's application runs on EC2 instances in a single Availability Zone. The operations team wants to improve resilience without redesigning the application. Which action is the MOST effective?

A.Use a larger instance type to handle more traffic.
B.Enable EC2 Auto Recovery to automatically restart the instance if it fails.
C.Deploy EC2 instances across multiple Availability Zones using an Auto Scaling group.
D.Place the instance in a placement group to ensure low latency.
AnswerC

Deploying EC2 instances across multiple Availability Zones with an Auto Scaling group is the standard pattern for high availability. If one AZ becomes unavailable, the load balancer routes traffic to instances in the remaining healthy AZs, and the Auto Scaling group maintains instance count across AZs to replace any that are terminated. This eliminates the single-AZ point of failure and ensures application availability during an AZ outage.

Why this answer

Deploying EC2 instances across multiple Availability Zones (AZs) using an Auto Scaling group is the most effective action because it eliminates the single point of failure at the AZ level. If one AZ experiences an outage, the Auto Scaling group automatically launches replacement instances in the remaining healthy AZs, ensuring application availability without requiring any application-level changes. This directly addresses the goal of improving resilience by leveraging AWS's fault-isolated infrastructure.

Exam trap

The trap here is that candidates often confuse instance-level recovery (Auto Recovery) with infrastructure-level resilience (multi-AZ deployment), mistakenly thinking that restarting a failed instance in the same AZ provides sufficient protection against the most common cause of downtime—an AZ outage.

How to eliminate wrong answers

Option A is wrong because using a larger instance type only increases compute capacity, not resilience; a single AZ failure still takes down all instances regardless of size. Option B is wrong because EC2 Auto Recovery only recovers an instance within the same AZ if the underlying hardware fails, but it does not protect against an entire AZ outage, which is the primary risk. Option D is wrong because a placement group is designed to reduce network latency by ensuring instances are in close proximity, but it actually increases the risk of correlated failures and does not improve resilience against AZ-level failures.

58
Multi-Selecteasy

A company wants to ensure that its application running on AWS can withstand the failure of an entire AWS Region. Which TWO strategies should the company implement?

Select 2 answers
A.Deploy the application in multiple AWS Regions using an active-active or active-passive pattern
B.Deploy the application across multiple Availability Zones in a single Region
C.Replicate data across Regions using services like DynamoDB global tables or RDS cross-Region replication
D.Use a single CloudFront distribution with multiple origins in the same Region
E.Configure RDS read replicas in the same Region
AnswersA, C

Running the application in multiple AWS Regions using an active-active or active-passive pattern is the foundational disaster recovery strategy for a regional outage. Active-active routes live traffic across Regions for automatic failover, while active-passive keeps a warm standby in another Region that can be promoted via Route 53 health checks and failover policies. This approach directly addresses the failure domain of an entire Region.

Why this answer

Option A is correct because deploying the application in multiple AWS Regions using an active-active or active-passive pattern ensures that if an entire Region fails, traffic can be served from another Region, providing true Region-level fault tolerance. Option C is correct because replicating data across Regions with services like DynamoDB global tables or RDS cross-Region replication ensures the application's data is available in the secondary Region, which is essential for the failover strategy in Option A to work. Option B is incorrect because multiple Availability Zones within a single Region only protect against AZ-level failures, not the failure of an entire Region.

Option D is incorrect because a single CloudFront distribution with multiple origins in the same Region still depends on that one Region and does not survive a Region-wide outage. Option E is incorrect because RDS read replicas in the same Region remain within that Region and would be lost if the entire Region failed.

Exam trap

DOP-C02 often tests the distinction between multi-AZ (single-Region HA) and multi-Region (disaster recovery); candidates select multi-AZ answers because they sound resilient but do not survive a Region outage.

59
MCQhard

A company runs a critical e-commerce platform on AWS. The architecture includes an Application Load Balancer (ALB) that distributes traffic to a fleet of EC2 instances in an Auto Scaling group across three Availability Zones. The instances run a Java application that connects to an Amazon RDS Multi-AZ MySQL database. The application also uses Amazon ElastiCache for Redis for session caching. The company recently experienced a severe outage where the ALB's 5xx error rate spiked to 100% for 45 minutes. The root cause was a combination of a slow-running query on the RDS primary instance and a subsequent failover that caused the application to lose connections to the database. The failover happened because the slow query caused the primary to become unresponsive, triggering a Multi-AZ failover. During the failover, the application's connection pool exhausted, and new connections failed. The application logs show a high rate of 'java.sql.SQLTimeoutException' and 'com.mysql.cj.exceptions.CJCommunicationsException'. The DevOps team needs to implement a long-term solution that minimizes the impact of similar incidents. The solution must be cost-effective and require minimal application changes. Which combination of actions should the DevOps team take?

A.Implement Amazon RDS Proxy to manage database connections and add read replicas to offload read traffic.
B.Use an Auto Scaling policy for EC2 based on RDS connection count and implement a read replica for the primary.
C.Configure Multi-AZ RDS with a synchronous standby and use Amazon RDS for MySQL with enhanced monitoring.
D.Increase the instance size of the RDS primary and enable Performance Insights to identify slow queries.
AnswerA

Amazon RDS Proxy sits between the application and the database, maintaining a warm connection pool that absorbs the spike in connection requests when EC2 instances reconnect during a failover. Because the proxy keeps connections to the RDS instance open and multiplexes client sessions, the primary no longer gets overwhelmed by thousands of short-lived connections. Adding read replicas moves read-heavy queries off the primary, reducing CPU/IO contention that can cause slow queries and cascading failovers. This directly addresses the root cause of connection exhaustion while preserving write consistency on the primary.

Why this answer

Amazon RDS Proxy is the correct solution because it efficiently manages database connection pooling, reducing the likelihood of connection exhaustion during failovers. By maintaining a warm connection pool and automatically reconnecting to the new primary after a Multi-AZ failover, RDS Proxy minimizes application-side connection timeouts and errors like SQLTimeoutException and CJCommunicationsException. Adding read replicas offloads read traffic, reducing the load on the primary and mitigating the risk of slow queries causing unresponsiveness.

This combination requires minimal application changes and is cost-effective compared to scaling the primary instance.

Exam trap

The trap here is that candidates often focus on scaling the database (e.g., increasing instance size or adding read replicas) to fix performance issues, but overlook the critical connection management problem that causes application-level timeouts during failover, which RDS Proxy directly addresses.

How to eliminate wrong answers

Option B is wrong because using an Auto Scaling policy based on RDS connection count does not address the root cause of connection exhaustion during failover; it only scales EC2 instances reactively, which may not prevent timeouts and adds complexity without solving the connection management issue. Option C is wrong because simply configuring Multi-AZ RDS with a synchronous standby and enhanced monitoring does not prevent connection pool exhaustion during failover; the application still needs to manage connections, and enhanced monitoring only provides visibility, not mitigation. Option D is wrong because increasing the instance size of the RDS primary and enabling Performance Insights addresses performance but does not solve the connection management problem during failover; it may delay the issue but does not prevent connection timeouts or exhaustion.

60
MCQhard

A company runs a critical microservice on Amazon ECS with AWS Fargate. The service must be highly available across multiple Availability Zones. The DevOps engineer configured the service with a desired count of 4 tasks spread across 2 Availability Zones. During a deployment, a new task fails to start due to a missing environment variable. The deployment fails, but the old tasks continue to run. What is the most likely cause of the deployment failure and how can the engineer ensure future deployments are resilient?

A.The deployment failed because the ECS service was using the rolling update deployment controller. Change to blue/green deployment.
B.The deployment failed because the ECS service did not have the deployment circuit breaker enabled. Enable the circuit breaker with rollback.
C.The deployment failed because the desired count was too low. Increase the desired count to 6.
D.The deployment failed because the health check grace period was too short. Increase the grace period.
AnswerB

The correct fix is to enable the ECS deployment circuit breaker with rollback. This feature monitors the deployment for indicators such as repeated task launch failures, container exits, and health check failures; when it detects that the deployment is failing beyond a configured threshold, it stops the deployment and automatically restores the service to the most recent successful task set definition. This preserves service availability by keeping the existing stable tasks running and eliminates the need for you to manually roll back a bad task definition. Without the circuit breaker, a deployment with continuously crashing tasks will eventually time out and leave the service in a FAILED state with no automatic recovery.

Why this answer

The deployment failed because the new task could not start due to a missing environment variable, and the ECS service did not have the deployment circuit breaker enabled. Without the circuit breaker, ECS continues to attempt the deployment indefinitely or until a timeout, but it does not automatically roll back to the previous stable task set. Enabling the deployment circuit breaker with rollback ensures that if a specified number of tasks fail to start (e.g., due to health checks or runtime errors), ECS automatically rolls back to the last successful deployment, maintaining service availability.

Exam trap

The trap here is that candidates may focus on the deployment controller type (rolling vs. blue/green) or task count, but the real issue is the lack of automatic rollback capability provided by the deployment circuit breaker, which is specifically designed to handle task startup failures during deployments.

How to eliminate wrong answers

Option A is wrong because the rolling update deployment controller is not the cause of the failure; it is the default and works correctly here by keeping old tasks running. Changing to blue/green deployment would not inherently fix the missing environment variable issue and adds complexity. Option C is wrong because the desired count of 4 tasks is sufficient for high availability across 2 AZs; increasing it to 6 does not address the root cause of task startup failure.

Option D is wrong because the health check grace period only delays the start of health checks, but the task failed to start entirely due to a missing environment variable, not because health checks failed prematurely.

61
MCQhard

A company runs a Stateful application on EC2 that requires sticky sessions. They use an ALB with duration-based stickiness. During a deployment, they want to drain existing connections gracefully before terminating instances. Which step is necessary?

A.Increase the deregistration delay on the target group.
B.Reduce the stickiness duration to zero.
C.Configure health checks to mark instances unhealthy.
D.Enable connection draining on the target group.
AnswerA

ALB implements graceful connection draining through the target group's deregistration delay. Increasing this delay allows in-flight sticky-session requests to complete before the instance is terminated.

Why this answer

ALB target groups use a 'deregistration delay' (formerly called connection draining) to allow in-flight requests to complete before an instance is terminated. This setting is configured on the target group, not the load balancer, and it works with sticky sessions by waiting for the delay period (default 300 seconds) for existing connections to finish, even if the stickiness cookie would otherwise route new requests to the same instance. During a deployment, increasing this delay or ensuring it is set appropriately is the necessary step to gracefully drain connections.

Exam trap

The trap here is that candidates confuse 'connection draining' (which is the deregistration delay on ALB target groups) with 'sticky session duration' or 'health check settings,' thinking that reducing stickiness or marking instances unhealthy alone will gracefully terminate connections, when in fact the deregistration delay is the specific mechanism that waits for in-flight requests to complete.

How to eliminate wrong answers

Option A is wrong because increasing the deregistration delay is not the step that enables draining; the deregistration delay is already the mechanism for connection draining (it is the same setting), but the question asks which step is necessary, and the correct answer is to enable connection draining, which is already the default behavior of the deregistration delay. Option B is wrong because reducing the stickiness duration to zero would disable sticky sessions entirely, breaking the application requirement for sticky sessions, and it does not drain existing connections—it simply stops new sessions from being sticky. Option C is wrong because configuring health checks to mark instances unhealthy would cause the ALB to stop sending new traffic to the instance, but it does not wait for existing connections to complete; the deregistration delay is what handles in-flight requests, not health check status.

62
MCQhard

Refer to the exhibit. A Lambda function uses the IAM role with the above policy. The function is configured to access a DynamoDB table MyTable and an RDS instance in a VPC. When invoked, the function fails with an error indicating it cannot describe VPC subnets. What is the MOST likely cause?

A.The Lambda function is missing permissions to describe VPC subnets and security groups.
B.The Lambda function does not have permission to write to DynamoDB.
C.The Lambda function cannot create network interfaces in the VPC.
D.The DynamoDB table's resource policy denies access from Lambda.
AnswerA

The Lambda execution role must explicitly allow ec2:DescribeSubnets and ec2:DescribeSecurityGroups so that the Lambda service can inspect your VPC and locate the subnets and security groups you specified. These read-only actions are prerequisites for Lambda to create an elastic network interface (ENI); without them, the invocation fails during the VPC provisioning step, even if the role allows creating ENIs. This is why the error specifically mentions missing subnet or security group describe permissions.

Why this answer

The Lambda execution role policy shown in the exhibit grants DynamoDB and RDS-related actions but omits the ec2:DescribeSubnets, ec2:DescribeSecurityGroups, and ec2:DescribeNetworkInterfaces permissions required when a function is attached to a VPC. When Lambda is configured with VPC access, the service must call these EC2 APIs to validate and place the ENIs, so the invocation fails with a subnet-description error. Adding the missing ec2:Describe* permissions to the role resolves the failure.

Exam trap

DOP-C02 often tests the misconception that VPC-attached Lambda only needs ENI creation permissions, when in fact DescribeSubnets, DescribeSecurityGroups, and DescribeNetworkInterfaces are also mandatory and their absence produces the exact error described.

How to eliminate wrong answers

Option B is wrong because DynamoDB write failures would surface as AccessDeniedException on PutItem/UpdateItem, not as a VPC subnet description error, and the exhibit policy already grants DynamoDB actions. Option C is wrong because ENI creation failures produce a different error (e.g., 'The provided execution role does not have permissions to call CreateNetworkInterface'), and the reported error is specifically about describing subnets. Option D is wrong because a DynamoDB resource policy denial would return an access-denied error on the DynamoDB call itself, unrelated to VPC subnet enumeration.

63
MCQeasy

A company runs a static website on Amazon S3 with public read access. The website content is stored in an S3 bucket and served through an Amazon CloudFront distribution for better performance and security. Recently, the company noticed that some users are accessing the S3 bucket directly via the S3 endpoint, bypassing CloudFront. This increases costs and exposes the bucket to potential attacks. The company wants to ensure that all access to the website goes through CloudFront only. Which solution should the company implement?

A.Set the S3 bucket policy to deny all requests that do not come from the CloudFront distribution's IP addresses.
B.Configure the S3 bucket to use AWS WAF to block requests that do not have a custom header set by CloudFront.
C.Create an origin access identity (OAI) in CloudFront and update the S3 bucket policy to allow only the OAI to read objects.
D.Change the S3 bucket to be private and use presigned URLs for all requests.
AnswerC

Creating an origin access identity (OAI) in CloudFront and updating the bucket policy to permit only that OAI to read objects is the correct approach. The OAI is a special CloudFront user that validates requests to the S3 origin with AWS Signature Version 4, and the bucket policy grants s3:GetObject permission exclusively to this principal. This blocks any request that does not come through the CloudFront distribution, while still allowing the static content to be publicly served to end users via CloudFront.

Why this answer

To restrict access to the S3 bucket only through CloudFront, use an origin access identity (OAI) and a bucket policy that allows only the OAI. This way, direct access via S3 URL is denied.

64
MCQmedium

A company runs a containerized application on Amazon EKS. They want to ensure that if a node fails, the pods are rescheduled on healthy nodes. Which configuration is necessary?

A.Configure a pod disruption budget to prevent too many pods from being terminated simultaneously.
B.Use a horizontal pod autoscaler to increase the number of pods during high load.
C.Configure the EKS managed node group with a health check and ensure that the Kubernetes control plane automatically reschedules pods from failed nodes.
D.Use a cluster autoscaler to automatically add new nodes when pods are pending.
AnswerC

An EKS managed node group is backed by an Auto Scaling group whose health checks include both Amazon EC2 status checks and Kubernetes node status (including the node-lost and NodeReady conditions). When a node fails or becomes unhealthy, the Auto Scaling group replaces the underlying instance, while Kubernetes' node controller marks the node as NotReady and evicts pods, leading them to be rescheduled onto healthy nodes. This combination of automatic instance replacement and control-plane-driven pod rescheduling directly addresses the failure of a node and is the expected solution for maintaining availability.

Why this answer

EKS managed node groups automatically register nodes with the Kubernetes control plane, and the Kubernetes node controller (part of the kube-controller-manager) monitors node health via the NodeLifecycleController. When a node fails (e.g., due to an EC2 instance termination or health check failure), the control plane marks the node as `NotReady` and, after the default pod eviction timeout (5 minutes), evicts pods from the failed node, rescheduling them on healthy nodes. This behavior is inherent to Kubernetes and does not require additional configuration beyond using a managed node group.

Exam trap

The trap here is that candidates confuse the Cluster Autoscaler (which adds nodes) with the node controller's pod rescheduling behavior, or they think a PodDisruptionBudget is needed for failure recovery when it only applies to voluntary disruptions.

How to eliminate wrong answers

Option A is wrong because a PodDisruptionBudget (PDB) controls voluntary disruptions (e.g., node drains during updates) and does not handle involuntary node failures; it would actually prevent pods from being rescheduled if the PDB's minAvailable or maxUnavailable constraints are violated. Option B is wrong because a HorizontalPodAutoscaler (HPA) scales the number of pod replicas based on CPU/memory metrics, not in response to node failures; it does not reschedule pods from failed nodes. Option D is wrong because the Cluster Autoscaler adds new nodes when pods are unschedulable due to resource constraints, but it does not reschedule pods from failed nodes—that is the responsibility of the Kubernetes node controller and kube-scheduler.

65
MCQhard

A company runs a production e-commerce platform on AWS. The architecture includes an Application Load Balancer (ALB) that distributes traffic to a fleet of Amazon EC2 instances running in an Auto Scaling group across three Availability Zones (AZs). The application stores session state in Amazon ElastiCache for Redis (cluster mode disabled) with a single node. The database is an Amazon Aurora MySQL DB cluster with one writer and two reader instances in different AZs. The platform experiences intermittent slowdowns and occasional timeouts during peak traffic hours. The CloudWatch metrics show that the ALB's TargetResponseTime is elevated, and the Redis CPU utilization is consistently above 80% during these periods. The Auto Scaling group is scaling out, but new instances take several minutes to become healthy. The DevOps team has been asked to improve the resilience and performance of the application with minimal changes to the application code. Which solution should the team implement?

A.Replace the ALB with a Network Load Balancer (NLB) to reduce latency, and use an Auto Scaling group with a step scaling policy based on Redis CPU utilization.
B.Increase the instance size of the ElastiCache for Redis node and the size of the Aurora writer instance. Also, increase the cooldown period for the Auto Scaling group to allow new instances to warm up.
C.Implement Amazon RDS Proxy in front of the Aurora cluster to reduce database connection overhead, and increase the size of the Redis instance to handle more connections.
D.Migrate ElastiCache for Redis to a cluster mode enabled configuration with multiple shards and enable Multi-AZ with automatic failover. Also, use an ElastiCache replication group with read replicas in different AZs.
AnswerD

Migrating to cluster mode enabled with multiple shards horizontally partitions the Redis keyspace across nodes, which directly lowers per-shard CPU utilization and enables linear scaling as traffic grows. Enabling Multi-AZ with automatic failover and placing read replicas in different AZs provides high availability and lets reads be served by replicas, reducing primary node load and cutting failover time from minutes to seconds—this tackles both the immediate CPU bottleneck and the resilience requirement for a production e-commerce platform.

Why this answer

The primary bottleneck is the single-node Redis instance (CPU > 80%), which cannot scale horizontally and lacks high availability. Migrating to cluster mode enabled with multiple shards distributes CPU load across shards, while Multi-AZ with automatic failover and read replicas in different AZs provides high availability and read scaling. This directly addresses the elevated ALB TargetResponseTime caused by Redis latency.

Note that the application will need to use a Redis Cluster-compatible client; since the stem allows minimal code changes, this solution is still appropriate.

Exam trap

The trap here is that candidates focus on scaling the database or load balancer (options A, B, C) instead of recognizing that the single-node Redis cache is the bottleneck and requires horizontal scaling and high availability to resolve both performance and resilience issues.

How to eliminate wrong answers

Option A is wrong because replacing the ALB with an NLB does not reduce application-layer latency (NLB operates at Layer 4, not Layer 7, and cannot offload TLS or inspect HTTP sessions), and a step scaling policy based on Redis CPU utilization does not fix the single-node Redis bottleneck or the slow instance warm-up. Option B is wrong because increasing the instance size of the single Redis node and the Aurora writer instance only vertically scales the existing bottlenecks, and increasing the Auto Scaling group cooldown period would delay scaling further, worsening the timeouts. Option C is wrong because RDS Proxy reduces database connection overhead but does not address the Redis CPU bottleneck (the primary cause of elevated response times), and increasing the Redis instance size alone does not provide the read scaling or high availability needed.

66
MCQhard

A company runs a high-traffic web application on a fleet of EC2 instances behind an Application Load Balancer (ALB) with Auto Scaling. The application uses an Amazon RDS for PostgreSQL database. Recently, during a traffic spike, the application became unresponsive. Investigation revealed that the database CPU utilization reached 100%, causing queries to timeout. The Auto Scaling group added more EC2 instances, which only increased the load on the database. The DevOps team needs to implement a solution that prevents the database from being overwhelmed during traffic spikes while maintaining application availability. The solution must be cost-effective and require minimal changes to the application code. Which solution should the DevOps team implement?

A.Implement read replicas for the RDS database and modify the application to use read replicas for read queries.
B.Increase the instance size of the RDS database to a larger instance type to handle more connections.
C.Use Amazon RDS Proxy between the application and the database to pool and reuse connections.
D.Configure Auto Scaling to launch EC2 instances based on a custom metric that tracks database CPU utilization, and throttle the number of instances.
AnswerC

Amazon RDS Proxy presents a single endpoint to the application while pooling and reusing database connections on the backend, dramatically cutting the CPU and memory load caused by connection handling. It is fully managed and transparent, and during Multi-AZ failovers it can keep connections warm for faster recovery. For a high-traffic web fleet, this directly mitigates CPU exhaustion due to connection volume, making it the most efficient and cost-effective option.

Why this answer

RDS Proxy manages database connections efficiently, reducing the number of connections and CPU overhead. It also provides connection pooling, which helps handle spikes without overwhelming the database.

67
MCQhard

A company runs a stateful web application on EC2 instances behind a Network Load Balancer (NLB) in a single Availability Zone. The application stores session state locally on the instance. The company wants to achieve high availability across multiple AZs with minimal application changes. What should the DevOps engineer do?

A.Add more AZs and configure the NLB with cross-zone load balancing.
B.Replace the NLB with an ALB and use ElastiCache for session storage.
C.Use a Multi-AZ RDS instance to store session state.
D.Replace the NLB with an ALB and enable sticky sessions (session affinity) using the ALB's cookie.
AnswerD

Replacing the NLB with an ALB and enabling sticky sessions via the ALB's load balancer-generated cookie is the correct minimal-change solution. The ALB inserts a stickiness cookie on the first response, and all subsequent requests from that client are routed to the same EC2 instance, preserving the locally stored session state without any application modifications. This leverages the ALB's Layer 7 capabilities to achieve session affinity while keeping the existing web application code unchanged.

Why this answer

Replacing the NLB with an ALB and enabling sticky sessions (session affinity) using the ALB's cookie allows the stateful web application to maintain session state across multiple AZs without modifying the application code. The ALB generates a cookie (AWSALB) that binds a client's session to a specific target instance, ensuring subsequent requests from the same client are routed to the same EC2 instance. This achieves high availability across AZs with minimal changes, as the application continues to store session state locally on the instance.

Exam trap

The trap here is that candidates often assume cross-zone load balancing or adding more AZs inherently solves high availability for stateful applications, but they overlook that session affinity is required to keep a client's requests directed to the same instance when session state is stored locally.

How to eliminate wrong answers

Option A is wrong because adding more AZs and configuring cross-zone load balancing with an NLB does not solve the session state problem; the NLB distributes traffic across instances without session affinity, so a client's requests may be routed to different instances in different AZs, breaking the locally stored session. Option B is wrong because replacing the NLB with an ALB and using ElastiCache for session storage requires application code changes to read/write session data to ElastiCache, which contradicts the requirement for minimal application changes. Option C is wrong because using a Multi-AZ RDS instance for session storage also requires significant application code changes to store and retrieve session data from the database, and it introduces unnecessary complexity and latency for session management.

68
Multi-Selecthard

A company's application uses Amazon DynamoDB as its primary data store. The application experiences occasional throttling errors during traffic spikes. The DevOps team needs to implement a solution that ensures consistent performance without manual intervention. Which TWO actions should the team take? (Choose TWO.)

Select 2 answers
A.Use eventually consistent reads for all queries.
B.Move the data to Amazon RDS with read replicas.
C.Implement DynamoDB Accelerator (DAX) to cache read requests.
D.Enable DynamoDB Auto Scaling for read and write capacity.
E.Switch DynamoDB to On-Demand capacity mode.
AnswersC, D

DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of a DynamoDB table, serving read-heavy workloads with microsecond latency while absorbing a large fraction of read requests. By caching frequently accessed items (including strongly consistent reads when DAX is enabled), DAX reduces the number of read requests that actually reach DynamoDB, thereby reducing the table's consumed read capacity and preventing read throttling. It is a native, fully managed solution specifically designed for this scenario, preserving the DynamoDB API and requiring no application rewrite beyond adding a DAX client endpoint.

Why this answer

To handle occasional throttling during traffic spikes without manual intervention, the team should use DynamoDB Accelerator (DAX) to cache read requests, reducing read load on the table, and enable DynamoDB Auto Scaling to automatically adjust read and write capacity based on traffic patterns. DAX absorbs spikey read traffic, while Auto Scaling ensures sufficient capacity for writes and uncached reads, together providing consistent performance without manual scaling. Option E (On-Demand) also handles spikes automatically but can be costlier for predictable workloads; the combination of DAX and Auto Scaling is often more cost-effective for read-heavy applications.

Exam trap

The trap here is that candidates may think On-Demand capacity mode (Option E) is the only way to handle spikes without manual intervention, but it ignores the cost implications and the fact that DAX plus Auto Scaling provides a more balanced and cost-effective solution for read-heavy workloads.

69
MCQhard

A company is implementing a disaster recovery strategy for its Amazon Aurora MySQL database. The primary database is in us-west-2. The company requires an RPO of less than 1 minute and an RTO of less than 5 minutes. Which solution meets these requirements?

A.Create a cross-Region read replica in the secondary Region and promote it during failover.
B.Use automated backups and restore to a new DB instance in the secondary Region.
C.Use Amazon Aurora Global Database with a secondary Region cluster.
D.Take manual snapshots of the DB instance and copy them to the secondary Region every hour.
AnswerC

Aurora Global Database replicates data from the primary Region to a secondary cluster using a dedicated storage-based replication channel with typical latency under one second, ensuring an RPO of under one minute. Failover can be initiated either manually or automatically, and a promoted secondary cluster becomes available in minutes without the need to restore from a backup or apply transaction logs. This is the only option that inherently satisfies both the 1-minute RPO and a recovery time objective measured in minutes.

Why this answer

Amazon Aurora Global Database is designed for low-latency cross-Region replication with a typical RPO of 1 second and RTO of 1 minute or less, meeting the <1 minute RPO and <5 minute RTO requirements. It uses a dedicated storage-level replication channel that keeps the secondary cluster fully synchronized without impacting primary performance, and failover involves promoting the secondary cluster to primary in under a minute.

Exam trap

The trap here is that candidates confuse a cross-Region read replica (Option A) with Aurora Global Database, assuming both provide similar failover speed, but the read replica's promotion process is slower and less reliable for meeting strict RTO/RPO targets.

How to eliminate wrong answers

Option A is wrong because a cross-Region read replica for Aurora MySQL uses asynchronous replication with a typical RPO of several seconds to minutes, but the promotion process can take longer than 5 minutes due to the need to apply remaining redo logs and reconfigure endpoints, failing the RTO requirement. Option B is wrong because automated backups are taken once per day (default retention of 1-35 days) and restoring to a new instance in a secondary Region requires copying the backup across Regions, which can take hours and far exceeds both the RPO and RTO limits. Option D is wrong because manual snapshots taken every hour provide an RPO of up to 60 minutes, which violates the <1 minute RPO requirement, and restoring from a snapshot in a secondary Region also takes significantly longer than 5 minutes.

70
Multi-Selectmedium

A company runs a stateful web application on EC2 instances that store session data locally. They want to migrate to a stateless architecture for better resilience. Which TWO actions should they take?

Select 2 answers
A.Use Amazon CloudFront to cache session data at the edge.
B.Use Amazon DynamoDB to store session data.
C.Use Amazon S3 to store session data as objects.
D.Use ElastiCache for Redis to store session data externally.
E.Use Amazon EFS to store session data as files.
AnswersB, D

Amazon DynamoDB is a fully managed NoSQL key-value database that delivers single-digit millisecond read/write performance at any scale, making it an excellent external session store. Its fine-grained access control, encryption, backup, and TTL support for automatic item expiration align well with session lifecycle needs. Because DynamoDB is serverless and horizontally scalable, EC2 instances can share session state without statefulness, enabling the application tier to scale freely behind a load balancer.

Why this answer

DynamoDB provides a fully managed, low-latency, highly available NoSQL database that is ideal for storing session state externally. By moving session data to DynamoDB, the EC2 instances become stateless, allowing any instance to handle any request without relying on local storage, which improves resilience and scalability.

Exam trap

The trap here is that candidates may confuse stateless session storage with caching or file storage, incorrectly choosing S3 or EFS because they are persistent, while overlooking the need for low-latency, high-throughput, and consistent access that only DynamoDB or ElastiCache can provide.

71
MCQeasy

A company runs a serverless application using AWS Lambda functions behind an Amazon API Gateway. The application processes user uploads stored in an S3 bucket. The Lambda function writes results to a DynamoDB table. Recently, the function started timing out when processing large files. What should the DevOps engineer do to improve resilience for large file processing?

A.Increase the Lambda function memory to improve CPU performance.
B.Use S3 event notifications to trigger an AWS Step Functions workflow that processes the file asynchronously.
C.Increase the Lambda function timeout to the maximum 15 minutes.
D.Add Amazon ElastiCache to cache processed results and reduce Lambda execution time.
AnswerB

S3 event notifications triggering Step Functions decouple processing from the API request, letting large files be handled asynchronously with retries and state tracking. This removes the Lambda timeout constraint imposed by synchronous invocation behind API Gateway.

Why this answer

Using S3 event notifications to trigger an AWS Step Functions workflow enables asynchronous processing of large files, decoupling the upload from the processing and avoiding Lambda's timeout limits. Option A (increasing memory) may improve CPU performance but does not address the timeout issue for large files. Option C (increasing timeout) can extend up to 15 minutes, but large files may still exceed this limit and it does not provide a resilient architecture.

Option D (ElastiCache) caches processed results but does not solve the initial timeout problem; it is irrelevant to the large file processing issue.

72
MCQeasy

A company uses AWS CodeDeploy to deploy a new version of an application to EC2 instances. They want to minimize downtime and roll back quickly if the deployment fails. Which deployment type should they use?

A.Canary deployment
B.Linear deployment
C.Blue/green deployment
D.In-place deployment
AnswerC

Blue/green deployment in AWS CodeDeploy is the only deployment type that provisions a new, separate environment (green) alongside the existing one (blue), then shifts production traffic to the green environment. This architecture enables instant rollback by simply switching traffic back to the blue environment if the deployment fails or exhibits issues, with no need to reinstall or reconfigure instances. CodeDeploy manages this through lifecycle hooks like AllowTraffic and redirecting traffic via the load balancer, providing the fastest and most reliable rollback path.

Why this answer

Blue/green deployment creates two separate environments (blue and green) and shifts traffic from the old to the new after testing. This minimizes downtime because traffic is switched instantly, and rollback is achieved by reverting traffic to the original environment. Option A (Canary) is a traffic shifting pattern used within blue/green deployments, not a standalone deployment type that offers immediate rollback.

Option B (Linear) is also a traffic shifting pattern for blue/green. Option D (In-place) updates existing instances, causing downtime during deployment and requiring a manual rollback process.

73
Multi-Selecthard

A company has a microservices architecture running on Amazon ECS with Fargate launch type. Each service is deployed in multiple Availability Zones. The services communicate via REST APIs. Recently, a downstream service experienced a partial outage, causing upstream services to time out and leading to cascading failures. The team wants to improve resilience against such failures. Which combination of actions should the DevOps engineer take? (Choose TWO.)

Select 2 answers
A.Increase the HTTP timeout values for all service-to-service calls.
B.Implement circuit breaker patterns in the service clients.
C.Remove all retry logic from service calls.
D.Adopt an asynchronous communication pattern using Amazon SQS or Amazon EventBridge.
E.Configure Auto Scaling for all services based on request count.
AnswersB, D

Circuit breakers in service clients detect repeated downstream failures and trip open, failing fast instead of holding threads until timeout. This halts the cascade at the caller, satisfying the resilience requirement by preventing upstream exhaustion when a downstream service partially fails.

Why this answer

Option B is correct because a circuit breaker in the service clients detects repeated failures from a downstream dependency and trips open, failing fast instead of letting upstream threads block on slow REST calls, which directly prevents the timeout propagation and cascading failures described. Option D is correct because moving to asynchronous communication with Amazon SQS or Amazon EventBridge decouples services: upstream services enqueue or publish events and return immediately, so a partial outage in a downstream consumer no longer blocks callers, and messages can be retried or buffered until the consumer recovers. Option A is not appropriate because increasing HTTP timeouts makes callers wait longer, consuming threads and connections and worsening cascading failures rather than containing them.

Option C is wrong because removing retry logic eliminates a useful resilience mechanism for transient errors; retries should be bounded and combined with circuit breakers and backoff, not deleted. Option E is not the right fix because scaling on request count does not address a downstream dependency that is failing or slow, and could even amplify load against the impaired service.

Exam trap

DOP-C02 often tests whether candidates confuse 'make the timeout longer' with resilience — longer timeouts amplify cascading failures rather than preventing them.

74
Multi-Selectmedium

A company runs a microservices application on Amazon ECS with Fargate. The application includes a service that processes orders and stores them in an RDS PostgreSQL database. The company wants to ensure that the order service is resilient to AZ failures and can handle a sudden increase in order volume. Which TWO actions should the DevOps engineer take? (Choose TWO.)

Select 2 answers
A.Increase the CPU and memory limits for the ECS task definition.
B.Place an Amazon CloudFront distribution in front of the order service.
C.Deploy the RDS instance in a Multi-AZ configuration.
D.Configure the ECS service to run tasks in multiple Availability Zones.
E.Use RDS Proxy to manage database connections.
AnswersC, D

Deploying RDS in a Multi-AZ configuration creates a synchronous standby replica in a different Availability Zone, and Amazon RDS automatically fails over to the standby if the primary instance becomes unhealthy or the AZ fails. The DNS name stays the same, so the ECS order service can reconnect without code changes, and the standby is continuously updated with synchronous replication to prevent data loss. This directly addresses the database as a single point of failure, which is necessary because the order service depends on durable transactions to record orders.

Why this answer

Deploying the RDS instance in a Multi-AZ configuration provides automatic failover to a standby replica in a different Availability Zone, ensuring database resilience to AZ failures. Option D is correct because configuring the ECS service to run tasks in multiple Availability Zones distributes the order processing workload across AZs, improving both fault tolerance and scalability during sudden traffic spikes.

Exam trap

The trap here is that candidates often confuse connection pooling (RDS Proxy) with high availability (Multi-AZ) or assume that vertical scaling (increasing task limits) is sufficient for both resilience and sudden load, when in fact horizontal distribution across AZs is required for fault tolerance and elasticity.

75
Multi-Selectmedium

A company is designing a highly available architecture for a stateless web application using AWS services. Which TWO steps should they take to achieve high availability?

Select 2 answers
A.Store session state in an EBS volume attached to each instance
B.Deploy EC2 instances in multiple Availability Zones
C.Use a single NAT instance in a public subnet
D.Use only M5 instance types for better performance
E.Use an Application Load Balancer to distribute traffic
AnswersB, E

Distributing EC2 instances across multiple Availability Zones ensures the application can tolerate a complete failure of one physical data center, because the remaining instances continue to serve traffic. Each AZ has independent power, cooling, and network connectivity, so an outage in one AZ does not affect the others. This redundancy is the fundamental building block of high availability on AWS and is a necessary condition for achieving a higher service-level agreement.

Why this answer

Option B is correct because deploying EC2 instances across multiple Availability Zones ensures the application survives an AZ-level failure, which is a fundamental requirement for high availability in AWS. Option E is correct because an Application Load Balancer distributes incoming traffic across healthy targets in multiple AZs, performs health checks, and automatically routes around failed instances, directly supporting high availability for a stateless web tier. Option A is incorrect because storing session state on an EBS volume tied to a single instance creates a single point of failure and is unnecessary for a stateless application.

Option C is incorrect because a single NAT instance is itself a single point of failure and cannot provide high availability. Option D is incorrect because choosing a specific instance type like M5 improves performance but does nothing to increase availability.

Exam trap

The trap is thinking that a single NAT instance or EBS-backed session state can provide HA; both introduce single points of failure.

Page 1 of 3 · 184 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Resilient Cloud questions.