Courseiva

CCNA Resilient Cloud Solutions Questions

34 of 184 questions · Page 3/3 · Resilient Cloud Solutions · Answers revealed

151
MCQeasy

A company's DevOps team is designing a disaster recovery plan for a critical application. The application runs on EC2 instances with an RDS MySQL database. The Recovery Time Objective (RTO) is 15 minutes, and the Recovery Point Objective (RPO) is 1 hour. Which approach BEST meets these requirements?

A.Use backup and restore with daily snapshots stored in S3 and cross-Region replication.
B.Use a multi-Region application with Route 53 latency-based routing and RDS read replicas in the DR Region.
C.Use a warm standby strategy with a scaled-down copy of the production environment in the DR Region, and replicate data using RDS Multi-AZ with synchronous replication.
D.Use a pilot light strategy with EC2 instances stopped and RDS snapshots copied to the DR Region.
AnswerB

Cross-Region RDS read replicas provide asynchronous replication with an RPO of seconds to minutes, meeting the 1-hour RPO. Promoting a read replica and redirecting traffic via Route 53 can be done within minutes, meeting the 15-minute RTO. This is a valid warm standby configuration.

Why this answer

The best approach for a multi-Region disaster recovery with RTO of 15 minutes and RPO of 1 hour. By deploying the application in multiple regions and using RDS cross-Region read replicas, data is asynchronously replicated with an RPO typically within seconds to minutes, well within 1 hour. In the event of a failure, the read replica can be promoted to a primary instance, and Route 53 routing (preferably failover routing, but latency-based routing can also redirect traffic) can shift traffic to the DR region.

This failover can be completed within a few minutes, meeting the 15-minute RTO. Option A fails because daily snapshots exceed the 1-hour RPO and restore times exceed the RTO. Option C incorrectly relies on RDS Multi-AZ, which is a single-region high-availability feature and does not provide cross-region replication; thus it cannot serve as a disaster recovery solution across regions.

Option D, pilot light with snapshots, has a longer RTO as it requires restoring instances from snapshots and starting them, likely exceeding 15 minutes.

Exam trap

A common trap is to assume that RDS Multi-AZ provides cross-region replication; however Multi-AZ is a high-availability feature within a single region. For cross-region disaster recovery, asynchronous cross-Region read replicas or other cross-region replication methods are required. A warm standby architecture can be combined with cross-region replication, but the key is the replication mechanism, not Multi-AZ.

How to eliminate wrong answers

Option A is wrong because daily snapshots with cross-Region replication result in an RPO of up to 24 hours, far exceeding the 1-hour requirement, and the restore process takes longer than 15 minutes. Option B is wrong because Route 53 latency-based routing is for active-active traffic distribution, not disaster recovery failover, and RDS read replicas are asynchronous, leading to potential data loss and RPO that can exceed 1 hour during a failure. Option D is wrong because a pilot light strategy with stopped EC2 instances and RDS snapshots copied to the DR Region requires provisioning and restoring from snapshots, which typically takes longer than 15 minutes to become fully operational, and the RPO is limited by snapshot frequency.

152
MCQeasy

Refer to the exhibit. A DevOps engineer applies the IAM policy shown to an S3 bucket to enforce server-side encryption. However, users report that some uploads succeed without encryption. What is the most likely reason?

A.The policy uses StringEquals instead of StringNotEquals.
B.The policy only allows the action but does not deny actions that do not meet the condition.
C.The resource ARN is incorrect; it should be the bucket ARN.
D.The action should be s3:PutEncryptedObject instead of s3:PutObject.
AnswerB

This is the core problem: IAM is default-deny, so this Allow statement only grants the upload when the condition is met; it does nothing to block uploads that fail the condition if another policy grants them. To enforce encryption, you must include an explicit Deny for s3:PutObject without the required encryption condition, because Deny always overrides Allow. Merely adding a condition to an Allow does not constrain other permissions.

Why this answer

The IAM policy only allows the s3:PutObject action when the encryption condition is met, but it does not include an explicit Deny statement to block uploads that do not satisfy the condition. In IAM, an Allow statement with a condition does not automatically deny requests that fail the condition; it simply does not apply the Allow. If there is another policy (e.g., a bucket policy or an identity-based policy) that grants s3:PutObject without the encryption condition, or if the default S3 behavior permits unencrypted uploads (since S3 does not require encryption by default), then unencrypted uploads can still succeed.

To enforce encryption, you must add a Deny statement with a condition like `StringNotEquals` on `s3:x-amz-server-side-encryption` to explicitly reject requests that lack the required encryption header.

Exam trap

The trap here is that candidates assume an Allow statement with a condition implicitly denies requests that don't meet the condition, but AWS IAM requires an explicit Deny to block non-compliant requests, and the absence of that Deny is the root cause of the enforcement failure.

How to eliminate wrong answers

Option A is wrong because using StringEquals is correct for allowing only requests with the specified encryption value; the issue is not the operator but the lack of a Deny statement. Option C is wrong because the resource ARN in the policy is correct for the bucket itself (e.g., `arn:aws:s3:::bucket-name`), and the action s3:PutObject applies to objects, but the policy's Resource field can be the bucket ARN or the bucket ARN with a wildcard for objects; the given ARN is not the root cause of the failure to enforce encryption. Option D is wrong because there is no such action as s3:PutEncryptedObject in AWS S3; encryption is controlled via request headers and conditions, not a separate API action.

153
MCQmedium

A company runs a critical batch processing workload on a fleet of EC2 instances in a single AWS Region. The workload can tolerate interruptions but must complete within 4 hours and must not lose progress if an instance is terminated. The company wants to minimize cost and operational overhead while ensuring resilience against a single Availability Zone failure. Which solution meets these requirements?

A.Use Reserved Instances in a single Availability Zone, and store intermediate results on instance store volumes.
B.Use Spot Instances with a diversified allocation strategy across multiple Availability Zones, and store intermediate results in Amazon S3.
C.Use On-Demand instances spread across multiple Availability Zones, and store intermediate results in an Amazon S3 bucket.
D.Use Spot Instances in a single Availability Zone, and store intermediate results on Amazon EBS volumes attached to each instance.
AnswerB

Spot Instances reduce cost by up to 90% and are ideal for interruptible workloads. A diversified allocation strategy across multiple Availability Zones reduces the chance of simultaneous interruptions. Storing progress in S3 ensures that if an instance is reclaimed, another instance can resume from the last checkpoint. This combination meets cost, resilience, and progress-preservation requirements.

Why this answer

The workload is interruptible, so Spot Instances are the most cost-effective choice. A diversified allocation strategy across multiple Availability Zones reduces correlated interruptions and provides AZ resilience. Storing intermediate results in Amazon S3 ensures progress is preserved because any instance can retrieve the latest checkpoint.

This solution satisfies cost, availability, and durability requirements with minimal operational overhead.

Exam trap

The trap here is assuming that Spot Instances alone provide resilience, when actually the allocation strategy and checkpoint storage are what ensure the workload survives interruptions and AZ failures.

154
Multi-Selecthard

A company is migrating a monolithic application to a microservices architecture on AWS. To improve resilience, which THREE design patterns should be implemented? (Select THREE.)

Select 3 answers
A.Synchronous communication between services to ensure consistency
B.Single shared database to maintain data consistency
C.Retry with exponential backoff for transient failures
D.Bulkhead pattern to isolate critical services from non-critical ones
E.Circuit breaker pattern to stop calls to a failing service
AnswersC, D, E

Retry with exponential backoff is a core resilience pattern that reattempts an operation after an exponentially increasing delay, typically with jitter to avoid synchronized retries, which prevents a thundering herd. It specifically targets transient failures, such as network timeouts, database connection drops, or throttling (e.g., API Gateway 429s or DynamoDB throttling), where the operation may succeed on a subsequent attempt if the original failure was short-lived. This pattern preserves system stability by giving the dependency time to recover and reducing continuous load, and it is widely supported in AWS SDKs and infrastructure.

Why this answer

Implementing retry with exponential backoff allows services to handle transient failures (e.g., network timeouts, throttling) by automatically retrying operations after increasing delays, reducing load on recovering systems. This pattern is essential for microservices on AWS, where services like DynamoDB or Lambda may throttle requests, and exponential backoff (e.g., using jitter as per AWS SDK defaults) prevents cascading failures.

Exam trap

The trap here is that candidates confuse synchronous communication (Option A) with resilience, but in microservices, synchronous calls increase failure propagation, while asynchronous patterns and the three selected patterns (retry, circuit breaker, bulkhead) are the correct resilience mechanisms.

155
MCQhard

A company runs a containerized application on Amazon ECS with Fargate launch type. The application experiences intermittent failures when the ECS service scheduler attempts to place tasks during a deployment. The DevOps engineer notices that tasks fail to start due to insufficient IP addresses in the VPC subnets. What is the MOST resilient solution to prevent this issue?

A.Create an ECS service-linked role with permissions to allocate IPs.
B.Increase the desired task count in the ECS service to pre-warm IP addresses.
C.Use VPC endpoints for ECS to reduce IP usage.
D.Configure the ECS service to use multiple subnets with larger CIDR blocks across multiple Availability Zones.
AnswerD

Deploying the ECS service across multiple subnets with larger CIDR blocks increases the total number of usable private IP addresses in different Availability Zones, allowing additional tasks to be launched. Since each awsvpc-mode task requires an ENI with its own private IP, a larger IP pool directly solves the exhaustion. Using multiple AZs also improves availability by spreading tasks across fault domains, making this the appropriate corrective action.

Why this answer

Using a larger CIDR block for subnets provides more IP addresses, and using multiple subnets across Availability Zones increases availability and capacity. Option A is wrong because increasing desired count does not solve IP shortage. Option B is wrong because ECS service-linked role does not affect IP allocation.

Option C is wrong because VPC endpoints do not provide IP addresses for tasks.

156
MCQeasy

A company runs a containerized application on Amazon ECS with Fargate. The application needs to store session state. Which service provides the MOST resilient and scalable solution?

A.Amazon ElastiCache for Redis
B.Amazon EFS
C.Ephemeral storage on the container instance
D.Amazon S3
AnswerA

Amazon ElastiCache for Redis is a fully managed in-memory data store that provides sub-millisecond read/write latencies and is ideal for storing transient session state. It supports replication across Availability Zones and automatic failover, so the session store remains available even if a container task is rescheduled or an Availability Zone fails. Because it is external to the ECS task, the session data persists independently of the container lifecycle, enabling stateless containers and horizontal scaling of the web tier.

Why this answer

Amazon ElastiCache for Redis provides a highly available, scalable, and low-latency in-memory data store ideal for session state management in a containerized environment. It supports replication and automatic failover, ensuring resilience. Option B (Amazon EFS) is a file storage service with higher latency and not designed for sub-millisecond session retrieval.

Option C (ephemeral storage on the container instance) is not durable; data is lost when the container stops or fails. Option D (Amazon S3) is object storage with higher latency and not optimized for frequent read/write operations required for session state.

157
MCQeasy

A company wants to ensure its Amazon RDS DB instance is highly available with automatic failover in case of an AZ failure. Which configuration should they use?

A.Multi-AZ deployment
B.Amazon RDS Proxy
C.Single-AZ with automated backups
D.Read replicas in multiple AZs
AnswerA

Multi-AZ deployment creates a primary DB instance and a standby replica in a different Availability Zone, using synchronous physical replication to keep the standby transactionally consistent. When an AZ failure, network outage, or instance health check failure occurs, Amazon RDS automatically flips the DNS CNAME to the standby within about 60–120 seconds, providing automatic failover without manual intervention. This is the only option that gives the DB instance true high availability with a single endpoint, which is why it is correct.

Why this answer

Amazon RDS Multi-AZ deployments maintain a synchronous standby replica in a different Availability Zone and automatically fail over to it if the primary fails, providing high availability with minimal downtime. The DNS endpoint remains the same, so applications reconnect without configuration changes. This is the standard AWS-recommended configuration for HA with automatic failover.

Exam trap

DOP-C02 often tests the confusion between Multi-AZ (synchronous HA with automatic failover) and read replicas (asynchronous read scaling, manual promotion) — candidates who pick read replicas for HA miss that they do not provide automatic failover in standard RDS.

How to eliminate wrong answers

Option B is wrong because Amazon RDS Proxy is a connection pooler that improves scalability and resilience for applications (especially Lambda), but it does not itself provide AZ-level failover — it complements Multi-AZ. Option C is wrong because Single-AZ with automated backups provides durability and point-in-time recovery but no automatic failover; recovery requires restoring from backup, causing significant downtime. Option D is wrong because read replicas are asynchronous copies used for read scaling and are not automatic failover targets in RDS (unlike Aurora, where read replicas can be promoted); promoting a read replica is a manual process and may lose recent transactions.

158
Multi-Selecteasy

A company is designing a highly available architecture for a web application that uses Amazon EC2 instances. The application must be resilient to the failure of a single instance and a single Availability Zone. Which TWO actions should the company take? (Choose TWO.)

Select 2 answers
A.Use an Auto Scaling group with a minimum of two instances spread across two Availability Zones.
B.Distribute EC2 instances across at least two Availability Zones.
C.Place all EC2 instances in a single Availability Zone and use a Network Load Balancer.
D.Use a single Application Load Balancer in one Availability Zone.
E.Use a single large EC2 instance in one Availability Zone.
AnswersA, B

An Auto Scaling group with a minimum of two instances across two Availability Zones is the most complete solution because it combines horizontal scaling with automated self-healing. The ASG continuously monitors instance health and automatically replaces failed instances, while distributing the workload across two AZs ensures that a single Availability Zone outage does not eliminate all capacity. This is the gold standard for highly available EC2 architectures.

Why this answer

An Auto Scaling group with a minimum of two instances spread across two Availability Zones ensures that if one instance or one entire AZ fails, the remaining instance(s) in the other AZ can continue serving traffic, and Auto Scaling will automatically launch a replacement instance in the healthy AZ to restore the desired count. Option B is correct because distributing EC2 instances across at least two Availability Zones is the fundamental requirement for AZ-level resilience, as it eliminates a single point of failure at the AZ boundary.

Exam trap

The trap here is that candidates often think a load balancer alone provides high availability, but they overlook that the load balancer itself must be deployed across multiple AZs (or be a Regional service like ALB with cross-zone load balancing enabled) and that instances must be in at least two AZs to survive an AZ failure.

159
MCQmedium

A DevOps engineer runs the above command and sees that one target is unhealthy with reason 'Target.Timeout'. The target is an EC2 instance running a web server on port 80. The security group for the instance allows inbound traffic on port 80 from the ALB's security group. What is the most likely cause of the health check failure?

A.The instance does not have a public IP address.
B.The instance's network ACL is blocking inbound traffic from the internet.
C.The web server on the instance is not responding to health check requests on port 80.
D.The security group for the ALB does not allow outbound traffic to the instance.
AnswerC

Timeout indicates no response from the web server.

Why this answer

The 'Target.Timeout' reason indicates that the health check request timed out, meaning the instance is not responding within the timeout period. The most common cause is that the web server is not running or is not listening on the correct port. Option C is correct because a non-responsive web server on port 80 would cause the timeout.

Option A is incorrect because health checks are performed over private IPs, not public IPs. Option B is incorrect because the network ACL applies at the subnet level and would affect all traffic; the ALB's security group is already allowed inbound, and the instance's security group allows inbound from the ALB. Option D is incorrect because the ALB's security group does not need to allow outbound to the instance; traffic flows from the ALB to the instance, and the instance's security group must allow inbound from the ALB.

160
MCQhard

A company uses an NLB to distribute traffic to a fleet of EC2 instances in a single Availability Zone. During a recent AWS outage in that zone, the application became completely unavailable. The company wants to achieve high availability without rearchitecting the application. Which change is MOST appropriate?

A.Use a larger instance type and enable detailed CloudWatch monitoring
B.Replace the NLB with an Application Load Balancer and enable cross-zone load balancing
C.Create an Auto Scaling group with a scheduled scaling policy to add instances during peak hours
D.Launch EC2 instances in a second Availability Zone and register them with the NLB target group
AnswerD

Launching additional EC2 instances in a second Availability Zone and registering them with the NLB target group creates a multi-AZ target fleet, allowing the load balancer to route new connections to healthy instances in the surviving zone when the original zone has an outage. To make this work you must also configure the NLB itself with subnets in at least two AZs so its nodes are geographically distributed; then the target group's health checks automatically remove failed instances and keep the service available. This is the correct architecture for fault tolerance because it eliminates the single AZ as a point of failure and aligns with AWS's region-based resiliency patterns.

Why this answer

Registering EC2 instances in a second Availability Zone with the NLB target group allows NLB to route traffic to healthy instances across zones, providing high availability during a zone outage. Option A is incorrect because using a larger instance type and enabling detailed CloudWatch monitoring does not add redundancy across zones. Option B is incorrect because replacing NLB with an ALB still requires multi-AZ configuration to achieve high availability; cross-zone load balancing is already available on NLB.

Option C is incorrect because scheduled scaling does not protect against zone failures; it only adjusts capacity predictably.

161
MCQeasy

A company is designing a multi-region active-active architecture for a stateless web application. The application uses a DynamoDB table as its data store. The company wants to minimize write latency and ensure that writes are accepted in any region with eventual consistency. Which DynamoDB feature should they use?

A.DynamoDB read replicas in each region.
B.Cross-region replication using AWS Lambda function.
C.DynamoDB global tables.
D.DynamoDB Accelerator (DAX) with multi-region endpoints.
AnswerC

Amazon DynamoDB global tables provide a fully managed multi-Region, multi-master replication solution that maintains two or more identical tables across AWS Regions. When a write is issued to any replica, DynamoDB automatically propagates that write to every other region using DynamoDB Streams, with last-writer-wins conflict resolution to reconcile concurrent updates. This makes global tables the appropriate choice for an active-active architecture because each region can accept both reads and writes while providing low-latency access for geographically distributed users. The service handles the replication plumbing, versioning, and conflict resolution internally, eliminating the need for custom code.

Why this answer

DynamoDB global tables provide a fully managed, multi-region, multi-active solution that automatically replicates data across AWS Regions. This enables low-latency writes in any region with eventual consistency, meeting the requirement for an active-active architecture without custom code or additional infrastructure.

Exam trap

The trap here is that candidates may confuse DynamoDB Accelerator (DAX) as a solution for multi-region writes, but DAX only caches reads in a single region and does not replicate writes across regions.

How to eliminate wrong answers

Option A is wrong because DynamoDB read replicas are not a native feature; DynamoDB supports global tables for multi-region replication, not read replicas. Option B is wrong because using AWS Lambda for cross-region replication introduces custom code, complexity, and potential latency, and is not a managed, native DynamoDB feature for active-active setups. Option D is wrong because DAX is an in-memory cache that reduces read latency but does not replicate writes across regions; it operates within a single region and does not provide cross-region write acceptance.

162
Multi-Selecthard

A company is designing a disaster recovery plan for a MySQL database running on Amazon RDS. The database is critical and must have an RPO of 5 minutes and an RTO of 1 hour. The primary Region is us-east-1, and the DR Region is us-west-2. Which TWO steps should the company take to meet these requirements? (Choose TWO.)

Select 2 answers
A.Create a cross-Region read replica in us-west-2.
B.Enable automated backups with a 5-minute backup window.
C.Configure cross-Region automated snapshot copy to us-west-2.
D.Set up a process to promote the read replica to a standalone instance in us-west-2 during a disaster.
E.Enable Multi-AZ deployment in us-east-1.
AnswersA, D

A cross-Region read replica continuously replicates asynchronously from us-east-1 to us-west-2, keeping data loss well within the 5-minute RPO. During failover, promoting it to a standalone primary completes in minutes, satisfying the 1-hour RTO without restoring from snapshots.

Why this answer

Option A is correct because a cross-Region read replica continuously replicates data asynchronously from the primary RDS instance in us-east-1 to us-west-2, which can keep the DR copy within the 5-minute RPO requirement. Option D is correct because, during a disaster, the cross-Region read replica must be promoted to a standalone DB instance in us-west-2 so it can accept writes and become the new primary, enabling recovery within the 1-hour RTO. Option B is not correct because a 5-minute backup window does not guarantee a 5-minute RPO; automated backups are periodic snapshots and the backup window only defines when backups occur, not continuous replication.

Option C is not correct because cross-Region automated snapshot copy is periodic and typically cannot achieve a 5-minute RPO, and restoring from a snapshot would likely exceed the 1-hour RTO. Option E is not correct because Multi-AZ in us-east-1 only provides high availability within the same Region and does not protect against a Region-wide disaster or provide a DR copy in us-west-2.

Exam trap

The trap here is that candidates confuse Multi-AZ (which provides automatic failover within a Region) with cross-Region disaster recovery, or assume that automated backups or snapshot copies can meet a low RPO/RTO when they actually require time-consuming restore operations.

163
MCQeasy

A company wants to ensure that its application can recover from an Amazon S3 service disruption. The application reads and writes data to S3. Which strategy should the application implement to achieve resilience?

A.Store all data in a single S3 bucket with versioning enabled
B.Implement application logic to fall back to an S3 bucket in a different Region if the primary bucket is unavailable
C.Enable S3 Cross-Region Replication with automatic failover
D.Use S3 Transfer Acceleration to improve data transfer speed
AnswerB

This pattern gives the application explicit control over failover by first attempting to read from the primary bucket and, on failure (e.g., throttling, regional outage, or S3 service disruption), switching to a pre-created bucket in another Region. It is a common active-passive architecture that avoids reliance on any AWS feature providing automatic DNS-level or data-plane failover. Because the application itself detects the failure, it can also manage consistency, replication lag, and write buffering appropriately. This satisfies the recovery requirement because data availability is maintained as long as at least one Region is operational.

Why this answer

Implementing application logic to fall back to an S3 bucket in a different Region provides resilience against a regional S3 service disruption. S3 buckets are regional resources, so if one Region experiences an outage, the application can redirect reads and writes to a bucket in another Region. This approach requires the application to handle errors from the primary bucket and switch to the secondary bucket, ensuring continued availability without relying on automatic failover mechanisms that may not be instantaneous.

Exam trap

The trap here is that candidates often confuse S3 Cross-Region Replication (CRR) with automatic failover, but CRR is asynchronous and does not provide built-in failover; the application must still implement its own fallback logic to achieve resilience.

How to eliminate wrong answers

Option A is wrong because storing all data in a single S3 bucket with versioning enabled protects against accidental deletion or overwrite, but it does not provide resilience against a regional S3 service disruption, as the bucket is still tied to a single Region. Option C is wrong because S3 Cross-Region Replication (CRR) replicates objects asynchronously to another Region, but it does not include automatic failover; the application must still implement logic to detect the primary bucket's unavailability and switch to the replicated bucket. Option D is wrong because S3 Transfer Acceleration improves data transfer speed over long distances by using AWS edge locations, but it does not provide any resilience or failover capability during a regional S3 service disruption.

164
Multi-Selectmedium

A company is using AWS CloudFormation to deploy a critical application stack. The company wants to ensure that the stack can be recovered quickly in case of a failure. Which THREE strategies should the company implement? (Choose THREE.)

Select 3 answers
A.Disable rollback on stack creation failure to preserve resources for debugging.
B.Use StackSets to deploy the stack across multiple Regions.
C.Define the entire application in a single CloudFormation template.
D.Use nested stacks to separate components into reusable templates.
E.Use change sets to review changes before updating the stack.
AnswersB, D, E

StackSets enable multi-Region deployment for resilience.

Why this answer

AWS CloudFormation StackSets allow you to deploy stacks across multiple AWS Regions and accounts from a single template, enabling multi-Region disaster recovery. By deploying the critical application stack in multiple Regions, you can quickly fail over to a secondary Region if the primary fails, meeting the requirement for rapid recovery.

Exam trap

The trap here is that candidates often confuse 'recovery' with 'debugging' and select disabling rollback (Option A) thinking it helps preserve resources, but it actually hinders recovery by leaving failed resources in place.

165
MCQmedium

A company runs a serverless application using AWS Lambda functions that process messages from an Amazon SQS queue. The function scales up to handle high traffic but sometimes experiences throttling errors (HTTP 429) from Lambda. The company wants to improve the resilience of the application by reducing throttling. The SQS queue is configured as a Lambda event source with a batch size of 10. The Lambda function has a reserved concurrency of 100. Which combination of actions will best reduce throttling? (Choose the single best answer.)

A.Change the SQS queue to use a FIFO queue to guarantee exactly-once processing.
B.Increase the SQS batch size to 50 to process more messages per invocation.
C.Use a dead-letter queue (DLQ) for unprocessed messages and set up a CloudWatch alarm to trigger a second Lambda function to reprocess them.
D.Increase the Lambda function's reserved concurrency to 500.
AnswerD

Raising the Lambda function's reserved concurrency to 500 is correct because it directly increases the maximum number of simultaneous executions, allowing the SQS event source mapping to scale out beyond the previous limit and process more messages in parallel. With a higher concurrency ceiling, incoming SQS messages are consumed faster, preventing the burst of throttling attempts that occur when the function is already running at its current cap. This aligns with Lambda's SQS scaling model, where the number of active pollers grows with the message volume until the reserved concurrency is exhausted.

Why this answer

Throttling errors (HTTP 429) occur when Lambda function invocations exceed the account-level concurrency limit or the function's reserved concurrency. By increasing the reserved concurrency from 100 to 500, the function can handle more concurrent invocations, reducing the likelihood of throttling when traffic spikes. This directly addresses the scaling bottleneck without changing the event source or message processing pattern.

Exam trap

The trap here is that candidates often confuse throttling with message processing failures and choose a dead-letter queue or batch size change, but the core issue is insufficient concurrency allocation, which only reserved concurrency adjustment can fix.

How to eliminate wrong answers

Option A is wrong because changing to a FIFO queue does not affect concurrency or throttling; FIFO queues guarantee exactly-once processing and message ordering but do not increase the invocation capacity of Lambda. Option B is wrong because increasing the batch size to 50 may reduce the number of invocations but does not prevent throttling if the reserved concurrency is still too low; it could even cause timeouts or processing delays if messages accumulate. Option C is wrong because a dead-letter queue and a second Lambda function handle failed messages after throttling occurs, but they do not prevent the initial throttling errors; they add complexity without addressing the root cause of insufficient concurrency.

166
Multi-Selectmedium

A company is designing a resilient architecture for a web application that uses Amazon RDS for MySQL. The application must be able to withstand the loss of an entire AWS Region. Which TWO actions should the company take?

Select 2 answers
A.Use RDS Proxy to pool database connections.
B.Configure automated backups to be copied to another Region.
C.Enable Multi-AZ deployment for the RDS instance.
D.Create a Cross-Region Read Replica.
E.Enable deletion protection on the RDS instance.
AnswersB, D

Copying automated backups to another Region stores point-in-time snapshots outside the primary Region, so data survives regional failure. It satisfies the requirement to withstand loss of an entire AWS Region by enabling restore of the database in the unaffected Region.

Why this answer

To withstand the loss of an entire AWS Region, the company must have a disaster recovery strategy that includes cross-region data replication. Option B is correct because copying automated backups to another Region ensures that a recoverable copy of the database exists in a different geographic area, allowing restoration in a separate Region if the primary Region fails. Option D is correct because a Cross-Region Read Replica provides a live, asynchronously replicated copy of the database in another Region, which can be promoted to a standalone primary instance during a regional outage, minimizing recovery time.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which provides high availability within a Region) with cross-region disaster recovery, leading them to incorrectly select Multi-AZ as a solution for regional failure.

167
MCQmedium

A company uses a third-party backup solution to back up its EC2 instances daily. The backups are stored in an S3 bucket with default settings. The company wants to ensure that backups are protected from accidental deletion and are available for at least one year. Which combination of S3 features should the DevOps engineer implement?

A.Enable MFA Delete and set a lifecycle policy to transition to S3 Glacier after 30 days.
B.Enable versioning and set a lifecycle policy to expire noncurrent versions after 365 days.
C.Enable cross-Region replication to a bucket with versioning enabled.
D.Enable S3 Object Lock with Governance mode and a retention period of 365 days, and set a lifecycle policy to transition to S3 Glacier Deep Archive after 30 days.
AnswerD

S3 Object Lock with Governance mode applies a Write-Once-Read-Many (WORM) policy that guarantees the backup objects cannot be modified or deleted by any ordinary user—even an account administrator with full S3 access—until the 365-day retention period expires. Governance mode does allow users with the s3:BypassGovernanceRetention permission to override the lock if needed, but a typical backup scenario doesn't grant that to regular IAM roles, so the data remains immutable for the entire year. Pairing this with a lifecycle rule that transitions the objects to S3 Glacier Deep Archive after 30 days satisfies both the need for protection and cost efficiency: the transition preserves Object Lock metadata and the object stays protected through the transition, after which Storage Costs drop to the lowest tier while retention still applies for the remaining 335 days.

Why this answer

S3 Object Lock with Governance mode prevents objects from being deleted or overwritten by any user (including the root user) for the specified retention period of 365 days, meeting the one-year availability requirement. The lifecycle policy to transition to S3 Glacier Deep Archive after 30 days reduces storage costs while still keeping the data accessible for retrieval within 12 hours, which is acceptable for backup retention. This combination ensures immutability and cost-effective long-term storage.

Exam trap

The trap here is that candidates often confuse versioning with immutability, assuming that versioning alone prevents deletion, but versioning only creates multiple versions and does not prevent the current version from being deleted (it becomes a delete marker), whereas S3 Object Lock provides true immutability by preventing any deletion or overwrite during the retention period.

How to eliminate wrong answers

Option A is wrong because MFA Delete only protects against accidental deletion of objects and versioning suspension, but it does not enforce a minimum retention period or prevent overwrites, so backups could still be deleted after the MFA-authenticated action. Option B is wrong because versioning with expiration of noncurrent versions after 365 days does not prevent deletion of the current version; a user could delete the current version (which becomes a delete marker), and the noncurrent versions would expire after 365 days, but the data could be lost before that if the delete marker is not handled. Option C is wrong because cross-Region replication to a bucket with versioning enabled provides redundancy but does not protect against accidental deletion in the source bucket; if an object is deleted in the source, the replication delete marker is replicated, and the destination bucket may also lose the object unless additional safeguards like S3 Object Lock are used.

168
MCQmedium

A DevOps team is designing a disaster recovery plan for a production RDS for PostgreSQL database. The RPO must be less than 5 minutes and the RTO less than 1 hour. The database size is 2 TB. Which solution is MOST cost-effective?

A.Enable cross-Region automated backups with a retention period of 1 day
B.Take manual snapshots every 5 minutes and copy them to another Region
C.Use AWS Database Migration Service (DMS) for continuous replication to another Region
D.Create a cross-Region read replica and promote it during disaster
AnswerD

A cross-Region read replica uses asynchronous replication with lag typically under 5 minutes and can be promoted quickly, meeting both RPO and RTO cost-effectively.

Why this answer

The most cost-effective solution that meets the RPO < 5 minutes and RTO < 1 hour for a 2 TB RDS PostgreSQL database. A cross-Region read replica uses asynchronous replication, typically with lag of seconds, ensuring RPO well under 5 minutes. Promoting the replica to a standalone instance takes minutes, satisfying the RTO.

It leverages existing RDS features without additional services like DMS, and the replica instance can be sized smaller than the primary if not used, minimizing cost. Option A (cross-Region automated backups) only copies daily backups, resulting in RPO up to 24 hours, failing the requirement. Option B (manual snapshots every 5 minutes) is impractical and costly.

Option C (DMS continuous replication) meets RPO but incurs extra compute and data transfer costs, making it less cost-effective than a read replica.

Exam trap

Candidates often overlook that cross-Region automated backups do not include transaction logs for point-in-time recovery, so RPO can be up to 24 hours, not minutes.

169
Multi-Selectmedium

A company runs a critical web application on Amazon EC2 instances behind an Application Load Balancer (ALB) across multiple Availability Zones. The application stores session data in a shared Amazon ElastiCache for Redis cluster. The operations team reports that during a recent AZ failure, users experienced session loss and application errors. Which combination of actions should the company take to improve resilience and maintain session state during an AZ failure? (Choose TWO.)

Select 2 answers
A.Configure the ALB with cross-zone load balancing enabled and connection draining set to a suitable timeout.
B.Deploy an Auto Scaling group with a dynamic scaling policy that adds instances in the remaining AZs.
C.Enable cluster mode for the ElastiCache for Redis cluster and configure replica nodes in different Availability Zones.
D.Configure the application to use a custom DNS name with a low TTL pointing to the ElastiCache cluster endpoint.
E.Enable Multi-AZ for the ElastiCache cluster to automatically fail over to a replica in another AZ.
AnswersA, C

Cross-zone load balancing on the ALB ensures that incoming traffic is distributed evenly across all registered targets in every Availability Zone, preventing any single AZ from being overloaded and allowing the ALB to continue serving requests even if one AZ is impaired. Connection draining gives in-flight requests a grace period to complete before an instance is deregistered or replaced, avoiding request interruption during rolling updates or failed health checks. Together, these features support seamless instance replacement without dropping active requests, though they do not on their own preserve stored session data — they protect the connection lifecycle while the application layer (e.g., ElastiCache) handles state.

Why this answer

Enabling cross-zone load balancing on the ALB ensures traffic is distributed evenly across all EC2 instances in all AZs, and connection draining with a suitable timeout allows in-flight requests to complete before instances are deregistered, preventing session loss during an AZ failure. Option C is correct because enabling cluster mode for ElastiCache for Redis with replica nodes in different AZs provides automatic sharding and replication, ensuring session data remains available and consistent even if a primary node in one AZ fails. Option E is incorrect because while ElastiCache for Redis supports Multi-AZ with automatic failover, it alone does not guarantee that replica nodes are placed in different Availability Zones for each shard; enabling cluster mode with replicas in different AZs (Option C) provides a more comprehensive solution for maintaining session state during an AZ failure.

Exam trap

Candidates may choose Multi-AZ (Option E) thinking it provides cross-AZ failover for ElastiCache for Redis, which is true. However, Multi-AZ with automatic failover requires replication groups with replicas in different AZs. In a cluster-mode setup, you must explicitly ensure replicas are in different AZs per shard.

Option C directly addresses this by enabling cluster mode and configuring replica nodes in different AZs, making Option C a more complete solution for the given scenario of a shared cluster.

170
MCQmedium

A company runs a stateful application on EC2 instances. The application stores session data locally. The instances are behind an ALB with sticky sessions enabled. A scaling event terminates an instance, causing loss of session data. How can the company prevent this while maintaining performance?

A.Use Amazon ElastiCache to store session data
B.Use a dedicated EC2 instance for sessions
C.Disable sticky sessions
D.Increase the sticky session duration
AnswerA

Storing sessions in ElastiCache moves session state off the instance into a shared, low-latency in-memory tier, so any instance behind the ALB can serve subsequent requests and instance termination no longer destroys session data, while preserving performance.

Why this answer

Storing session data in Amazon ElastiCache (Redis or Memcached) externalizes the state from the EC2 instances, so when an instance is terminated during scaling, the session data persists and remains accessible to all instances. This maintains performance because ElastiCache provides sub-millisecond latency and high throughput, and it also allows any instance to handle any request, eliminating the need for sticky sessions. The application must be modified to read/write session data to ElastiCache, but this is the standard AWS-recommended approach for stateful web applications.

Exam trap

DOP-C02 often tests the misconception that sticky sessions alone can solve session persistence, but they only tie a user to an instance; they do not prevent data loss when that instance is terminated. Candidates may also think that increasing session duration or using a dedicated instance is sufficient, but these are not scalable or highly available solutions.

How to eliminate wrong answers

Option B is wrong because a dedicated EC2 instance for sessions still represents a single point of failure and does not solve the scaling problem—if that instance fails, all sessions are lost, and it also becomes a performance bottleneck. Option C is wrong because disabling sticky sessions without externalizing session state would cause session data to be lost or inconsistent across instances, as each instance would have its own local session store. Option D is wrong because increasing the sticky session duration only delays the problem; it does not prevent session loss when an instance is terminated, and it can lead to uneven load distribution.

171
Multi-Selectmedium

A company is building a multi-tier web application on AWS. The application must be resilient to the failure of an entire Availability Zone. The architecture includes an Application Load Balancer (ALB), EC2 instances in an Auto Scaling group, and an Amazon RDS for MySQL database. Which TWO actions should be taken to achieve this resilience? (Choose two.)

Select 2 answers
A.Configure an RDS read replica in a different Availability Zone.
B.Use a Single-AZ RDS for MySQL database to keep costs low.
C.Place all EC2 instances in the same Availability Zone to reduce cross-AZ data transfer costs.
D.Configure the Auto Scaling group to launch EC2 instances in at least two Availability Zones.
E.Deploy the RDS for MySQL database in a Multi-AZ configuration.
AnswersD, E

Spreading EC2 instances across at least two Availability Zones means an AZ outage leaves capacity in the surviving zone, so the Auto Scaling group keeps serving traffic behind the ALB. This directly satisfies the requirement to survive failure of an entire Availability Zone.

Why this answer

Option D is correct because an Auto Scaling group configured to span at least two Availability Zones ensures that if one AZ fails, the remaining AZ can continue serving traffic behind the ALB, maintaining compute capacity and application availability. Option E is correct because an RDS for MySQL Multi-AZ deployment maintains a synchronous standby replica in a different AZ and automatically fails over to it if the primary AZ becomes unavailable, providing database resilience. Option A is incorrect because a read replica is asynchronous and is intended for read scaling, not automatic failover for high availability.

Option B is incorrect because a Single-AZ RDS database has no standby in another AZ and cannot survive an AZ failure. Option C is incorrect because placing all EC2 instances in one AZ concentrates the workload in a single failure domain, defeating the resilience requirement.

Exam trap

The trap here is that candidates often confuse read replicas (asynchronous, for read scaling) with Multi-AZ deployments (synchronous, for high availability), and mistakenly think placing all resources in one AZ reduces costs without recognizing the critical single point of failure it introduces.

172
Multi-Selecthard

A company runs a microservices architecture on Amazon ECS. They want to ensure that if a service fails, it does not cascade to other services. Which TWO design patterns should they implement?

Select 2 answers
A.Cache-aside pattern
B.Saga pattern
C.Circuit breaker pattern
D.Throttling pattern
E.Bulkhead pattern
AnswersC, E

Circuit breaker pattern monitors calls to a remote service and maintains three states—closed, open, and half-open—progressing to open when failure thresholds are exceeded, at which point subsequent calls fail fast without attempting the network operation. This prevents a failing service from being overwhelmed and stops the same repeated errors from saturating caller resources, thereby breaking the chain of cascading failures and giving the dependency time to recover.

Why this answer

The circuit breaker pattern (C) is correct because it wraps calls to a failing downstream microservice and, after a threshold of failures, trips open to fail fast instead of repeatedly invoking the unhealthy service, preventing a localized failure from cascading through the call chain. The bulkhead pattern (E) is correct because it isolates resources such as thread pools, connection pools, or ECS tasks per service or dependency, so exhaustion or failure in one component cannot consume shared capacity and take down other services. Together these patterns directly address fault isolation and cascade prevention in a microservices architecture on Amazon ECS.

The cache-aside pattern (A) only improves read performance and reduces backend load by populating a cache on demand; it does not stop failure propagation. The Saga pattern (B) manages distributed transaction consistency across services via compensating actions, which is about data integrity rather than preventing cascading failures. The throttling pattern (D) limits request rates to protect a service from overload, but it is a rate-control mechanism and not the primary isolation pattern for stopping cross-service failure cascades.

173
MCQhard

A company uses AWS CloudFormation to deploy a multi-tier application. The stack includes an RDS DB instance with Multi-AZ enabled. The database experiences a failover during maintenance. The application reports connection errors for several minutes. What is the MOST likely cause and solution?

A.The RDS failover took longer than expected; increase the Multi-AZ timeout
B.The read replica was promoted incorrectly; recreate the read replica
C.The RDS proxy is misconfigured; disable the proxy for Multi-AZ
D.The application does not implement connection retry logic; implement exponential backoff and retry
AnswerD

During an RDS Multi-AZ failover, the existing TCP connections to the primary instance are forcibly closed, and the database endpoint's DNS CNAME is updated to point to the new primary. If the application does not implement connection retry with exponential backoff, it will never re-establish those dropped connections, causing an outage even after the failover completes. Retry logic with backoff and jitter is the standard, required practice for any production application using RDS, especially when combined with a connection pool or RDS Proxy.

Why this answer

During an RDS Multi-AZ failover, the DNS CNAME for the DB instance is updated to point to the standby, and existing connections are dropped. Applications that do not implement connection retry logic will continue to fail until they reconnect, causing several minutes of errors. Implementing exponential backoff and retry logic allows the application to reconnect after the failover completes.

Exam trap

DOP-C02 often tests whether candidates blame the AWS service (failover time, read replica, proxy) instead of recognizing that application-level retry logic is required to handle transient connection failures during failover.

How to eliminate wrong answers

Option A is wrong because there is no 'Multi-AZ timeout' parameter to increase; failover time is determined by AWS and is typically 60-120 seconds, and the issue is application reconnection, not the failover duration. Option B is wrong because Multi-AZ failover promotes the standby, not a read replica, and recreating a read replica does not address connection errors during failover. Option C is wrong because RDS Proxy is not misconfigured by default and disabling it would not fix the issue; in fact, RDS Proxy can help by masking failovers, but the root cause is missing retry logic.

174
MCQeasy

A company uses Amazon Route 53 for DNS and wants to ensure high availability for a web application hosted on two EC2 instances in different Availability Zones. The application uses an Application Load Balancer. What is the simplest way to achieve resilience if one Availability Zone becomes unavailable?

A.Launch both instances in the same Availability Zone.
B.Place the instances in different Availability Zones behind the ALB.
C.Configure Route 53 latency-based routing to each instance.
D.Configure Route 53 failover routing with health checks pointing to each instance.
AnswerB

An Application Load Balancer is a regional, layer-7 service that can route traffic to healthy targets across multiple Availability Zones. By placing EC2 instances in different AZs and registering them in the ALB's target group, the ALB automatically performs health checks and reroutes traffic away from unhealthy instances or entirely failed AZs. This design ensures that if one AZ becomes unavailable, the ALB continues to serve requests using instances in the other AZ. The combination of cross-zone load balancing and health-based target removal makes this the industry-standard resilient architecture.

Why this answer

Placing EC2 instances in different Availability Zones behind an Application Load Balancer (ALB) is the simplest and most effective way to achieve high availability. The ALB automatically distributes traffic across healthy targets in multiple AZs, and if one AZ becomes unavailable, the ALB stops routing requests to instances in that AZ, ensuring continued service from the remaining AZ.

Exam trap

The trap here is that candidates overcomplicate the solution by choosing DNS-level failover (Option D) when the ALB already provides built-in cross-AZ failover, making the simpler architecture the correct answer.

How to eliminate wrong answers

Option A is wrong because launching both instances in the same Availability Zone creates a single point of failure; if that AZ goes down, the application becomes completely unavailable. Option C is wrong because Route 53 latency-based routing directs traffic based on lowest latency, not availability; it does not automatically failover when an AZ becomes unavailable, and it requires additional health check configuration to be effective. Option D is wrong because Route 53 failover routing with health checks is more complex than necessary; the ALB already provides health checks and automatic failover across AZs, making this an overly complicated solution that adds unnecessary DNS-level complexity.

175
MCQeasy

A company wants to ensure its data in Amazon S3 is protected against accidental deletion. The bucket stores critical documents. Which approach provides the HIGHEST level of resilience?

A.Apply a bucket policy that denies s3:DeleteObject for all users.
B.Enable S3 lifecycle policies to archive objects to Glacier.
C.Enable versioning and MFA delete on the bucket.
D.Configure cross-region replication (CRR) to another bucket.
AnswerC

Enabling versioning on the bucket creates a new version for every object put, overwrite, or delete, so a delete operation only inserts a delete marker that hides old versions rather than purging them; with versioning enabled, you can permanently recover any object version. MFA Delete fortifies this by requiring an MFA code (using root credentials) to permanently delete an object version, suspend versioning, or change the versioning state—thus preventing attackers or accidental admin actions from irreversibly destroying data.

Why this answer

Enabling versioning and MFA delete provides protection against both accidental overwrites and malicious deletions. Versioning allows recovery of deleted or overwritten objects, while MFA delete adds an extra layer of security by requiring multi-factor authentication for permanent deletions. Option A is incorrect because a bucket policy that denies s3:DeleteObject can prevent deletions but does not allow recovery if the policy is bypassed or changed.

Option B is incorrect because lifecycle policies archive objects to Glacier, which reduces costs but does not prevent or recover from accidental deletion. Option D is incorrect because cross-region replication protects against regional failures but does not protect against accidental deletion within the source bucket.

176
MCQhard

A company runs a containerized microservices application on Amazon EKS. The application includes a critical service that processes real-time financial transactions. This service must be highly available and resilient to node failures. The current setup uses a Deployment with 3 replicas and a ClusterIP service. During a recent node failure, the application experienced a brief period of unavailability. Which action should the DevOps engineer take to improve resilience without changing the underlying infrastructure?

A.Change the service type from ClusterIP to NodePort and configure an external load balancer.
B.Increase the number of replicas to 10 and use a node selector to schedule all pods on the largest instance type.
C.Configure a PodDisruptionBudget with a maxUnavailable of 1, and add pod anti-affinity rules to spread pods across different nodes.
D.Enable HorizontalPodAutoscaler with a target CPU utilization of 50% to automatically scale the Deployment.
AnswerC

A PodDisruptionBudget with maxUnavailable:1 guarantees that at most one Pod is unavailable during voluntary evictions such as node drains, and pod anti-affinity rules (preferably with topologyKey kubernetes.io/hostname) force the scheduler to place replicas on distinct nodes. This means an involuntary node failure can kill only one replica, and the remaining replicas continue to serve traffic. Combined, these mechanisms directly address both failure classes—involuntary hardware failures and voluntary maintenance—by ensuring the application always has at least N-1 replicas available across different failure domains.

Why this answer

A PodDisruptionBudget with maxUnavailable=1 ensures that at most one pod is unavailable during voluntary disruptions, while pod anti-affinity rules force the scheduler to distribute pods across different nodes. This combination prevents a single node failure from taking down all replicas, maintaining service availability without altering the underlying infrastructure.

Exam trap

The trap here is that candidates often confuse scaling (HPA or more replicas) with resilience, failing to realize that without proper pod distribution and disruption budgets, scaling alone cannot prevent downtime from node failures.

How to eliminate wrong answers

Option A is wrong because changing to NodePort with an external load balancer adds network complexity and does not address pod distribution or node failure resilience; the ClusterIP service already provides internal load balancing. Option B is wrong because increasing replicas to 10 and using node selector to pin pods to the largest instance type actually reduces resilience by creating a single point of failure on that node. Option D is wrong because HorizontalPodAutoscaler scales based on CPU utilization, which does not protect against node failures; it may even exacerbate the problem by scaling pods onto the same failing nodes.

177
Multi-Selecthard

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application generates logs that are sent to Amazon CloudWatch Logs. The DevOps team needs to configure a metric filter to monitor for error patterns and trigger an alarm when the error rate exceeds 5% of total requests over a 5-minute period. Which TWO steps should the team take? (Choose TWO.)

Select 2 answers
A.Create a CloudWatch Logs log group for the error metric.
B.Create a metric filter on the log group to count occurrences of the error pattern.
C.Create a CloudWatch Logs subscription filter to stream errors to a Lambda function that calculates the error rate.
D.Create a CloudWatch metric for the error count.
E.Create a CloudWatch alarm that uses a math expression to calculate the error rate (error count / total request count) and compare it to the threshold of 5%.
AnswersB, E

CloudWatch metric filters scan log events for a pattern and publish a numeric metric each time it matches. Counting error-pattern occurrences produces the error count metric that the subsequent math expression and alarm consume to compute the error rate.

Why this answer

Option B is correct because a CloudWatch Logs metric filter is the mechanism that scans log events in a log group for a specified pattern (for example, "ERROR") and increments a custom metric each time the pattern matches, which is exactly how the error count is derived from the application logs. Option E is correct because the requirement is an error rate (a percentage), not a raw count, so the alarm must use a CloudWatch metric math expression that divides the error-count metric by the total-request-count metric and compares the result against the 5% threshold over the 5-minute evaluation period. Option A is not needed because metric filters are applied to an existing log group that already receives the application logs; you do not create a separate log group for the metric.

Option C is wrong because subscription filters stream log data to destinations such as Lambda or Kinesis for real-time processing, which is unnecessary when CloudWatch metric filters and metric math can compute the rate natively. Option D is wrong because the metric is created automatically by the metric filter defined in Option B, so a separate manual metric creation step is redundant.

Exam trap

DOP-C02 often tests the difference between metric filters (which auto-publish metrics from log patterns) and manually created metrics or subscription filters, catching candidates who think they must create the metric separately or route logs through Lambda.

178
Drag & Dropmedium

Drag and drop the steps to perform a disaster recovery failover from a primary region to a secondary region using AWS Route 53 and RDS.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First configure health checks, then lower TTL, then promote RDS, then update DNS, then verify.

179
MCQeasy

A company is designing a resilient architecture for a web application using AWS Global Accelerator and two Application Load Balancers in different AWS Regions. The application is stateless and uses a global DynamoDB table for data. What is the primary benefit of using Global Accelerator in this architecture?

A.It replaces the need for an Application Load Balancer.
B.It provides static IP addresses and automatically routes traffic to the closest healthy ALB, improving availability and performance.
C.It provides DNS-based failover between Regions.
D.It caches static content at AWS edge locations.
AnswerB

Global Accelerator uses two static Anycast IP addresses at AWS edge locations, which are announced from multiple points of presence. It monitors the health of your ALB or NLB endpoints and automatically routes each user's traffic to the closest endpoint that is healthy, using the AWS backbone instead of the public internet. This improves availability by failing over to a healthy endpoint within seconds and improves performance by reducing latency and jitter, but it still depends on the ALB to handle the actual HTTP/HTTPS load balancing.

Why this answer

Global Accelerator provides static IP addresses and directs traffic to the nearest healthy endpoint, improving resilience and performance. Option A is wrong because Global Accelerator does not cache content. Option C is wrong because DNS routing is not the primary benefit; Global Accelerator uses anycast.

Option D is wrong because Global Accelerator does not replace ALB; it works with ALBs.

180
MCQhard

A company uses AWS Lambda with Amazon DynamoDB to process orders. During peak hours, the Lambda function sometimes fails with throttling errors from DynamoDB. The system must be resilient and cost-effective. What should a DevOps engineer do?

A.Use Amazon SQS to buffer the requests and have Lambda pull from the queue with a reserved concurrency limit.
B.Increase the DynamoDB provisioned read and write capacity units to a high fixed value.
C.Provision DynamoDB Accelerator (DAX) to cache reads and reduce throttling.
D.Configure DynamoDB auto scaling and implement a dead-letter queue in Lambda to retry failed events.
AnswerD

DynamoDB auto scaling adjusts provisioned capacity based on actual usage, preventing most throttling, but it cannot anticipate sudden one-off spikes because it relies on trends. A Lambda dead-letter queue, combined with the function's built-in retries and exponential backoff, ensures that any event which still fails due to a throttle is safely captured for manual or automated replay rather than silently dropped. This two-tier approach balances elasticity with data durability, which is why it is the recommended solution for unpredictable write spikes.

Why this answer

Configuring DynamoDB auto scaling allows the table to adjust its provisioned capacity based on actual traffic patterns, preventing throttling during peak hours while remaining cost-effective during low usage. Implementing a dead-letter queue (DLQ) in Lambda ensures that failed events (e.g., due to transient throttling) are captured and can be retried or investigated, providing resilience without manual intervention.

Exam trap

The trap here is that candidates may confuse read caching solutions (DAX) or queue-based decoupling (SQS) with the direct need to scale write capacity and handle retries, overlooking the combination of auto scaling and DLQ as the most resilient and cost-effective approach for write-throttling scenarios.

How to eliminate wrong answers

Option A is wrong because using Amazon SQS to buffer requests and having Lambda pull from the queue with a reserved concurrency limit does not directly address DynamoDB throttling; it only controls Lambda concurrency, not the underlying DynamoDB capacity, and could still result in throttling if the database cannot handle the aggregate write volume. Option B is wrong because increasing DynamoDB provisioned read and write capacity units to a high fixed value is not cost-effective; it leads to over-provisioning during off-peak hours and does not adapt to variable traffic, contradicting the requirement for a cost-effective solution. Option C is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for read operations only; it does not mitigate write throttling errors, which are the primary issue described in the scenario.

181
MCQeasy

A DevOps engineer needs to ensure that an application running on EC2 can automatically recover from an underlying hardware failure without manual intervention. Which AWS feature should be enabled?

A.Enable termination protection
B.Configure EC2 Auto Recovery with a CloudWatch alarm
C.Configure an Auto Scaling group with a minimum size of 1
D.Enable CloudWatch detailed monitoring
AnswerB

EC2 Auto Recovery uses a CloudWatch alarm on the StatusCheckFailed_System metric to detect underlying host or network failures. When the alarm triggers, EC2 automatically relaunches the same instance on a new physical host while preserving the instance ID, private and Elastic IP addresses, and attached EBS volumes. This in-place recovery requires the instance to be EBS-backed and in a VPC, making it the correct choice for restoring availability after hardware failure.

Why this answer

EC2 Auto Recovery automatically recovers an instance when it becomes impaired due to underlying hardware failure, preserving its instance ID, private IP, and Elastic IP. Option A is incorrect because termination protection only prevents accidental deletion, not recovery. Option C is incorrect because an Auto Scaling group with a minimum size of 1 can replace a failed instance but launches a new instance rather than recovering the same one.

Option D is incorrect because CloudWatch detailed monitoring provides more frequent metrics but does not trigger recovery actions.

182
MCQmedium

A company has a production environment that uses Amazon Route 53 for DNS and an Application Load Balancer (ALB) to distribute traffic to EC2 instances. The company wants to implement a disaster recovery plan that automatically fails over to a secondary region in case the primary region becomes unavailable. Which configuration should be used?

A.Use Route 53 weighted routing policy with equal weights for both regions.
B.Use Route 53 geolocation routing policy to route users based on their location.
C.Use Route 53 failover routing policy with primary and secondary records and health checks.
D.Use Route 53 latency routing policy to route to the region with lowest latency.
AnswerC

Failover routing is the Route 53 policy explicitly designed for active-passive disaster recovery. It lets you mark one record (or record set within a group) as the primary and another as secondary, and associate each with a health check. Route 53 monitors the primary endpoint's health and, when the health check fails after the configured threshold, automatically returns the secondary record's answer for all DNS queries. This deterministic, health-driven promotion of the standby region matches the company's requirement for automatic failover.

Why this answer

Route 53 failover routing is purpose-built for active-passive DR: you designate a PRIMARY record and a SECONDARY record, attach a health check to the primary, and Route 53 automatically returns the secondary record's value when the primary health check fails. This gives deterministic, automatic regional failover without manual DNS changes. The other policies distribute traffic but do not implement a primary/secondary failover decision.

Exam trap

DOP-C02 often tests whether candidates confuse traffic-distribution policies (weighted, latency, geolocation) with true active-passive failover — only failover routing with health checks provides automatic primary/secondary switching.

How to eliminate wrong answers

Option A is wrong because weighted routing distributes traffic proportionally by weight and has no concept of a primary/secondary relationship — equal weights would send ~50% of traffic to the DR region permanently, not fail over on failure. Option B is wrong because geolocation routing selects answers based on the resolver's geographic location, which is for localization/compliance, not health-based failover. Option D is wrong because latency routing picks the lowest-latency region and would keep sending traffic to a degraded region as long as it responds, and it has no primary/secondary semantics.

183
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer. To improve resilience, they want to automatically replace failed instances. Which AWS service should they use?

A.EC2 Instance Recovery
B.AWS Systems Manager Automation
C.CloudFormation Stack update
D.Auto Scaling group with health checks
AnswerD

An Auto Scaling group with health checks automatically replaces failed EC2 instances by monitoring them using EC2 status checks and optionally Elastic Load Balancing health checks. When an instance is marked unhealthy, the ASC terminates it and launches a new instance in an attempt to maintain the desired capacity, and it does so across multiple Availability Zones to increase resilience. This meets the requirement for automatic replacement based on instance health, and because it is a continuous, managed process, it is the correct solution for keeping critical applications available.

Why this answer

Auto Scaling groups with health checks automatically replace unhealthy instances based on ELB health check integration. Option D is correct. Option A (EC2 Instance Recovery) recovers instances on the same host but does not replace them if the host fails.

Option B (Systems Manager Automation) requires manual or scheduled automation, not automatic health-based replacement. Option C (CloudFormation Stack update) does not provide auto-replacement; it manages infrastructure updates.

184
MCQeasy

A company runs a web application on EC2 instances behind an Application Load Balancer. The application experiences intermittent failures due to a single Availability Zone failing. Which solution is MOST resilient and cost-effective?

A.Use a larger instance type in the same Availability Zone
B.Use an Auto Scaling group with a single instance in each of three Availability Zones and a Network Load Balancer
C.Migrate to a single larger instance in a different region
D.Deploy EC2 instances across two Availability Zones and configure the ALB to distribute traffic
AnswerD

Deploying EC2 instances across two Availability Zones behind an Application Load Balancer provides both fault isolation and active load balancing. The ALB performs health checks, routes traffic only to healthy instances, and can distribute requests across both AZs, so if one entire AZ becomes unavailable, the other continues to serve traffic. This is the standard architecture for achieving regional high availability with an HTTP web application.

Why this answer

Most resilient and cost-effective. Deploying EC2 instances across two Availability Zones and configuring the ALB to distribute traffic ensures high availability by handling a single AZ failure without over-provisioning. Option A is wrong because using a larger instance in the same AZ does not address AZ failure.

Option B is wrong because using three AZs with a single instance each and a Network Load Balancer is more expensive and not necessary for this scenario; ALB already supports cross-zone load balancing. Option C is wrong because migrating to a different region adds latency, complexity, and cost.

← PreviousPage 3 of 3 · 184 questions total

Ready to test yourself?

Try a timed practice session using only Resilient Cloud Solutions questions.