Courseiva

CCNA Resilient Cloud Solutions Questions

75 of 184 questions · Page 2/3 · Resilient Cloud Solutions · Answers revealed

76
MCQmedium

An e-commerce platform uses Amazon DynamoDB as its primary database. The platform experiences occasional read throttling during flash sales. The operations team needs to ensure that read traffic is handled without errors, while keeping costs low. What should a DevOps engineer recommend?

A.Enable DynamoDB Accelerator (DAX) to cache frequently read data.
B.Increase the read capacity units for the table during flash sale events.
C.Use DynamoDB Streams to replicate reads to a separate table.
D.Implement Global Tables to distribute read traffic across multiple regions.
AnswerA

DAX is a fully managed, in-memory caching service placed in front of DynamoDB, returning cached items with microsecond latency. By writing through and caching the frequently read flash-sale items, it absorbs the burst of read traffic before it reaches the table, which directly reduces consumed read capacity units and throttling events. This requires no costly rearchitecture or constantly adjusting provisioned throughput, making it the most appropriate solution for unpredictable read spikes.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that reduces read load on the table by serving repeated read requests from cache. During flash sales, read traffic spikes on popular items; DAX absorbs these reads, preventing throttling and improving latency, while keeping costs low because it reduces the need to over-provision read capacity. This is the most cost-effective solution for read-heavy, repetitive access patterns.

Exam trap

The trap is thinking that increasing RCUs or using Global Tables is the best fix for read throttling, when the question emphasizes cost and handling read traffic without errors; DAX is the purpose-built caching solution.

How to eliminate wrong answers

Option B is wrong because increasing read capacity units (RCUs) during flash sales is a manual, reactive approach that can be costly if over-provisioned and still may not handle sudden spikes quickly enough; it also does not reduce costs. Option C is wrong because DynamoDB Streams capture item-level changes for replication or triggers, not for serving reads, and replicating to another table does not offload read traffic from the original table. Option D is wrong because Global Tables are for multi-region active-active replication, which increases cost and complexity, and does not directly solve read throttling in a single region unless reads are distributed across regions, which may introduce latency.

77
MCQeasy

A company is deploying a critical application on Amazon EC2 instances behind an Application Load Balancer (ALB) across multiple Availability Zones. The application must be resilient to the failure of an entire Availability Zone. Which design should the company implement?

A.Launch EC2 instances in at least two Availability Zones and place them behind an Application Load Balancer with cross-zone load balancing enabled.
B.Use one EC2 instance in a single Availability Zone behind a Network Load Balancer.
C.Launch EC2 instances in one Availability Zone and use an Application Load Balancer to distribute traffic.
D.Deploy EC2 instances in two Availability Zones but use a single Application Load Balancer in one AZ.
AnswerA

This is the correct approach. An Application Load Balancer is a regional service; by enabling subnets in at least two Availability Zones, you create redundant ALB nodes. The ALB performs health checks and automatically routes traffic to healthy EC2 instances across both AZs. Cross-zone load balancing ensures each instance receives an equal share of requests, so the architecture tolerates an AZ failure and even an instance failure without manual intervention.

Why this answer

Deploying EC2 instances across at least two Availability Zones (AZs) behind an Application Load Balancer (ALB) with cross-zone load balancing enabled ensures that if an entire AZ fails, the ALB can route traffic to healthy instances in the remaining AZs. Cross-zone load balancing allows the ALB to distribute incoming requests evenly across all registered instances in all enabled AZs, which improves fault tolerance and resource utilization. This design meets the requirement for resilience to an AZ failure by eliminating a single point of failure at the AZ level.

Exam trap

The trap here is that candidates often assume that simply placing instances in multiple AZs behind a load balancer is sufficient, but they overlook the critical requirement that the load balancer itself must be deployed across multiple AZs to avoid being a single point of failure.

How to eliminate wrong answers

Option B is wrong because using a single EC2 instance in one AZ behind a Network Load Balancer (NLB) does not provide resilience to an AZ failure; if that AZ goes down, the application becomes unavailable. Option C is wrong because launching EC2 instances in only one AZ behind an ALB still creates a single point of failure at the AZ level; the ALB cannot route traffic to healthy instances if the entire AZ fails. Option D is wrong because deploying EC2 instances in two AZs but using a single ALB in one AZ means the ALB itself is a single point of failure; if that AZ fails, the ALB becomes unavailable, and traffic cannot be distributed to instances in the other AZ.

78
MCQmedium

A DevOps engineer is designing a multi-Region active-active architecture for a stateless web application using Route 53 latency-based routing and DynamoDB global tables. The application must continue to serve traffic even if an entire AWS Region becomes unavailable. Which additional step is MOST critical for resilience?

A.Use an Auto Scaling group with a scheduled scaling policy
B.Enable DynamoDB Accelerator (DAX) in each Region
C.Place a CloudFront distribution in front of the application
D.Configure Route 53 health checks and associate them with the latency records
AnswerD

Route 53 health checks associated with the latency records let DNS stop returning the failed Region's endpoints, redirecting users to a healthy Region. Without this, latency records would continue resolving to the unavailable Region, defeating the active-active resilience requirement.

Why this answer

In a multi-Region active-active architecture using Route 53 latency-based routing, the routing policy alone does not detect or react to a Region failure. Route 53 health checks must be created and associated with each latency record so that unhealthy endpoints are removed from DNS responses. Without health checks, Route 53 would continue directing traffic to a failed Region, defeating the resilience goal.

Exam trap

The trap is that candidates assume latency-based routing automatically fails over when a Region goes down, when in fact Route 53 requires explicitly configured health checks to remove unhealthy endpoints from DNS responses.

How to eliminate wrong answers

Option A is wrong because scheduled scaling adjusts capacity based on time, not health or demand, and does nothing to redirect traffic away from a failed Region. Option B is wrong because DAX is an in-memory cache for DynamoDB that improves read latency but does not provide cross-Region failover or traffic redirection. Option C is wrong because CloudFront caches content at edge locations but does not perform origin health-based failover across Regions unless paired with origin groups and health checks — and it does not replace Route 53 health checks for latency-based routing.

79
MCQhard

An organization runs a critical application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application requires that all traffic be encrypted in transit. The security team mandates the use of TLS 1.2 or higher and specific ciphers. What is the MOST efficient way to enforce this requirement?

A.Use a Network Load Balancer with TLS listeners and target groups.
B.Place a CloudFront distribution in front of the ALB and configure the origin protocol policy.
C.Install a self-signed certificate on each EC2 instance and configure the web server.
D.Configure the ALB with a security policy that enforces TLS 1.2 and the required ciphers.
AnswerD

ALB security policies terminate TLS at the load balancer and enforce the minimum protocol version and cipher suite list, satisfying the TLS 1.2 requirement without modifying each EC2 instance. This centralises enforcement at the single ingress point.

Why this answer

The Application Load Balancer (ALB) supports predefined security policies that enforce TLS version and cipher requirements at the load balancer level, providing centralized control. Option D is correct because configuring the ALB with a security policy that mandates TLS 1.2 and specific ciphers meets the requirement efficiently without modifying individual instances. Option A is incorrect because a Network Load Balancer (NLB) with TLS listeners does not offer the same granular security policy options as ALB and is less efficient for this requirement.

Option B is incorrect because CloudFront is a CDN service; placing it in front of the ALB does not directly enforce TLS settings on the ALB itself and adds unnecessary complexity. Option C is incorrect because installing self-signed certificates on each EC2 instance is inefficient, not centralized, and does not enforce consistent TLS version and cipher requirements across all traffic.

80
Multi-Selecteasy

A company wants to design a highly available and fault-tolerant architecture for a stateless web application on AWS. Which TWO actions should they take? (Choose two.)

Select 2 answers
A.Use a single large EC2 instance to simplify management
B.Deploy multiple Application Load Balancers in each AZ
C.Launch EC2 instances in at least two Availability Zones
D.Use an RDS Multi-AZ deployment for the web server fleet
E.Use an Auto Scaling group to replace failed instances automatically
AnswersC, E

Launching EC2 instances in at least two Availability Zones is the foundational principle for highly available web server fleets. Availability Zones are physically separate data centers with independent power, cooling, and networking, so a failure in one AZ does not affect the other. Combined with a load balancer that spans those same AZs, traffic automatically continues to be served by healthy instances if one AZ loses capacity. This design eliminates the single-AZ dependency and is a core requirement for fault-tolerant compute tiers.

Why this answer

To achieve high availability and fault tolerance for a stateless web application, you should deploy EC2 instances in at least two Availability Zones (C) to eliminate a single point of failure, and use an Auto Scaling group (E) to automatically replace failed instances and maintain desired capacity. Option A is incorrect because a single large instance is a single point of failure and does not provide fault tolerance. Option B is incorrect because multiple Application Load Balancers per AZ are unnecessary; a single ALB can route traffic across multiple AZs.

Option D is incorrect because RDS Multi-AZ is a database feature, not for the web server fleet.

81
MCQmedium

A company runs a stateful application on EC2 instances. They want to distribute traffic evenly and maintain session stickiness. Which AWS service should they use?

A.Network Load Balancer
B.Application Load Balancer with sticky sessions
C.Amazon Route 53 weighted routing policy
D.Amazon CloudFront with origin failover
AnswerB

An Application Load Balancer with sticky sessions is the correct solution because it operates at Layer 7 and can inspect HTTP/HTTPS traffic. ALB uses either the AWSALB cookie with a configurable duration or an application-controlled cookie to send all requests from the same client session to the same EC2 instance. This preserves in-memory session state, which is exactly what a stateful application requires.

Why this answer

An Application Load Balancer with sticky sessions (session affinity) distributes HTTP/HTTPS traffic across targets while binding a client to the same target for the duration of a session via a cookie. This satisfies both even distribution and session persistence for a stateful application.

Exam trap

DOP-C02 often tests whether candidates know that only ALB provides cookie-based sticky sessions; picking NLB because it is 'high performance' ignores the session-affinity requirement.

How to eliminate wrong answers

Option A is wrong because Network Load Balancer operates at Layer 4 and, while it supports source-IP affinity, it does not offer cookie-based sticky sessions and is not ideal for HTTP-layer session state. Option C is wrong because Route 53 weighted routing distributes DNS answers across endpoints but provides no session stickiness and is not a load balancer. Option D is wrong because CloudFront with origin failover is a CDN feature for content delivery and origin redundancy, not session-aware load balancing across EC2 targets.

82
MCQhard

A company runs a critical web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application frequently experiences high latency during peak hours. The DevOps team needs to implement a solution that automatically adds capacity based on demand and reduces cost during off-peak hours. Which combination of AWS services should the team use?

A.Use an AWS Auto Scaling group with scheduled scaling policies that add instances during known peak hours and remove them during off-peak hours.
B.Implement Amazon Route 53 weighted routing policies to distribute traffic to multiple ALBs, each fronting a fixed set of EC2 instances.
C.Use an AWS Auto Scaling group with simple scaling policies based on CPU utilization and attach it to the ALB target group.
D.Use an AWS Auto Scaling group with target tracking scaling policies based on the ALB's request count per target, and attach it to the ALB target group.
AnswerD

This configuration uses the ALB's per-target request count as a scaling metric, so Auto Scaling continuously adjusts the EC2 fleet to keep that value near the chosen target. The ALB target group integration ensures newly launched instances are immediately registered to receive traffic, and the policy can scale both out and in based on real observed load. It is designed for variable workloads like critical web apps and responds faster and more accurately than scheduled or simple scaling policies.

Why this answer

Target tracking scaling policies allow the Auto Scaling group to automatically adjust capacity based on a specific metric, such as ALB request count per target, which directly reflects the load on each instance. This ensures that capacity is added during high latency periods and removed during off-peak hours, optimizing both performance and cost. The ALB target group integration ensures that new instances are automatically registered and start receiving traffic.

Exam trap

The trap here is that candidates often choose scheduled scaling (Option A) because it seems straightforward for known peak hours, but they overlook the requirement to handle unpredictable high latency during peak hours, which demands a dynamic, metric-based scaling solution like target tracking.

How to eliminate wrong answers

Option A is wrong because scheduled scaling policies only add or remove instances at predefined times, which cannot react to real-time demand fluctuations or unexpected traffic spikes, leading to either over-provisioning or under-provisioning. Option B is wrong because Route 53 weighted routing policies distribute traffic across multiple ALBs but do not dynamically scale the underlying EC2 instances; each fixed set of instances would still suffer from high latency during peak hours. Option C is wrong because simple scaling policies based on CPU utilization require manual configuration of thresholds and cooldown periods, which can cause slow reaction to sudden load changes and may not directly correlate with application latency as effectively as request count per target.

83
MCQeasy

A development team wants to ensure that their application can continue serving traffic even if an entire AWS Availability Zone (AZ) becomes unavailable. The application runs on Amazon EC2 instances in an Auto Scaling group and uses an Application Load Balancer (ALB). Which configuration should the team implement to meet this requirement?

A.Configure the Auto Scaling group to launch EC2 instances across multiple AZs, and ensure the ALB is enabled for cross-zone load balancing.
B.Use a launch template with multiple instance types to ensure diversity across the fleet.
C.Use a single AZ but configure EC2 Auto Scaling to replace unhealthy instances automatically.
D.Launch all EC2 instances in the same AZ to minimize latency, and configure the Auto Scaling group to maintain a minimum of two instances.
AnswerA

By configuring the Auto Scaling group to span multiple Availability Zones and enabling cross-zone load balancing on the ALB, the application can survive an entire AZ failure. If one AZ becomes unavailable, the ALB routes traffic only to the remaining healthy targets, while the ASG automatically launches replacement instances in the other AZs to maintain capacity. Cross-zone load balancing ensures that traffic is distributed evenly across all instances, even when one AZ has fewer instances. This design provides both redundancy and optimal utilization.

Why this answer

Deploying EC2 instances across multiple Availability Zones (AZs) ensures that if one AZ fails, the remaining AZs continue to serve traffic. Enabling cross-zone load balancing on the ALB distributes incoming requests evenly across all healthy instances in all AZs, preventing traffic from being sent only to instances in the same AZ as the client. This architecture meets the requirement for high availability and fault tolerance at the AZ level.

Exam trap

The trap here is that candidates often confuse instance-level resilience (e.g., replacing unhealthy instances) with AZ-level resilience, or they think that multiple instances in a single AZ provide sufficient fault tolerance, ignoring the fact that an AZ failure takes down all instances in that AZ.

How to eliminate wrong answers

Option B is wrong because using multiple instance types in a launch template addresses instance diversity and spot instance interruption resilience, not AZ-level failure; it does not protect against an entire AZ becoming unavailable. Option C is wrong because using a single AZ means all instances are in one failure domain; even with automatic replacement, the application cannot serve traffic if that AZ fails, as the new instances would also be launched in the same unavailable AZ. Option D is wrong because launching all instances in the same AZ and maintaining a minimum of two instances does not provide AZ-level redundancy; if that AZ fails, all instances become unavailable, and the application cannot serve traffic.

84
MCQeasy

A DevOps engineer is designing a resilient architecture for a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application experiences occasional spikes in traffic that cause Lambda function throttling and increased error rates. What is the MOST effective way to improve resilience and reduce throttling?

A.Increase the Lambda function memory to the maximum allowed.
B.Enable DynamoDB auto scaling for the table to handle traffic spikes.
C.Set API Gateway throttling limits to match the expected peak traffic.
D.Reserve concurrency for the Lambda function to ensure it always has available capacity.
AnswerD

Reserved concurrency sets a hard upper limit on the number of simultaneous executions for a specific function and, more importantly, guarantees that this amount of capacity is reserved exclusively for that function from the account's total concurrency pool. This prevents other functions from exhausting the shared pool and ensures the critical function can always handle its peak load without being throttled. It also makes the function's behavior predictable during traffic spikes, because the reserved capacity is always available for that function's invocations.

Why this answer

The stem specifies that traffic spikes cause Lambda function throttling. Reserving concurrency for the Lambda function (Option D) guarantees a dedicated portion of the account-level concurrency limit, preventing other functions from exhausting capacity and directly mitigating throttling. Option B (DynamoDB auto scaling) would only help if the throttling were caused by DynamoDB capacity issues, but the stem makes no mention of database errors.

Option A (increasing memory) does not address concurrency limits, and Option C (API Gateway throttling) can limit incoming requests but does not guarantee that Lambda has capacity to process them.

Exam trap

The trap is to assume that downstream services like DynamoDB are the bottleneck causing Lambda throttling. However, the stem clearly states that traffic spikes directly cause Lambda throttling, indicating that the concurrency limit is the primary constraint.

How to eliminate wrong answers

Option A is wrong because increasing Lambda memory also increases CPU and network allocation, but it does not resolve throttling caused by DynamoDB capacity limits or Lambda concurrency limits; it only improves execution speed for compute-bound functions. Option C is wrong because setting API Gateway throttling limits to match expected peak traffic would cap requests at that level, rejecting legitimate traffic during spikes rather than improving resilience. Option D is wrong because reserving concurrency for the Lambda function guarantees a fixed number of concurrent executions, but if the DynamoDB table lacks sufficient capacity, those executions will still fail due to database throttling, and reserved concurrency can also waste capacity during low traffic.

85
MCQeasy

A DevOps team is designing a disaster recovery plan for an RDS MySQL database. The database must be recoverable with minimal data loss in case of a regional failure. Which solution provides the LOWEST Recovery Point Objective (RPO)?

A.Configure a Cross-Region Read Replica.
B.Take daily automated snapshots and copy them to another Region.
C.Use a Multi-AZ deployment with synchronous standby.
D.Use RDS Proxy to cache database writes.
AnswerA

A Cross-Region Read Replica uses Amazon RDS's asynchronous replication to continuously apply transactions from the primary DB instance to a replica in a different AWS Region. This typically achieves an RPO in the low seconds (usually 1–5 seconds) and a short RTO by promoting the replica to a standalone primary via the console or API. Because replication is ongoing, it minimizes data loss far more effectively than any interval-based snapshot strategy.

Why this answer

A Cross-Region Read Replica provides the lowest Recovery Point Objective (RPO) because it uses asynchronous replication to continuously replicate data changes from the primary region to a replica in another region. In the event of a regional failure, you can promote the replica to a standalone primary database, typically losing only a few seconds to minutes of data, depending on replication lag. This minimizes data loss compared to snapshot-based or batch replication methods.

Exam trap

The trap here is that candidates often confuse Multi-AZ deployments (which provide high availability within a region with zero RPO) with cross-region disaster recovery, mistakenly thinking synchronous replication across regions is possible, but AWS RDS does not support synchronous replication across regions, and Multi-AZ does not protect against regional outages.

How to eliminate wrong answers

Option B is wrong because daily automated snapshots have an RPO of up to 24 hours, as they are taken only once per day, and copying them to another region adds additional latency, resulting in significantly higher potential data loss. Option C is wrong because a Multi-AZ deployment with synchronous standby provides high availability within a single region but does not protect against a regional failure; it offers zero RPO within the region but cannot recover data if the entire region goes down. Option D is wrong because RDS Proxy caches database writes to improve connection pooling and reduce failover time, but it does not replicate data to another region or provide any disaster recovery capability; it does not reduce RPO for regional failures.

86
MCQmedium

A company runs a high-traffic e-commerce application on EC2 instances in an Auto Scaling group behind an ALB. The application uses an in-memory cache on the EC2 instances. During a recent deployment, the Auto Scaling group terminated an instance that had active user sessions, causing users to lose their cart data and leading to a poor customer experience. The company wants to prevent this in future deployments. They need a solution that allows existing sessions to complete before instance termination, without manual intervention. Which solution should they use?

A.Increase the Auto Scaling group's cooldown period and health check grace period.
B.Enable connection draining on the ALB target group and increase the deregistration delay.
C.Implement an Auto Scaling lifecycle hook that puts the instance in a 'terminating:wait' state, and have a script on the instance that signals completion after draining sessions.
D.Change the health check type to ELB and mark instances unhealthy before deployment.
AnswerC

A lifecycle hook pauses the Auto Scaling termination process by moving the instance into the 'terminating:wait' state, giving you a configurable period to perform custom actions before the instance is finally terminated. An in-instance script or agent can monitor and gracefully drain active user sessions, commit any required state, and then call complete-lifecycle-action with the appropriate lifecycle hook token to signal that termination may proceed. This is the only option that actually holds the termination process itself and coordinates with application-level work rather than merely affecting network connections or scaling timers.

Why this answer

An Auto Scaling lifecycle hook puts the instance into a 'terminating:wait' state when the ASG decides to terminate it, pausing the termination process. A script on the instance can then drain active sessions (e.g., finish processing requests, flush cart data to a persistent store) and call CompleteLifecycleAction to signal completion, after which the ASG proceeds with termination. This is the only option that provides a programmable, instance-level mechanism to gracefully complete sessions before termination without manual intervention.

Exam trap

DOP-C02 often tests the misconception that ALB connection draining alone handles graceful instance termination — candidates forget that connection draining only covers in-flight requests, not in-memory application state or session persistence.

How to eliminate wrong answers

Option A is wrong because cooldown periods and health check grace periods control scaling timing and health check behavior, not graceful session draining during termination — they do not pause termination to let sessions finish. Option B is wrong because ALB connection draining (deregistration delay) only handles in-flight HTTP requests at the load balancer level; it does not address in-memory session state on the instance or allow the application to persist cart data before the instance is terminated. Option D is wrong because changing health check type to ELB and manually marking instances unhealthy is a manual, error-prone process that does not provide a programmable completion signal and does not guarantee session draining.

87
Multi-Selecteasy

A company wants to protect its application from DDoS attacks. Which THREE AWS services should they use?

Select 3 answers
A.Amazon Inspector
B.AWS WAF
C.AWS Shield Advanced
D.Amazon CloudFront
E.Amazon GuardDuty
AnswersB, C, D

AWS WAF is a web application firewall that filters and monitors HTTP(S) requests using rules for IP reputation, geographic origin, URI patterns, SQL injection, and cross-site scripting. Its rate-based rules automatically block IPs that exceed configured request thresholds, making it effective against HTTP floods and the low-and-slow application-layer DDoS attacks that target web endpoints. WAF integrates with CloudFront, Application Load Balancer, and API Gateway, allowing it to enforce Web ACLs at the edge or origin.

Why this answer

AWS Shield Advanced, WAF, and CloudFront provide layered DDoS protection.

88
MCQhard

A company runs a critical microservices architecture on Amazon ECS with Fargate. They want to ensure that if a task fails, it is automatically restarted, and the service remains available across multiple Availability Zones. How should they configure the ECS service?

A.Place all tasks in the same Availability Zone to reduce latency
B.Run a standalone Fargate task and use a CloudWatch alarm to restart it
C.Use an EC2 launch type with a single instance to reduce complexity
D.Define an ECS service with a task definition, set desired count across multiple Availability Zones, and use Service Auto Scaling
AnswerD

An ECS service with a task definition and a multi-AZ placement strategy is the correct pattern because the service scheduler continuously maintains the desired count, automatically replacing tasks that fail, become unhealthy, or lose connectivity. Distributing tasks across multiple Availability Zones ensures the service remains available even if an entire AZ goes down, as the remaining AZs still host task copies. Service Auto Scaling independently adjusts the desired count based on CloudWatch metrics or target tracking, providing elasticity and capacity management without sacrificing the resilience built into the service model.

Why this answer

An ECS service configured with a task definition, desired count across multiple Availability Zones, and Service Auto Scaling ensures resilience. The ECS service scheduler automatically restarts failed tasks, and distributing tasks across AZs provides high availability. Option A is wrong because placing all tasks in a single AZ creates a single point of failure.

Option B is wrong because a standalone Fargate task does not have automatic restart; a CloudWatch alarm can restart it but lacks the built-in resilience of an ECS service. Option C is wrong because using a single EC2 instance is a single point of failure and does not provide multi-AZ resilience.

89
MCQhard

A company uses AWS CloudFormation to deploy a multi-tier application. During an update, the stack fails and rolls back. The rollback also fails, leaving the stack in UPDATE_ROLLBACK_FAILED state. The operations team needs to resolve this with minimal disruption. What is the MOST efficient approach?

A.Use the 'ContinueUpdateRollback' API or AWS Management Console to retry the rollback.
B.Manually modify the resources to match the previous stack state.
C.Delete the stack and recreate it from the original template.
D.Execute a change set to update the stack to the desired configuration.
AnswerA

The Correct answer: The `ContinueUpdateRollback` API or the AWS Management Console action is specifically designed for stacks stuck in `UPDATE_ROLLBACK_FAILED`. It resumes the rollback from the point where it failed, skips resources that were already successfully rolled back, and can optionally skip resources you explicitly specify. This recovers the stack to its last known good state with minimal disruption, and it is the only built-in mechanism that directly resolves this status.

Why this answer

When a CloudFormation stack is stuck in UPDATE_ROLLBACK_FAILED, the supported recovery path is to call ContinueUpdateRollback (via API, CLI, or console) to resume the rollback, optionally skipping specific resources that are blocking it. This preserves the stack and its resources with minimal disruption.

Exam trap

DOP-C02 often tests whether candidates know that ContinueUpdateRollback is the only supported way out of UPDATE_ROLLBACK_FAILED — deleting the stack or running a change set sounds reasonable but is either destructive or blocked.

How to eliminate wrong answers

Option B is wrong because manually modifying resources to match the previous state does not update CloudFormation's internal state and can cause drift or further failures; it is not a supported recovery mechanism. Option C is wrong because deleting and recreating the stack destroys all resources and causes significant disruption, violating the 'minimal disruption' requirement. Option D is wrong because a change set cannot be executed while the stack is in UPDATE_ROLLBACK_FAILED — CloudFormation rejects updates until the rollback is resolved.

90
MCQhard

A company deploys the above CloudFormation stack. They want to enforce HTTPS for all requests to the S3 bucket. After deployment, users are still able to make HTTP requests. What is the problem?

A.The condition key 'aws:SecureTransport' is misspelled; it should be 'aws:SecureTransport' with a capital 'T'
B.The bucket is not versioned, so the policy does not apply to object versions
C.The policy uses Deny, but an Allow policy from another statement overrides it
D.The Deny statement's Resource specifies only the objects, not the bucket itself
AnswerD

The Resource does not include the bucket ARN, so bucket-level operations like ListBucket are not denied.

Why this answer

The Deny statement in the bucket policy uses `arn:aws:s3:::example-bucket/*` as the Resource, which applies only to objects within the bucket, not to the bucket itself. To enforce HTTPS for all requests, including those to the bucket endpoint (e.g., `GET /` or `PUT /`), the Resource must also include the bucket ARN without the `/*` suffix. Without it, HTTP requests targeting the bucket itself (such as listing objects or configuring website hosting) are not denied, allowing HTTP access to bypass the policy.

Exam trap

The trap here is that candidates assume a Deny statement on `/*` covers all requests, but they overlook that the bucket ARN itself must be explicitly included to enforce HTTPS on bucket-level operations, not just object operations.

How to eliminate wrong answers

Option A is wrong because `aws:SecureTransport` is correctly spelled with a capital 'S' and capital 'T' — the condition key is case-sensitive and must be exactly `aws:SecureTransport`. Option B is wrong because versioning is irrelevant to enforcing HTTPS; bucket policies apply to all object versions regardless of versioning status, and the Deny statement would still apply to `/*` resources. Option C is wrong because an explicit Deny in a bucket policy always overrides any Allow, regardless of other statements, per IAM policy evaluation logic (Deny is evaluated first and is definitive).

91
MCQmedium

A DevOps team uses AWS CodePipeline to deploy a web application. The pipeline has a deploy stage that uses CodeDeploy to deploy to an Auto Scaling group. During deployment, the new instances fail health checks and the deployment rolls back. However, the rollback also fails because the old instances have been terminated. What should the team do to avoid this issue?

A.Increase the health check grace period in the Auto Scaling group.
B.Add a manual approval step before the deploy stage.
C.Configure the pipeline to deploy to a new Auto Scaling group each time.
D.Use a blue/green deployment strategy in CodeDeploy to keep the old instances running until the new ones pass health checks.
AnswerD

A blue/green deployment creates a fresh set of green instances beside the original blue instances, and traffic is shifted to the green fleet only after it successfully passes the configured health checks. If the green instances fail, CodeDeploy can keep the blue fleet available and immediately route traffic back to it, giving you a deterministic rollback path that does not require re-deploying the old revision. This directly solves the problem in the question, because the old instances remain running and intact until the new ones are proven healthy, making rollback both possible and rapid.

Why this answer

A blue/green deployment strategy in CodeDeploy keeps the old (blue) instances running while the new (green) instances are provisioned and validated, so if the new instances fail health checks, the deployment can roll back to the still-running old instances. This avoids the scenario where in-place deployment terminates old instances before new ones are confirmed healthy, leaving nothing to roll back to. Blue/green also supports traffic shifting controls like all-at-once, linear, or canary.

Exam trap

DOP-C02 often tests the misconception that increasing health check grace periods or adding approvals solves rollback failures, when the real fix is preserving old instances via blue/green deployment.

How to eliminate wrong answers

Option A is wrong because increasing the health check grace period only delays health check evaluation; it does not prevent old instances from being terminated during an in-place deployment, so rollback would still fail if the new instances never become healthy. Option B is wrong because a manual approval step adds a gate before deployment but does not change the deployment mechanics — once approved, the in-place deployment still terminates old instances, so rollback failure remains possible. Option C is wrong because deploying to a new Auto Scaling group each time is essentially what blue/green does, but configuring the pipeline to do this manually is not the standard CodeDeploy feature and does not by itself provide the traffic shifting and rollback guarantees of blue/green; the correct answer is to use CodeDeploy's blue/green capability.

92
MCQhard

A company has a serverless application using AWS Lambda functions that process messages from an Amazon SQS queue. The Lambda function sometimes fails due to transient errors. The company wants to ensure that failed messages are retried and eventually processed or sent to a dead-letter queue after 3 retries. What is the correct configuration?

A.Set the Lambda function's retry policy to Maximum retries: 3 and configure a DLQ on the Lambda function.
B.Set the Lambda function's DLQ to an SQS queue and configure the event source mapping to use that DLQ after 3 retries.
C.Configure the SQS queue's redrive policy with maxReceiveCount: 3 and a dead-letter queue.
D.Create an AWS Step Functions workflow that polls the SQS queue, processes messages, and retries failures up to 3 times before moving to a DLQ.
AnswerC

This is the correct pattern because SQS itself owns the retry and dead-letter behavior when Lambda consumes from a queue. A redrive policy with maxReceiveCount: 3 instructs SQS to allow a message to be received up to three times; if the Lambda function fails to process it each time, SQS automatically moves the message to the configured dead-letter queue. This is a native, serverless-friendly mechanism that avoids unnecessary compute and precisely matches the requirement for '3 retries before a DLQ'.

Why this answer

For Lambda functions that poll an SQS queue, the retry behavior and dead-letter queue are configured on the SQS queue itself using a redrive policy. The redrive policy specifies the maxReceiveCount (e.g., 3) and the ARN of the dead-letter queue. After a message is received the specified number of times without being deleted, SQS moves it to the DLQ.

This is the correct way to handle retries and DLQ for SQS-triggered Lambda.

Exam trap

The trap is assuming Lambda's retry settings apply to SQS-triggered invocations; candidates often confuse asynchronous invocation retries with poll-based event source mapping retries, leading them to pick Lambda DLQ options instead of SQS redrive policy.

How to eliminate wrong answers

Option A is wrong because Lambda's own retry policy (Maximum retries) applies to asynchronous invocations, not to poll-based event source mappings like SQS. Option B is wrong because the Lambda function's DLQ is for asynchronous invocations, and the event source mapping does not have a DLQ configuration; the DLQ must be on the SQS queue. Option D is wrong because using Step Functions adds unnecessary complexity and does not leverage the native SQS redrive policy; it is not the simplest or correct configuration for this requirement.

93
MCQhard

A company uses AWS CloudFormation to deploy infrastructure. The stack creation fails with the error: 'Resource handler returned message: 'The security group does not exist in VPC'.' The template references a security group by name. What is the MOST likely cause?

A.The security group name is misspelled or uses incorrect case
B.The IAM role used for CloudFormation does not have permissions to describe security groups
C.The stack is being created in a Region where the security group does not exist
D.The template uses a parameter that resolves to the default VPC security group
AnswerA

Security group names in EC2 are case-sensitive, and CloudFormation's lookup by name uses an exact, literal string match against the security groups within the specified VPC. If you reference a group by name and the template contains a typo, wrong case, or an unintended trailing space, the API returns no matching group, which CloudFormation reports as a 'security group not found' error. This is the most frequent root cause in practice because names like 'web-SG' and 'web-sg' are considered distinct, and the error message often appears immediately after a small edit or a manual copy-paste from a different source.

Why this answer

The error 'Resource handler returned message: The security group does not exist in VPC' occurs when CloudFormation cannot find a security group with the specified name in the target VPC. The most likely cause is a misspelling or case sensitivity issue (Option A), as CloudFormation matches security group names exactly. While region scoping (Option C) can also cause a similar error, the combination of a name-based reference and the specific error wording points to a name mismatch as the most common issue.

IAM permission errors (Option B) would typically return an authorization error, not a 'does not exist' error. Option D is incorrect because a default VPC security group exists and would not cause this error unless it is missing, which is unlikely.

Exam trap

Candidates often overlook that security group name resolution is case-sensitive and exact. While region scoping can cause a similar error, the most frequent cause is a typo or case mismatch in the security group name referenced in the template.

How to eliminate wrong answers

Option B is wrong because the error message specifically indicates the security group does not exist in the VPC, not a permissions issue; an IAM permissions error would produce a different message such as 'AccessDenied' or 'Unauthorized operation'. Option D is wrong because referencing a parameter that resolves to the default VPC security group would not cause this error; the default security group exists in every VPC and would be found, so the error would not occur unless the VPC itself is missing or the parameter value is invalid.

94
MCQmedium

A company is designing a disaster recovery strategy for a critical application that requires a Recovery Time Objective (RTO) of 15 minutes and a Recovery Point Objective (RPO) of 1 hour. The application runs on EC2 with data stored in Amazon RDS Multi-AZ. Which approach meets these requirements?

A.Use a pilot light strategy with RDS cross-Region read replicas and automated backups
B.Use backup and restore with daily snapshots to another Region
C.Use a warm standby with a scaled-down production environment in another Region
D.Use a Multi-AZ deployment in the same Region for DR
AnswerA

A pilot light strategy keeps a minimal core, such as the RDS instance, running in the DR Region via a cross-Region read replica while other services remain off. The read replica provides continuous asynchronous replication, so a promotion on failover can complete in roughly 15 minutes, and automated backups in the DR Region give point-in-time restore support. Since the database is already warm and data is being copied, the worst-case data loss stays within the 1-hour RPO.

Why this answer

A pilot light strategy with RDS cross-Region read replicas and automated backups meets the RTO of 15 minutes and RPO of 1 hour. The cross-Region read replica provides near-synchronous replication with an RPO typically under 5 seconds, and automated backups enable point-in-time recovery within the 1-hour RPO. The pilot light approach allows rapid promotion of the replica to a primary instance, achieving the 15-minute RTO by keeping minimal core services running in the DR Region.

Exam trap

The trap here is that candidates often confuse Multi-AZ (high availability within a Region) with cross-Region disaster recovery, assuming Multi-AZ alone provides DR, but it does not protect against Region-wide outages.

How to eliminate wrong answers

Option B is wrong because daily snapshots to another Region result in an RPO of up to 24 hours, far exceeding the required 1-hour RPO, and restoring from snapshots takes longer than 15 minutes. Option C is wrong because a warm standby with a scaled-down production environment typically has an RTO of minutes but requires continuous replication and failover orchestration; while it could meet the RPO, it is over-engineered and more costly than necessary, and the question asks for an approach that meets requirements, not the most optimal. Option D is wrong because a Multi-AZ deployment in the same Region does not provide disaster recovery across Regions; it only protects against Availability Zone failures, not regional disasters, and thus fails to meet the DR requirement.

95
MCQmedium

A DevOps engineer ran the above command and saw this output. What is the MOST likely cause of the stack creation failure?

A.The key pair specified in the launch template does not exist.
B.The IAM role does not have permission to create the Auto Scaling group.
C.The AMI ID specified in the launch template is not available in this Region.
D.The launch template name specified in the CloudFormation template is incorrect or does not exist.
AnswerD

When you reference a launch template by name in an Auto Scaling group's LaunchTemplateSpecification, CloudFormation and the Auto Scaling API require that the launch template already exists in the same account and Region. If the name is misspelled, or the launch template was created under a different account/Region or never created, the API returns an error indicating the launch template parameter is invalid or not found. The error message in the output is consistent with this cause, as it specifically flags the launch template name reference, not the AMI, key pair, or IAM permissions.

Why this answer

The error message indicates that the launch template name specified in the CloudFormation template does not match any existing launch template in the account and Region. CloudFormation resolves the launch template name at stack creation time; if the name is incorrect or the template does not exist, the Auto Scaling group creation fails with a validation error. This is the most direct cause because the launch template name is a required parameter that must reference a pre-existing resource.

Exam trap

The trap here is that candidates confuse launch template validation errors with instance-level errors (like missing AMI or key pair), but CloudFormation validates the launch template name at the Auto Scaling group resource level before any EC2 instances are launched.

How to eliminate wrong answers

Option A is wrong because a missing key pair would cause an EC2 instance launch failure, not a stack creation failure at the Auto Scaling group level; the error message would reference 'InvalidKeyPair.NotFound' or similar. Option B is wrong because an IAM role lacking permissions to create an Auto Scaling group would produce an 'AccessDenied' or authorization error, not a validation error about a missing launch template. Option C is wrong because an unavailable AMI ID would cause an instance launch failure with an 'InvalidAMIID.NotFound' error, not a stack creation failure related to the launch template name.

96
MCQeasy

A company is building a multi-tier web application on AWS. The web tier runs on EC2 instances behind an ALB. The application tier runs on EC2 instances that are not publicly accessible. The database tier runs on RDS MySQL. Which design provides the HIGHEST level of resilience for the database tier?

A.Deploy a single RDS DB instance in one Availability Zone.
B.Deploy an RDS DB instance with a cross-region read replica.
C.Deploy an RDS DB instance with a read replica in the same region.
D.Deploy an RDS DB instance in a Multi-AZ configuration.
AnswerD

Deploying an RDS DB instance in a Multi-AZ configuration creates a synchronous standby in a different Availability Zone and enables automatic failover. RDS monitors the primary instance and, if a failure is detected, automatically switches the endpoint to the standby, typically in 60–120 seconds. This is the standard RDS pattern for high availability within a single region.

Why this answer

Multi-AZ RDS automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary DB instance fails, RDS automatically fails over to the standby, providing high availability with minimal downtime. This design ensures the database tier remains resilient against AZ-level failures without manual intervention.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and disaster recovery) with Multi-AZ deployments (which are for high availability and automatic failover), leading them to choose a read replica option for resilience.

How to eliminate wrong answers

Option A is wrong because a single RDS DB instance in one Availability Zone is a single point of failure; any AZ outage or instance failure will cause database downtime. Option B is wrong because a cross-region read replica provides disaster recovery and read scaling, but it is asynchronous and does not support automatic failover for high availability; it requires manual promotion to become the primary. Option C is wrong because a read replica in the same region is designed for offloading read traffic, not for automatic failover; it is asynchronous and cannot be used as a synchronous standby for high availability.

97
Multi-Selecthard

A company runs a containerized application on Amazon EKS. The application must be highly available across multiple Availability Zones and must automatically recover from node failures. Which THREE steps should be taken?

Select 3 answers
A.Use Pod Disruption Budgets to ensure a minimum number of pods are available during voluntary disruptions.
B.Configure the Cluster Autoscaler to add nodes when pods are unschedulable.
C.Deploy worker nodes across multiple Availability Zones.
D.Deploy worker nodes in a single Availability Zone to reduce cross-AZ data transfer costs.
E.Use a single large instance type for all worker nodes to simplify management.
AnswersA, B, C

Pod Disruption Budgets (PDBs) constrain voluntary disruptions such as node drains during cluster upgrades, node-group updates, or Cluster Autoscaler scale-in by specifying minAvailable or maxUnavailable for a pod selector. During a voluntary disruption, the Kubernetes eviction API rejects requests that would cause the number of available pods to fall below the budget. PDBs do not protect against involuntary failures like an EC2 instance crash, so they are a complement, not a substitute, for multi-AZ node deployment.

Why this answer

Pod Disruption Budgets (PDBs) are correct because they allow you to specify the minimum number of pods that must remain available during voluntary disruptions, such as node drains or cluster upgrades. This ensures that the application maintains high availability even when Kubernetes performs planned maintenance, preventing all replicas from being taken down simultaneously.

Exam trap

The trap here is that candidates often think deploying in a single AZ or using a single instance type simplifies management and reduces costs, but the DOP-C02 exam specifically tests the principle of designing for failure across multiple AZs and instance diversity to achieve true high availability.

98
MCQhard

A company runs a critical application on Amazon ECS with the Fargate launch type. The application is deployed across three Availability Zones. Each service has its own Application Load Balancer. The company wants to implement a blue/green deployment strategy to reduce risk. They currently use AWS CodeDeploy for ECS deployments. During a recent deployment, the company noticed that the new version (green) was not receiving any traffic even after passing all health checks. The CodeDeploy configuration uses a 'Linear10PercentEvery3Minutes' traffic shifting configuration. What is the most likely reason that the green tasks are not receiving traffic?

A.The CodeDeploy deployment group is not associated with the correct ECS service.
B.The green target group's health check is misconfigured, causing CodeDeploy to consider the green tasks unhealthy and not route traffic.
C.The blue target group is still set as the production target group in the load balancer listener.
D.The green tasks are in a different VPC than the load balancer.
AnswerB

In CodeDeploy's ECS blue/green deployment, the green target group's health checks are the gating mechanism for traffic shifting. If the health check path, port, or interval is misconfigured, the green tasks are marked unhealthy even though the containerized application is running normally, so CodeDeploy never receives the signal that the green fleet is ready. As a result, the listener remains pointed at the blue target group, and the deployment times out waiting for a healthy green target.

Why this answer

The green target group's health check is misconfigured, causing CodeDeploy to consider the green tasks unhealthy. With a 'Linear10PercentEvery3Minutes' traffic shifting configuration, CodeDeploy gradually shifts traffic in 10% increments every 3 minutes, but only if the green target group passes health checks. If the health check fails, CodeDeploy stops traffic shifting, leaving the green tasks with zero traffic despite the tasks themselves being healthy.

Exam trap

The trap here is that candidates assume health checks passing on the ECS tasks means traffic will automatically route, but CodeDeploy relies on the target group's health check configuration, not the task's health status, to determine when to shift traffic.

How to eliminate wrong answers

Option A is wrong because if the CodeDeploy deployment group were not associated with the correct ECS service, the deployment would fail entirely or target the wrong service, but the green tasks would still be created and potentially receive traffic if health checks passed. Option C is wrong because CodeDeploy automatically updates the load balancer listener rules to point to the green target group during the traffic shifting process; the blue target group being set as production is the initial state, but CodeDeploy changes it as traffic shifts. Option D is wrong because ECS Fargate tasks and the load balancer must be in the same VPC for the service to function; if they were in different VPCs, the service would not register targets or pass health checks at all, not just fail to receive traffic after health checks pass.

99
Multi-Selecteasy

A company is designing a highly available architecture for a web application using AWS. Which TWO of the following design principles should be applied? (Select TWO.)

Select 2 answers
A.Run all resources in a single Availability Zone to reduce complexity
B.Store session data on EC2 instances to improve performance
C.Deploy resources across multiple Availability Zones
D.Use loosely coupled components, such as queues and asynchronous processing
E.Use tightly coupled components to reduce latency
AnswersC, D

Deploying across multiple Availability Zones (AZs) is the primary AWS design pattern for high availability because each AZ is an independent failure domain with separate power, cooling, and networking. By placing resources (e.g., application servers behind an Application Load Balancer, and a Multi-AZ database) in at least two AZs, the workload can continue serving traffic if one AZ suffers an outage, as the load balancer automatically routes requests only to healthy instances in the remaining AZs. This approach directly satisfies the requirement for fault tolerance, and when combined with Auto Scaling, it also provides capacity to absorb increased load in the surviving AZs.

Why this answer

Correct answers: C and D. Deploying resources across multiple Availability Zones (C) ensures high availability by tolerating an AZ failure. Using loosely coupled components like queues (D) improves resilience by decoupling components, preventing cascading failures and allowing independent scaling.

Option A is wrong because running in a single AZ creates a single point of failure. Option B is wrong because storing session data on EC2 instances is not recommended for high availability; session data should be stored externally (e.g., ElastiCache or DynamoDB). Option E is wrong because tightly coupled components increase dependency and reduce fault tolerance.

100
Multi-Selectmedium

A company is designing a resilient architecture for a critical application. Which TWO strategies improve resilience?

Select 2 answers
A.Deploy resources across multiple Availability Zones
B.Use a single large instance instead of multiple smaller ones
C.Use health checks to automatically replace unhealthy resources
D.Disable automated backups to reduce latency
E.Deploy resources in a single Availability Zone
AnswersA, C

Deploying across multiple Availability Zones (AZs) provides infrastructure-level fault isolation because each AZ has independent power, cooling, and network access. This architecture ensures that an outage in one AZ does not take down the entire workload, supporting a higher availability SLA. For example, running EC2 instances in two or more AZs with an Application Load Balancer allows traffic to continue to healthy AZs even if one becomes isolated.

Why this answer

Multi-AZ deployments and health checks with auto-remediation improve resilience by handling failures automatically.

101
MCQhard

A DevOps engineer runs the above command and sees that one target is unhealthy with a 503 error. The application is a web server running on port 80. The health check is configured to hit the root path '/'. Which action should the engineer take to resolve the issue?

A.Change the health check port to 443 and use HTTPS
B.Verify that the application on the unhealthy instance is configured to respond to '/' with a 200 status code
C.Increase the health check interval and timeout settings
D.Check the security group rules for the target group to ensure port 80 is open
AnswerB

A 503 status code from the health check endpoint indicates that the target instance is reachable at the HTTP layer, but the application logic serving the root path '/' is returning a Service Unavailable error. Elastic Load Balancing requires a 2xx or 3xx response for the health check to mark the instance healthy; any 4xx or 5xx status counts as unhealthy. Verify that the web server or application is configured to serve '/' with a 200 OK during normal operation and that there are no authentication, redirect, or maintenance-mode rules that cause a 503 on that specific path. This is the correct troubleshooting step because the health check is working as designed—it is detecting an application-level failure.

Why this answer

The health check is configured to hit the root path '/' and expects a 200 status code. A 503 error indicates the application on the unhealthy instance is not serving the correct response for that path. Verifying that the application responds with a 200 status code on '/' directly addresses the root cause of the health check failure.

Exam trap

The trap here is that candidates often confuse network-level issues (like security groups) with application-level HTTP errors (like 503), leading them to check connectivity instead of the application's response logic.

How to eliminate wrong answers

Option A is wrong because changing the health check port to 443 and using HTTPS does not fix a 503 error; it changes the protocol and port, which may not match the application's actual configuration and could cause further failures. Option C is wrong because increasing the health check interval and timeout settings only delays detection or reduces false positives, but does not resolve the underlying issue of the application returning a 503 error. Option D is wrong because a 503 error is an application-level HTTP status code, not a network connectivity issue; security group rules for port 80 would cause a timeout or connection refused, not a 503 response.

102
Multi-Selecthard

A company runs a microservices architecture on Amazon ECS with Fargate. Services communicate via an internal Application Load Balancer (ALB). The operations team notices that occasional traffic spikes cause increased latency and timeouts. The team wants to improve resilience without over-provisioning. Which THREE steps should be taken? (Choose THREE.)

Select 3 answers
A.Increase the CPU and memory limits in the task definitions.
B.Enable ECS Service Connect for inter-service communication to manage traffic distribution.
C.Configure ECS service auto scaling with a target tracking policy based on ALB request count per target.
D.Implement a graceful shutdown handler in the application to handle SIGTERM.
E.Use EC2 launch type with Spot Instances to reduce cost.
AnswersB, C, D

ECS Service Connect provides a resilient service mesh that simplifies inter-service communication by giving each service a stable DNS name and managing traffic distribution at the application layer. It enables fine-grained traffic splitting, automatic retries, and connection draining, which collectively reduce latency and improve reliability between microservices. This is a direct remedy for latency caused by inefficient service-to-service calls, as it optimizes the network path and avoids overloading individual task instances.

Why this answer

B is correct because ECS Service Connect provides built-in traffic management for inter-service communication, including load balancing, retries, and circuit breaking. This helps distribute traffic more evenly during spikes, reducing latency and timeouts without requiring over-provisioning.

Exam trap

The trap here is that candidates often confuse vertical scaling (increasing task resources) with horizontal scaling (adding more tasks), and they may overlook the importance of application-level resilience patterns like graceful shutdowns and service mesh features for traffic management.

103
MCQhard

An application running on Amazon ECS with Fargate experiences intermittent failures. The task definition includes a single container with a health check command. Despite the health check passing, the application occasionally returns HTTP 500 errors. The application logs are sent to CloudWatch Logs. What is the MOST likely root cause?

A.The health check command only checks the process status, not the application's ability to serve requests.
B.The application is missing environment variables that are required for certain requests.
C.The ECS service is configured with a target tracking scaling policy that reacts too slowly.
D.The container port and host port in the task definition do not match the ALB target group port.
AnswerA

Using a process-only health check (e.g., checking that the PID exists) means the ECS/ALB considers the container healthy whenever the runtime is alive, regardless of whether the app can actually serve HTTP traffic. As a result, intermittent 500 errors caused by a hung worker, exhausted connection pool, or an unhandled exception inside request handling remain invisible to the health check, so the task stays in service and keeps receiving traffic. A valid health check should send a real request to the application's endpoint and verify the response code.

Why this answer

A health check command in an ECS task definition typically checks the container process status (e.g., via a shell command or a simple TCP check), not the application's HTTP layer. If the health check passes but the application returns HTTP 500 errors, it indicates the container is running but the application logic is failing (e.g., unhandled exceptions, database connection issues). This mismatch between process-level health and application-level health is a common cause of intermittent failures in containerized applications.

Exam trap

The trap here is that candidates assume a passing health check guarantees the application is fully functional, but AWS specifically tests the distinction between container-level health (process running) and application-level health (HTTP response correctness), which is a key concept for the DOP-C02 exam.

How to eliminate wrong answers

Option B is wrong because missing environment variables would cause consistent failures for specific requests, not intermittent HTTP 500 errors; the application would fail predictably when those variables are accessed. Option C is wrong because a target tracking scaling policy that reacts too slowly would cause performance degradation or timeouts under load, not intermittent HTTP 500 errors when the health check passes; scaling latency affects capacity, not application logic. Option D is wrong because mismatched container port and ALB target group port would cause the ALB health check to fail entirely, not allow the health check to pass while the application intermittently returns HTTP 500 errors; the ALB would mark the target as unhealthy.

104
MCQmedium

A company's application uses Amazon SQS to decouple microservices. During peak hours, the SQS queue backlog grows significantly, causing processing delays. The DevOps team wants to reduce latency without increasing costs unnecessarily. What should the team do?

A.Increase the visibility timeout to allow consumers more time to process messages.
B.Use an SQS queue with priority settings to process high-priority messages first.
C.Increase the SQS queue's throughput by requesting a quota increase.
D.Configure Auto Scaling for the consumer fleet based on the ApproximateNumberOfMessagesVisible metric.
AnswerD

Scaling consumers on ApproximateNumberOfMessagesVisible directly matches fleet capacity to the actual backlog, so added workers drain the queue during peaks and terminate when it empties. This satisfies the latency constraint without over-provisioning, since capacity tracks demand rather than a fixed schedule or CPU proxy.

Why this answer

Scaling the consumer fleet based on the ApproximateNumberOfMessagesVisible metric directly addresses the backlog by adding more processing capacity when the queue grows. This approach reduces latency dynamically without incurring unnecessary costs during off-peak hours, as it only scales up when needed. Auto Scaling with SQS metrics is a cost-effective, elastic solution for handling variable workloads.

Exam trap

The trap here is that candidates may confuse SQS's throughput capabilities with consumer-side scaling, assuming that increasing queue throughput (Option C) solves backlog, when in fact SQS already handles high throughput and the bottleneck is the consumer processing rate.

How to eliminate wrong answers

Option A is wrong because increasing the visibility timeout does not reduce backlog; it only gives consumers more time to process a message, which can actually increase latency if consumers fail or take longer, as messages remain hidden longer. Option B is wrong because standard SQS queues do not support priority settings; FIFO queues offer ordering but not priority-based message selection, and SQS has no built-in priority feature. Option C is wrong because SQS queues already offer virtually unlimited throughput by default (up to 3,000 messages per second for FIFO with batching, and unlimited for standard), so requesting a quota increase is unnecessary and does not address consumer-side processing capacity.

105
MCQmedium

An application on EC2 instances in an Auto Scaling group uses an ALB. The ALB health checks are failing for some instances, but the instances are healthy from the OS perspective. What is the most likely cause?

A.The ALB idle timeout is too low
B.The security group for the instances does not allow traffic from the ALB
C.The Auto Scaling group cooldown period is too short
D.The ALB cross-zone load balancing is disabled
AnswerB

If the instance security group does not allow inbound traffic from the ALB's security group on the health-check port, the ALB's TCP or HTTP health-check probes are silently dropped. Because the instance never completes the health-check handshake, the target fails the required number of consecutive checks and the ALB marks it unhealthy. This is a classic misconfiguration that often appears right after adding the ALB, and it also blocks normal client traffic routed by the load balancer.

Why this answer

ALB health checks originate from the ALB's nodes and require the instance's security group to permit inbound traffic on the health check port and protocol from the ALB's security group. If the security group does not allow this traffic, the health check fails even though the OS and application are healthy. This is the most common cause of ALB health check failures when instances are otherwise reachable.

Exam trap

DOP-C02 often tests whether candidates jump to scaling or timeout settings when the real issue is a security group rule blocking the ALB's health check traffic, so candidates must check network ACLs and security groups first.

How to eliminate wrong answers

Option A is wrong because the ALB idle timeout affects long-lived connections, not health check success; a low idle timeout would drop idle connections but would not cause health checks to fail. Option C is wrong because the Auto Scaling group cooldown period controls how quickly the ASG launches or terminates instances after a scaling activity; it does not affect ALB health check results. Option D is wrong because cross-zone load balancing affects traffic distribution across AZs, not whether health checks succeed; disabling it does not cause health checks to fail.

106
MCQmedium

A media company runs a video processing pipeline on AWS. Raw videos are uploaded to an S3 bucket, which triggers a Lambda function to start an AWS Batch job for transcoding. The Batch job reads the source video from S3, processes it, and writes the output to another S3 bucket. Recently, the company has seen an increase in processing failures. Investigation shows that the Batch jobs are being terminated with a 'TIMEOUT' status after running for exactly 30 minutes. The video files are large, and some jobs legitimately take up to 45 minutes. The Batch job definition has a 'timeout' setting configured. Which action should be taken to resolve this issue?

A.Modify the Batch job definition to increase the 'timeout' value to 3600 seconds (60 minutes).
B.Increase the S3 bucket lifecycle policy to retain videos longer.
C.Increase the Lambda function timeout to 60 minutes.
D.Change the Batch job queue to a different compute environment.
AnswerA

The AWS Batch job definition includes a `timeout` field that sets the maximum duration a job attempt is allowed to run before Batch forcibly terminates it as a timeout. Increasing this value to 3600 seconds (60 minutes) directly accommodates longer video processing tasks, preventing premature termination when the workload legitimately needs more than the current limit. This is the intended control for adjusting how long Batch permits a single job attempt to execute.

Why this answer

The Batch job is being terminated with TIMEOUT after exactly 30 minutes, and the job definition has a timeout setting — this is the Batch job attempt timeout (default 30 minutes if not explicitly set, or set to 1800 seconds). Since legitimate jobs take up to 45 minutes, the fix is to raise the job definition's timeout to at least 3600 seconds (60 minutes) to accommodate the longest jobs with headroom.

Exam trap

DOP-C02 often tests whether candidates can pinpoint the exact configuration parameter responsible for a symptom — here, confusing the Lambda trigger timeout with the Batch job definition timeout, or blaming the compute environment, is the trap.

How to eliminate wrong answers

Option B is wrong because S3 lifecycle policies govern object retention/transition/deletion and have no effect on Batch job execution timeouts. Option C is wrong because the Lambda function only triggers the Batch job submission; the Lambda timeout (max 15 minutes anyway) is unrelated to the Batch job's 30-minute termination, and Lambda cannot be set to 60 minutes. Option D is wrong because changing the job queue or compute environment does not alter the job definition's timeout — the TIMEOUT status comes from the job definition's timeout parameter, not from the compute environment.

107
MCQhard

A company runs a critical web application on EC2 instances in an Auto Scaling group. The application uses an Application Load Balancer (ALB) with health checks pointing to /health. Recently, the application experienced intermittent failures where the ALB would mark instances as unhealthy and route traffic away, causing a reduction in capacity. The development team noticed that the /health endpoint occasionally returns HTTP 503 when the application is under heavy load, but the application can recover quickly. The team wants to avoid unnecessary instance replacements while ensuring availability. Which solution should the DevOps engineer implement?

A.Implement a custom health check using Lambda that ignores 503 responses
B.Decrease the unhealthy threshold to mark instances unhealthy faster
C.Increase the health check interval and increase the unhealthy threshold
D.Decrease the health check interval and decrease the healthy threshold
AnswerC

In a target group's health-check settings, increasing the interval spaces out probe requests, while increasing the unhealthy threshold demands more consecutive failures before an instance is marked unhealthy. Together, these settings build a longer smoothing window: sporadic errors such as brief 503s during a rolling deploy or dependency hiccup will not immediately cause a healthy instance to be replaced. The instance remains in service and in rotation until the failures are sustained over an extended period, which is exactly the desired behavior when the goal is to reduce replacement churn. The trade-off is that genuinely dead instances take longer to detect, but that is acceptable in many critical applications that favor stability over instantaneous failover.

Why this answer

Increasing the health check interval and increasing the unhealthy threshold makes the health check less sensitive to transient errors, such as occasional 503 responses under heavy load. This prevents unnecessary instance replacements while maintaining availability. Option A is incorrect because implementing a custom Lambda health check that ignores 503 responses would not leverage the built-in ALB health check tuning and adds complexity.

Option B is incorrect because decreasing the unhealthy threshold would make instances more easily marked unhealthy, worsening the problem. Option D is incorrect because decreasing the health check interval increases the frequency of checks, which might cause more frequent detections of transient errors, and decreasing the healthy threshold does not address the issue of avoiding unnecessary replacements.

108
MCQhard

A company runs a microservices application on Amazon EKS. The application's frontend service needs to communicate with the backend service. The DevOps team wants to implement service-to-service authentication using AWS IAM. Which method should the team use?

A.Configure the backend service as an Amazon RDS database with IAM database authentication.
B.Use AWS App Mesh with mTLS for authentication between services.
C.Create an IAM user with access keys and store them as Kubernetes secrets.
D.Use IAM roles for service accounts (IRSA) to associate an IAM role with each service's Kubernetes service account.
AnswerD

IRSA associates a Kubernetes ServiceAccount with an IAM role by annotating the ServiceAccount with the role ARN and configuring an OIDC trust policy. A projected service account token is exchanged via STS AssumeRoleWithWebIdentity for short-lived AWS credentials, and the AWS SDK automatically reads AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE. This gives each microservice a distinct IAM role with scoped permissions and no long-lived keys stored in the cluster.

Why this answer

IAM roles for service accounts (IRSA) allows each Kubernetes service account to assume an IAM role with fine-grained permissions, enabling secure service-to-service authentication without managing long-lived credentials. The frontend service can use its associated IAM role to sign AWS API requests (e.g., STS AssumeRole) to authenticate to the backend service, which validates the role via IAM policies. This approach integrates natively with EKS and follows AWS best practices for workload identity.

Exam trap

The trap here is that candidates may confuse mTLS (which provides encryption and certificate-based authentication) with IAM-based authentication, or assume that static IAM users with secrets are acceptable in Kubernetes, when IRSA is the recommended AWS-native approach for pod-level IAM integration.

How to eliminate wrong answers

Option A is wrong because Amazon RDS IAM database authentication is designed for database access, not for service-to-service authentication between microservices on EKS; it does not provide a mechanism for frontend-to-backend communication. Option B is wrong because AWS App Mesh with mTLS provides transport-layer encryption and mutual TLS authentication, but it does not use AWS IAM for authentication; it relies on X.509 certificates, not IAM roles or policies. Option C is wrong because creating an IAM user with access keys and storing them as Kubernetes secrets introduces long-lived static credentials, which violates security best practices (e.g., no automatic rotation, risk of exposure) and does not leverage IAM roles for dynamic, scoped access.

109
Multi-Selecthard

Which THREE components are required to implement a global application that can withstand the failure of an entire AWS Region? (Select THREE.)

Select 3 answers
A.An Application Load Balancer in the primary Region.
B.Amazon CloudFront with multiple origins and origin failover.
C.Amazon DynamoDB Global Tables.
D.Amazon RDS with a single-AZ deployment.
E.Amazon Route 53 with health checks and failover routing policy.
AnswersB, C, E

Amazon CloudFront is a global content delivery network that terminates connections at edge locations and can be configured with multiple origins, including an origin group where the primary origin's failure triggers automatic failover to a secondary origin. This origin failover is initiated when the primary returns specific HTTP error codes or fails connection attempts, making it a key component for routing requests to a healthy Region. Additionally, edge caching shields the origin and improves performance, which is essential for a globally resilient application.

Why this answer

B is correct because Amazon CloudFront with multiple origins and origin failover automatically routes requests to a healthy origin in another Region when the primary origin becomes unavailable, providing a global entry point that survives a full Region failure. C is correct because DynamoDB Global Tables replicate data across Regions in a multi-active configuration, so the application can continue reading and writing to a replica table in a surviving Region. E is correct because Route 53 health checks detect an unhealthy Region endpoint and the failover routing policy then directs DNS queries to the standby Region's endpoint, enabling regional failover.

A is not required because an Application Load Balancer is Region-scoped and cannot by itself provide cross-Region resilience; it would only be part of a single-Region design. D is not required because a single-AZ RDS deployment has no cross-Region (or even cross-AZ) redundancy and would not survive a Region failure.

Exam trap

The trap is selecting Regional services (ALB, single-AZ RDS) as if they provide cross-Region resilience — candidates must recognize that only global services (Route 53, CloudFront, DynamoDB Global Tables) or multi-Region replicated services can survive a full Region failure.

110
Multi-Selecteasy

Which TWO actions can help ensure that an application running on EC2 instances can survive the loss of an entire Availability Zone?

Select 2 answers
A.Deploy all instances in a single Availability Zone for consistency
B.Use an Auto Scaling group with multiple Availability Zones
C.Deploy EC2 instances in at least two Availability Zones
D.Use a larger instance type to handle more load
E.Use CloudWatch alarms to monitor instance health
AnswersB, C

An Auto Scaling group configured across multiple Availability Zones automatically distributes instances among those AZs and enforces the desired capacity by replacing any instance that fails its health checks. This provides both high availability and operational automation: if an entire AZ becomes impaired, the ASG launches replacement instances in the remaining healthy AZs, and it also performs capacity rebalancing when AZs become imbalanced. Combined with an Elastic Load Balancer, this is the standard AWS pattern for building a fault-tolerant, self-healing application tier.

Why this answer

Deploying instances in multiple Availability Zones (AZs) ensures that if one AZ fails, instances in other AZs continue to run. Using an Auto Scaling group with multiple AZs automatically distributes instances across AZs and replaces failed instances, further enhancing resilience. Options B and C are both correct because they achieve multi-AZ deployment.

Option A is incorrect because a single AZ is a single point of failure. Option D is incorrect because instance type does not provide AZ resilience. Option E is incorrect because CloudWatch alarms can detect issues but do not distribute instances across AZs.

111
MCQmedium

A company runs a stateful application on EC2 instances in an Auto Scaling group. The application stores state on local instance storage. During a scaling event, users lose session data. How can the company make the application resilient without modifying the application code?

A.Reduce the Auto Scaling group cooldown period.
B.Enable sticky sessions on the Application Load Balancer.
C.Increase the instance size to reduce scaling events.
D.Use Elastic Block Store (EBS) volumes instead of instance store.
AnswerD

EBS volumes are network-attached block storage that exist independently of the EC2 instance lifecycle, so when an instance terminates the volume can be detached and reattached to a replacement instance. Data written to an EBS volume survives Auto Scaling scale-in as long as the volume's DeleteOnTermination attribute is set to false and the application writes to that mount point. This approach makes application state durable across instance replacements without requiring a distributed storage architecture or application code changes, because the application still sees a standard block device.

Why this answer

Using EBS volumes instead of instance store provides persistent storage that survives instance termination. By attaching an EBS volume to the EC2 instances and configuring the application to store state on that volume (e.g., through the same file path), the data is preserved even when instances are scaled in. This does not require modifying application code, only infrastructure configuration.

Option A is incorrect because reducing the cooldown period does not prevent data loss; it only affects scaling speed. Option B is incorrect because sticky sessions route users to the same instance but do not preserve session data when that instance is terminated during scale-in. Option C is incorrect because larger instances reduce the frequency of scaling events but data loss still occurs when instances are terminated.

Exam trap

The trap is assuming that sticky sessions (session affinity) provide resilience by routing users to the same instance. However, sticky sessions do not prevent data loss when that instance is terminated during scale-in events.

112
MCQeasy

A DevOps team uses the above CloudFormation template to create an S3 bucket. What does the bucket policy accomplish?

A.It denies all S3 operations on the bucket unless the request uses HTTPS.
B.It denies all read access to the bucket for anonymous users.
C.It prevents anyone from deleting objects in the bucket.
D.It allows only HTTPS requests to the bucket and denies all HTTP requests.
AnswerA

This statement uses an explicit Deny on all s3:* actions, conditioned on aws:SecureTransport being false. Any request made to the bucket over plain HTTP is therefore blocked regardless of the principal or whether an Allow policy exists; HTTPS requests are not affected by this Deny. The effect is to require TLS for every S3 API operation on the bucket.

Why this answer

The bucket policy uses a Deny effect with a condition that the request must use HTTPS (SecureTransport: false). This denies all S3 operations on the bucket unless the request is sent over HTTPS. Option B is incorrect because the policy does not target anonymous users specifically; it applies to all principals.

Option C is incorrect because the policy denies all actions, not just delete. Option D is incorrect because it allows HTTP requests when the condition is not met, but the Deny overrides; the policy explicitly denies non-HTTPS requests.

113
MCQeasy

A company wants to automate the recovery of an Amazon RDS DB instance in a different region if the primary region becomes unavailable. Which service should they use?

A.RDS Multi-AZ deployment.
B.RDS cross-region automated backups.
C.RDS read replicas.
D.AWS CloudFormation custom resource.
AnswerB

RDS cross-region automated backups continuously copy automated backups and applicable transaction logs from the source DB instance to a destination AWS Region without manual intervention. This allows the database to be restored to a specific point in time, or to the latest restorable time, in the secondary region, making recovery from a regional outage repeatable and largely automated. Because the backups are managed by RDS and restored through the standard restore workflow, this option directly supports automated cross-region recovery.

Why this answer

RDS cross-region automated backups can be restored to a different region. Option A is incorrect because RDS Multi-AZ only provides failover within the same region. Option C is incorrect because read replicas can be promoted but require manual intervention.

Option D is incorrect because RDS does not support CloudFormation for automated recovery across regions.

114
MCQhard

A company has a multi-region application with an RDS for MySQL database in us-east-1. They want to minimize downtime if the primary region fails. They set up a cross-region read replica in us-west-2. What additional step is needed for automated failover?

A.Create a second read replica in the secondary region
B.Use a custom automation to monitor the primary and promote the replica
C.Configure automatic backup retention on the replica
D.Enable Multi-AZ on the read replica
AnswerB

RDS does not natively perform cross-region automatic failover to a read replica. You must implement custom health monitoring, for example using Route 53 health checks or a Lambda function, to detect primary DB instance or AZ failure; upon detection, the automation invokes PromoteReadReplica to convert the replica into a standalone primary and then updates application DNS or database endpoints to redirect traffic. This is the only way to achieve the stated automated failover objective.

Why this answer

B is correct because RDS cross-region read replicas do not support automatic failover; you must use custom automation (e.g., AWS Lambda, Amazon Route 53 health checks, or a custom script) to detect primary region failure and promote the read replica to a standalone instance. This promotion breaks the replication link and makes the replica a writable primary, enabling failover.

Exam trap

The trap here is that candidates confuse Multi-AZ (which provides automatic failover within a region) with cross-region read replicas (which require manual or custom automation for failover).

How to eliminate wrong answers

Option A is wrong because creating a second read replica in the same secondary region does not enable automated failover; it only adds another read target and does not handle promotion or detection logic. Option C is wrong because automatic backup retention on the replica is a backup configuration, not a failover mechanism; it does not monitor the primary or promote the replica. Option D is wrong because Multi-AZ on a read replica provides high availability within a single region, not cross-region automated failover; it does not detect primary region failure or promote the replica.

115
MCQeasy

A company is designing a disaster recovery strategy for its primary RDS for PostgreSQL database in us-east-1. The RTO is 15 minutes and RPO is 1 minute. Which solution meets these requirements?

A.Create a cross-Region read replica in the secondary Region and promote it during failover.
B.Use AWS Backup to copy automated backups to the secondary Region every hour.
C.Deploy a Multi-AZ RDS instance and failover to the standby in the same region.
D.Take manual snapshots of the database every 5 minutes and copy them to the secondary Region.
AnswerA

A cross-Region read replica in RDS for MySQL, MariaDB, or PostgreSQL uses engine-native asynchronous replication to continuously stream changes from the primary in the source Region to a replica instance in the secondary Region. During a failover, you simply promote the replica, which stops replication and makes it a standalone primary database; promotion typically completes in minutes, giving you an RTO in minutes and an RPO of only seconds (or sub-second in low-write workloads). Because replication is ongoing rather than backup-based, this option gives the lowest RPO and RTO among the choices and is the correct DR strategy for cross-Region recovery.

Why this answer

A cross-Region read replica can be promoted to a primary instance in a matter of minutes, meeting the 15-minute RTO, and replication lag is typically under 1 minute, satisfying the 1-minute RPO. Option B (AWS Backup every hour) provides an RPO of up to 1 hour, exceeding the requirement. Option C (Multi-AZ in same region) does not provide cross-Region failover.

Option D (manual snapshots every 5 minutes) has a higher RPO than 1 minute and may not be promotable quickly enough.

116
MCQeasy

A company runs a production web application on EC2 instances behind an Application Load Balancer. The application experiences intermittent high latency. The operations team needs to identify the root cause without affecting live traffic. Which approach is the MOST efficient?

A.Deploy a separate test environment with identical configuration and run load tests
B.Enable EC2 detailed monitoring and SSH into each instance to run top and iostat
C.Enable detailed CloudWatch metrics on the ALB and analyze ALB access logs
D.Run tcpdump on all EC2 instances and analyze packet captures
AnswerC

Enabling detailed CloudWatch metrics on the ALB provides high-frequency measurements such as TargetResponseTime, RequestCount, and TargetConnectionErrorCount, and analyzing ALB access logs offers per-request timestamps, target processing time, request time, client IP, and target status codes. This combination yields a historical, request-level view of latency from the client through the ALB to the target without adding any agents or scripts to the EC2 instances. It also lets you filter for slow requests, group by URL or target, and determine whether delay occurs in the ALB-to-target hop or in the target application itself.

Why this answer

Enabling detailed CloudWatch metrics on the ALB and analyzing ALB access logs provides visibility into request latency patterns without affecting live traffic. Option A is wrong because setting up a separate test environment does not directly help diagnose the current intermittent issue. Option B is wrong because SSHing into instances and running commands can impact production performance and does not provide historical latency data.

Option D is wrong because tcpdump generates large packet captures that can degrade performance and requires significant analysis effort.

117
MCQeasy

A company wants to automatically recover an Amazon RDS DB instance if the underlying hardware fails. Which feature should the DevOps engineer enable?

A.Multi-AZ deployment.
B.Deletion protection.
C.Read replicas in a different Region.
D.Automated backups with a retention period of 35 days.
AnswerA

Multi-AZ maintains a synchronous standby in a different Availability Zone and automatically fails over to it when the primary's hardware fails, satisfying the stem's automatic recovery requirement. A single-AZ instance would remain unavailable until AWS or manual intervention restored it.

Why this answer

Multi-AZ deployment maintains a synchronous standby replica in a different Availability Zone and automatically fails over to it if the primary instance's hardware fails, the AZ becomes unavailable, or the instance is rebooted with failover. This is the only RDS feature designed for automatic hardware-failure recovery with minimal downtime (typically 60–120 seconds). Deletion protection, cross-Region read replicas, and automated backups do not provide automatic failover on hardware failure.

Exam trap

DOP-C02 often tests the distinction between high availability (Multi-AZ, automatic failover) and disaster recovery (cross-Region read replicas, backups), so candidates who see 'different Region' or 'backups' and assume automatic recovery pick the wrong option.

How to eliminate wrong answers

Option B is wrong because deletion protection only prevents accidental deletion of the DB instance via the console, CLI, or API — it has no role in detecting or recovering from hardware failure. Option C is wrong because cross-Region read replicas are asynchronous copies used for disaster recovery or read scaling; they do not automatically promote on primary hardware failure and require manual intervention (or custom automation) to fail over. Option D is wrong because automated backups with a 35-day retention period only enable point-in-time restore to a new instance after a failure — recovery is manual, takes time, and does not provide automatic failover.

118
MCQeasy

A DevOps engineer is designing a highly available web application using Amazon Route 53. The application is deployed in two AWS Regions. The engineer wants to route traffic to the nearest healthy endpoint. Which routing policy should be used?

A.Failover routing
B.Weighted routing
C.Geolocation routing
D.Latency routing
AnswerD

Latency routing directs users to the Region with the lowest measured latency, satisfying the nearest-healthy-endpoint requirement. Route 53 latency records support health checks, so unhealthy endpoints are excluded automatically. Unlike geolocation, which maps users by geography, latency uses actual network performance between resolver and endpoint.

Why this answer

Latency routing in Amazon Route 53 directs users to the AWS Region that provides the lowest network latency, which matches the requirement to route traffic to the nearest healthy endpoint. Route 53 uses latency measurements between the user and each AWS Region to select the optimal record. Health checks can be associated with latency records to ensure traffic only goes to healthy endpoints.

Exam trap

DOP-C02 often tests the distinction between latency and geolocation routing; candidates may incorrectly choose geolocation because they equate 'nearest' with geographic proximity, but latency routing is specifically for lowest network latency.

How to eliminate wrong answers

Option A is wrong because failover routing is designed for active-passive configurations, not for selecting the lowest-latency endpoint among multiple active regions. Option B is wrong because weighted routing distributes traffic based on assigned weights, not on network latency or proximity. Option C is wrong because geolocation routing directs traffic based on the geographic location of the user (e.g., country or continent), which does not guarantee the lowest latency and may not align with 'nearest healthy endpoint' if latency differs from geographic distance.

119
MCQhard

A DevOps engineer applies this S3 bucket policy to an S3 bucket. What is the effect of this policy?

A.All objects uploaded must be encrypted with SSE-C.
B.All uploads to the bucket are blocked.
C.All objects uploaded must use server-side encryption with Amazon S3 managed keys (SSE-S3).
D.All objects uploaded must be encrypted with SSE-KMS.
AnswerC

The policy grants s3:PutObject only when the s3:x-amz-server-side-encryption request header is exactly AES256. In S3, that header value maps to SSE-S3, where Amazon S3 manages the encryption keys using AES-256. Consequently, any object uploaded must be encrypted with SSE-S3, and uploads using no encryption or any other encryption method are denied.

Why this answer

The S3 bucket policy in question denies uploads unless the `x-amz-server-side-encryption` header is set to `AES256`, which is the value for SSE-S3 (Amazon S3 managed keys). This ensures that all objects uploaded to the bucket must be encrypted using server-side encryption with S3-managed keys (SSE-S3). Option C correctly identifies this requirement.

Exam trap

The trap here is that candidates confuse the `x-amz-server-side-encryption` header values: `AES256` is specific to SSE-S3, not SSE-C or SSE-KMS, leading to incorrect selections of A or D.

How to eliminate wrong answers

Option A is wrong because SSE-C requires the `x-amz-server-side-encryption-customer-algorithm` header, not `x-amz-server-side-encryption: AES256`. Option B is wrong because the policy does not block all uploads; it only denies uploads that do not meet the encryption requirement, so uploads with the correct encryption header are allowed. Option D is wrong because SSE-KMS requires the `x-amz-server-side-encryption` header set to `aws:kms`, not `AES256`.

120
MCQmedium

A company has deployed a multi-tier application on AWS. The web tier uses an Auto Scaling group of EC2 instances behind an Application Load Balancer. The application tier uses another Auto Scaling group of EC2 instances that process messages from an Amazon SQS queue. The database tier uses Amazon RDS Multi-AZ. Recently, the application experienced a complete outage when the SQS queue became overwhelmed with messages due to a sudden spike in traffic. The application tier could not process messages fast enough, causing the queue to grow indefinitely and eventually exceed the visibility timeout, leading to message loss and degraded performance. The DevOps engineer needs to improve the resilience of the architecture to handle traffic spikes without losing messages. Which solution should be implemented?

A.Limit the maximum message size and set a queue size limit to prevent overflow
B.Replace the standard SQS queue with a FIFO SQS queue to ensure exactly-once processing
C.Increase the visibility timeout in the SQS queue to allow more time for processing
D.Configure a dead-letter queue for unprocessed messages and implement Auto Scaling based on SQS queue depth
AnswerD

A dead-letter queue (DLQ) captures messages that remain unprocessed after multiple attempts, preserving them for later analysis or manual intervention, while Auto Scaling based on SQS queue depth dynamically adjusts the number of consumers to match incoming load. This combination ensures that spikes in traffic are handled by adding more processing capacity, and messages that cannot be processed are not permanently lost, providing both reliability and scalability.

Why this answer

It addresses both the message loss and processing bottleneck. A dead-letter queue (DLQ) captures messages that cannot be processed successfully after a specified number of attempts, preventing them from being lost when the visibility timeout expires. Implementing Auto Scaling based on SQS queue depth (using a CloudWatch alarm on the ApproximateNumberOfMessagesVisible metric) dynamically adds more application-tier EC2 instances when the queue grows, ensuring the processing rate scales with demand and prevents the queue from being overwhelmed.

Exam trap

The trap here is that candidates often focus on increasing visibility timeout or switching queue types to fix message loss, but they miss the core issue: the application tier lacks elasticity to scale with queue depth, and a DLQ is needed to safely capture messages that exceed processing attempts.

How to eliminate wrong answers

Option A is wrong because limiting maximum message size and setting a queue size limit does not prevent message loss; SQS does not have a configurable queue size limit (it can hold an unlimited number of messages), and reducing message size does not address the processing speed bottleneck. Option B is wrong because replacing a standard queue with a FIFO queue does not improve resilience to traffic spikes; FIFO queues have lower throughput (3000 messages per second with batching) and are designed for exactly-once ordering, not for handling high-volume bursts, and they do not prevent message loss from visibility timeout expiration. Option C is wrong because increasing the visibility timeout only delays when unprocessed messages become visible again; it does not prevent the queue from growing indefinitely or stop message loss if the application tier cannot keep up, and it can actually worsen the problem by hiding messages longer, leading to higher latency and potential starvation.

121
MCQhard

A company runs a production application on Amazon ECS with Fargate, fronted by an Application Load Balancer (ALB). The application experiences periodic latency spikes and occasional 502 errors. The ECS service is configured with a desired count of 2 tasks, and the ALB health check is set to /health with a 30-second interval and 2 consecutive failures threshold. The team uses CloudWatch Container Insights and has noticed that CPU and memory utilization of tasks remain below 50%. However, the ALB TargetGroup's HealthyHostCount metric occasionally drops to 0 for a few minutes before recovering. The deployment strategy is rolling update with a minimum healthy percent of 50% and maximum percent of 200%. The team recently updated the task definition to increase memory and CPU, but the issue persists. What is the MOST likely cause of the problem?

A.The ECS service role lacks permissions to register targets with the ALB.
B.The ALB health check is too aggressive, causing tasks to be marked unhealthy during brief initialization or deployment.
C.The ALB target group is configured with only one Availability Zone, causing loss of all targets when that AZ fails.
D.The task's CPU or memory limits are set too low, causing the container to be throttled.
AnswerB

The health check interval (30 seconds) and failure threshold (2) mean it takes up to 60 seconds to mark a task unhealthy. During deployments, the rolling update may temporarily have only 1 healthy task (minimum 50% of 2 = 1), and if that task becomes unhealthy, HealthyHostCount drops to 0.

Why this answer

The health check interval (30 seconds) and failure threshold (2) mean it takes up to 60 seconds to mark a task unhealthy. During deployments, the rolling update may temporarily have only 1 healthy task (minimum 50% of 2 = 1), and if that task becomes unhealthy, HealthyHostCount drops to 0. Option A is wrong because CPU and memory are below 50%, so resource limits are not the issue.

Option C is wrong because a target group with 2 tasks and 2 AZs is fine; the problem is not AZ-specific. Option D is wrong because ECS service-linked role does not affect health checks.

122
MCQeasy

A company is designing a disaster recovery strategy for its on-premises database to AWS using AWS Elastic Disaster Recovery (AWS DRS). The recovery time objective (RTO) is 15 minutes, and the recovery point objective (RPO) is 1 minute. Which configuration should they use?

A.Use AWS CloudEndure Disaster Recovery (now AWS DRS) with periodic replication every hour.
B.Use AWS Backup to take snapshots every 5 minutes and restore in the target region.
C.Use RDS cross-region automated backups with a 5-minute backup window.
D.Configure AWS DRS to continuously replicate data to a staging area in AWS, and launch instances in the target region on failover.
AnswerD

AWS DRS provides continuous block-level replication from source servers to a low-cost staging area in the AWS target Region, capturing every data change and maintaining a nearly synchronous copy. On failover, DRS automatically launches EC2 instances from the most recent consistent recovery point, achieving an RPO of seconds (well under the required 1 minute) and an RTO that can be measured in minutes. This directly meets the stated DR requirements, making it the correct choice.

Why this answer

AWS DRS (formerly CloudEndure) provides continuous block-level replication with sub-second RPO, which meets the 1-minute RPO requirement. By replicating to a staging area and launching instances in the target region on failover, it can achieve an RTO of 15 minutes or less, as the instances are pre-configured and ready for rapid launch.

Exam trap

The trap here is that candidates may confuse periodic backup solutions (like AWS Backup or RDS automated backups) with continuous replication, failing to recognize that only AWS DRS can achieve sub-minute RPO and rapid RTO for on-premises databases.

How to eliminate wrong answers

Option A is wrong because periodic replication every hour would result in an RPO of up to 60 minutes, far exceeding the required 1-minute RPO. Option B is wrong because AWS Backup snapshots every 5 minutes cannot achieve a 1-minute RPO, and restoring from snapshots typically takes longer than 15 minutes, failing the RTO. Option C is wrong because RDS cross-region automated backups have a minimum backup window of 5 minutes, which does not meet the 1-minute RPO, and this option is specific to RDS, not the on-premises database being migrated.

123
Multi-Selecthard

A company is designing a disaster recovery plan for an Amazon S3 data lake. The data lake stores sensitive data that must be replicated to a secondary Region with an RPO of 15 minutes. Which THREE actions should the company take? (Choose THREE.)

Select 3 answers
A.Enable S3 Versioning on both the source and destination buckets.
B.Enable S3 Replication Time Control (RTC) for the replication rule.
C.Configure cross-Region replication (CRR) from the source bucket to the destination bucket.
D.Configure S3 Event Notifications to trigger a Lambda function that copies objects to the secondary Region.
E.Enable S3 Transfer Acceleration on the source bucket.
AnswersA, B, C

S3 Versioning must be enabled on both the source and destination buckets because S3 Cross-Region Replication (CRR) requires versioned buckets to function. Versioning allows every object version, including overwrites and delete markers, to be captured and replicated, ensuring a complete and recoverable history in the DR Region. Without versioning, CRR cannot replicate updates or track deletes consistently, and the RPO guarantee would be unattainable.

Why this answer

Enabling S3 Versioning on both the source and destination buckets is a prerequisite for S3 Replication. Without versioning, replication cannot track object versions, which is essential for meeting the 15-minute RPO with consistency guarantees.

Exam trap

The trap here is that candidates may confuse S3 Event Notifications with Lambda as a viable replication method, overlooking that native CRR with RTC provides guaranteed RPO and versioning consistency without custom code overhead.

124
MCQhard

A company runs a critical application on EC2 instances behind an Application Load Balancer. The application uses an Amazon RDS for PostgreSQL Multi-AZ DB instance. During a recent failover test, the application experienced a 5-minute downtime. The RDS failover completed within 30 seconds. What is the most likely cause of the prolonged downtime?

A.The application caches DNS resolutions, causing it to connect to the old writer endpoint
B.The RDS Multi-AZ failover took longer than expected due to a large transaction log
C.The Application Load Balancer health checks marked all instances as unhealthy during the failover
D.The application was using read replicas for writes, which failed during failover
AnswerA

After a Multi-AZ failover, the RDS DNS record for the writer endpoint is updated to point to the new primary instance, but an application that caches DNS resolutions continues sending write connections to the old IP. That old IP belongs to the former primary, which is now promoted to standby (or replaced) and will not accept write connections, so all new database requests fail until the TTL expires or the application is restarted. This directly causes the five-minute outage because the application never re-resolves the endpoint until the cached entry times out.

Why this answer

The most likely cause is that the application caches DNS resolutions, causing it to continue connecting to the old writer endpoint after failover. When an RDS Multi-AZ failover occurs, the DNS record for the writer endpoint is updated to point to the new primary instance, but the application's cached DNS entry still points to the old IP address. Since the old primary is now a standby and no longer accepts connections, the application experiences downtime until the DNS cache expires (typically 5–60 seconds) or the application refreshes the DNS resolution.

The 5-minute downtime suggests the application uses a long DNS TTL or a custom caching layer that delays reconnection.

Exam trap

The trap here is that candidates assume the 5-minute downtime must be caused by the database failover itself, but the question explicitly states the failover completed in 30 seconds, so the real issue is application-side DNS caching or stale connection handling.

How to eliminate wrong answers

Option B is wrong because the scenario explicitly states the RDS failover completed within 30 seconds, so a large transaction log did not cause the prolonged downtime. Option C is wrong because the Application Load Balancer health checks are independent of RDS failover; even if the database is briefly unavailable, the ALB does not mark instances unhealthy unless the application itself fails health checks due to database connectivity issues. Option D is wrong because read replicas are not used for writes in a standard RDS Multi-AZ setup; writes always go to the primary instance, and read replicas are read-only, so this scenario does not apply.

125
MCQhard

A company uses DynamoDB global tables with two regions. They notice that writes in one region are not replicating to the other region after a brief network partition. Which configuration will ensure replication resumes automatically?

A.Use DynamoDB Streams with a Lambda function to manually replicate writes.
B.No action needed; DynamoDB automatically resumes replication when connectivity is restored.
C.Manually fail over the table to the other region.
D.Delete the replica table and recreate it.
AnswerB

No operator action is required because DynamoDB global tables are designed to tolerate temporary partition failures and network interruptions by persisting replication metadata and resuming from the correct point in the replication stream once connectivity is restored. Each write is durably committed in the source region first, and then the service replays it to other replicas with de-duplication and ordering, so the system automatically reconciles the data across regions. This resilience is a core feature, so no manual intervention is necessary.

Why this answer

DynamoDB global tables use a last-writer-wins (LWW) conflict resolution mechanism and are designed to handle temporary network partitions. When connectivity is restored, DynamoDB automatically resumes replication of pending writes using its internal replication engine, which is built on DynamoDB Streams. No manual intervention is required because the service manages eventual consistency across regions.

Exam trap

The trap here is that candidates may assume network partitions require manual recovery actions, but DynamoDB global tables are designed to automatically resume replication without intervention, testing the understanding of the service's built-in resilience and eventual consistency model.

How to eliminate wrong answers

Option A is wrong because DynamoDB global tables already use DynamoDB Streams internally for replication; adding a separate Lambda function would introduce unnecessary complexity, latency, and potential data inconsistency, and it does not leverage the built-in conflict resolution. Option C is wrong because manual failover is not a replication resumption mechanism; it is used for disaster recovery to change the write region, but it does not address the automatic resumption of replication after a network partition. Option D is wrong because deleting and recreating the replica table would cause data loss and downtime, and it is an extreme, unnecessary action when DynamoDB automatically recovers replication upon reconnection.

126
MCQeasy

A company is running a production database on Amazon RDS for PostgreSQL with Multi-AZ deployment. The database experiences a failover due to an AZ outage. What happens to the existing database connections during the failover?

A.Existing connections are automatically redirected to the standby without interruption.
B.The RDS endpoint IP address changes, and the application must update its configuration.
C.Existing connections are dropped, and applications must reconnect to the new primary using the same endpoint.
D.The primary DB instance is promoted to standby and connections remain active.
AnswerC

A Multi-AZ failover promotes the standby instance to primary and updates the RDS DNS record to the new primary's IP, but existing connections are not preserved. Because the old primary is no longer reachable, active TCP sessions and open database connections are dropped, forcing client applications to re-establish them. Applications must reconnect to the same CNAME endpoint—ideally with retry/backoff logic—to transparently recover once DNS propagation completes. This is the expected and documented behavior.

Why this answer

During an RDS Multi-AZ failover, the DNS CNAME for the DB instance is updated to point to the standby, which is promoted to primary. However, existing TCP connections are severed because the standby has a different IP and the failover breaks the session; applications must reconnect using the same endpoint hostname, which now resolves to the new primary.

Exam trap

DOP-C02 often tests whether candidates know that the RDS endpoint hostname stays the same during failover (only the IP changes) — many incorrectly believe the application must update its configuration with a new endpoint.

How to eliminate wrong answers

Option A is wrong because RDS does not transparently redirect existing connections — the failover drops them, and the application must re-establish connections. Option B is wrong because the RDS endpoint hostname (CNAME) does not change; only the underlying IP it resolves to changes, so the application does not need to update its configuration. Option D is wrong because the primary is not 'promoted to standby' — the standby is promoted to primary, and the old primary becomes the new standby (or is replaced).

127
MCQmedium

A company runs a critical REST API on Amazon ECS Fargate behind an Application Load Balancer. The operations team wants the service to automatically recover from a failed task without manual intervention. They have configured a target group health check, but the ECS service still does not replace unhealthy tasks. What should the DevOps engineer do to ensure automatic task replacement?

A.Enable the ECS service to use the load balancer target group health check by specifying the target group ARN in the service definition.
B.Create an Amazon CloudWatch alarm based on the UnhealthyHostCount metric and configure it to invoke an AWS Lambda function that calls the ECS StopTask API.
C.Set the deployment configuration to use the rolling update strategy with a minimum healthy percent of 100 and a maximum percent of 200.
D.Configure the ECS service to use a container health check with the HEALTHCHECK instruction in the Dockerfile.
AnswerA

When an ECS service is configured with a load balancer, ECS uses the target group health check to determine task health. If a task fails the health check, ECS will stop the task and start a new one. Specifying the target group ARN in the service definition is required for ECS to monitor and replace unhealthy tasks. Without this, ECS does not know which target group to use for health checks.

Why this answer

For ECS services with a load balancer, ECS uses the target group health check to determine if a task is healthy. When a task fails the health check, ECS automatically stops it and starts a new one. This behavior requires that the service definition includes the load balancer configuration with the target group ARN.

Without this, ECS does not monitor task health via the load balancer and will not replace unhealthy tasks.

Exam trap

The trap here is assuming that a container health check alone is sufficient for ECS to replace unhealthy tasks, when in fact the service must be integrated with the load balancer target group health check.

128
MCQmedium

Your company runs a multi-tier web application on AWS. The web tier consists of EC2 instances behind an Application Load Balancer (ALB) in an Auto Scaling group across three Availability Zones. The application tier runs on a separate Auto Scaling group of EC2 instances that process requests from the web tier. The database tier uses an Amazon RDS for PostgreSQL Multi-AZ deployment. All application servers write logs to Amazon CloudWatch Logs. Recently, the operations team reported that during peak hours, the web tier experiences intermittent 503 errors. The ALB access logs show that the errors occur when the target group's healthy host count drops to zero momentarily. The Auto Scaling group's minimum and desired capacity is 6, with a maximum of 12. The scaling policy is based on average CPU utilization, with a target of 60%. The health check grace period is 300 seconds. The application health check endpoint returns a 200 status when healthy. The DevOps engineer suspects that the scaling policy is too slow to react to traffic spikes. The engineer wants to implement a more proactive scaling approach. Which solution should the engineer implement?

A.Implement a predictive scaling policy combined with dynamic scaling to proactively adjust capacity based on forecasted traffic.
B.Implement a scheduled scaling policy that increases capacity 30 minutes before the expected peak.
C.Increase the health check grace period to 600 seconds to give new instances more time to become healthy.
D.Switch to a step scaling policy with a lower cooldown period and a greater scaling adjustment.
AnswerA

Predictive scaling uses the Auto Scaling group's historical load data with machine learning to forecast future demand, generating scheduled scaling actions before the predicted load arrives. By pairing it with dynamic scaling (target tracking or step), you cover both the expected trend and real-time deviations, ensuring that instances are already available before traffic peaks. This proactive capacity prevents the healthy host count from ever collapsing to zero, unlike purely reactive policies that still have to wait for alarms and instance startup.

Why this answer

Predictive scaling uses machine learning to forecast future traffic based on historical patterns and proactively provisions capacity ahead of predicted demand, which directly addresses the slow reaction of reactive CPU-based scaling. Combining predictive scaling with dynamic scaling provides both proactive baseline capacity and reactive adjustments for unexpected spikes. This is the AWS-recommended approach for cyclical traffic patterns with intermittent 503 errors caused by insufficient capacity during peaks.

Exam trap

The trap is assuming that tuning reactive scaling parameters (cooldown, step adjustments, grace period) will solve a proactive capacity problem, when the real fix is forecasting-based scaling.

How to eliminate wrong answers

Option B is wrong because scheduled scaling only works for predictable, fixed-time peaks and does not adapt to variable or forecasted traffic; it also requires manual schedule maintenance and may over- or under-provision. Option C is wrong because increasing the health check grace period only delays health check evaluation for new instances; it does not speed up scaling and could actually delay detection of unhealthy instances, worsening the issue. Option D is wrong because step scaling with a lower cooldown and larger adjustment is still reactive — it responds after CPU thresholds are breached, so it cannot prevent the momentary zero healthy hosts during rapid spikes.

129
Multi-Selectmedium

A company is building a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application is expected to have unpredictable traffic patterns. The DevOps team needs to ensure that the application can handle sudden spikes in traffic without throttling. Which TWO actions should the team take? (Choose TWO.)

Select 2 answers
A.Use DynamoDB on-demand capacity mode for the table.
B.Configure Lambda provisioned concurrency to keep a set number of execution environments warm.
C.Configure DynamoDB auto scaling with a minimum capacity of 10 read/write capacity units.
D.Increase the Lambda function timeout to the maximum (15 minutes).
E.Set API Gateway throttling limits to a high value to prevent throttling.
AnswersA, B

On-demand instantly scales to handle spikes.

Why this answer

DynamoDB on-demand capacity mode automatically scales to handle unpredictable traffic spikes without requiring capacity planning or throttling. This mode charges per request and can accommodate sudden bursts of traffic up to the table's previous peak, making it ideal for serverless applications with variable workloads.

Exam trap

The trap here is that candidates often confuse DynamoDB auto scaling with on-demand capacity, thinking auto scaling can handle sudden spikes as effectively as on-demand, but auto scaling has a lag time and can still throttle during rapid bursts.

130
MCQmedium

A company runs a web application on AWS that uses Amazon SQS to decouple the frontend from the backend processing. The application experiences sudden spikes in traffic, causing the SQS queue to accumulate a large number of messages. The backend workers are unable to process messages fast enough, leading to increased latency. What solution can the company implement to improve the resilience and scalability of the backend?

A.Reduce the receive message wait time (long polling) to poll the queue more frequently.
B.Increase the visibility timeout of the SQS queue to allow more time for processing.
C.Use an SQS FIFO queue instead of a standard queue to ensure ordered processing.
D.Configure an Auto Scaling group for the backend workers with a scaling policy based on the SQS queue depth.
AnswerD

The correct solution is to attach an Auto Scaling policy to the SQS queue depth metric (e.g., ApproximateNumberOfMessagesVisible) and configure the backend workers as an Auto Scaling group. As the number of available messages grows, the policy launches additional EC2 workers to increase aggregate polling and processing throughput; as the queue drains, it terminates excess workers. This directly ties compute capacity to the ingested message volume, which is the standard pattern for decoupled, event-driven autoscaling with SQS.

Why this answer

Configuring an Auto Scaling group for the backend workers with a scaling policy based on the SQS queue depth (ApproximateNumberOfMessagesVisible) directly addresses the sudden traffic spikes. This approach dynamically adds more worker instances when the queue depth increases, improving processing throughput and reducing latency. It ensures the backend scales in response to demand, enhancing both resilience and scalability.

Exam trap

The trap here is that candidates often confuse operational fixes (like adjusting polling or visibility timeout) with architectural scalability solutions, failing to recognize that only dynamic scaling of compute resources can handle unpredictable traffic spikes.

How to eliminate wrong answers

Option A is wrong because reducing the receive message wait time (long polling) to poll more frequently would increase the number of empty responses and API calls, potentially throttling the workers without improving processing capacity; long polling (wait time up to 20 seconds) is actually more efficient for reducing latency and empty receives. Option B is wrong because increasing the visibility timeout only gives workers more time to process a single message, but does not address the root cause of insufficient worker capacity; it can even cause message processing delays if workers fail and messages become visible again after the timeout. Option C is wrong because using an SQS FIFO queue ensures exactly-once processing and message ordering, but does not improve throughput or scalability; FIFO queues have a lower throughput limit (300 transactions per second without batching) compared to standard queues, which would worsen the backlog during spikes.

131
Multi-Selectmedium

Which TWO AWS services can be used to distribute incoming traffic across multiple AWS resources in different Availability Zones within a single region?

Select 2 answers
A.AWS Global Accelerator
B.Amazon Route 53
C.Amazon CloudFront
D.AWS Direct Connect
E.Application Load Balancer
AnswersA, E

AWS Global Accelerator uses static anycast IPs at AWS edge locations to route traffic over the AWS global network to endpoint groups, which can contain NLB or ALB endpoints in multiple Availability Zones. It performs health checks and weight-based traffic distribution across endpoints, so incoming TCP/UDP connections are deliberately distributed across AZs rather than being resolved once by DNS. This edge-to-backbone path also improves latency while keeping the same IP addresses during failover.

Why this answer

AWS Global Accelerator uses the AWS global network and Anycast static IP addresses to route incoming traffic to the optimal endpoint, such as an Application Load Balancer or EC2 instance, across multiple Availability Zones within a single region. It improves performance and reliability by directing traffic to the healthiest endpoint and automatically rerouting in case of failure, making it a valid service for distributing traffic across AZs.

Exam trap

The trap here is that candidates often think only Elastic Load Balancers (like ALB) can distribute traffic across AZs, but AWS Global Accelerator also performs this function at the network layer, and the question asks for TWO services, so both ALB and Global Accelerator are correct.

132
MCQhard

A company uses AWS Lambda functions to process events from an Amazon SQS queue. The Lambda function occasionally fails due to a transient downstream service error. The DevOps team wants to ensure that failed messages are not lost and can be retried later. The team also wants to reduce the number of invocations on the downstream service. Which configuration should the team use?

A.Configure a dead-letter queue (DLQ) on the SQS queue and set the Lambda function's reserved concurrency to 1.
B.Configure an Amazon SNS topic as a Lambda destination for failure events and subscribe the SQS queue to it.
C.Configure a dead-letter queue (DLQ) on the Lambda function and set the function's maximum retry attempts to 2.
D.Configure the Lambda function to write failed messages to an Amazon DynamoDB table and set up a scheduled Lambda to retry.
AnswerA

Setting a dead-letter queue (DLQ) on the SQS queue ensures that messages which exhaust their retry attempts are preserved for later inspection, while assigning reserved concurrency of 1 to the Lambda function caps the maximum number of concurrent invocations to exactly one. This hard limit prevents Lambda from scaling out to hundreds of executions when the downstream service is slow or failing, because the SQS event source mapping can only invoke one function at a time, thereby throttling the rate of calls to the downstream service. As a result, the downstream service receives at most one in‑flight request, avoiding overload and allowing it to recover gracefully. The queue‑level DLQ captures messages that ultimately fail after all retries, so no data is lost while the concurrency limit protects the bottleneck.

Why this answer

Configuring a dead-letter queue (DLQ) on the SQS queue ensures that messages that exhaust their retries (due to Lambda failures) are preserved for later reprocessing, preventing data loss. Setting the Lambda function's reserved concurrency to 1 throttles the function to a single concurrent invocation, which naturally reduces the rate of downstream service calls and allows the SQS queue's visibility timeout and redrive policy to manage retry timing, thereby reducing pressure on the downstream service.

Exam trap

The trap here is that candidates often confuse a Lambda function's DLQ (which captures invocation records) with an SQS queue's DLQ (which captures the original messages), and they overlook that reserved concurrency is a direct way to throttle invocation rate, not just a capacity planning tool.

How to eliminate wrong answers

Option B is wrong because using an SNS topic as a Lambda destination for failure events and subscribing the SQS queue to it would create an asynchronous loop where failed events are re-sent to the same SQS queue, potentially causing infinite retries without a controlled retry mechanism or throttling to protect the downstream service. Option C is wrong because a dead-letter queue on the Lambda function (via Lambda destinations) only captures invocation records, not the original SQS messages; setting maximum retry attempts to 2 on the Lambda function does not reduce downstream service invocations—it actually increases them by retrying immediately without backoff. Option D is wrong because writing failed messages to DynamoDB and using a scheduled Lambda to retry adds unnecessary complexity and latency, and does not inherently reduce downstream service invocations; it also bypasses SQS's built-in retry and DLQ mechanisms, which are simpler and more reliable for transient failures.

133
MCQhard

A company uses Amazon DynamoDB with global tables for a multi-region active-active application. They notice that occasionally, concurrent updates to the same item in different regions cause data inconsistency. How can they resolve this?

A.Disable global tables and use a single region
B.Use DynamoDB read replicas instead of global tables
C.Use conditional writes and design the application to handle conflicts
D.Use DynamoDB Accelerator (DAX) to cache writes
AnswerC

Conditional writes use attribute-based conditions (e.g., version numbers or timestamps) to ensure a write only succeeds if the item hasn't changed since the last read, preventing blind overwrites. Since DynamoDB global tables use last-writer-wins by default, application-level conflict resolution—such as storing version metadata and retrying with reconciliation logic—gives you control over which update wins. This approach aligns with DynamoDB's recommended pattern for multi-region applications that require strong consistency for business-critical updates.

Why this answer

DynamoDB global tables use an eventually consistent model for multi-region replication, meaning concurrent updates to the same item in different regions can lead to conflicts. Conditional writes allow the application to enforce a last-writer-wins (LWW) strategy or custom conflict resolution logic, ensuring data consistency by only applying updates that meet specified conditions (e.g., a version number or timestamp check). This approach aligns with the recommended practice for handling concurrent writes in an active-active global table setup.

Exam trap

The trap here is that candidates often assume DynamoDB global tables automatically resolve all write conflicts, but the exam tests the understanding that without conditional writes or custom conflict resolution, concurrent updates can cause data inconsistency due to eventual consistency.

How to eliminate wrong answers

Option A is wrong because disabling global tables and using a single region eliminates multi-region active-active capability, which is a core requirement of the scenario, and does not resolve the underlying conflict issue—it just avoids it by sacrificing availability and latency benefits. Option B is wrong because DynamoDB read replicas (via global tables or otherwise) are designed for read scaling, not for handling concurrent writes; they do not address write conflicts or provide write consistency across regions. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for read-heavy workloads that reduces read latency, but it does not manage write conflicts or provide cross-region write consistency; caching writes does not resolve the fundamental issue of concurrent updates to the same item in different regions.

134
MCQeasy

A company wants to ensure its RDS Multi-AZ deployment automatically fails over to a standby instance in a different Availability Zone. Which additional step is required?

A.Create an Amazon Route 53 health check to update the DNS record.
B.No additional step; RDS Multi-AZ handles automatic failover.
C.Configure a read replica in another AZ.
D.Deploy the standby instance in a different VPC.
AnswerB

No additional step is needed because RDS Multi-AZ is designed to handle automatic failover entirely within the managed service. When the primary instance becomes unavailable (due to hardware failure, AZ outage, or even a patching event), RDS automatically promotes the standby in the other Availability Zone and updates the DNS endpoint to point to the new primary within minutes. Applications using the existing RDS endpoint will reconnect automatically once the DNS change propagates, without requiring any manual intervention or external tooling. This is the inherent advantage of a Multi-AZ deployment for high availability.

Why this answer

Amazon RDS Multi-AZ automatically handles failover to the standby instance in a different Availability Zone without any additional configuration. When the primary DB instance becomes unavailable, RDS automatically detects the failure and flips the DNS record to point to the standby instance, typically within 60–120 seconds. No manual intervention or extra services like Route 53 health checks are required for this built-in failover mechanism.

Exam trap

The trap here is that candidates overthink the need for external DNS management (like Route 53 health checks) or confuse read replicas with Multi-AZ standby instances, not realizing that RDS Multi-AZ is a fully managed, automatic failover solution that handles DNS updates internally.

How to eliminate wrong answers

Option A is wrong because Amazon RDS Multi-AZ already manages DNS record updates automatically during failover; creating a separate Route 53 health check is unnecessary and could introduce additional complexity or failover delays. Option C is wrong because a read replica is designed for read scaling and can be promoted to a primary, but it does not provide automatic synchronous replication or automatic failover as part of a Multi-AZ deployment; Multi-AZ uses a dedicated standby instance with synchronous replication. Option D is wrong because the standby instance must be in a different Availability Zone within the same VPC to maintain low-latency synchronous replication and automatic failover; deploying in a different VPC would break network connectivity and the replication mechanism.

135
MCQmedium

A company is deploying a stateful application on Amazon EKS. The application requires persistent storage that can be reattached to a new pod if the original pod fails. The cluster spans multiple Availability Zones. Which storage solution provides the BEST resilience and meets these requirements?

A.Amazon S3 bucket with a mountpoint.
B.Amazon EBS with gp3 volume type.
C.EC2 instance store volumes.
D.Amazon EFS file system.
AnswerD

Amazon EFS is a regional, elastic, fully managed NFS file system that is accessible from all Availability Zones in the region. It supports the ReadWriteMany access mode, allowing multiple pods across different nodes and AZs to share the same file system simultaneously. EFS integrates with the EKS CSI driver and provides strong consistency and durability, making it an ideal persistent storage solution for stateful applications deployed on EKS.

Why this answer

Amazon EFS provides a fully managed, regional NFS file system that can be mounted concurrently by multiple pods across different Availability Zones. It is designed for high availability and durability, automatically replicating data across multiple AZs, and supports automatic reattachment to a new pod if the original pod fails, making it the best choice for stateful applications requiring resilient, shared persistent storage on Amazon EKS.

Exam trap

The trap here is that candidates often assume EBS is the default persistent storage for Kubernetes because of its common use with single-node stateful workloads, but they overlook the multi-AZ requirement that makes EBS unsuitable due to its zonal scope, while EFS's regional nature provides the necessary cross-AZ resilience.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service, not a file system; using a mountpoint (e.g., s3fs) introduces POSIX compatibility issues, performance overhead, and does not provide the native file locking or consistent read-after-write semantics required for a stateful application's persistent storage. Option B is wrong because Amazon EBS volumes are tied to a single Availability Zone and cannot be reattached to a pod in a different AZ; if the original pod fails and a replacement pod is scheduled in another AZ, the EBS volume cannot be mounted, breaking resilience across the multi-AZ cluster. Option C is wrong because EC2 instance store volumes are ephemeral and data is lost if the instance stops, terminates, or fails; they do not provide persistent storage that survives pod or node failures.

136
MCQmedium

A company is designing a disaster recovery strategy for a critical application. They need a Recovery Time Objective (RTO) of 15 minutes and a Recovery Point Objective (RPO) of 1 minute. Which AWS database service configuration meets these requirements?

A.RDS MySQL with Multi-AZ and cross-region read replica
B.DynamoDB global tables
C.Aurora Global Database
D.RDS PostgreSQL with cross-region read replica
AnswerC

Aurora Global Database is the correct DR choice because it uses dedicated storage-level replication across regions with a typical RPO of sub-second and an RTO of under 1 minute when you promote a secondary region. The primary and secondary remain fully readable during normal operation, and promotion is a single API call with automatic DNS update, giving the best RTO/RPO among the options.

Why this answer

Aurora Global Database provides a fully managed cross-region disaster recovery solution with typical RPO of 1 second and RTO of 1 minute for regional failover, which meets the required RTO of 15 minutes and RPO of 1 minute. It uses storage-level replication that is asynchronous but very low-latency, and failover can be promoted to the secondary region in under a minute.

Exam trap

Common misconception: Any cross-region read replica (like RDS MySQL or PostgreSQL) can achieve sub-minute RPO and RTO. In reality, manual promotion steps and asynchronous replication lag make them unsuitable for strict 15-minute RTO and 1-minute RPO requirements. Aurora Global Database's storage-level replication achieves RPO of 1 second and RTO under 1 minute, meeting the requirements.

How to eliminate wrong answers

Option A is wrong because RDS MySQL Multi-AZ provides high availability within a single region, not cross-region DR, and cross-region read replicas have asynchronous replication with typical RPO of seconds to minutes but failover requires manual promotion and DNS changes, often exceeding 15 minutes RTO. Option B is wrong because DynamoDB global tables are designed for multi-region active-active workloads with eventual consistency, and while RPO is typically sub-second, RTO for regional failover can be minutes but requires application-side retry logic and does not guarantee 1-minute RPO under all failure scenarios. Option D is wrong because RDS PostgreSQL cross-region read replicas have similar limitations to MySQL: asynchronous replication with variable RPO and manual promotion steps that make achieving 15-minute RTO unreliable.

137
MCQeasy

A company is designing a highly available architecture for a web application. The application runs on Amazon EC2 instances in an Auto Scaling group across three Availability Zones. The instances are behind an Application Load Balancer (ALB). Which additional step should the team take to ensure that traffic is evenly distributed across all healthy instances in all Availability Zones?

A.Use Amazon Route 53 weighted routing policy to distribute traffic to each AZ.
B.Configure health checks on the target group to mark instances as unhealthy if they are in an AZ with fewer instances.
C.Enable cross-zone load balancing on the ALB.
D.Configure the ALB to use least outstanding requests routing algorithm.
AnswerC

This ensures even distribution across all instances in all AZs.

Why this answer

By default, ALB distributes traffic evenly across AZs, but cross-zone load balancing must be enabled to distribute traffic evenly across all instances regardless of AZ. Option A is wrong because Route 53 weighted routing is not needed for internal load balancing; the ALB already distributes traffic across AZs. Option B is wrong because configuring health checks to mark instances unhealthy based on AZ instance count would not help with even distribution and could cause unnecessary failures.

Option D is wrong because the least outstanding requests routing algorithm optimizes for request queue depth but does not enable cross-zone load balancing; cross-zone load balancing must be explicitly enabled.

138
Multi-Selectmedium

Which TWO strategies can be used to improve the resilience of an application running on Amazon ECS with Fargate? (Select TWO.)

Select 2 answers
A.Use a single subnet for all tasks to simplify networking.
B.Configure the ECS service to place tasks in multiple Availability Zones.
C.Increase the task memory reservation to handle peak load.
D.Implement a circuit breaker pattern for downstream dependencies.
E.Use scheduled scaling to adjust task count based on historical patterns.
AnswersB, D

Placing tasks in multiple Availability Zones is a core high-availability strategy because each AZ is an isolated, independent failure domain. The ECS service scheduler spreads tasks across the chosen subnets, so when one AZ is impacted by an outage, the tasks in the other AZs continue serving traffic and the service can still meet its desired count. This addresses resilience by ensuring no single infrastructure failure can take down the entire service.

Why this answer

Configuring the ECS service to place tasks in multiple Availability Zones distributes the application across physically separate data centers, so if one AZ fails, the tasks in other AZs continue to run. Option D is correct because implementing a circuit breaker pattern for downstream dependencies prevents cascading failures by detecting faults and failing fast, allowing the system to recover gracefully. Option A is incorrect; using a single subnet for all tasks typically places them in a single Availability Zone, reducing fault tolerance.

Option C is incorrect; increasing task memory reservation helps handle peak load but does not improve resilience against failures. Option E is incorrect; scheduled scaling adjusts capacity based on historical patterns and does not handle unexpected spikes or failures.

139
MCQeasy

A company runs a web application on EC2 instances behind an ALB. To improve resilience, they want to automatically replace failed instances and maintain a minimum number of instances. Which AWS service should be used?

A.Amazon EC2 Auto Scaling
B.AWS CloudFormation
C.AWS Elastic Beanstalk
D.AWS Systems Manager
AnswerA

Amazon EC2 Auto Scaling is the service that continuously monitors the health of EC2 instances using EC2 status checks and, when configured, Elastic Load Balancing health checks. If an instance fails these checks, Auto Scaling automatically terminates it and launches a replacement instance to maintain the desired or minimum fleet size. This health-based replacement is an inherent capability of an Auto Scaling group, making it the correct answer for automatically replacing unhealthy instances.

Why this answer

Amazon EC2 Auto Scaling is purpose-built to maintain a desired/minimum number of instances and automatically replace unhealthy ones via health checks integrated with ELB. It continuously monitors instance health and launches replacements when instances fail, ensuring the minimum capacity is preserved. This directly satisfies the resilience requirement without manual intervention.

Exam trap

DOP-C02 often tests the misconception that CloudFormation or Elastic Beanstalk 'does' auto scaling, when in fact Auto Scaling is the underlying service and the others merely orchestrate it — candidates must pick the service that directly provides the capability.

How to eliminate wrong answers

Option B is wrong because AWS CloudFormation is an infrastructure-as-code provisioning service — it can create an Auto Scaling group but does not itself perform health-based replacement or maintain minimum instance counts at runtime. Option C is wrong because Elastic Beanstalk is a PaaS layer that abstracts deployment; while it uses Auto Scaling under the hood, it is not the direct service for configuring automatic instance replacement and minimum capacity. Option D is wrong because AWS Systems Manager is for operational management (patching, run commands, parameter store) and does not provide automatic scaling or health-based instance replacement.

140
Multi-Selectmedium

A company runs a stateful web application on EC2 instances behind an ALB. The application stores session data in memory. The company wants to make the application stateless to improve resilience. Which TWO changes should the company make?

Select 2 answers
A.Increase the instance memory to store more sessions
B.Disable sticky sessions on the ALB
C.Enable sticky sessions (session affinity) on the ALB
D.Store session data in Amazon ElastiCache for Redis
E.Use an NLB instead of an ALB
AnswersB, D

Disabling sticky sessions on the ALB is a necessary precondition for a horizontally scalable, fault-tolerant design. With stickiness off, the ALB can route any request to any healthy target, so if an instance fails, the next request can be served by a different instance — assuming the session state is stored externally (for example, in ElastiCache or DynamoDB). This makes the application effectively stateless at the instance level, which also allows Auto Scaling to add or remove instances without worrying about breaking client sessions on a particular host.

Why this answer

To make the application stateless, the company should disable sticky sessions on the ALB (option B) and store session data in Amazon ElastiCache for Redis (option D). Disabling sticky sessions ensures that requests can be routed to any instance, and storing session data externally removes the dependency on in-memory state on individual instances, improving resilience. Option A is incorrect because increasing instance memory does not solve the statefulness issue.

Option C is incorrect because enabling sticky sessions would maintain state on instances. Option E is incorrect because using an NLB does not address session state management.

141
MCQeasy

A company wants to protect its S3 bucket data from accidental deletion or overwrite. Which feature should be enabled?

A.Enable cross-region replication
B.Apply a bucket policy that denies DeleteObject
C.Enable S3 Versioning
D.Enable MFA Delete
AnswerC

Enabling S3 Versioning is the correct first-line protection because it retains every version of an object, including the original, whenever an overwrite (PUT) or delete (DELETE) occurs. With versioning, an overwritten object's previous version is preserved as a non-current version, and a DELETE action only inserts a deletemarker, leaving all prior versions intact. You can recover accidental changes by simply fetching a previous version, making it the foundational mechanism that enables other features like lifecycle rules, MFA Delete, and point-in-time restores.

Why this answer

S3 Versioning preserves every prior version of an object, so an accidental overwrite creates a new version while the original remains recoverable, and an accidental delete inserts a delete marker rather than removing data. This directly satisfies the requirement to protect against both accidental deletion and overwrite without blocking legitimate writes. Versioning is the foundational control that MFA Delete and replication build upon.

Exam trap

The trap is assuming MFA Delete or a deny-DeleteObject bucket policy is the primary protection — both are secondary controls that depend on versioning already being enabled.

How to eliminate wrong answers

Option A is wrong because cross-region replication copies objects to another bucket but does not protect against deletion or overwrite in the source bucket — a deleted source object can also be deleted at the destination depending on replication configuration. Option B is wrong because a bucket policy denying s3:DeleteObject blocks all deletions, including legitimate ones, and does nothing to prevent overwrites; it is a blunt instrument, not a protection mechanism. Option D is wrong because MFA Delete only adds an extra authentication requirement for permanently deleting versions or suspending versioning — it requires versioning to already be enabled and does not by itself preserve overwritten data.

142
MCQmedium

A company uses AWS CodeDeploy for blue/green deployments to an Auto Scaling group. The deployment fails because the new instances do not pass health checks. The DevOps engineer discovers that the health check URL returns a 503 error. What is the MOST likely cause?

A.The target group health check path is '/health' but the application does not serve that endpoint
B.The CodeDeploy agent on the new instances is not running
C.The security group for the ALB does not allow inbound traffic on port 80
D.The Auto Scaling group health check type is set to EC2 instead of ELB
AnswerA

A 503 response from the ALB health check means the target instance accepted the TCP connection and returned an HTTP response, but the response status code was not a success (2xx/3xx). If the health check path is '/health' and the application does not define that route, the web server returns a 503 error because no handler matches the request. To resolve this, the health check path must be changed to an existing endpoint or the application must implement an endpoint that returns 200 OK on '/health'.

Why this answer

The health check URL returning a 503 error indicates that the application is not responding to the health check endpoint. Since the target group health check path is configured as '/health' but the application does not serve that endpoint, the ALB considers the instances unhealthy, causing CodeDeploy to fail the deployment. This is the most direct cause because the health check is failing at the application layer, not due to infrastructure issues.

Exam trap

The trap here is that candidates may confuse a 503 error with a network-level failure (like a security group blocking traffic) rather than recognizing it as an application-layer response indicating the health check endpoint is missing or misconfigured.

How to eliminate wrong answers

Option B is wrong because if the CodeDeploy agent were not running, the deployment would likely fail earlier (e.g., during the Install event) or the agent would not report success, but the health check failure (503) specifically indicates the application is running but not responding correctly. Option C is wrong because if the security group for the ALB did not allow inbound traffic on port 80, the health check would likely time out or return a connection refused error, not a 503 (Service Unavailable) which is an HTTP response from the application. Option D is wrong because the Auto Scaling group health check type (EC2 vs ELB) affects how ASG replaces unhealthy instances, but it does not directly cause the health check URL to return a 503; the 503 error is a symptom of the application not serving the correct endpoint.

143
Matchingmedium

Match each AWS CloudFormation concept to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Collection of AWS resources managed as a single unit

JSON or YAML document describing AWS resources

Preview of changes before applying to a stack

Enables stack creation across multiple accounts and regions

Identifies differences between stack and actual resource configurations

Why these pairings

Change Set previews stack changes; Stack is a resource collection; Stack Set manages stacks across accounts; Template is a JSON/YAML description.

144
MCQeasy

A company uses Amazon Route 53 to route traffic to an Application Load Balancer. They want to improve availability by routing traffic to multiple ALBs in different AWS Regions. Which routing policy should they use?

A.Latency-based routing policy
B.Weighted routing policy
C.Geolocation routing policy
D.Simple routing policy
AnswerA

Latency-based routing policy uses AWS-measured round-trip time data to direct each user to the endpoint that offers the fastest response, and it supports health checks on each record. When the primary endpoint fails its health check, Route 53 automatically selects the next lowest-latency healthy endpoint, providing seamless active-passive failover without manual intervention. This combination of performance optimization and health-aware failover makes it the correct choice for an application with multiple regional endpoints.

Why this answer

Latency-based routing policy uses AWS-measured round-trip time data to direct each user to the AWS Region offering the fastest response, and each record can have an associated Route 53 health check; when the primary region's endpoint fails its health check, Route 53 automatically routes to the next lowest-latency healthy region. This combination of performance optimization and health-aware failover, without manual intervention, makes it the best fit for improving availability and performance across multiple regional ALBs. (Note: weighted routing policy also supports per-record health checks and will automatically exclude an unhealthy endpoint from rotation, redistributing traffic among the remaining healthy weighted records -- it is not correct that weighted routing requires manual updates to fail over.)

Exam trap

The trap here is that candidates often confuse latency-based routing with geolocation routing, mistakenly thinking that geographic proximity equals low latency, but latency-based routing uses actual network measurements rather than fixed geographic boundaries.

How to eliminate wrong answers

Option B (Weighted routing policy) is wrong because it distributes traffic based on assigned weights (e.g., 80% to one region, 20% to another) and does not consider real-time latency or availability; it is designed for load balancing or testing, not for optimizing user-perceived performance across regions. Option C (Geolocation routing policy) is wrong because it routes traffic based on the geographic location of the user (e.g., country or continent), not on actual network latency or regional health; it can cause traffic to be sent to a distant region if the user's location is mapped there, even if that region is degraded. Option D (Simple routing policy) is wrong because it only supports a single record with multiple values (e.g., multiple IPs) and returns all values in a random order without any health checking or latency awareness, making it unsuitable for active-active multi-region failover.

145
Multi-Selecthard

A company uses DynamoDB global tables for a multi-region application. They notice that write conflicts are occurring. Which TWO strategies can reduce write conflicts?

Select 2 answers
A.Reduce read capacity units to limit concurrent reads
B.Enable DynamoDB Streams with last writer wins
C.Use conditional writes in the application code
D.Increase write capacity units on the table
E.Implement application-level conflict resolution
AnswersC, E

Conditional writes enable optimistic concurrency by allowing the application to assert a precondition—such as an item version or updated timestamp—before the write commits. If the condition evaluates to false because another concurrent write modified the item, DynamoDB rejects the request without overwriting, forcing the application to re-read and retry. This prevents silent data loss from last-writer-wins and is the appropriate DynamoDB-native way to enforce a safe update workflow in a multi-region setup.

Why this answer

Conditional writes prevent overwriting data unless a specified condition is met, thereby reducing write conflicts by ensuring that updates are only applied when the data is in a known state. Application-level conflict resolution allows the application to handle conflicts when they occur, using custom logic to merge or resolve differences, which reduces the impact of conflicts on the database. Option D (increasing write capacity) does not reduce conflicts; it only increases throughput capacity.

Option A (reducing read capacity) is unrelated to write conflicts. Option B (DynamoDB Streams with last writer wins) is the default behavior and does not reduce conflicts; it may cause data loss.

146
MCQmedium

A company uses AWS Lambda to process messages from an SQS queue. They need to ensure that if the Lambda function fails, the message is not lost and can be processed again. Which configuration is required?

A.Set the visibility timeout to less than the Lambda function timeout.
B.Enable SQS redrive policy to retry messages.
C.Configure a dead-letter queue (DLQ) on the SQS queue.
D.Set the Lambda event source mapping to not delete messages from the queue on failure.
AnswerD

The Lambda event source mapping for an SQS queue governs message deletion: it only deletes a message after the function returns a success response. By ensuring the mapping does not delete messages on failure (which is the standard behavior unless the function reports success), the failed message remains in the source queue. On subsequent polls, the event source mapping receives it again and invokes the function, giving the desired retry behavior. This is the correct way to let Lambda retry processing of a failed message.

Why this answer

The Lambda event source mapping for SQS can be configured to not delete messages from the queue if the function fails. This ensures that the message remains in the queue and becomes visible again after the visibility timeout expires, allowing it to be retried. Without this setting, Lambda automatically deletes messages upon successful processing, but on failure, the default behavior is to delete them as well, which would cause message loss.

Exam trap

The trap here is that candidates often confuse the dead-letter queue (DLQ) as the mechanism for retrying messages, when in fact it only stores messages after all retry attempts are exhausted, and the key to ensuring retries on failure is the event source mapping's delete behavior.

How to eliminate wrong answers

Option A is wrong because setting the visibility timeout to less than the Lambda function timeout would cause the message to become visible again before the function finishes, leading to duplicate processing, not preventing message loss. Option B is wrong because an SQS redrive policy moves messages to a dead-letter queue after a specified number of receive attempts, but it does not retry messages; it only redirects them after exhaustion of retries. Option C is wrong because configuring a dead-letter queue on the SQS queue is a best practice for capturing messages that cannot be processed after all retries, but it does not ensure that the message is retried on failure; it only stores failed messages after retries are exhausted.

147
MCQeasy

A company uses AWS CloudFormation to deploy infrastructure. During a recent deployment, the stack failed to create an Amazon RDS DB instance because of a parameter validation error. The DevOps engineer fixed the parameter and wants to resume the stack creation without recreating the resources that were already successfully created. The stack template is parameterized and uses nested stacks. What is the MOST efficient way to resume the stack creation?

A.Use the CloudFormation stack update operation with the corrected parameter.
B.Manually create the RDS instance with the corrected parameter and update the stack to import it.
C.Delete the entire stack and redeploy with the corrected parameter.
D.Use the 'ContinueUpdateRollback' feature to rollback the failed stack and then redeploy.
AnswerA

Calling the UpdateStack API (or the update operation in the console) on a stack in CREATE_FAILED state is the correct recovery path. CloudFormation only re-evaluates the resources that depend on the changed parameter, so the failed RDS instance is created while already-created resources remain untouched and no downtime is introduced. Any other approach either destroys existing resources or is not applicable to a creation failure.

Why this answer

CloudFormation stack updates can be used to fix the issue. By updating the stack with the corrected parameters, CloudFormation will only modify the failed resource and not recreate already created resources.

148
MCQeasy

A company uses Amazon DynamoDB as the database for a mobile application. The application requires single-digit millisecond read and write latency and must be resilient to the failure of an entire AWS Region. Which DynamoDB feature should the company use?

A.DynamoDB point-in-time recovery (PITR)
B.DynamoDB global tables
C.DynamoDB Accelerator (DAX)
D.DynamoDB on-demand capacity mode
AnswerB

DynamoDB global tables create a fully managed, multi-Region, multi-master replicated database using DynamoDB Streams and a backend replication engine. Changes made in any replica table are replicated to all other selected Regions with sub-second latency, enabling active-active failover: if one Region fails, traffic can be redirected to a remaining Region with minimal or zero downtime. This architecture directly addresses the requirement for low-latency access and resilience to Regional outages, making it the correct answer.

Why this answer

DynamoDB global tables provide multi-Region, multi-active replication, ensuring the application can withstand an entire AWS Region failure while maintaining single-digit millisecond read and write latency in each Region. This is achieved through DynamoDB Streams and a last-writer-wins conflict resolution mechanism, making it the correct choice for cross-Region resilience.

Exam trap

The trap here is that candidates often confuse high-availability features like DAX (caching) or PITR (backup) with true disaster recovery and multi-Region resilience, failing to recognize that only global tables replicate data across Regions for active-active failover.

How to eliminate wrong answers

Option A is wrong because point-in-time recovery (PITR) protects against accidental writes or deletions by enabling restoration to any point within the last 35 days, but it does not provide cross-Region resilience or continuous availability during a Region outage. Option C is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read latency but operates within a single Region and does not replicate data across Regions, offering no protection against a full Region failure. Option D is wrong because on-demand capacity mode handles traffic spikes automatically but is a scaling feature within a single Region, not a disaster recovery or multi-Region replication solution.

149
MCQeasy

A company uses Amazon Route 53 for DNS. They want to ensure that if their primary website endpoint fails, traffic is automatically routed to a secondary endpoint in a different Region. Which routing policy should be used?

A.Latency routing
B.Simple routing
C.Failover routing
D.Weighted routing
AnswerC

Failover routing is the correct choice because it explicitly pairs a primary record with a secondary record and uses Route 53 health checks to decide which to return. When the health check for the primary endpoint fails, Route 53 automatically returns the secondary record's value. This works in an active-passive configuration and gives the deterministic failover behavior the company requires.

Why this answer

Failover routing policy allows you to configure an active-passive failover setup.

150
MCQhard

Refer to the exhibit. A DevOps engineer runs the describe-target-health command and receives the output shown. The ALB target group has two instances. One instance is healthy, and the other is unhealthy with a 502 error. What is the most likely cause of the 502 error?

A.The security group for the instance does not allow inbound traffic on port 80 from the ALB.
B.The application running on the instance is not responding correctly or has crashed.
C.The instance's route table does not have a route to the internet gateway.
D.The health check path is configured to return a 404 status code.
AnswerB

When the ALB forwards a request to a healthy EC2 instance, the target must send a complete, valid HTTP response within the configured timeout. If the application logic crashes, the process closes the socket, or the web server returns a malformed response, the ALB cannot construct a valid HTTP response for the client and surfaces as a 502 Bad Gateway. Thus, a faulty or crashed application is the direct underlying cause of a 502 error.

Why this answer

A 502 Bad Gateway error from an ALB indicates that the target (EC2 instance) is not responding correctly or has closed the connection prematurely. This is commonly caused by the application or web server on the instance crashing or being unable to handle the request. Option A is incorrect because security group issues typically lead to connection timeouts (504) or refused connections, not 502.

Option C is incorrect because a missing route to the internet gateway would cause network unreachability, resulting in a different error. Option D is incorrect because the health check path returning a 404 would cause the target to fail health checks, not return a 502 during actual traffic.

← PreviousPage 2 of 3 · 184 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Resilient Cloud Solutions questions.