Courseiva

CCNA Reliability and Business Continuity Questions

75 of 205 questions · Page 1/3 · Reliability and Business Continuity · Answers revealed

1
MCQmedium

A company runs an Amazon RDS for MySQL DB instance in us-east-1. The SysOps administrator needs to implement a disaster recovery solution that can recover from a regional outage with a Recovery Point Objective (RPO) of less than 1 second and a Recovery Time Objective (RTO) of less than 1 minute. Which solution should the administrator use?

A.Multi-AZ deployment
B.Cross-region read replica
C.Aurora Global Database
D.Automated snapshot copy to another region
AnswerC

Aurora Global Database is specifically designed for cross-region disaster recovery. It uses a dedicated replication channel, typically with sub-second replication lag between the primary Region and up to five secondary Regions, which gives an RPO of usually under a second. A secondary Region can be promoted to primary in about a minute, providing a low RTO. This makes Aurora Global Database the correct choice when the requirement is for rapid, cross-region failover while losing almost no committed transactions.

Why this answer

Aurora Global Database is the correct choice because it provides a fully managed cross-region replication solution with a typical RPO of less than 1 second and an RTO of less than 1 minute during a regional failover. It uses a primary cluster in one region and up to five secondary clusters in other regions, with asynchronous replication that is optimized for low latency, meeting the stringent RPO/RTO requirements.

Exam trap

The trap here is that candidates often confuse Multi-AZ deployments (which are for high availability within a region) with cross-region disaster recovery, and they underestimate the replication lag and failover time of standard cross-region read replicas versus the optimized architecture of Aurora Global Database.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability within a single region, not cross-region disaster recovery, and its failover RTO is typically 1-2 minutes, exceeding the required 1 minute. Option B is wrong because a cross-region read replica for RDS MySQL uses asynchronous replication with a typical RPO of seconds to minutes, not less than 1 second, and promoting a read replica to a primary instance can take several minutes, failing the RTO requirement. Option D is wrong because automated snapshot copy to another region has an RPO of at least 5 minutes (the minimum snapshot interval) and restoring from a snapshot can take tens of minutes, both far exceeding the required RPO and RTO.

2
MCQhard

A company has a production DynamoDB table with on-demand capacity. They need to ensure business continuity with a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour in case of a regional outage. What is the MOST cost-effective solution?

A.Use AWS Backup to schedule daily backups and restore in another region.
B.Enable DynamoDB global tables for the table.
C.Enable point-in-time recovery (PITR) on the table.
D.Configure cross-region read replicas for the table.
AnswerB

DynamoDB global tables replicate your table automatically and asynchronously across multiple AWS Regions, typically with sub-second latency, which easily satisfies a 5-minute RPO. Because every replica accepts reads and writes, you can fail over instantly by rerouting application traffic (e.g., via Route 53 health checks) to a healthy region, meeting the 1-hour RTO without needing to restore data. Global tables also handle conflict resolution using last-writer-wins, ensuring data consistency across regions and making them the correct managed solution for regional disaster recovery.

Why this answer

DynamoDB global tables provide multi-region, multi-active replication with changes typically propagated within seconds, meeting an RPO of 5 minutes and an RTO of 1 hour by allowing traffic to be redirected to a replica in another region. This is the most cost-effective solution because it eliminates the need for separate backup storage or compute resources, and on-demand capacity scales automatically without provisioning.

Exam trap

The trap here is that candidates confuse point-in-time recovery (PITR) with cross-region disaster recovery, not realizing PITR is single-region only and cannot meet the RTO for a regional outage.

How to eliminate wrong answers

Option A is wrong because AWS Backup daily backups cannot achieve an RPO of 5 minutes (backups are at most daily) and restoring to another region would exceed the 1-hour RTO due to restore time. Option C is wrong because point-in-time recovery (PITR) only protects against accidental writes/deletes within the same region, not a regional outage, and restoring a table from PITR to another region would take longer than 1 hour. Option D is wrong because cross-region read replicas are read-only and cannot accept write traffic during a regional outage, failing the RTO requirement for write continuity.

3
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer across three Availability Zones. The application stores session data in an RDS MySQL database. To improve reliability, the company wants to ensure that a single Availability Zone failure does not impact the application's availability. Which combination of actions should the SysOps administrator take?

A.Configure the ALB to use only healthy instances and enable detailed CloudWatch metrics.
B.Increase the EC2 instance size to handle more traffic in a single AZ.
C.Increase the Auto Scaling group's desired capacity to a larger number.
D.Deploy RDS in Multi-AZ configuration with automatic failover, and enable cross-zone load balancing on the ALB.
AnswerD

Deploying RDS in a Multi-AZ configuration with automatic failover gives the database a synchronous standby replica in a different Availability Zone, so an AZ outage triggers a failover to the standby with minimal downtime. Enabling cross-zone load balancing on the ALB distributes incoming traffic across EC2 instances in multiple AZs, ensuring that if one AZ fails, the ALB continues to route requests to healthy instances in the remaining AZ. Together these actions provide both application-tier and database-tier high availability, which is the core requirement for running the workload across AZs and surviving an AZ failure.

Why this answer

Deploying RDS in Multi-AZ with automatic failover ensures database resilience against an AZ failure by maintaining a synchronous standby replica in a different AZ, while enabling cross-zone load balancing on the ALB distributes traffic across instances in all AZs, preventing a single AZ failure from taking down the entire application. Together, these actions address both the session data persistence and compute layer availability requirements.

Exam trap

The trap here is that candidates often think increasing instance count or size alone provides high availability, but they overlook the need for both database redundancy (Multi-AZ RDS) and cross-AZ traffic distribution (cross-zone load balancing) to survive an entire AZ failure.

How to eliminate wrong answers

Option A is wrong because configuring the ALB to use only healthy instances and enabling detailed CloudWatch metrics only improves monitoring and traffic routing to healthy targets, but does not provide redundancy for the RDS database or ensure compute capacity survives an AZ failure. Option B is wrong because increasing the EC2 instance size to handle more traffic in a single AZ does not eliminate the single point of failure; if that AZ fails, all instances are lost regardless of size. Option C is wrong because increasing the Auto Scaling group's desired capacity to a larger number does not guarantee instances are distributed across multiple AZs; without cross-zone load balancing and Multi-AZ RDS, a single AZ failure can still take down all instances and the database.

4
MCQmedium

A company runs a stateful web application on a single Amazon EC2 instance. The application stores session state in memory and writes critical data to an Amazon EBS volume. The SysOps administrator needs to implement a highly available architecture that can tolerate an Availability Zone (AZ) failure. The administrator plans to use an Auto Scaling group and an Application Load Balancer (ALB). Which combination of steps is required to make the application highly available while preserving session and data durability across AZ failures?

A.Create an AMI of the current instance, configure an Auto Scaling group with a launch template that uses the AMI, and attach the existing EBS volume to new instances.
B.Create a multi-AZ Auto Scaling group and use sticky sessions (session affinity) on the ALB to tie users to specific instances.
C.Use an Auto Scaling group across multiple AZs, migrate session storage to Amazon ElastiCache (multi-AZ), and migrate application data from EBS to Amazon EFS (file system mounted across AZs).
D.Use an Auto Scaling group in a single AZ and use a Multi-AZ RDS instance for data storage.
AnswerC

ElastiCache provides a shared, cross-AZ in-memory session store. EFS provides a shared, cross-AZ file system. The Auto Scaling group launches instances in multiple AZs, and the ALB distributes traffic. This architecture survives an AZ failure.

Why this answer

It addresses both session state and data durability across AZ failures. Migrating session storage to ElastiCache (multi-AZ) ensures session data survives instance failure, and migrating application data from EBS to EFS provides a shared, multi-AZ file system that persists independently of any single EC2 instance. This combination allows the Auto Scaling group to launch new instances in any AZ and immediately access both session and application data.

Exam trap

The trap here is that candidates often assume sticky sessions (session affinity) alone are sufficient for high availability, but they fail to realize that sticky sessions do not replicate session state across instances, so an instance failure still loses the session data.

How to eliminate wrong answers

Option A is wrong because attaching the existing EBS volume to new instances is not possible across AZs (EBS volumes are AZ-scoped) and does not provide a shared, durable data layer; it also fails to address session state persistence. Option B is wrong because sticky sessions alone do not preserve session data if the instance fails; they only route traffic to the same instance, and if that instance goes down, the session is lost. Option D is wrong because using a single AZ for the Auto Scaling group cannot tolerate an AZ failure, and while Multi-AZ RDS handles database durability, it does not address the application's in-memory session state or EBS-stored data.

5
MCQeasy

A company runs a web application on EC2 instances in an Auto Scaling group. The application is behind an Application Load Balancer. The company wants to ensure that the application can handle a sudden spike in traffic without downtime. What should the SysOps administrator do?

A.Use a scheduled scaling policy to add instances during business hours.
B.Configure a target tracking scaling policy based on average CPU utilization.
C.Reduce the number of Availability Zones to lower latency.
D.Manually increase the desired capacity of the Auto Scaling group when traffic increases.
AnswerB

A target tracking scaling policy based on average CPU utilisation automatically adjusts the desired capacity of the Auto Scaling group to maintain a predefined CPU target, such as 50%. This satisfies the requirement to handle a sudden traffic spike without downtime because the policy proactively adds EC2 instances as CPU load increases, preventing performance degradation before the application becomes overwhelmed.

Why this answer

A target tracking scaling policy based on average CPU utilization allows the Auto Scaling group to automatically adjust capacity in response to real-time demand spikes. This dynamic scaling approach maintains a target metric (e.g., 50% CPU) by adding or removing instances, ensuring the application can handle sudden traffic bursts without downtime.

Exam trap

The trap here is that candidates often confuse scheduled scaling (for predictable patterns) with dynamic scaling (for unpredictable spikes), or they mistakenly think reducing Availability Zones improves performance when it actually harms reliability.

How to eliminate wrong answers

Option A is wrong because scheduled scaling policies are designed for predictable traffic patterns (e.g., business hours), not for sudden, unplanned spikes; they cannot react to real-time changes. Option C is wrong because reducing the number of Availability Zones actually decreases fault tolerance and increases the risk of downtime during a zone failure, contradicting the goal of handling spikes without downtime. Option D is wrong because manually increasing desired capacity requires human intervention, which is too slow to respond to sudden spikes and defeats the purpose of automated elasticity.

6
Multi-Selecthard

A company runs a stateless web application on EC2 instances behind an Application Load Balancer. The SysOps Administrator needs to ensure the application can withstand the loss of an entire Availability Zone. Which THREE steps should be taken? (Choose THREE.)

Select 3 answers
A.Enable cross-zone load balancing on the ALB.
B.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
C.Ensure the ALB is configured to route traffic to all enabled AZs.
D.Configure the Auto Scaling group to use a dynamic scaling policy based on CPU utilization.
E.Use an Elastic IP address for each EC2 instance.
AnswersB, C, D

Specifying subnets in at least two Availability Zones within the Auto Scaling group causes Amazon EC2 Auto Scaling to launch instances across those zones, ensuring that if one Availability Zone fails, the remaining zones continue serving traffic. The Auto Scaling group also uses health checks to detect failed instances and automatically launches replacements in the other healthy Availability Zones, maintaining capacity and availability. This is the foundational design for a fault-tolerant, stateless web tier and directly addresses the requirement to survive an Availability Zone outage.

Why this answer

Configuring the Auto Scaling group to launch instances in at least two Availability Zones (AZs) ensures that if one entire AZ fails, the remaining AZ(s) still have running instances to serve traffic. This is a fundamental pattern for high availability and fault tolerance in AWS, as it distributes the application across physically separate data centers within a region.

Exam trap

The trap here is that candidates often confuse cross-zone load balancing (which optimizes traffic distribution within healthy AZs) with multi-AZ deployment (which ensures instance availability across AZs), leading them to incorrectly select Option A as a solution for AZ failure.

7
Multi-Selectmedium

A company is designing a disaster recovery strategy for a production RDS for MySQL database. The database is currently single-AZ. The recovery point objective (RPO) is 1 hour, and the recovery time objective (RTO) is 15 minutes. Which steps should the SysOps administrator take to meet these requirements? (Choose THREE.)

Select 3 answers
A.Disable automated backups to reduce performance impact.
B.Take manual DB snapshots every hour.
C.Enable automated backups with a retention period of at least 1 day.
D.Modify the DB instance to be Multi-AZ.
E.Create a read replica in a different Availability Zone.
AnswersB, C, D

Taking manual DB snapshots every hour creates a discrete recovery point every 60 minutes, so in the worst case you lose at most one hour of data, satisfying the stated RPO. Restoring from a manual snapshot provisions a new DB instance from that exact point; the RTO includes the time to restore the snapshot and update connection strings, which can be several minutes but is often acceptable for a recovery plan. Manual snapshots are stored in S3 and persist until you delete them, giving you a reliable backup mechanism independent of automated backup retention settings.

Why this answer

Manual DB snapshots can be taken on demand, and taking them every hour ensures that the recovery point objective (RPO) of 1 hour is met. In the event of a failure, you can restore the database from the latest manual snapshot, which provides a point-in-time recovery point within the RPO window. However, manual snapshots alone do not meet the 15-minute recovery time objective (RTO), so they must be combined with other measures like Multi-AZ and automated backups.

Exam trap

The trap here is that candidates often think a read replica or manual snapshots alone can meet both RPO and RTO, but they fail to recognize that Multi-AZ is required for the low RTO and automated backups are needed for the granular RPO.

8
Multi-Selecthard

A company runs a web application on EC2 instances in an Auto Scaling group. The application uses an Amazon RDS Multi-AZ DB instance. The SysOps administrator notices that during a recent failover test, the application became unresponsive for several minutes. The administrator wants to improve the application's resilience during failover. Which three actions should the administrator take? (Choose THREE.)

Select 3 answers
A.Configure the Application Load Balancer health checks to have a low threshold (e.g., 2 consecutive failures) and a short interval (e.g., 5 seconds).
B.Implement retry logic in the application to handle transient database connection failures.
C.Change the RDS DB instance to use asynchronous replication instead of synchronous replication.
D.Increase the EC2 instance size to handle more connections during failover.
E.Configure an Amazon RDS Proxy in front of the RDS database to pool and share database connections.
AnswersA, B, E

Configuring ALB health checks with a low threshold (e.g., 2 consecutive failures) and a short interval (e.g., 5 seconds) enables the load balancer to detect and deregister unhealthy instances within about 10 seconds, instead of the default ~60-90 seconds. This rapid detection prevents the ALB from continuing to route user traffic to instances whose database connections have been terminated during a failover. By quickly redirecting requests to healthy instances, the overall impact on users is minimized, and the application appears more resilient during RDS failover events.

Why this answer

Configuring the Application Load Balancer (ALB) health checks with a low threshold (e.g., 2 consecutive failures) and a short interval (e.g., 5 seconds) allows the ALB to quickly detect unhealthy EC2 instances and stop routing traffic to them. This reduces the time the application spends trying to serve requests through failing instances during an RDS failover, improving overall responsiveness.

Exam trap

The trap here is that candidates often assume increasing instance size (Option D) or changing replication mode (Option C) will improve failover resilience, but neither addresses the core issue of connection handling and rapid health check detection during a database failover.

9
MCQhard

A company runs a web application on EC2 instances in a private subnet. The application needs to connect to an RDS database in a different VPC. The VPCs are peered. The SysOps Administrator is troubleshooting connectivity issues. The RDS security group allows inbound traffic from the EC2 security group, but connections still fail. What could be the issue?

A.The RDS instance does not have public DNS resolution enabled.
B.The network ACL for the private subnet is blocking inbound traffic.
C.The route tables in each VPC do not have routes to the peered VPC CIDR.
D.The security group outbound rules on the EC2 instance are blocking traffic.
AnswerC

For VPC peering to function, each VPC's route table must contain a route to the CIDR block of the peered VPC, with the peering connection (e.g., pcx-xxxx) as the target. Without these routes, traffic destined for the other VPC is not directed to the peering connection and instead falls back to the local route, making the connection unreachable. This is the most common cause of failed peering connectivity and directly explains why the EC2 instance cannot reach the RDS database.

Why this answer

For traffic to flow between peered VPCs, each VPC's route table must have a route pointing to the CIDR block of the other VPC. Without these routes, packets from the EC2 instance in VPC A destined for the RDS database in VPC B will be dropped, even if security groups and network ACLs are permissive. The SysOps Administrator must add a route in the private subnet's route table for the RDS VPC's CIDR, and a corresponding route in the RDS VPC's route table for the EC2 VPC's CIDR, both pointing to the VPC peering connection.

Exam trap

The trap here is that candidates often assume security groups or network ACLs are the sole cause of connectivity failures in peered VPCs, overlooking the mandatory route table entries required for traffic to traverse the peering connection.

How to eliminate wrong answers

Option A is wrong because public DNS resolution is irrelevant for private connectivity within a VPC peering; RDS instances in a VPC use private DNS names that resolve to private IP addresses, and the EC2 instance can connect using the RDS endpoint without public DNS. Option B is wrong because network ACLs are stateless and must allow both inbound and outbound traffic; however, the question states connections fail, and the most common cause is missing routes, not NACL rules, and NACLs are evaluated before security groups. Option D is wrong because security group outbound rules on the EC2 instance are stateful; if the EC2 security group allows outbound traffic (which it does by default), responses from the RDS database are automatically allowed regardless of outbound rules, so this would not cause a failure.

10
MCQhard

A company runs a critical e-commerce application on Amazon ECS with Fargate launch type, fronted by an Application Load Balancer. The application uses an Amazon ElastiCache for Redis cluster for session state and an Amazon RDS for MySQL Multi-AZ database for persistent data. Recently, during a deployment of a new service version, the application became unresponsive for 15 minutes. The SysOps administrator discovered that the deployment updated the task definition with a new environment variable that pointed to an incorrect ElastiCache endpoint. The ECS service was configured with a rolling update, minimum healthy percent of 50%, and maximum percent of 200%. After the deployment, all tasks failed health checks due to a connection timeout to the wrong Redis endpoint. What is the MOST effective way to prevent this issue in future deployments?

A.Configure a CloudWatch alarm that triggers an automatic rollback if the error rate exceeds 10%.
B.Update the ECS service to use a canary deployment by updating one task at a time.
C.Implement a blue/green deployment strategy using AWS CodeDeploy and test the new task definition before shifting traffic.
D.Enable ECS deployment circuit breaker and set the rollback configuration to automatically roll back failed deployments.
AnswerC

CodeDeploy's blue/green deployment on ECS creates a new 'green' task set running the proposed task definition while the existing 'blue' task set continues to serve full production traffic. This lets you run smoke tests, endpoint checks, or a Lambda-based validation hook against the green task set before shifting any traffic from the blue to the green target group using the ALB's weighted routing. Only after you explicitly confirm the new task definition works can you shift traffic (optionally incrementally), and if issues are detected, you can re-shift back to the original high-fidelity task set with minimal disruption. This pre-validation is exactly what prevents the misconfigured variable from ever affecting end users.

Why this answer

Implementing a blue/green deployment with AWS CodeDeploy allows testing the new task definition in a separate target group before shifting traffic. If the new tasks fail health checks (e.g., due to incorrect ElastiCache endpoint), traffic remains on the blue environment, preventing application downtime. Option A is incorrect because a CloudWatch alarm only triggers an alert or rollback after the issue occurs; it does not prevent the deployment from impacting users.

Option B is incorrect because updating one task at a time (canary) still exposes tasks to the wrong configuration, and since the minimum healthy percent is 50%, at least half the tasks would fail before detection. Option D is incorrect because the ECS deployment circuit breaker rolls back only after the deployment fails, but it does not prevent the initial impact during the rolling update.

11
MCQhard

A company runs a critical application on an EC2 instance that stores data on an EBS volume. The SysOps administrator needs to implement a backup strategy that provides the ability to restore the volume to a specific point in time within the last 24 hours, with a recovery time objective (RTO) of less than 15 minutes. Which solution meets these requirements?

A.Configure a RAID 1 mirror of the EBS volume across two Availability Zones.
B.Enable automated backups on the EC2 instance.
C.Use AWS Backup to create backup plans for the EBS volume.
D.Schedule EBS snapshots every hour and keep them for 24 hours.
AnswerD

Scheduling EBS snapshots hourly provides point-in-time recovery points no more than one hour apart, satisfying an RPO of up to one hour, and retaining 24 snapshots covers a full day of rollback options. Restoring is done by creating a new EBS volume from the desired snapshot and attaching it to the instance, which typically takes only minutes and meets the required RTO. EBS snapshots are incremental and stored in Amazon S3, making them a durable, low-cost backup mechanism.

Why this answer

Scheduling EBS snapshots every hour and retaining them for 24 hours provides point-in-time recovery granularity within the last 24 hours. EBS snapshots are incremental, stored in Amazon S3, and can be used to create a new volume or restore an existing one; restoring from a snapshot typically takes only a few minutes, meeting the RTO of less than 15 minutes.

Exam trap

The trap here is that candidates may confuse AWS Backup (Option C) as a service that automatically provides hourly snapshots with 24-hour retention, but AWS Backup requires explicit configuration of a backup plan with the desired schedule and retention—it is not a default behavior, and the question tests whether you know the specific implementation (scheduled snapshots) rather than the service name.

How to eliminate wrong answers

Option A is wrong because RAID 1 mirroring across Availability Zones is not a native EBS feature—it would require software RAID on the EC2 instance, which adds complexity, does not provide point-in-time snapshots, and cannot guarantee a restore to a specific time within 24 hours. Option B is wrong because EC2 instances do not have a native 'automated backups' feature; the term is ambiguous and likely refers to EBS snapshots or AMI backups, but without a defined schedule and retention policy, it cannot ensure point-in-time recovery within 24 hours. Option C is wrong because AWS Backup can create backup plans for EBS volumes, but it does not inherently provide the required granularity of hourly snapshots with 24-hour retention unless explicitly configured; the question specifies a solution that meets the requirements, and AWS Backup is a service that can be used to orchestrate snapshots, but the correct answer is the specific action of scheduling hourly snapshots with 24-hour retention.

12
MCQmedium

A company runs a critical production database on Amazon RDS for MySQL with Multi-AZ deployment. The SysOps administrator needs to be automatically notified when a failover event occurs, and also capture the exact time and reason for the failover for compliance purposes. Which AWS service or feature should be used to capture the failover event details with the least operational overhead?

A.Create an Amazon CloudWatch Events rule that matches the 'RDS DB Instance Event' for 'failover' and sends the event to an Amazon SNS topic for notification and logging.
B.Enable detailed monitoring on the RDS instance and stream the logs to Amazon CloudWatch Logs where a metric filter can detect failover patterns.
C.Configure AWS CloudTrail to log all RDS API calls and analyze the logs for the 'Failover' event type.
D.Use AWS Config to create a config rule that evaluates whether the 'DBInstanceStatus' changes to 'failover' and then trigger a remediation action.
AnswerA

Amazon CloudWatch Events (now part of Amazon EventBridge) natively integrates with RDS event notifications, emitting a structured event whenever a DB instance experiences a failover. By creating a rule that matches the 'RDS DB Instance Event' source and the specific detail type for failover, you can route that event to an SNS topic in near-real time, enabling automated alerting, logging, and downstream remediation. This is the intended, low-overhead approach because RDS already publishes these lifecycle events, and no polling or custom detection logic is required.

Why this answer

Amazon CloudWatch Events (now part of Amazon EventBridge) can match RDS DB Instance events, including 'failover', and route them to an SNS topic for notification and to CloudWatch Logs for logging. This approach requires no custom scripting or polling, providing the least operational overhead while capturing the exact time and reason for the failover directly from the RDS event stream.

Exam trap

The trap here is that candidates confuse CloudTrail (which logs API calls) with RDS events (which log internal service events), leading them to choose CloudTrail even though automatic failovers are not API-driven and thus not recorded by CloudTrail.

How to eliminate wrong answers

Option B is wrong because detailed monitoring on RDS provides enhanced metrics (e.g., CPU, memory) but does not generate failover events or detect failover patterns; metric filters on CloudWatch Logs would require RDS to log failover details to CloudWatch Logs, which RDS does not do by default. Option C is wrong because AWS CloudTrail logs API calls (e.g., FailoverDBInstance), not internal failover events triggered by AWS; a Multi-AZ failover is an automatic process, not an API call, so CloudTrail will not capture it. Option D is wrong because AWS Config evaluates resource configuration changes (e.g., DBInstanceStatus) but does not natively detect a 'failover' status change; the DBInstanceStatus transitions through multiple states (e.g., 'creating', 'available', 'resetting-master-credentials') and 'failover' is not a valid status—Config rules would require custom logic and still not capture the exact reason for the failover.

13
MCQeasy

A company uses Amazon S3 to store critical data. The SysOps administrator needs to protect against accidental deletion of objects. Which combination of actions should the administrator take? (Choose the best answer.)

A.Apply a bucket policy that denies s3:DeleteObject for all principals.
B.Set a lifecycle policy to expire objects after 30 days.
C.Enable S3 Versioning and MFA Delete on the bucket.
D.Configure cross-Region replication to a different bucket.
AnswerC

Enabling S3 Versioning preserves every version of an object, so a deletion request creates a delete marker instead of erasing the underlying data, allowing you to restore the object at any time. Adding MFA Delete requires a valid MFA code to permanently delete object versions or to suspend versioning, which prevents accidental or malicious purges even by the root user. Together, these features provide comprehensive, built-in protection against both overwrites and deletes, making them the correct choice for safeguarding critical data.

Why this answer

Enabling S3 Versioning preserves all versions of an object, allowing recovery from accidental deletion or overwrite. MFA Delete adds an additional authentication layer, requiring multi-factor authentication to permanently delete object versions or suspend versioning, thus preventing unauthorized or accidental deletions.

Exam trap

The trap here is that candidates often choose a bucket policy denying s3:DeleteObject (Option A) thinking it prevents accidental deletion, but they overlook that it does not protect against overwrites or that versioning with MFA Delete is the only comprehensive solution that also covers permanent deletion and versioning state changes.

How to eliminate wrong answers

Option A is wrong because a bucket policy that denies s3:DeleteObject for all principals would also block legitimate administrative deletions, and it does not protect against accidental overwrites (PUT operations) or deletion of versioned objects without versioning enabled. Option B is wrong because a lifecycle policy to expire objects after 30 days would actually delete objects automatically, increasing the risk of data loss rather than protecting against accidental deletion. Option D is wrong because cross-Region replication replicates objects to another bucket but does not prevent deletion in the source bucket; deletions are replicated as well unless a delete marker replication rule is configured, and it does not protect against accidental deletion in the source.

14
MCQmedium

A company runs a file-sharing application on AWS. Users upload files to an S3 bucket, which triggers a Lambda function to process the files and store metadata in a DynamoDB table. Recently, users have reported that some uploaded files are never processed. The SysOps Administrator checks the CloudWatch logs and finds no errors from the Lambda function. The S3 bucket is configured to send events to the Lambda function. The DynamoDB table has sufficient write capacity. The administrator suspects that the event notifications are being lost. Which action should the SysOps Administrator take to ensure that every file upload triggers a Lambda function and that the function processes the file successfully?

A.Configure an SQS queue as the event destination for the S3 bucket, and have the Lambda function process messages from the queue.
B.Use DynamoDB Streams to capture file metadata changes instead of Lambda invocation.
C.Increase the Lambda function's reserved concurrency to handle more invocations.
D.Increase the write capacity of the DynamoDB table to avoid throttling.
AnswerA

S3 event notifications can be delivered to SQS, providing a durable buffer. Lambda polls the queue, so messages are not lost if the function is throttled or busy. SQS retains messages until processed, and Lambda event source mapping handles retries and batching. This decouples the upload rate from Lambda's invocation capacity.

Why this answer

Using an SQS queue as the event destination for S3 bucket events provides a durable, reliable mechanism to capture every event. S3 sends event notifications to the SQS queue, and if the Lambda function fails or is throttled, the message remains in the queue for later processing. This decouples the event source from the function and ensures no events are lost.

Option B is incorrect because DynamoDB Streams capture changes to DynamoDB items, not S3 events. Option C is wrong because increasing reserved concurrency only helps with Lambda scaling but does not address potential event loss due to failures or delivery issues. Option D is wrong because DynamoDB write capacity is already sufficient; the issue is with S3 event delivery, not database writes.

15
MCQmedium

A company has a production Amazon RDS for MySQL DB instance in a single Availability Zone. The SysOps administrator needs to improve database availability to ensure automatic failover in the event of a database failure or an Availability Zone outage. Which configuration should the administrator enable?

A.Enable Multi-AZ deployment
B.Create a read replica in another Availability Zone
C.Enable automated backups
D.Change the DB instance to a larger instance class
AnswerA

Enabling Multi-AZ on an RDS for MySQL instance provisions a standby replica in a separate Availability Zone and uses synchronous replication to keep it current. If the primary fails or its AZ becomes unavailable, Amazon RDS automatically flips the DNS endpoint to the standby, delivering failover in typically 60–120 seconds with no manual intervention and minimal data loss. This is the only option that provides automatic failover and meets the high availability requirement for a production database.

Why this answer

Enabling a Multi-AZ deployment for Amazon RDS for MySQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of a database failure or an AZ outage, Amazon RDS automatically fails over to the standby replica, typically within 60–120 seconds, without requiring manual intervention. This configuration meets the requirement for automatic failover and improved availability.

Exam trap

The trap here is that candidates often confuse a read replica with a Multi-AZ standby, mistakenly believing that a read replica can provide automatic failover, but read replicas require manual promotion and do not maintain synchronous replication.

How to eliminate wrong answers

Option B is wrong because a read replica is an asynchronous copy used for offloading read traffic, not for automatic failover; while it can be promoted to a standalone instance, this requires manual action and does not provide automatic failover. Option C is wrong because automated backups only enable point-in-time recovery and do not provide any failover capability or high availability. Option D is wrong because changing the DB instance to a larger instance class improves performance and scalability but does not provide redundancy or automatic failover across Availability Zones.

16
MCQeasy

A company is designing a disaster recovery plan for its on-premises database. They need to replicate the database to AWS with low latency. Which AWS service should they use?

A.Amazon S3 with Cross-Region Replication.
B.AWS Storage Gateway with volume gateway.
C.AWS Database Migration Service (DMS) with ongoing replication.
D.AWS Direct Connect to establish a dedicated network connection.
AnswerC

AWS Database Migration Service (DMS) with ongoing replication is the correct choice because DMS can perform continuous change data capture (CDC) from the source database and apply those changes to a target Amazon RDS instance in another Region. This enables a near-real-time replica of the database with low RPO, supporting disaster recovery by keeping the target transactionally consistent. DMS supports both homogeneous and heterogeneous migrations, and ongoing replication is a key feature for DR scenarios when native replication options are not available.

Why this answer

AWS DMS with ongoing replication (change data capture, CDC) is the correct choice because it can continuously replicate changes from an on-premises database to a target database in AWS with low latency, supporting heterogeneous and homogeneous migrations. This meets the requirement for a disaster recovery plan that keeps the AWS copy nearly synchronized with the on-premises source.

Exam trap

The trap here is confusing network connectivity services (like Direct Connect) or storage replication (like S3 CRR or Storage Gateway) with database-level replication, which requires transaction-consistent change capture and application.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Cross-Region Replication is an object-level replication mechanism for S3 buckets, not designed for database replication; it cannot capture transactional changes or maintain database consistency. Option B is wrong because AWS Storage Gateway with volume gateway provides block-level storage volumes that can be backed up to S3, but it does not offer ongoing database replication with low latency; it is intended for hybrid storage, not for replicating live database transactions. Option D is wrong because AWS Direct Connect establishes a dedicated network connection for improved bandwidth and latency, but it is a connectivity service, not a replication service; it does not replicate the database itself.

17
MCQeasy

A company uses Amazon Route 53 for DNS resolution. The company wants to ensure that if a web server becomes unhealthy, traffic is automatically routed to a healthy server in another Availability Zone. Which routing policy should be used?

A.Latency routing policy
B.Weighted routing policy
C.Geolocation routing policy
D.Failover routing policy
AnswerD

Failover routing policy is specifically designed for active-passive configurations: you create a primary record and a secondary record, and Route 53 uses health checks on the primary to determine its status. When the primary fails its health check, Route 53 automatically returns the secondary record's IP address, enabling DNS-level failover without manual intervention. This matches the company's need to route to a secondary endpoint whenever the primary is unhealthy, making it the correct answer.

Why this answer

The Failover routing policy (D) is correct because it is specifically designed to route traffic to a primary resource (e.g., a web server) and automatically redirect to a secondary resource in a different Availability Zone when health checks fail. Route 53 uses health checks to monitor the primary endpoint; if the primary becomes unhealthy, it returns the secondary record in DNS responses, ensuring automatic failover.

Exam trap

The trap here is that candidates often confuse Failover routing policy with Weighted routing policy, mistakenly thinking that weights can be adjusted dynamically to simulate failover, but Route 53 does not automatically adjust weights based on health—only Failover routing policy provides automatic, health-check-driven failover between a primary and secondary resource.

How to eliminate wrong answers

Option A is wrong because Latency routing policy routes traffic based on the lowest network latency to the client, not on health status or failover. Option B is wrong because Weighted routing policy distributes traffic across multiple resources based on assigned weights, but it does not automatically reroute all traffic to a healthy endpoint when one fails—it continues to send a portion of traffic to unhealthy endpoints unless health checks are manually configured to remove them. Option C is wrong because Geolocation routing policy routes traffic based on the geographic location of the DNS resolver, not on the health of the resources, and it does not provide automatic failover between Availability Zones.

18
MCQmedium

A SysOps administrator needs to implement a backup strategy for an Amazon RDS for PostgreSQL database. The database is 500 GB and experiences heavy write traffic. Which solution provides the most cost-effective backup with the least impact on database performance?

A.Enable automated backups with a retention period of 7 days.
B.Create a Multi-AZ deployment and use the standby for backups.
C.Use AWS Database Migration Service to continuously replicate data to an S3 bucket.
D.Take manual DB snapshots daily during off-peak hours.
AnswerA

Automated backups are the native RDS backup mechanism: daily snapshots of the database volume are captured, and transaction logs are continuously uploaded to S3 so you can restore to any point within the 7-day retention window. These backup operations are low overhead because they leverage the EBS snapshot facility, so the performance impact is minimal and no long-duration downtime is required. With a 7-day retention, you automatically meet a standard RPO (point-in-time) without manual effort or risk of expiring snapshots.

Why this answer

Automated backups are enabled by default with minimal performance impact and include transaction logs for point-in-time recovery. Option B is wrong; while using a standby for backups in Multi-AZ can reduce impact, automated backups already have minimal impact and are more cost-effective. Option C is wrong because AWS DMS replication to S3 is not a backup solution and adds complexity and cost.

Option D is wrong because manual snapshots cause a brief I/O suspension and are less automated than automated backups.

19
Multi-Selectmedium

A company is designing a disaster recovery plan for its critical applications. The plan must minimize data loss and recovery time. Which TWO measures should the SysOps administrator implement?

Select 2 answers
A.Perform regular backups to Amazon S3.
B.Set a recovery time objective (RTO) of 24 hours.
C.Use manual procedures to restore from backups.
D.Run all workloads in a single AWS region.
E.Replicate data to another AWS region.
AnswersA, E

Perform regular backups to Amazon S3 is a valid DR component because S3 provides 11 nines of durability, versioning for point-in-time recovery, and lifecycle policies to archive to S3 Glacier. Backups create immutable, restorable copies of critical data, protecting against accidental deletion, corruption, or ransomware. However, backups alone are not a complete DR plan; they establish an RPO tied to backup frequency and must be restored before workloads can resume, which elongates RTO. Therefore, this action is correct as a foundational data-protection measure, not as a full DR strategy.

Why this answer

Regular backups to Amazon S3 are a foundational data protection measure because S3 provides 99.999999999% durability and supports lifecycle policies for cost-effective long-term retention. This directly addresses the requirement to minimize data loss by ensuring point-in-time recovery copies exist independently of the primary infrastructure.

Exam trap

The trap here is that candidates often confuse RTO/RPO definitions with actual implementation measures, or they assume that a single region with backups is sufficient for disaster recovery, ignoring the need for geographic separation to survive a region-wide outage.

20
MCQmedium

A company is running a web application on EC2 instances behind an Application Load Balancer. They want to ensure that if an entire Availability Zone fails, the application remains available. Which configuration should they implement?

A.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
B.Use an RDS Multi-AZ deployment for the application.
C.Use a larger EC2 instance type.
D.Enable detailed monitoring on the EC2 instances.
AnswerA

Configuring the Auto Scaling group to span multiple Availability Zones is the correct approach because it distributes the EC2 instances across independent failure domains. If one AZ becomes unavailable, the instances in the other AZs continue to serve traffic, and the ASG automatically replaces the unhealthy instances to maintain desired capacity. This architecture provides high availability for the application tier.

Why this answer

To ensure application availability during an entire Availability Zone (AZ) failure, the Auto Scaling group must be configured to launch EC2 instances across multiple AZs. This distributes the workload so that if one AZ becomes unavailable, the remaining AZs continue serving traffic. The Application Load Balancer (ALB) automatically routes requests only to healthy instances in the surviving AZs, maintaining application uptime.

Exam trap

The trap here is that candidates often confuse high availability at the database layer (RDS Multi-AZ) with compute layer fault tolerance, or they mistakenly believe that scaling vertically (larger instances) or increasing monitoring granularity can compensate for a full AZ outage.

How to eliminate wrong answers

Option B is wrong because RDS Multi-AZ provides high availability for the database layer, not for the compute layer (EC2 instances) handling the web application; it does not address EC2 instance distribution across AZs. Option C is wrong because using a larger EC2 instance type increases compute capacity within a single AZ but does not protect against an AZ failure; the instance would still be unavailable if its AZ fails. Option D is wrong because enabling detailed monitoring on EC2 instances provides more granular CloudWatch metrics (1-minute intervals) but does not affect instance placement or fault tolerance across AZs.

21
Multi-Selectmedium

A SysOps administrator is designing a highly available architecture for a web application using an Application Load Balancer (ALB) with EC2 instances in an Auto Scaling group. Which TWO configurations are required to ensure high availability? (Choose TWO.)

Select 2 answers
A.Launch all EC2 instances in a single Availability Zone to reduce latency
B.Use t2.micro instances to reduce cost
C.Configure the ALB with health checks for the target group
D.Disable health checks to reduce load on the ALB
E.Configure the Auto Scaling group to launch instances in at least two Availability Zones
AnswersC, E

The ALB continuously sends health-check requests (e.g., HTTP GET to a specified path) to each instance in its target group; an instance that fails a set number of consecutive checks is marked unhealthy and automatically deregistered, so the ALB stops forwarding new traffic to it and reroutes incoming requests to healthy instances. Configuring health checks with appropriate interval, timeout, and threshold values detects underlying application or instance failures early, enabling rapid failover and improving overall fault tolerance. This is a critical control point; without meaningful health checks, even a perfectly scaled fleet will serve errors to a portion of requests.

Why this answer

Health checks allow the ALB to monitor the status of each EC2 instance in the target group. If an instance fails health checks, the ALB automatically stops routing traffic to it, preventing user requests from reaching a failed instance. This is essential for maintaining application availability and is a core feature of the ALB's high-availability design.

Exam trap

The trap here is that candidates often confuse cost-saving measures (like using smaller instance types) or performance optimizations (like single-AZ deployment for lower latency) with high-availability requirements, but the exam specifically tests the understanding that high availability requires redundancy across Availability Zones and active health monitoring.

22
MCQeasy

A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which action should be taken?

A.Launch a second instance in a different Availability Zone.
B.Assign an Elastic IP address to the instance.
C.Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
D.Place the instance in an Auto Scaling group with a min size of 1.
AnswerC

Creating a CloudWatch alarm on the StatusCheckFailed metric and configuring the recovery action is the correct method because EC2 instance recovery automatically restarts the instance on new hardware when the underlying host fails. This recovery action preserves the instance ID, private IP address, Elastic IP address, and all EBS volumes, so the instance's identity and configuration are maintained. The alarm must monitor the StatusCheckFailed_System metric (or the aggregate StatusCheckFailed metric) and invoke the 'recover' action to trigger the automatic recovery.

Why this answer

A CloudWatch alarm on the StatusCheckFailed metric can be configured with the 'recover' action to automatically restart the EC2 instance on a new underlying host if a hardware failure is detected. This recovery action preserves the instance ID, private IP, Elastic IP, and instance metadata, ensuring minimal disruption. The StatusCheckFailed metric specifically monitors the instance's system status checks, which detect AWS hardware issues.

Exam trap

The trap here is that candidates often confuse Auto Scaling recovery (which replaces the instance) with CloudWatch alarm recovery (which recovers the same instance), leading them to choose Option D despite the requirement to preserve the original instance.

How to eliminate wrong answers

Option A is wrong because launching a second instance in a different Availability Zone does not automatically recover the original instance from hardware failure; it creates a separate instance that requires manual or automated traffic redirection. Option B is wrong because assigning an Elastic IP address only provides a static public IP, but does not trigger any recovery mechanism when the underlying hardware fails. Option D is wrong because placing the instance in an Auto Scaling group with a min size of 1 will replace a failed instance with a new one, but it does not preserve the original instance ID, private IP, or Elastic IP, and the replacement is not a recovery of the same instance.

23
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer. The database is an RDS MySQL instance with Multi-AZ enabled. The application experiences intermittent 5xx errors that correlate with database failover events. What is the MOST likely cause and solution?

A.Configure the application to use the RDS reader endpoint.
B.Use a read replica to offload read traffic and reduce failover impact.
C.Increase the database connection pool size to handle retries.
D.Enable DNS caching with a low TTL in the application and use the RDS instance endpoint with a retry mechanism.
AnswerD

The RDS instance endpoint is a CNAME record that AWS automatically points to the current primary instance's IPv4 address. By configuring the application's DNS resolver to cache this record with a low TTL (e.g., 5–10 seconds), the app will quickly re-resolve to the new primary after a failover. Pairing this with a retry loop that catches transient connection errors ensures that any packets sent during the failover window are re-attempted once DNS has propagated. This combination directly minimizes downtime and is the officially recommended client-side failover approach.

Why this answer

During an RDS Multi-AZ failover, the DNS record for the primary instance endpoint is updated to point to the standby instance. If the application caches the old DNS resolution with a high TTL, it continues to send connections to the unreachable primary, causing 5xx errors. Enabling DNS caching with a low TTL (e.g., 5 seconds) and implementing a retry mechanism ensures the application quickly resolves the new endpoint and reconnects, minimizing downtime.

Exam trap

The trap here is that candidates often confuse the reader endpoint with the writer endpoint, or assume that read replicas or connection pooling can mitigate failover errors, when the real issue is DNS caching and the need for a retry mechanism.

How to eliminate wrong answers

Option A is wrong because the RDS reader endpoint is used for read-only traffic from read replicas, not for handling failover of the primary writer instance; during a failover, the writer endpoint is the one that updates its DNS. Option B is wrong because read replicas do not participate in Multi-AZ failover; they are for scaling read traffic and do not provide automatic failover for the primary database. Option C is wrong because increasing the connection pool size does not address the root cause of stale DNS resolution; it only adds more connections that will still fail until the DNS cache is refreshed.

24
Multi-Selecteasy

Which TWO actions should a SysOps administrator take to ensure high availability of a web application running on EC2 instances? (Choose two.)

Select 2 answers
A.Enable termination protection on all EC2 instances.
B.Launch all EC2 instances in a single Availability Zone.
C.Use a larger instance type for all EC2 instances.
D.Configure an Auto Scaling group with a health check to replace unhealthy instances.
E.Deploy EC2 instances across multiple Availability Zones.
AnswersD, E

An Auto Scaling group with health checks continuously monitors instance state and automatically terminates and replaces unhealthy instances, maintaining the desired capacity. This removes failed nodes from service, satisfying the availability requirement without manual intervention.

Why this answer

Option D is correct because an Auto Scaling group with health checks (EC2 status checks and optionally ELB health checks) automatically detects and replaces unhealthy instances, maintaining the desired capacity and thus high availability. Option E is correct because deploying EC2 instances across multiple Availability Zones ensures the application survives an AZ-level failure, since each AZ has independent power, cooling, and networking. Option A is incorrect because termination protection only prevents accidental instance termination; it does not improve availability or recover failed instances.

Option B is incorrect because placing all instances in a single AZ creates a single point of failure, reducing availability. Option C is incorrect because a larger instance type only increases capacity, not redundancy or fault tolerance.

Exam trap

The trap here is that candidates often confuse termination protection (a safety feature) with high availability, or think that larger instance types inherently provide fault tolerance, when in fact only redundancy across multiple Availability Zones and automated health-based replacement ensure high availability.

25
MCQhard

A company uses Amazon MQ (RabbitMQ) for messaging between microservices. The SysOps administrator needs to ensure the message broker is highly available with automatic failover and no data loss. Which deployment mode should be used?

A.Single-instance broker
B.Active/standby broker
C.Cluster deployment
D.Multi-AZ broker with read replicas
AnswerB

An active/standby broker deploys two broker instances in different Availability Zones, with synchronous replication of messages and metadata from the active broker to the standby. On failure of the active broker, automatic failover promotes the standby to active with minimal downtime, ensuring messages are not lost because they are already replicated synchronously. This is the correct Amazon MQ deployment mode for meeting high-availability and data-loss-prevention requirements.

Why this answer

Amazon MQ for RabbitMQ supports an active/standby deployment mode that provides automatic failover and no data loss. In this mode, one broker instance is active and a second is a synchronous standby; if the active fails, the standby takes over without losing messages because all data is replicated synchronously across both instances. This meets the high availability and data durability requirements specified in the question.

Exam trap

The trap here is that candidates confuse Amazon MQ's cluster deployment (which is for scaling) with active/standby (which is for high availability and data durability), or they incorrectly apply RDS Multi-AZ concepts to Amazon MQ.

How to eliminate wrong answers

Option A is wrong because a single-instance broker has no redundancy or automatic failover; if it fails, all messages are lost until manual recovery. Option C is wrong because RabbitMQ cluster deployment in Amazon MQ is designed for horizontal scaling and throughput, not for automatic failover with zero data loss; it uses asynchronous replication and can lose messages during a node failure. Option D is wrong because Amazon MQ does not support Multi-AZ brokers with read replicas; that concept applies to Amazon RDS, not to message brokers.

26
MCQmedium

A company runs a critical application on Amazon EC2 instances with data stored on Amazon EBS volumes. The SysOps administrator needs to implement a backup strategy that supports point-in-time recovery with a Recovery Point Objective (RPO) of 1 hour and a Recovery Time Objective (RTO) of 4 hours. Which solution meets these requirements with the least operational overhead?

A.Use AWS Backup to schedule hourly EBS snapshots and restore to a new volume when needed.
B.Use Amazon Data Lifecycle Manager (DLM) to take hourly snapshots and create an AWS CloudFormation template to launch a new instance from the snapshot.
C.Use custom scripts to copy snapshots to an Amazon S3 bucket and restore from there.
D.Use Amazon S3 Lifecycle policies to transition data to Amazon S3 Glacier.
AnswerA

AWS Backup offers a fully managed, policy-based backup service that can create hourly EBS snapshots automatically. It provides centralized backup governance, retention management, and lifecycle policies, with the ability to restore a snapshot to a new EBS volume quickly. This minimizes RTO because restore is a native AWS operation and does not require custom scripting or additional orchestration.

Why this answer

AWS Backup provides a fully managed, policy-based backup service that can schedule EBS snapshots hourly, meeting the 1-hour RPO. Restoring from an AWS Backup snapshot to a new EBS volume and attaching it to an EC2 instance can be completed within the 4-hour RTO, with minimal operational overhead as it eliminates the need for custom scripts or lifecycle management.

Exam trap

The trap here is that candidates may choose DLM (Option B) because it can schedule snapshots, but they overlook the operational overhead of manually creating a CloudFormation template for recovery, whereas AWS Backup provides a fully managed restore workflow that meets the least operational overhead requirement.

How to eliminate wrong answers

Option B is wrong because Amazon Data Lifecycle Manager (DLM) can schedule hourly snapshots, but requiring a CloudFormation template to launch a new instance from the snapshot adds unnecessary operational overhead and complexity, whereas AWS Backup can directly restore the volume and instance. Option C is wrong because using custom scripts to copy snapshots to S3 introduces additional complexity, potential for errors, and does not leverage native AWS backup services, increasing operational overhead. Option D is wrong because Amazon S3 Lifecycle policies are designed for object lifecycle management in S3, not for EBS snapshots or point-in-time recovery of EC2 instances, and S3 Glacier is for archival, not rapid recovery with a 4-hour RTO.

27
Multi-Selectmedium

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application stores data on an Amazon EBS volume. The SysOps administrator needs to implement a backup strategy that ensures data can be recovered in the event of an AZ failure. Which TWO actions should be taken? (Choose TWO.)

Select 2 answers
A.Increase the EBS volume size to maximize I/O performance.
B.Configure automated snapshots using Amazon Data Lifecycle Manager.
C.Create a lifecycle policy to automatically take snapshots of the EBS volume and copy them to another region.
D.Enable EBS encryption using AWS KMS.
E.Enable EBS Multi-Attach to allow the volume to be attached to instances in another AZ.
AnswersB, C

Amazon Data Lifecycle Manager (DLM) is the native service for automating EBS snapshot creation, retention, and deletion according to a policy-defined schedule. By configuring a DLM policy, you ensure consistent point-in-time backups of the EBS volume without manual intervention, and you can set retention rules to age out old snapshots to control costs. This directly satisfies the backup requirement for a critical application and is the correct answer to the scenario.

Why this answer

Amazon Data Lifecycle Manager (DLM) automates the creation, retention, and deletion of EBS snapshots, providing a scheduled backup mechanism that protects against data loss. Option C is correct because copying snapshots to another region ensures data is recoverable even if an entire AWS Availability Zone (AZ) fails, as the snapshots are stored independently in a different geographic region.

Exam trap

The trap here is that candidates often confuse EBS Multi-Attach (which provides high availability within an AZ) with a cross-AZ backup strategy, or they mistakenly think increasing volume size or enabling encryption alone constitutes a backup plan.

28
MCQmedium

A company runs a stateless web application on Amazon EC2 instances in an Auto Scaling group with a minimum of 2 and maximum of 10 instances. The instances are behind an Application Load Balancer (ALB). The SysOps administrator needs to ensure that the application can survive the failure of an entire AWS Availability Zone (AZ) in the region. Which configuration is necessary?

A.Configure the Auto Scaling group with subnets in at least two Availability Zones and ensure the ALB has subnets in the same AZs.
B.Increase the Auto Scaling group minimum to 10 instances to absorb the failure.
C.Use larger instance types to handle the load of a failed AZ.
D.Use multiple Application Load Balancers in different AZs.
AnswerA

Correctly designed for failure domain isolation: an Auto Scaling group spanning subnets in at least two Availability Zones (AZs) lets EC2 instances be provisioned across independent infrastructure, and an ALB with subnets in those same AZs can route traffic to healthy instances in any AZ. If one AZ becomes unavailable, the ALB continues distributing requests to instances in the remaining AZs, while Auto Scaling replaces failed instances in the other AZs. This provides high availability because the application is stateless and can serve all traffic from a single AZ when needed.

Why this answer

Deploying the Auto Scaling group across multiple Availability Zones (AZs) and ensuring the ALB has subnets in the same AZs allows the application to continue serving traffic even if one entire AZ fails. The ALB can route requests to healthy instances in the remaining AZs, and the Auto Scaling group will replace failed instances in other AZs as needed, maintaining the minimum instance count. This architecture is a fundamental pattern for high availability in AWS.

Exam trap

The trap here is that candidates often think increasing instance count or size alone provides high availability, but without multi-AZ distribution, a single AZ failure can still cause total application downtime.

How to eliminate wrong answers

Option B is wrong because simply increasing the minimum to 10 instances does not provide AZ resilience; all instances could still be in a single AZ, and a failure of that AZ would take down all 10 instances. Option C is wrong because using larger instance types only increases compute capacity per instance, but does not distribute instances across AZs; a single AZ failure would still eliminate all instances if they are all in that AZ. Option D is wrong because using multiple ALBs in different AZs is unnecessary and adds complexity; a single ALB can already distribute traffic across multiple AZs, and multiple ALBs would require additional DNS routing logic (e.g., Route 53) and do not inherently improve AZ failure survival.

29
MCQhard

A company is running a stateful web application on a single EC2 instance in a public subnet. The instance stores user sessions locally. The company wants to improve availability without rewriting the application. Which design should they use?

A.Create a second EC2 instance in a different AZ and use Route 53 with health checks.
B.Use an Auto Scaling group across multiple AZs but keep sessions on instance.
C.Deploy an Application Load Balancer across multiple AZs, move session storage to ElastiCache, and use an Auto Scaling group.
D.Use an Application Load Balancer with sticky sessions and an Auto Scaling group in a single AZ.
AnswerC

This is correct because an ALB in multiple AZs distributes traffic across healthy instances, while ElastiCache stores session data independently of any individual EC2 instance, making the app stateless. If an instance fails or is terminated by the ASG, other instances can immediately serve users because their sessions are still in ElastiCache. The ASG handles capacity and automatically replaces unhealthy instances based on ALB health checks, giving both elasticity and high availability.

Why this answer

It addresses the core issue of stateful sessions without rewriting the application. By moving session storage to ElastiCache (a centralized, external data store), the application becomes stateless from the instance's perspective, allowing an Auto Scaling group across multiple Availability Zones (AZs) and an Application Load Balancer (ALB) to distribute traffic seamlessly. This design improves availability by enabling horizontal scaling and fault tolerance, as any instance can handle any request since sessions are stored externally.

Exam trap

The trap here is that candidates often assume sticky sessions (session affinity) alone solve the stateful application problem, but they fail to recognize that sticky sessions do not protect against instance failure or AZ outages, and they still require local session storage, which is lost when an instance is replaced.

How to eliminate wrong answers

Option A is wrong because simply adding a second EC2 instance in a different AZ with Route 53 health checks does not solve the session state problem; user sessions stored locally on the original instance would be lost if traffic fails over to the new instance, breaking the stateful application. Option B is wrong because keeping sessions on the instance while using an Auto Scaling group across multiple AZs means that if an instance is terminated or replaced, all local session data is lost, and new instances cannot serve existing sessions, leading to user disruption. Option D is wrong because using an ALB with sticky sessions and an Auto Scaling group in a single AZ still creates a single point of failure at the AZ level; if that AZ goes down, the entire application becomes unavailable, and sticky sessions alone do not persist session data across instance replacements.

30
MCQmedium

A team of developers is deploying a new microservice that uses Amazon DynamoDB as its data store. The SysOps administrator must ensure that the application can handle a sudden spike in read traffic without throttling. Which DynamoDB feature can be used to automatically handle increases in read capacity?

A.DynamoDB Global Tables
B.DynamoDB Time to Live (TTL)
C.DynamoDB Auto Scaling
D.DynamoDB Accelerator (DAX)
AnswerC

DynamoDB Auto Scaling is the correct solution because it continuously adjusts a table's provisioned read and write capacity in response to actual traffic using the Application Auto Scaling target-tracking policy. By monitoring CloudWatch metrics like consumed capacity, it scales out before throttling occurs and scales in when traffic declines, all within user-defined minimum and maximum limits. This gives the microservice a fully automated way to handle variable load without manual capacity planning.

Why this answer

DynamoDB Auto Scaling is the correct feature because it automatically adjusts the provisioned read and write capacity based on actual traffic patterns, preventing throttling during sudden spikes. The SysOps administrator can define a target utilization percentage (e.g., 70%), and DynamoDB Auto Scaling uses the Application Auto Scaling service to increase capacity before requests are throttled, ensuring consistent performance without manual intervention.

Exam trap

The trap here is that candidates often confuse DynamoDB Accelerator (DAX) with a scaling solution, thinking its caching automatically handles capacity increases, but DAX only reduces read load and does not modify provisioned capacity or prevent throttling from a sudden spike in uncached reads.

How to eliminate wrong answers

Option A is wrong because DynamoDB Global Tables provide multi-region replication for disaster recovery and low-latency reads, but they do not automatically handle sudden read traffic spikes within a single region; they require provisioned capacity management per replica. Option B is wrong because DynamoDB Time to Live (TTL) automatically deletes expired items to reduce storage costs, but it has no effect on read capacity or throttling prevention. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read latency for repeated queries, but it does not automatically increase provisioned read capacity; it only reduces the load on the underlying table by caching frequently accessed data.

31
MCQhard

A company uses S3 to store critical data. They need to ensure that data can be recovered in the event of accidental deletion or overwriting by users. Which combination of actions should they take?

A.Enable S3 Cross-Region Replication and S3 Transfer Acceleration.
B.Enable S3 Versioning and S3 Transfer Acceleration.
C.Enable S3 Versioning and MFA Delete.
D.Enable S3 Server Access Logging and S3 Object Lock.
AnswerC

S3 Versioning maintains every version of an object, including the original before any overwrite or deletion, so a deleted or replaced object can be restored by reverting to an older version ID. MFA Delete requires the AWS account root's MFA code to permanently delete a version or to change the bucket's versioning state, preventing an attacker who compromises IAM credentials from executing an unrecoverable purge. Together these features provide both a recovery mechanism and a hardened authorization boundary, directly addressing the requirement to ensure critical data cannot be lost through accidental or malicious deletion.

Why this answer

Enabling S3 Versioning preserves all versions of an object, allowing recovery from accidental deletion or overwriting. MFA Delete adds an extra layer of protection by requiring multi-factor authentication to permanently delete object versions or suspend versioning, preventing unauthorized or accidental permanent data loss.

Exam trap

The trap here is that candidates often think S3 Cross-Region Replication (CRR) or S3 Object Lock alone can prevent accidental deletion, but CRR replicates delete markers and does not protect the source, while Object Lock without Versioning cannot recover overwritten data; the correct combination requires both Versioning and MFA Delete to enable recovery and prevent permanent deletion.

How to eliminate wrong answers

Option A is wrong because S3 Cross-Region Replication (CRR) replicates objects to another region for disaster recovery or compliance, but it does not protect against accidental deletion or overwriting within the source bucket; deleted objects are also replicated as delete markers. S3 Transfer Acceleration speeds up uploads over long distances but provides no data recovery capabilities. Option B is wrong because while S3 Versioning enables recovery, S3 Transfer Acceleration is irrelevant for data recovery; it only improves upload performance.

Option D is wrong because S3 Server Access Logging records requests for auditing but does not enable recovery of deleted or overwritten objects; S3 Object Lock prevents objects from being deleted or overwritten for a fixed retention period, but without Versioning, it cannot recover objects that were overwritten before the lock was applied, and it does not address accidental deletion by users with sufficient permissions.

32
MCQmedium

A SysOps administrator is designing a disaster recovery plan for a web application that runs on EC2 instances with data stored in an RDS MySQL database. The application requires a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour. Which solution meets these requirements most cost-effectively?

A.Use a single RDS instance with daily snapshots and EC2 instance store.
B.Use EC2 Auto Scaling across Regions with an RDS standby instance.
C.Use RDS Multi-AZ with read replicas in another Region.
D.Use RDS Multi-AZ with automated backups and EC2 AMI backups.
AnswerD

RDS Multi-AZ automatically fails over to a standby instance in a different Availability Zone within minutes, satisfying the 1-hour RTO for the database tier. Automated backups on RDS enable point-in-time recovery to any second within the backup retention period, achieving an RPO of approximately 5 minutes. EC2 AMI backups provide a consistent, restorable image of the application servers, allowing rapid relaunch of the compute tier. Together, these components form a practical, cost-effective disaster recovery plan that meets both the RPO and RTO targets without relying on ephemeral storage or manual cross-region promotion.

Why this answer

RDS Multi-AZ provides automatic synchronous replication to a standby instance in a different Availability Zone, ensuring minimal data loss and a fast failover that meets the 5-minute RPO and 1-hour RTO. Automated backups enable point-in-time recovery within the retention period, and EC2 AMI backups allow quick restoration of the application tier, making this the most cost-effective solution that satisfies both objectives without requiring cross-Region resources.

Exam trap

The trap here is that candidates confuse Multi-AZ with cross-Region disaster recovery, assuming that a read replica in another Region is required for a 5-minute RPO, but Multi-AZ within a single Region with automated backups is sufficient and more cost-effective for the given RPO/RTO targets.

How to eliminate wrong answers

Option A is wrong because daily snapshots cannot achieve a 5-minute RPO, as data loss could be up to 24 hours, and EC2 instance store is ephemeral and does not persist data across stops or terminations, failing the RTO/RPO requirements. Option B is wrong because EC2 Auto Scaling across Regions introduces significant complexity and cost for cross-Region data replication, and RDS standby instances in another Region (Multi-Region) are not natively supported without additional replication mechanisms like cross-Region read replicas, which increase latency and cost beyond what is needed. Option C is wrong because RDS Multi-AZ with read replicas in another Region provides read scaling and disaster recovery but incurs higher costs for cross-Region data transfer and replica maintenance, and the read replica is asynchronous, potentially exceeding the 5-minute RPO during a failover scenario.

33
Multi-Selecthard

A company runs a stateless web application on EC2 instances behind an Application Load Balancer. The application is deployed in an Auto Scaling group with a minimum of 2 and maximum of 10 instances. During a traffic spike, the Auto Scaling group launches new instances, but the new instances are immediately marked as unhealthy by the ALB and terminated. What could be the cause? (Choose TWO.)

Select 2 answers
A.The health check path is misconfigured.
B.The Auto Scaling group does not have sufficient capacity in the target AZ.
C.The instances do not have the required IAM role to register with the ALB.
D.The security group for the instances does not allow inbound traffic from the ALB.
E.The instances are launched with a larger instance type than expected.
AnswersA, D

The ALB health check sends HTTP(S) requests to a configured path and expects a 2xx or 3xx response within a set timeout. If the path is incorrect (e.g., a missing endpoint or a route that returns 404), the health check fails, causing the ALB to mark the instance unhealthy and eventually terminate it. This is the most common cause of healthy-appearing instances being deregistered, and correcting the path to a verified reachable endpoint resolves the issue.

Why this answer

Option A is correct because if the ALB health check path is misconfigured (for example, pointing to a non-existent URL or wrong port), the target group health checks will fail and the ALB will mark the newly launched instances as unhealthy, causing the Auto Scaling group to terminate them. Option D is correct because the ALB must be able to reach the instances on the health check and traffic ports; if the instances' security group does not allow inbound traffic from the ALB's security group (or from the ALB subnet CIDRs), the health checks will time out and the instances will be marked unhealthy. Option B is not correct because insufficient capacity in a target AZ would prevent instances from launching at all, not cause them to launch and then be marked unhealthy by the ALB.

Option C is not correct because EC2 instances do not need an IAM role to register with an ALB; target registration is handled by the Auto Scaling group or ELB service itself, not by instance-level IAM permissions. Option E is not correct because a larger instance type does not inherently cause ALB health checks to fail; instance size is unrelated to health check success.

Exam trap

The trap here is that candidates often overlook the security group requirement for inbound traffic from the ALB, assuming that the ALB can always reach instances, or they confuse IAM roles with network-level registration requirements.

34
MCQmedium

A company runs a critical web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The application uses session stickiness (sticky sessions) to maintain user sessions. The SysOps administrator notices that when instances are replaced during a scale-in or failure event, users lose their session data. The administrator needs to preserve session data across instance failures without losing stickiness benefits. What should the administrator do?

A.Disable sticky sessions on the ALB and configure the application to store session data in an external session store like Amazon ElastiCache for Redis.
B.Increase the stickiness duration to a very high value so that sessions are not lost during brief interruptions.
C.Change the Auto Scaling group to use a larger instance type to handle more sessions per instance, reducing the likelihood of session loss.
D.Configure the Auto Scaling group to use a larger minimum size and a lower maximum, so instances are less likely to be terminated.
AnswerA

Disabling sticky sessions and moving session state to an external service like ElastiCache for Redis decouples user session data from individual EC2 instance lifecycles. When an ALB routes requests to any healthy instance, the instance can retrieve the session from Redis, so a failed or terminated instance does not lose state. Because ElastiCache replicates across AZs, sessions also survive single-cache-node failures, making the app tier effectively stateless and highly resilient.

Why this answer

It eliminates the dependency on stickiness by storing session data externally in Amazon ElastiCache for Redis. This way, if an instance fails or is scaled in, any other instance can retrieve the session data from the shared cache, preserving the user session. Disabling sticky sessions is necessary because with external storage, stickiness is no longer needed and can cause uneven load distribution.

Exam trap

The trap is that candidates may think they need to keep stickiness active, but the correct solution is to remove stickiness and store session data externally. Stickiness only provides routing affinity, not data persistence, and with external storage, any instance can serve any session.

How to eliminate wrong answers

Option B is wrong because increasing the stickiness duration does not preserve session data when an instance is terminated or fails; it only controls how long the ALB remembers the routing cookie, but the session data stored locally on the instance is still lost. Option C is wrong because using a larger instance type does not solve the fundamental problem of session data being stored locally; it only reduces the frequency of scale-in events but does not protect against instance failures or replacements. Option D is wrong because adjusting the Auto Scaling group's minimum and maximum sizes does not prevent session loss during scale-in or failure events; it only changes the number of instances running, but any instance that is terminated or replaced will still lose its locally stored session data.

35
MCQeasy

A company uses Amazon Route 53 for DNS. They want to ensure that if the primary web server fails, traffic is automatically routed to a secondary server in another region. Which routing policy should be used?

A.Simple routing policy
B.Failover routing policy
C.Latency routing policy
D.Weighted routing policy
AnswerB

Failover routing is designed specifically for active-passive failover: you create two (or more) records with the same name and type, designate one as primary and one as secondary, and attach health checks. Route 53 monitors the health of the primary and automatically routes traffic to the secondary if the primary becomes unhealthy. This behavior directly matches the requirement to ensure DNS-based failover when an endpoint fails.

Why this answer

The Failover routing policy in Amazon Route 53 is specifically designed for active-passive failover configurations. When the primary endpoint fails a health check, Route 53 automatically returns the secondary record in DNS responses, ensuring traffic is routed to the secondary server in another region. This directly meets the requirement for automatic failover between primary and secondary web servers.

Exam trap

The trap here is that candidates often confuse Failover routing policy with Weighted routing policy, mistakenly thinking weights can be set to 100/0 for failover, but Weighted routing does not automatically fail over based on health checks—it requires manual intervention or custom automation to adjust weights.

How to eliminate wrong answers

Option A is wrong because Simple routing policy only returns a single record (e.g., one IP) and does not support health checks or automatic failover; if the primary server fails, DNS continues to return the same IP, causing downtime. Option C is wrong because Latency routing policy routes traffic based on the lowest network latency to the client, not on the health or availability of endpoints; it does not provide failover between primary and secondary servers. Option D is wrong because Weighted routing policy distributes traffic across multiple endpoints based on assigned weights, but it does not automatically fail over to a secondary endpoint when the primary fails unless combined with health checks and manual weight adjustment, which is not the intended use for active-passive failover.

36
MCQmedium

A company runs a critical production database on Amazon RDS for MySQL with a Multi-AZ deployment. The database experiences a primary instance failure. The SysOps administrator needs to understand exactly how the failover process worked and why the application experienced a longer-than-expected downtime. Which AWS service or feature should the administrator use to review detailed events and actions during the failover?

A.AWS Personal Health Dashboard
B.Amazon RDS Performance Insights
C.Amazon CloudWatch Logs
D.AWS CloudTrail
AnswerA

The AWS Personal Health Dashboard (PHD) is the correct resource because it surfaces service health events that are specific to your AWS account and resources. For an RDS Multi-AZ failover, PHD provides a detailed event with the exact time, date, affected database instance, and the cause of the failover (e.g., infrastructure maintenance, hardware degradation, or patching). PHD also includes a timeline of activity and often links to related operational guidance, making it the authoritative source for reviewing automated failover details. Unlike generic service health dashboards, PHD filters events down to your particular resources, ensuring you see the actual failover incident that occurred.

Why this answer

AWS Personal Health Dashboard provides a personalized view of the health of AWS services and resources, including detailed event logs for RDS Multi-AZ failovers. It surfaces the exact sequence of actions (e.g., DNS record update, failover initiation, completion) and any underlying AWS infrastructure issues that caused the extended downtime, such as degraded hardware or network latency. This is the correct tool because it gives the administrator a chronological, AWS-side account of the failover process, which is not available through other services.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (which records API calls) with the ability to view internal service events, but CloudTrail does not capture automatic failover processes or infrastructure health events that are only available through AWS Personal Health Dashboard.

How to eliminate wrong answers

Option B is wrong because Amazon RDS Performance Insights focuses on database performance metrics (e.g., CPU, memory, SQL query load) and does not log failover events or infrastructure-level actions. Option C is wrong because Amazon CloudWatch Logs can capture RDS log files (e.g., error logs, slow query logs) but does not inherently record the failover process steps or AWS-side infrastructure events; it would require custom agent configuration to capture such data. Option D is wrong because AWS CloudTrail records API calls made to the RDS service (e.g., ModifyDBInstance) but does not capture internal failover events or DNS propagation details that occur automatically during a Multi-AZ failover.

37
MCQhard

A company has a production RDS for PostgreSQL instance. They need to recover from a logical corruption that occurred 2 hours ago. Which recovery method will minimize data loss?

A.Restore from the latest automated snapshot taken 1 hour ago.
B.Use pg_dump to export the database and restore it.
C.Fail over to the read replica in another AZ.
D.Perform a point-in-time recovery to a time just before the corruption occurred.
AnswerD

Point-in-Time Recovery (PITR) leverages automated backups and WAL transaction logs to restore a new database instance to any time within the backup retention window, typically with accuracy to a second. By selecting a time immediately before the corrupt transaction executed, you recover a clean dataset while retaining all legitimate transactions that occurred up to that moment. This is the intended RDS technique for logical corruption incidents.

Why this answer

Point-in-time recovery (PITR) for RDS PostgreSQL allows you to restore the database to any second within the backup retention period, using automated backups and transaction logs. By restoring to a time just before the logical corruption occurred (2 hours ago), you can recover the database to its state before the corruption, minimizing data loss to only transactions that happened after that point. This is the only option that can target a specific moment before the corruption, unlike a snapshot which is a fixed point in time.

Exam trap

The trap here is that candidates often assume a read replica or a recent snapshot can protect against logical corruption, but both replicate the corruption because they are copies of the same data, whereas point-in-time recovery leverages transaction logs to rewind to a clean state.

How to eliminate wrong answers

Option A is wrong because restoring from the latest automated snapshot taken 1 hour ago would recover data from that snapshot time, which is after the corruption occurred (2 hours ago), so the corruption would be included in the restored data, resulting in data loss of the entire 2-hour window. Option B is wrong because pg_dump exports the current state of the database, which already includes the logical corruption; restoring from that dump would simply reapply the corruption. Option C is wrong because failing over to a read replica in another AZ promotes an asynchronous replica that replicates the same corrupted data from the primary instance, so it does not provide a point-in-time recovery before the corruption.

38
MCQhard

A company has a production RDS for PostgreSQL instance with Multi-AZ enabled. During a recent failover test, the application experienced a 5-minute downtime. The company requires that failover be completed within 2 minutes. Which action should be taken to meet this requirement?

A.Migrate the database to Amazon Aurora with Multi-AZ.
B.Enable automated backups with a short retention period.
C.Increase the DB instance class to a larger size.
D.Configure an RDS Proxy in front of the database.
AnswerD

Configuring an RDS Proxy in front of the database is the correct, targeted fix because the proxy maintains persistent outbound connections to the DB endpoints and pools inbound client connections. During a failover, RDS Proxy automatically establishes connections to the new primary and keeps the client-side connections open, making the failover appear essentially transparent to the application. This eliminates the need for clients to implement complex reconnect logic and reduces the overall failover time by avoiding a cold-start connection establishment storm.

Why this answer

RDS Proxy reduces failover time by maintaining database connections and connection pooling, allowing the application to reconnect quickly without waiting for DNS propagation or new connection setup. In a Multi-AZ failover, RDS Proxy can typically complete failover in under 60 seconds, meeting the 2-minute requirement by minimizing connection disruption.

Exam trap

The trap here is that candidates assume increasing instance size or enabling backups improves failover speed, but the real bottleneck is connection management and DNS propagation, which RDS Proxy addresses directly.

How to eliminate wrong answers

Option A is wrong because Amazon Aurora with Multi-AZ already provides fast failover (typically under 30 seconds), but migrating to Aurora is not the simplest action and does not directly address the application's connection handling issue that caused the 5-minute downtime. Option B is wrong because enabling automated backups with a short retention period does not affect failover speed; backups are for point-in-time recovery, not for reducing failover time. Option C is wrong because increasing the DB instance class improves performance but does not reduce failover time; failover duration depends on DNS propagation and connection re-establishment, not instance size.

39
MCQmedium

A company runs a stateful web application on a single Amazon EC2 instance. The SysOps administrator needs to implement a high availability architecture that can tolerate an Availability Zone (AZ) failure. The application stores session state in memory and also writes critical data to an Amazon EBS volume. The administrator wants to use an Auto Scaling group and an Application Load Balancer (ALB). Which combination of steps is required to make the application highly available?

A.Create an Auto Scaling group that spans at least two Availability Zones, attach the existing EBS volume to the new instances, and use an ALB to distribute traffic.
B.Migrate session state to Amazon ElastiCache for Redis, store critical data in Amazon EFS, create an Auto Scaling group across multiple AZs, and place it behind an ALB.
C.Place the EC2 instance in an Auto Scaling group with a minimum and maximum of 1 in the same AZ, and attach an Elastic IP to the instance.
D.Use an ALB with the existing single instance as the target, and enable cross-zone load balancing.
AnswerB

This option makes the application stateless at the compute layer by externalizing session state to ElastiCache for Redis, which all instances can access, and storing critical application data on Amazon EFS, a shared regional file system. An Auto Scaling group spanning multiple Availability Zones ensures that an instance failure or entire AZ outage triggers replacement, while the ALB distributes traffic only to healthy instances and performs health checks. This architecture achieves both high availability and horizontal scalability because no unique state is tied to any individual EC2 instance.

Why this answer

It addresses both the stateless requirement for horizontal scaling and the persistence of critical data across AZ failures. Migrating session state to ElastiCache for Redis removes the dependency on local instance memory, allowing any instance to handle any request. Storing critical data on Amazon EFS provides a shared, NFS-based file system that is accessible from all instances across multiple AZs, unlike EBS which is tied to a single AZ.

Combining these with a multi-AZ Auto Scaling group and an ALB ensures the application can survive an entire AZ outage.

Exam trap

The trap here is that candidates assume EBS volumes can be shared across instances or AZs, or that a single-instance setup with an ALB provides high availability, when in fact EBS is a single-AZ resource and the ALB requires multiple healthy targets to tolerate failures.

How to eliminate wrong answers

Option A is wrong because EBS volumes are AZ-scoped and cannot be attached to instances in a different AZ; attaching the existing EBS volume to new instances in another AZ is impossible without snapshotting and recreating, which defeats high availability. Option C is wrong because keeping a single instance in one AZ with an Elastic IP does not provide fault tolerance for an AZ failure; the Auto Scaling group with min/max of 1 cannot replace the instance in a different AZ automatically, and the Elastic IP does not reroute traffic to a healthy instance. Option D is wrong because using an ALB with a single instance as the target and enabling cross-zone load balancing does not add redundancy; if the instance or its AZ fails, the ALB has no other targets to route traffic to, so the application becomes unavailable.

40
MCQhard

A company uses a Multi-AZ RDS for MySQL instance for its production database. During a maintenance window, the primary instance fails and a failover occurs. However, the application experiences a 5-minute downtime. The application uses a DNS CNAME record pointing to the RDS endpoint. What is the MOST likely cause of the downtime?

A.The application was using a cached DNS resolution for the RDS endpoint.
B.The application was not configured to retry connections after a failover.
C.The RDS endpoint changed after failover and the application did not update.
D.The failover process took longer than expected due to a large transaction log.
AnswerA

The RDS endpoint is a DNS name, and during a Multi-AZ failover, Amazon RDS updates the DNS record to point to the new primary instance's private IP address. If the application caches the DNS resolution longer than the TTL (or ignores TTL), it continues to attempt connections to the old, now-unavailable IP, causing persistent failures even though the endpoint name itself is correct. This is the classic root cause for post-failover connection errors: the client is not using the updated DNS mapping.

Why this answer

During a Multi-AZ RDS failover, the RDS DNS CNAME record is updated to point to the new primary instance in a different Availability Zone. However, the application's DNS resolver may have cached the previous IP address (TTL-based). If the application does not flush its DNS cache or the TTL is long, it continues to connect to the old (failed) IP, causing connection timeouts until the cache expires.

This is the most likely cause of the 5-minute downtime, as RDS failovers typically complete within 1-2 minutes.

Exam trap

The trap here is that candidates assume the RDS endpoint changes after failover (Option C), but AWS explicitly states the CNAME remains the same; the real issue is client-side DNS caching, which is a common oversight in high-availability architectures.

How to eliminate wrong answers

Option B is wrong because even if the application retries connections, it will still fail if it keeps resolving the old cached IP address; retries alone do not fix a stale DNS cache. Option C is wrong because the RDS endpoint (CNAME) does not change after a failover — it remains the same; only the underlying IP address changes. Option D is wrong because a large transaction log can delay failover completion, but the question states the failover occurs and the downtime is 5 minutes, which is longer than typical failover time; the primary issue is DNS caching, not transaction log size.

41
MCQhard

A company runs a critical web application on Amazon EC2 instances that are part of an Auto Scaling group. The application receives unpredictable traffic spikes. The SysOps administrator needs to ensure that when a scale-out event occurs, new instances are ready to serve traffic quickly to minimize latency spikes. Currently, the instance launch and configuration process (including software installs and cache warming) takes about 5 minutes. The administrator wants to reduce the time it takes for new instances to start serving traffic. Which combination of Auto Scaling features should be used?

A.Use a launch template that includes a pre-warmed Amazon Machine Image (AMI) with all software pre-installed, and configure the Auto Scaling group to use a larger instance type to reduce initialization time.
B.Implement an Auto Scaling warm pool with a minimum number of pre-initialized instances in a 'Stopped' state. Configure the scaling policy to move instances from the warm pool to the Auto Scaling group when needed.
C.Use scheduled scaling to predictively launch instances before the traffic spikes based on historical patterns.
D.Configure lifecycle hooks to add a wait time during instance launch so that the instance is fully configured before it is placed behind the load balancer.
AnswerB

A warm pool maintains instances that have been fully launched and configured but are stopped or in a standby state. When scale-out occurs, instances from the warm pool are started or moved into service quickly, drastically reducing the time to handle traffic.

Why this answer

An Auto Scaling warm pool maintains a pool of pre-initialized instances in a 'Stopped' state that are fully configured (software installed, cache warmed) and ready to serve traffic. When a scale-out event occurs, instances from the warm pool are moved to the Auto Scaling group and transitioned to 'Running' state, bypassing the 5-minute launch and configuration delay, thereby minimizing latency spikes.

Exam trap

The trap here is that candidates often confuse warm pools with lifecycle hooks or pre-warmed AMIs, assuming that reducing software install time alone is sufficient, when the real bottleneck is the entire instance initialization process that warm pools bypass.

How to eliminate wrong answers

Option A is wrong because using a pre-warmed AMI reduces software installation time but does not eliminate the instance launch and initialization overhead (e.g., kernel boot, network setup, cache warming), and using a larger instance type does not inherently reduce initialization time—it may even increase it due to more hardware resources to initialize. Option C is wrong because scheduled scaling relies on predictable traffic patterns and cannot handle unpredictable traffic spikes; it would either over-provision or under-provision for unexpected demand. Option D is wrong because lifecycle hooks add a wait time during instance launch, which would increase the time before the instance is ready to serve traffic, contradicting the goal of reducing latency spikes.

42
MCQeasy

A company runs a critical application on an EC2 instance backed by Amazon EBS. To protect against data loss, the company wants to create a backup strategy that allows for point-in-time recovery. Which solution should be used?

A.Configure an S3 Lifecycle policy to move data to Glacier.
B.Create an Amazon Machine Image (AMI) of the instance.
C.Use Amazon EFS to store data.
D.Create automated EBS snapshots.
AnswerD

Automated EBS snapshots use Amazon Data Lifecycle Manager (DLM) or AWS Backup to take scheduled, point-in-time copies of your EC2 instance's volumes. Each snapshot is incremental—only the blocks that changed since the previous snapshot are stored—and snapshots are stored redundantly in S3 for high durability. By setting a regular schedule, you get an ongoing backup that can be restored to a new volume or used to rebuild the instance, making it the correct choice for protecting a critical application's data.

Why this answer

Automated EBS snapshots provide point-in-time backups of the EBS volume, enabling granular recovery to a specific moment. Snapshots are stored in Amazon S3 and can be used to restore the volume or create new instances, directly addressing the requirement for point-in-time recovery against data loss.

Exam trap

The trap here is that candidates may confuse AMIs (which are used for instance-level recovery and launching new instances) with EBS snapshots (which are volume-level backups designed for granular point-in-time recovery), leading them to select Option B instead of D.

How to eliminate wrong answers

Option A is wrong because an S3 Lifecycle policy to move data to Glacier is for archiving objects in S3, not for backing up an EC2 instance's EBS volume; it does not provide point-in-time recovery of the instance or its data. Option B is wrong because an AMI captures the entire instance configuration (including attached volumes) but is typically used for launching new instances, not for granular point-in-time recovery of individual EBS volumes; AMIs are less frequent and more heavyweight than snapshots for backup purposes. Option C is wrong because Amazon EFS is a separate network file system that must be mounted to the instance; it does not back up the existing EBS root or data volumes, and it introduces additional complexity without addressing the requirement for point-in-time recovery of the EBS-backed instance.

43
MCQmedium

A SysOps administrator receives an alert that an EC2 instance in an Auto Scaling group is unhealthy. The instance fails the EC2 status check. What is the BEST course of action to restore availability automatically?

A.Use AWS Systems Manager to replace the underlying host.
B.Manually reboot the instance from the EC2 console.
C.Create a CloudWatch alarm that triggers an SNS notification to the administrator.
D.Configure the Auto Scaling group to use EC2 status checks for health checks and set the health check grace period appropriately.
AnswerD

Configuring the Auto Scaling group to use EC2 status checks as its health check type allows the ASG to automatically detect when an instance fails either a system or instance status check. Once marked unhealthy, the ASG terminates the instance and launches a replacement, providing self-healing without manual intervention. Setting the health check grace period appropriately ensures the instance is given enough time to initialize and pass status checks before being evaluated, preventing premature termination during boot or application startup.

Why this answer

Configuring the Auto Scaling group to use EC2 status checks for health checks allows it to automatically detect when an instance fails the EC2 status check and replace it with a new one, ensuring high availability without manual intervention. The health check grace period prevents premature termination during initial instance bootstrapping. This is the most automated and resilient approach for restoring availability in response to EC2 status check failures.

Exam trap

The trap here is that candidates may think a CloudWatch alarm with SNS notification is sufficient for automatic recovery, but it only provides notification, not automated remediation, whereas the Auto Scaling group's health check configuration directly triggers instance replacement without manual steps.

How to eliminate wrong answers

Option A is wrong because AWS Systems Manager does not have a capability to replace the underlying host of an EC2 instance; it is used for operational management like patching and configuration, not for host replacement. Option B is wrong because manually rebooting the instance from the EC2 console requires human intervention and does not provide automatic recovery, which contradicts the requirement to restore availability automatically. Option C is wrong because creating a CloudWatch alarm that triggers an SNS notification only alerts the administrator but does not take any automated action to replace or recover the unhealthy instance, leaving the availability restoration dependent on manual response.

44
MCQhard

A company runs a critical database workload on an Amazon RDS for MySQL DB instance with Multi-AZ deployment in the us-east-1 region. The SysOps administrator must design a disaster recovery strategy that can recover from a complete regional outage. The Recovery Time Objective (RTO) is 2 hours and the Recovery Point Objective (RPO) is 1 hour. Which solution meets these requirements at the lowest cost?

A.Create manual snapshots of the DB instance every hour and copy them to another AWS Region.
B.Enable automated backups with a retention period of 35 days and restore to a different Region when needed.
C.Create a cross-Region read replica in another Region and promote it to a standalone DB instance during a disaster.
D.Use AWS Database Migration Service (DMS) to continuously replicate data to a DB instance in another Region.
AnswerC

A cross-Region read replica provides continuous asynchronous replication with low lag (typically seconds). In a disaster, promoting the replica to a primary instance takes only minutes, meeting the RTO and RPO requirements with minimal cost.

Why this answer

A cross-Region read replica continuously replicates data from the primary RDS MySQL instance to another Region with minimal lag, typically achieving an RPO of seconds to minutes, well within the 1-hour requirement. Promoting the replica to a standalone instance during a disaster can be done in minutes, meeting the 2-hour RTO. This approach is the lowest cost among the viable options as it uses existing replication infrastructure without additional data transfer fees for snapshots or DMS replication instances.

Exam trap

The trap here is that candidates often choose Option B (automated backups) because they assume backups can be restored cross-Region, but automated backups are Region-specific and do not support cross-Region restore without additional snapshot copy configuration, which is not mentioned in the option.

How to eliminate wrong answers

Option A is wrong because manual snapshots taken every hour would incur significant storage costs for storing and copying snapshots across Regions, and the copy process can take longer than 1 hour, potentially exceeding the RPO. Option B is wrong because automated backups with a 35-day retention period are stored only in the source Region and cannot be restored to a different Region; cross-Region snapshot copy must be explicitly configured and is not part of automated backups. Option D is wrong because AWS DMS incurs additional costs for a replication instance and data transfer, making it more expensive than a cross-Region read replica, and it adds operational complexity for continuous replication that is unnecessary when native MySQL replication can achieve the same RPO/RTO.

45
MCQeasy

A company has a fleet of EC2 instances that need to be patched monthly. The SysOps administrator must ensure that the patching process does not affect the availability of the application. Which strategy should the administrator use?

A.Patch one instance at a time manually by stopping and starting.
B.Use an Auto Scaling group with a rolling update strategy.
C.Use AWS Systems Manager Patch Manager to patch all instances at once.
D.Stop all instances, apply patches, then start them.
AnswerB

An Auto Scaling group with a rolling update strategy replaces a small number of instances at a time by updating the launch template with a patched AMI or user data and then performing an instance refresh (or manually incrementing the desired count). Each new instance must pass the group's health checks before the next batch is terminated, so the fleet maintains its desired capacity and continuous availability throughout the patching process. This approach is the correct one because it combines automation, health verification, and controlled blast radius, ensuring that no single point of failure or full downtime occurs.

Why this answer

Using an Auto Scaling group with a rolling update strategy allows the administrator to replace instances in batches, ensuring that a minimum number of instances remain in service throughout the patching process. This maintains application availability by avoiding simultaneous disruption to all instances, which is a key requirement for high-availability architectures.

Exam trap

The trap here is that candidates often assume AWS Systems Manager Patch Manager can be configured to patch instances in a rolling fashion, but by default it runs on all targeted instances simultaneously unless explicitly orchestrated with a maintenance window and rate control, which is not the same as an Auto Scaling group rolling update.

How to eliminate wrong answers

Option A is wrong because manually patching one instance at a time by stopping and starting is error-prone, lacks automation, and does not integrate with health checks or lifecycle hooks to ensure the instance is fully operational before proceeding, risking downtime if the manual process is slow or fails. Option C is wrong because using AWS Systems Manager Patch Manager to patch all instances at once would apply patches simultaneously, potentially causing all instances to reboot at the same time and resulting in complete application downtime. Option D is wrong because stopping all instances, applying patches, and then starting them creates a total outage window where the application is unavailable, violating the requirement to not affect availability.

46
MCQeasy

An application uploads files to an S3 bucket. The SysOps administrator needs to ensure that the files are automatically replicated to another bucket in a different AWS Region for disaster recovery. Which action should be taken?

A.Enable Cross-Region Replication on the source bucket.
B.Use S3 Transfer Acceleration for faster uploads.
C.Enable versioning on the source bucket.
D.Configure a lifecycle policy to transition objects to Glacier.
AnswerA

Cross-Region Replication (CRR) is the correct mechanism because it asynchronously copies objects from your source S3 bucket to a separate destination bucket in a different AWS Region. To use CRR, you must enable versioning on both the source and destination buckets, and S3 creates an IAM role with permissions to read from the source and write to the destination. Once configured, every new upload is automatically replicated, including metadata and object versions, giving you a second copy in another region for disaster recovery and compliance.

Why this answer

Cross-Region Replication (CRR) is the correct AWS feature to automatically replicate objects from a source S3 bucket to a destination bucket in a different AWS Region. CRR requires versioning to be enabled on both the source and destination buckets, and it copies every object uploaded to the source bucket asynchronously to the destination bucket, providing disaster recovery across regions.

Exam trap

The trap here is that candidates often confuse enabling versioning (a prerequisite for CRR) with the actual replication action, or they mistakenly think Transfer Acceleration or lifecycle policies can achieve cross-region replication.

How to eliminate wrong answers

Option B is wrong because S3 Transfer Acceleration speeds up uploads over long distances using AWS edge locations, but it does not replicate data to another region. Option C is wrong because enabling versioning alone does not replicate objects; it only preserves multiple versions of objects within the same bucket. Option D is wrong because a lifecycle policy transitions objects to Amazon S3 Glacier for cost optimization, not for cross-region replication or disaster recovery.

47
Multi-Selecthard

Which THREE measures help protect an S3 bucket from accidental data loss? (Choose 3)

Select 3 answers
A.Enable MFA Delete on the bucket.
B.Create a lifecycle policy to transition objects to S3 Glacier.
C.Enable server-side encryption on the bucket.
D.Configure cross-region replication to a destination bucket.
E.Enable versioning on the bucket.
AnswersA, D, E

MFA Delete requires a valid multi-factor authentication code to permanently delete an object version or suspend versioning, meaning an accidental delete request without the temporary code will be rejected. Because the MFA device is separate from AWS credentials, even a compromised access key cannot complete destructive actions like version removal. This adds a human/physical factor that effectively prevents costly accidental deletions and serves as a last line of defense for critical data.

Why this answer

Enabling MFA Delete on an S3 bucket requires multi-factor authentication for any delete operations, adding an extra layer of protection against accidental or unauthorized deletion of objects. This helps prevent data loss by ensuring that even if credentials are compromised, a delete action cannot be performed without the MFA token.

Exam trap

The trap here is that candidates often confuse data protection features like encryption or lifecycle policies with data durability and accidental deletion prevention, leading them to select options that secure data but do not prevent loss from deletion.

48
MCQhard

A company runs a critical application on EC2 instances in an Auto Scaling group. The application stores state information locally on the instance. The SysOps administrator needs to ensure that if an instance fails, the state is not lost. What should the administrator do?

A.Move the state data to an external data store such as ElastiCache or RDS.
B.Attach an EBS volume and set the 'DeleteOnTermination' flag to false.
C.Use instance store volumes for the state data.
D.Use Amazon SQS to store the state data.
AnswerA

Externalizing state to a managed service like ElastiCache or RDS decouples application data from the EC2 instance lifecycle, so any instance failure, termination, or replacement has no impact on data availability. ElastiCache provides extremely low-latency in-memory state ideal for session/state caching, while RDS offers durable, ACID-compliant storage with automated backups and multi-AZ failover. This design pattern makes the application stateless at the instance tier, enabling Auto Scaling, rolling deployments, and immediate recovery because new instances simply reconnect to the same external state store.

Why this answer

State stored locally on an EC2 instance is ephemeral and lost if the instance fails or is terminated. The most reliable solution is to externalize state to a durable, shared data store such as ElastiCache (for session state) or RDS (for relational state). This decouples the application from the instance lifecycle and allows any instance in the Auto Scaling group to serve requests.

Exam trap

SOA-C02 often tests the misconception that EBS volumes with DeleteOnTermination=false provide high availability for state; candidates must recognize that shared, external state stores are required for Auto Scaling groups.

How to eliminate wrong answers

Option B is wrong because attaching an EBS volume with DeleteOnTermination=false preserves the volume but does not automatically make the state available to a replacement instance; the new instance would need to attach and mount the volume, and EBS volumes are tied to an Availability Zone, complicating Auto Scaling. Option C is wrong because instance store volumes are ephemeral and lose data on instance stop, termination, or hardware failure, which is the opposite of what is needed. Option D is wrong because SQS is a message queue, not a data store for application state; it is designed for asynchronous messaging, not for storing and retrieving session state.

49
MCQeasy

A company wants to ensure that its EC2 instances automatically recover from an instance failure. Which feature should be used?

A.Create a CloudWatch alarm that sends an email when the instance status check fails.
B.Configure an Auto Scaling group with a launch configuration.
C.Attach the instance to an Elastic Load Balancer.
D.Enable EC2 Auto Recovery on the instance.
AnswerD

EC2 Auto Recovery, when enabled and attached to a CloudWatch alarm on the StatusCheckFailed_System metric, automatically restores the instance on new underlying hardware if the host is degraded. It preserves the instance ID, private IP address, Elastic IP address, and instance metadata, allowing applications to continue with minimal downtime. This is the only option that actually performs the instance recovery automation described in the scenario.

Why this answer

EC2 Auto Recovery is a feature that automatically recovers an EC2 instance when it becomes impaired due to an underlying hardware or system issue. When enabled, CloudWatch monitors the instance's status checks and, upon detecting a failure, automatically stops and starts the instance on a new healthy host, preserving its private IP, Elastic IP, and instance metadata.

Exam trap

The trap here is that candidates often confuse Auto Scaling groups (which replace instances) with EC2 Auto Recovery (which recovers the same instance), leading them to choose option B instead of D.

How to eliminate wrong answers

Option A is wrong because creating a CloudWatch alarm that sends an email only notifies an administrator of the failure; it does not automatically recover the instance. Option B is wrong because an Auto Scaling group with a launch configuration can replace a failed instance by launching a new one, but it does not recover the original instance's state, such as its private IP or Elastic IP, and is not designed for automatic recovery of a single instance. Option C is wrong because attaching the instance to an Elastic Load Balancer only distributes traffic and can route around a failed instance, but it does not recover the instance itself.

50
MCQhard

A company uses AWS Backup to back up its Amazon EFS file system daily. The backup retention policy is set to 30 days. Recently, a user accidentally deleted a critical directory. The company wants to restore the directory as it existed 2 days ago. What is the MOST cost-effective and quickest way to achieve this?

A.Use the EFS console to recover the directory from the .Trash folder.
B.Enable EFS replication to another region and then fail back.
C.Use AWS Backup to restore the entire file system to an on-premises server, then copy the directory back.
D.Restore the backup from 2 days ago to a new EFS file system, then copy the directory to the original file system.
AnswerD

AWS Backup for EFS captures point-in-time snapshots accessible as recovery points; you can restore the recovery point from two days ago to a brand-new EFS file system, which is a fast, fully managed operation. Once the restored file system is mounted (e.g., on a temporary EC2 instance), you can selectively copy the missing directory from it to the original file system using cp or rsync or AWS DataSync. This provides a clean, isolated recovery path that does not alter the current file system until you explicitly copy the target directory, so it is the correct way to achieve selective point-in-time recovery.

Why this answer

AWS Backup creates point-in-time snapshots of EFS file systems. Restoring a backup from 2 days ago to a new EFS file system allows you to mount that new file system, copy the specific directory back to the original file system, and then delete the temporary file system. This is the most cost-effective and quickest approach since it avoids moving data to on-premises servers and uses native AWS services without additional replication costs.

Exam trap

The trap here is that candidates may assume AWS Backup can restore directly into the original EFS file system, but AWS Backup for EFS always creates a new file system during restore, requiring a manual copy step to recover specific data.

How to eliminate wrong answers

Option A is wrong because EFS does not have a '.Trash' folder; that concept is specific to certain desktop operating systems, not Amazon EFS. Option B is wrong because EFS replication is a continuous, region-level feature designed for disaster recovery, not for restoring a specific point-in-time backup; enabling replication and failing back would be slow, expensive, and does not target a specific backup from 2 days ago. Option C is wrong because restoring an entire EFS file system to an on-premises server requires significant network bandwidth, time, and infrastructure, and is not the quickest or most cost-effective method for recovering a single directory.

51
MCQeasy

A company wants to create a disaster recovery (DR) strategy for its RDS for PostgreSQL database. The primary database is in us-east-1. The company needs a recovery point objective (RPO) of less than 5 minutes and a recovery time objective (RTO) of less than 1 hour. Which solution meets these requirements?

A.Enable Multi-AZ in us-east-1 and create a standby in a different Availability Zone.
B.Create a cross-region Read Replica in us-west-2 and promote it during a disaster.
C.Use AWS Database Migration Service (DMS) to continuously replicate to an EC2 instance.
D.Take daily automated snapshots and copy them to us-west-2.
AnswerB

A cross-region Read Replica in us-west-2 receives asynchronous replication from the primary in us-east-1, keeping a warm standby in a separate geographic region. During a disaster, you can promote the replica to a standalone primary in minutes, providing a low RPO (typically seconds or minutes of data loss) and a low RTO (often under 15 minutes). This option directly addresses region failure while also offloading read traffic before promotion.

Why this answer

A cross-region Read Replica for Amazon RDS PostgreSQL maintains an asynchronous replication lag typically under 5 seconds, easily meeting the RPO of less than 5 minutes. During a disaster, promoting the Read Replica to a standalone instance can be completed in minutes, satisfying the RTO of less than 1 hour. This approach provides both low RPO and fast recovery without the complexity of additional services.

Exam trap

The trap here is that candidates often confuse Multi-AZ (high availability within a region) with cross-region disaster recovery, assuming Multi-AZ provides regional fault tolerance, but it does not protect against a region-wide outage.

How to eliminate wrong answers

Option A is wrong because Multi-AZ provides high availability within a single region, not disaster recovery across regions; it cannot protect against a regional failure, and failover to a standby in a different Availability Zone does not meet the cross-region DR requirement. Option C is wrong because AWS DMS continuous replication to an EC2 instance introduces additional management overhead, potential licensing costs, and does not guarantee the sub-5-minute RPO or sub-1-hour RTO as reliably as a managed Read Replica; DMS is better suited for heterogeneous migrations or ongoing replication, not as a primary DR mechanism for RDS. Option D is wrong because daily automated snapshots have an RPO of up to 24 hours, far exceeding the required 5-minute RPO, and restoring from a snapshot in another region can take longer than 1 hour, failing the RTO requirement.

52
Multi-Selecthard

A company wants to implement a disaster recovery solution for its on-premises database using AWS. The solution must have an RPO of less than 1 hour and an RTO of less than 4 hours. Which THREE steps should the SysOps administrator take? (Choose THREE.)

Select 3 answers
A.Set up a cross-Region read replica for the RDS instance.
B.Launch an EC2 instance with the database software and configure replication.
C.Use AWS Database Migration Service (DMS) to replicate data to an RDS instance.
D.Use AWS DataSync to sync the database files to Amazon S3.
E.Configure the RDS instance with Multi-AZ.
AnswersA, B, C

A cross-Region read replica creates a continuously updated, asynchronous replicate of an RDS database in a different AWS Region, with typical replication lag well under the 1-hour RPO. Promoting the replica makes it a standalone writable instance, a process that generally completes within minutes and satisfies the RTO. This is a solid DR approach, but it presupposes that the database is already running on RDS; for an on-premises source, you would need an initial migration into RDS before this option becomes viable.

Why this answer

A cross-Region read replica for an RDS instance provides asynchronous replication to a secondary Region. After promoting the replica, the RTO can be under 4 hours, and the RPO is typically less than 1 hour. Option B is correct by launching an EC2 instance with the same database software and configuring continuous replication (e.g., log shipping or mirroring) from the on-premises database.

This allows failover to the EC2 instance within the RPO and RTO targets. Option C is correct as AWS DMS can perform ongoing replication from the on-premises database to an RDS instance, meeting the RPO requirement with minimal data loss. Together, these steps form a multi-layered DR strategy: DMS for continuous replication, EC2 as a standby, and a cross-Region replica for regional resilience.

Exam trap

Candidates often assume Multi-AZ (Option E) is a valid DR solution. However, Multi-AZ only provides high availability within a single Region, with synchronous replication and automatic failover. It does not protect against Region-wide outages or on-premises failures, and does not meet the cross-Region disaster recovery requirement implied by the need for an RPO < 1 hour and RTO < 4 hours for an on-premises database.

53
MCQhard

A company runs a critical database on an RDS for PostgreSQL instance in a single Availability Zone. The database experiences high write latency. The SysOps Administrator needs to improve the database's reliability and performance without downtime. Which solution meets these requirements?

A.Modify the RDS instance to be Multi-AZ with a standby in another Availability Zone.
B.Create a Multi-AZ deployment in the same Availability Zone.
C.Increase the allocated storage for the RDS instance.
D.Create a read replica in another Availability Zone and redirect read traffic.
AnswerA

Modifying the RDS instance to a Multi-AZ deployment provisions a synchronous standby replica in a different Availability Zone, and Amazon RDS automatically fails over to that standby if an AZ outage or primary instance failure occurs. This change can typically be applied without downtime, as it only requires a metadata modification and provisioning of the standby. This gives the database the high availability and automatic failover that the company needs.

Why this answer

Enabling Multi-AZ for an RDS for PostgreSQL instance provisions a standby replica in a different Availability Zone and synchronously replicates data to it. This eliminates the single point of failure, improving reliability. The modification is performed as a zero-downtime operation via a DNS update, meeting the requirement for no downtime.

Note that Multi-AZ improves availability but does not reduce write latency; performance improvement may come from offloading backups and other administrative tasks to the standby.

Exam trap

The trap here is that candidates confuse Multi-AZ (synchronous replication for high availability) with read replicas (asynchronous replication for read scaling), assuming a read replica can improve write performance or reliability when it only helps with read traffic.

How to eliminate wrong answers

Option B is wrong because Multi-AZ requires the standby to be in a different Availability Zone; deploying in the same AZ provides no fault isolation and does not improve reliability. Option C is wrong because increasing allocated storage addresses capacity or IOPS limits but does not improve reliability through redundancy or reduce write latency caused by synchronous replication overhead. Option D is wrong because a read replica is asynchronous and does not improve write latency or reliability for the primary database; it only offloads read traffic, leaving the primary as a single point of failure.

54
MCQmedium

A SysOps administrator is reviewing the reliability of a production system that uses Amazon DynamoDB as its primary data store. The table has on-demand capacity and a single partition key. The application experiences occasional throttling errors during peak hours. Which action would most effectively improve reliability?

A.Switch to provisioned capacity and set high read/write units.
B.Enable auto-scaling and increase the maximum capacity.
C.Review and optimize the partition key design to avoid hot partitions.
D.Enable DynamoDB Accelerator (DAX) to reduce read latency.
AnswerC

Reviewing and optimizing the partition key design is the correct fix because DynamoDB distributes data across partitions based on the partition key's hash value, and on-demand or provisioned table capacity does not guarantee uniform load per partition. Each partition has its own throughput ceiling (3000 RCU / 1000 WCU), so a skewed access pattern—such as one overly popular item or a timestamp prefix—creates a hot partition that gets throttled even when the table's overall capacity is underutilized. Adding high-cardinality suffixes like random or calculated bits to the partition key spreads the writes and reads across many partitions, eliminating the bottleneck while preserving query access if the design accounts for it. This approach directly addresses the root cause rather than masking symptoms.

Why this answer

Throttling errors in DynamoDB with a single partition key are most often caused by uneven access patterns creating hot partitions. Optimizing the partition key design (e.g., using a composite key or adding a suffix to distribute writes) directly addresses the root cause by ensuring requests are spread evenly across partitions, which on-demand capacity alone cannot fix. This improves reliability by preventing throttling at the partition level, regardless of the table's total capacity.

Exam trap

The trap here is that candidates assume throttling is always a capacity issue (solved by increasing RCU/WCU or enabling auto-scaling), when in reality it is often a data modeling problem where a single partition key creates a hot spot that no amount of capacity scaling can fix.

How to eliminate wrong answers

Option A is wrong because switching to provisioned capacity with high read/write units does not solve hot partition issues; if a single partition exceeds its 3,000 RCU or 1,000 WCU limit, throttling still occurs even with high provisioned capacity. Option B is wrong because enabling auto-scaling and increasing maximum capacity only adjusts total table throughput, not per-partition limits; it cannot prevent throttling caused by a hot key hammering one partition. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that reduces read latency for eventually consistent reads, but it does not eliminate throttling errors caused by write-heavy hot partitions or partition-level throughput limits.

55
Matchingmedium

Match each AWS storage service to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Object storage for any data

Block storage for EC2 instances

File storage for Linux instances

Managed file system for Windows or Lustre

Low-cost archival storage

Why these pairings

Amazon S3 is object storage, EBS is block storage, EFS is file storage, and Glacier is archival. Common confusions include mixing S3 and EBS definitions.

56
MCQmedium

An application running on EC2 instances in an Auto Scaling group uses an SQS queue for decoupling. The application experiences increased latency when the queue has a high number of messages. The SysOps Administrator needs to maintain responsiveness. Which solution is the most cost-effective?

A.Increase the desired capacity of the Auto Scaling group.
B.Configure a CloudWatch alarm on the queue depth to trigger Auto Scaling policies.
C.Use a larger instance type for the EC2 instances.
D.Increase the visibility timeout of the SQS queue.
AnswerB

Create a CloudWatch alarm on the SQS metric ApproximateNumberOfMessagesVisible, the number of messages waiting in the queue, and attach it to a scaling policy for the ASG. When the backlog exceeds a threshold, the alarm enters ALARM state and adds instances to consume messages faster; when the queue drains, it removes instances. This directly couples consumer fleet size to actual demand, providing cost-efficient elasticity that avoids both under-provisioning and idle over-provisioning.

Why this answer

Using a CloudWatch alarm on the SQS queue depth (ApproximateNumberOfMessagesVisible) to trigger Auto Scaling policies allows the Auto Scaling group to dynamically add EC2 instances only when the queue grows, directly addressing increased latency by scaling out compute capacity. This is the most cost-effective approach as it scales resources based on actual demand, avoiding over-provisioning.

Exam trap

The trap here is that candidates often confuse static scaling (Option A) or vertical scaling (Option C) with dynamic, demand-based scaling, or mistakenly think that increasing the visibility timeout (Option D) will reduce queue depth, when in fact it only delays message reprocessing.

How to eliminate wrong answers

Option A is wrong because increasing the desired capacity of the Auto Scaling group statically raises the number of running instances regardless of queue depth, leading to unnecessary cost when the queue is not deep. Option C is wrong because using a larger instance type increases per-instance cost and does not automatically scale with queue depth; it may still suffer from latency if the queue grows beyond the capacity of a single larger instance. Option D is wrong because increasing the visibility timeout of the SQS queue does not reduce the number of messages or processing time; it only delays when a message becomes visible again after a consumer fails, which can actually increase latency by hiding unprocessed messages longer.

57
MCQmedium

A company runs a critical web application on Amazon EC2 instances in an Auto Scaling group across three Availability Zones in us-east-1. The application stores data in an Amazon RDS for MySQL DB instance with Multi-AZ deployment. The SysOps administrator needs to design a disaster recovery strategy that can recover from a complete regional outage. The Recovery Time Objective (RTO) is 2 hours and the Recovery Point Objective (RPO) is 1 hour. Which solution should the administrator implement?

A.Create a read replica of the RDS instance in a second region. Configure an Amazon CloudFront distribution with the ALB as origin. Use Route53 failover routing policy to route traffic to the CloudFront distribution.
B.Take daily manual snapshots of the RDS instance and copy them to a second region. Store the AWS CloudFormation template for the infrastructure in an S3 bucket with cross-region replication. In the event of a disaster, manually deploy the stack and restore the snapshot.
C.Configure cross-region automated backups for the RDS instance with a backup window. Deploy an identical infrastructure stack in a second region using AWS CloudFormation StackSets. Create an Amazon Route53 DNS failover record set with health checks to automatically fail over to the second region.
D.Use AWS Database Migration Service (DMS) to continuously replicate data to a second region. Use an Application Load Balancer in the primary region and a Network Load Balancer in the secondary region. Create a Route53 weighted routing policy to distribute traffic.
AnswerC

Cross-region automated backups for RDS, when configured with a backup window, copy backups to a second region automatically, meeting the 1-hour RPO by enabling point-in-time recovery to within 5 minutes of the last transaction. Deploying an identical infrastructure stack in the second region using AWS CloudFormation StackSets ensures that compute, networking, and application resources are pre-provisioned and consistently configured, eliminating manual deployment delays. Amazon Route53 DNS failover with health checks continuously monitors the primary region and automatically routes traffic to the secondary region when the primary fails, providing the automated failover needed to meet the 2-hour RTO.

Why this answer

It meets both the RTO of 2 hours and RPO of 1 hour. Cross-region automated backups for RDS provide an RPO of 1 hour or less by continuously backing up transaction logs to a secondary region. Deploying an identical infrastructure stack via CloudFormation StackSets ensures rapid provisioning in the secondary region, and Route53 DNS failover with health checks automates traffic redirection within the RTO window.

Exam trap

The trap here is that candidates often confuse a read replica with a Multi-AZ standby. While a cross-region read replica can be promoted to a primary instance, the process requires manual intervention and may take longer than the 2-hour RTO. Additionally, manual snapshots cannot meet a 1-hour RPO due to the time needed to take and copy snapshots across regions.

Option C uses automated cross-region backups and CloudFormation StackSets to meet both RTO and RPO automatically.

How to eliminate wrong answers

Option A is wrong because a read replica in a second region does not support failover to become a standalone writer; it is read-only and cannot be promoted in a disaster scenario, and CloudFront with an ALB origin does not provide regional failover. Option B is wrong because daily manual snapshots cannot achieve an RPO of 1 hour (snapshots are taken at most once per day), and manual deployment of CloudFormation stacks in a disaster exceeds the 2-hour RTO. Option D is wrong because AWS DMS continuous replication can meet RPO but the use of a Network Load Balancer in the secondary region (which does not support path-based routing or health checks for HTTP applications) and weighted routing policy (which is not designed for automatic failover) fails to meet the RTO requirement.

58
MCQmedium

A company runs a production RDS for PostgreSQL instance with Multi-AZ enabled. The database experiences a failover due to an AZ outage. After the failover, the application experiences high latency on write operations. What is the most likely cause?

A.The application is now reading from the standby instance, which has higher read latency.
B.Synchronous replication to the standby instance in the other AZ is causing additional latency.
C.The failover switched to a read replica in a different AZ.
D.The failover switched to asynchronous replication mode.
AnswerB

Multi-AZ deployments use synchronous replication between the primary and the standby in a different Availability Zone. Every write transaction must be committed on the primary and then acknowledged by the standby before the primary returns success to the application. This adds at least one full network round-trip across AZs per write, which increases commit latency and write response times compared to a single-AZ deployment.

Why this answer

With Multi-AZ enabled, RDS for PostgreSQL uses synchronous replication to the standby instance in a different Availability Zone. After a failover, the new primary continues to use synchronous replication to the new standby, which adds latency to write operations because each write must be acknowledged by the standby before the primary commits. This synchronous replication overhead is the most likely cause of the increased write latency.

Exam trap

The trap here is that candidates confuse Multi-AZ standby with read replicas, assuming the standby can serve reads or that failover switches to a read replica, when in fact Multi-AZ uses a passive standby that only handles failover and synchronous replication.

How to eliminate wrong answers

Option A is wrong because after a failover, the application reads from the new primary, not the standby; the standby is used only for replication and failover, not for read traffic. Option C is wrong because a read replica is a separate instance used for read scaling, not for failover; Multi-AZ failover promotes the standby, not a read replica. Option D is wrong because Multi-AZ always uses synchronous replication; failover does not change the replication mode to asynchronous.

59
MCQeasy

A company wants to back up its on-premises file server to AWS. The backup must be encrypted in transit and at rest. Which AWS service should the company use to meet these requirements?

A.AWS Storage Gateway (File Gateway) backed by Amazon S3.
B.Amazon EBS volumes attached to an EC2 instance acting as a file server.
C.AWS CloudFormation to replicate the file server configuration.
D.Amazon S3 with server-side encryption and a custom script to upload files.
AnswerA

AWS Storage Gateway (File Gateway) is a managed hybrid cloud service that presents an SMB/NFS file share to on-premises servers while storing the underlying data as Amazon S3 objects. It automatically handles encryption in transit and at rest, local caching for frequently accessed files, and asynchronous uploads with resumable transfer, making it a turnkey backup solution without custom scripts or infrastructure management.

Why this answer

AWS Storage Gateway File Gateway provides a native NFS/SMB interface that allows on-premises file servers to back up data directly to Amazon S3. It encrypts data in transit using TLS (HTTPS) and at rest using S3 server-side encryption (SSE-S3 or SSE-KMS), meeting both requirements without custom scripting or additional infrastructure.

Exam trap

The trap here is that candidates often choose Option D (S3 with custom script) because they focus only on encryption at rest and in transit, overlooking the requirement for a managed, integrated backup solution that eliminates the operational burden of writing and maintaining custom upload scripts.

How to eliminate wrong answers

Option B is wrong because Amazon EBS volumes attached to an EC2 instance acting as a file server require the company to manage the file server OS, backup scripts, and encryption configuration themselves, and do not provide a native on-premises backup integration. Option C is wrong because AWS CloudFormation is an infrastructure-as-code service for provisioning resources, not a backup service; it cannot handle file-level backup or encryption in transit/at rest. Option D is wrong because while Amazon S3 with server-side encryption meets at-rest encryption, using a custom script to upload files does not guarantee encryption in transit unless the script explicitly enforces HTTPS, and it lacks the seamless on-premises integration and lifecycle management that Storage Gateway provides.

60
MCQmedium

A company uses AWS Backup to back up its Amazon EFS file systems. The SysOps administrator needs to ensure that backups are retained for 7 years to meet compliance requirements. What should the administrator do?

A.Create a backup plan with a lifecycle policy that retains backups for 7 years.
B.Manually delete backups older than 7 years every month.
C.Increase the backup frequency to daily.
D.Configure cross-region backup to copy backups to another region.
AnswerA

AWS Backup lifecycle policies let you specify a retention period, such as 7 years (2555 days), after which recovery points are automatically expired and deleted. This automated, policy-driven approach ensures backups are retained for the required duration without manual intervention and provides an auditable record of compliance. Setting the retention period in the backup plan is the definitive way to satisfy a 7-year retention mandate.

Why this answer

AWS Backup allows you to define backup plans that include lifecycle policies to automatically transition backups to cold storage and expire them after a specified retention period. By setting the retention period to 7 years (2557 days) in the backup plan, AWS Backup will automatically manage the deletion of backups after that time, ensuring compliance without manual intervention.

Exam trap

The trap here is that candidates confuse backup frequency (how often backups are taken) with retention (how long backups are kept), leading them to select option C, or they mistakenly think cross-region backup (option D) automatically handles retention, when in fact both require a lifecycle policy to expire backups.

How to eliminate wrong answers

Option B is wrong because manually deleting backups is error-prone, not scalable, and violates the principle of automated compliance; AWS Backup provides automated lifecycle management to avoid human error. Option C is wrong because increasing backup frequency only affects how often backups are taken, not how long they are retained; retention is controlled by the lifecycle policy, not the schedule. Option D is wrong because cross-region backup copies data to another region for disaster recovery or geographic redundancy, but it does not control the retention period; the copied backups still need a lifecycle policy to expire after 7 years.

61
MCQeasy

A company stores critical data in an S3 bucket. To ensure data durability and availability, the company wants to automatically replicate objects to a bucket in a different AWS Region. Which S3 feature should be used?

A.Enable S3 Standard storage class on the bucket.
B.Use S3 One Zone-IA storage class.
C.Configure S3 Cross-Region Replication.
D.Enable S3 Versioning on the bucket.
AnswerC

S3 Cross-Region Replication (CRR) asynchronously copies every uploaded object to a destination bucket in a different AWS Region, providing the geographic redundancy required for disaster recovery. CRR requires versioning to be enabled on both the source and destination buckets, and it can replicate to a different storage class. This ensures critical data remains available even if the entire source Region becomes unavailable.

Why this answer

S3 Cross-Region Replication (CRR) is the correct feature because it automatically and asynchronously replicates objects from a source S3 bucket in one AWS Region to a destination bucket in a different AWS Region. This ensures data durability and availability by maintaining a copy in a separate geographic location, meeting the requirement for cross-region replication. CRR requires versioning to be enabled on both source and destination buckets.

Exam trap

The trap here is that candidates often confuse enabling S3 Versioning (which is a prerequisite for CRR but does not itself replicate data) with the actual replication feature, leading them to select Option D instead of the correct CRR option.

How to eliminate wrong answers

Option A is wrong because enabling S3 Standard storage class only defines the storage tier for durability and availability within a single region, it does not replicate data to a different AWS Region. Option B is wrong because S3 One Zone-IA stores data in a single Availability Zone, which does not provide cross-region replication and actually reduces availability compared to multi-AZ storage classes. Option D is wrong because enabling S3 Versioning alone only preserves multiple versions of objects within the same bucket and region, it does not replicate objects to a different AWS Region.

62
Multi-Selecthard

Which TWO steps should a SysOps administrator take to ensure data durability for an Amazon S3 bucket that stores critical documents? (Choose two.)

Select 2 answers
A.Enable default encryption with SSE-S3.
B.Enable S3 Versioning.
C.Use S3 Transfer Acceleration.
D.Enable MFA Delete.
E.Configure cross-region replication (CRR).
AnswersB, E

S3 Versioning is a correct step because it preserves every version of an object, including all overwrites and the original copy before a delete operation. When a DELETE is issued, versioning inserts a delete marker rather than physically removing the object, allowing straightforward rollback to a prior state. This directly mitigates permanent data loss caused by accidental user actions or application bugs, thus strengthening durability.

Why this answer

S3 Versioning (B) protects against accidental deletion and overwrites by preserving every version of an object, including deletions as delete markers. This ensures data durability by allowing recovery of previous versions. Cross-Region Replication (E) provides durability by asynchronously replicating objects to a different AWS region, protecting against region-wide failures.

Exam trap

The trap here is that candidates confuse data durability (protection against loss) with data security (encryption or access control), leading them to select SSE-S3 or MFA Delete instead of versioning and replication.

63
MCQmedium

A company has a production RDS for MySQL database. The SysOps administrator receives an alert that the database instance is running out of storage. The company requires high availability and minimal downtime during any modifications. What should the administrator do?

A.Add a read replica and use it for read traffic to reduce load on the primary.
B.Modify the RDS instance to increase the allocated storage. Since the instance is Multi-AZ, the modification will be applied with minimal downtime.
C.Create a CloudWatch alarm to notify when storage is low, then manually clean up old data.
D.Create a new RDS instance with larger storage and migrate the data using AWS Database Migration Service.
AnswerB

Modifying the allocated storage on an existing Multi-AZ RDS for MySQL instance is the correct approach because RDS supports in-place storage scaling without a full rebuild. For Multi-AZ deployments, Amazon performs the modification with a brief, automatic failover to the standby, resulting in typically less than a minute of downtime rather than hours-long migrations. The KEY keyword 'production' and 'minimal downtime' align with RDS's native ModifyDBInstance operation, which can also enable Storage Auto Scaling as a proactive measure, but the immediate fix is to increase the allocated storage to accommodate the data growth.

Why this answer

Modifying the allocated storage on a Multi-AZ RDS for MySQL instance can be done with minimal downtime. When you modify storage settings, Amazon RDS performs the update in the background, and for Multi-AZ deployments, the modification is applied to the standby first, then a failover occurs to minimize any interruption. This approach satisfies the high availability requirement and keeps downtime to a few seconds or less.

Exam trap

The trap here is that candidates assume any storage modification requires significant downtime, but for Multi-AZ RDS instances, the modification is applied to the standby first with a controlled failover, resulting in minimal disruption.

How to eliminate wrong answers

Option A is wrong because adding a read replica does not increase the available storage on the primary instance; it only offloads read traffic, leaving the storage exhaustion issue unresolved. Option C is wrong because creating a CloudWatch alarm only provides notification, and manually cleaning up old data is not a scalable or automated solution for production systems, nor does it guarantee minimal downtime or high availability. Option D is wrong because creating a new RDS instance and migrating data using AWS DMS introduces significant downtime and complexity, and does not leverage the existing Multi-AZ configuration for minimal disruption.

64
MCQeasy

A SysOps administrator is designing a disaster recovery plan for a web application. The application runs on EC2 instances in a single Availability Zone. What is the FIRST step to improve availability?

A.Deploy EC2 instances in at least two Availability Zones.
B.Use an Application Load Balancer to distribute traffic.
C.Enable Multi-AZ for the RDS database.
D.Create an Amazon CloudFront distribution for the application.
AnswerA

Distributing EC2 instances across at least two Availability Zones ensures that the application layer remains operational even if an entire AZ experiences an outage, because each AZ is an isolated failure domain with independent power, networking, and cooling. This architecture allows a load balancer or DNS failover to route traffic to healthy instances in the remaining AZs, making it the foundational step for a resilient disaster recovery design that protects compute resources.

Why this answer

The application currently runs on EC2 instances in a single Availability Zone, which creates a single point of failure. The first step to improve availability is to eliminate this zone-level failure by deploying EC2 instances in at least two Availability Zones, as this provides fault isolation against an AZ outage. This foundational change directly addresses the root cause of the availability risk before adding other components like load balancers or database replication.

Exam trap

The trap here is that candidates often jump to adding a load balancer or database redundancy first, overlooking that the most fundamental step to improve availability is to eliminate the single point of failure at the compute layer by distributing instances across multiple Availability Zones.

How to eliminate wrong answers

Option B is wrong because an Application Load Balancer distributes traffic but does not itself provide high availability if all backend instances remain in a single Availability Zone; the ALB cannot route traffic to healthy instances if that entire AZ fails. Option C is wrong because enabling Multi-AZ for RDS improves database availability but does not address the application tier's single-AZ EC2 instances, which are the primary bottleneck described in the question. Option D is wrong because creating a CloudFront distribution caches content at edge locations but does not resolve the underlying single-AZ failure risk for the origin EC2 instances; CloudFront does not provide active-active failover for compute resources.

65
MCQhard

A company runs a stateful application on a single Amazon EC2 instance with an attached EBS volume. The SysOps administrator needs to ensure that in the event of an instance failure, a new instance can be launched quickly with the same data. The Recovery Point Objective (RPO) is 15 minutes and the Recovery Time Objective (RTO) is 30 minutes. Which strategy should the administrator implement?

A.Configure an Amazon EC2 automatic recovery action using a CloudWatch alarm
B.Schedule EBS snapshots every 15 minutes and use a Lambda function to launch a new instance from the latest snapshot
C.Use an Auto Scaling group with a custom AMI that is updated every 15 minutes
D.Use an Application Load Balancer with health checks to redirect traffic to a standby instance
AnswerA

An EC2 automatic recovery action driven by a CloudWatch alarm monitors the system status check and, if a failure is detected, restarts the instance on new underlying hardware while preserving its instance ID, private IP, Elastic IP addresses, and all attached EBS volumes. Because the root device is EBS-backed and the attached volumes remain intact, there is no data loss, achieving a zero recovery point objective (RPO), and the restart typically completes in only a few minutes, well within a 30-minute RTO. This makes it the only option that inherently satisfies both recovery targets without requiring separate backup or replication infrastructure.

Why this answer

Amazon EC2 automatic recovery, triggered by a CloudWatch alarm based on status checks, can restart the instance on new hardware while preserving the attached EBS volume and its data. This meets the RPO of 15 minutes (data is current on the EBS volume) and the RTO of 30 minutes (recovery is typically within a few minutes). The stateful application remains intact because the same EBS volume is reattached to the replacement instance.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing snapshot-based or AMI-based recovery strategies, failing to recognize that EC2 automatic recovery directly addresses instance failure while preserving the existing EBS volume and its stateful data without any data loss or manual intervention.

How to eliminate wrong answers

Option B is wrong because scheduling EBS snapshots every 15 minutes and launching a new instance from the latest snapshot introduces significant latency: snapshot creation is not instantaneous, and restoring a volume from a snapshot can take several minutes, potentially exceeding the 30-minute RTO. Option C is wrong because using an Auto Scaling group with a custom AMI updated every 15 minutes does not preserve the stateful application's live data; AMIs capture the root volume at a point in time, but any data written between updates is lost, and the RPO cannot be guaranteed. Option D is wrong because an Application Load Balancer with health checks and a standby instance requires a second EC2 instance with its own EBS volume, which would not have the same data unless continuous replication is configured, and the question does not mention replication; this approach also fails to address the single-instance failure scenario without additional complexity.

66
Multi-Selecteasy

A SysOps administrator wants to back up an Amazon EBS volume that is attached to an EC2 instance running a production database. The backup must be crash-consistent and should not cause any downtime. Which TWO steps should the administrator take? (Choose two.)

Select 2 answers
A.Stop the EC2 instance before taking the snapshot.
B.Take a snapshot directly from the attached volume without any preparation.
C.Detach the volume from the instance before taking a snapshot.
D.Take a snapshot of the EBS volume after freezing.
E.Freeze the filesystem and flush I/O operations using a tool like fsfreeze.
AnswersD, E

Taking a snapshot immediately after freezing the filesystem ensures the on-disk state is a stable, point-in-time representation. The freeze command flushes all dirty buffers and suspends new write operations, so the snapshot includes every block in a coherent relationship with the filesystem journal. This produces a crash-consistent backup without any interruption to the instance or application, making it the recommended backup method.

Why this answer

Taking a snapshot after freezing the filesystem ensures that the snapshot captures a crash-consistent state of the EBS volume. Option E is correct because using a tool like fsfreeze flushes all pending I/O operations and freezes the filesystem, which prevents data inconsistencies without stopping the EC2 instance or detaching the volume.

Exam trap

The trap here is that candidates may think stopping the instance or detaching the volume is necessary for a crash-consistent backup, but AWS allows crash-consistent snapshots without downtime by freezing the filesystem and flushing I/O operations.

67
MCQeasy

A company wants to ensure that its EC2 instances receive patches automatically to maintain security compliance. Which AWS service can be used to automate patch management?

A.Amazon CloudWatch
B.AWS Systems Manager
C.AWS Config
D.AWS CloudTrail
AnswerB

AWS Systems Manager Patch Manager automates the process of patching managed EC2 instances and on-premises servers. It uses the Systems Manager Agent (SSM Agent) to discover missing patches, download them from configured patch baselines, and install them according to maintenance window schedules. Patch Manager supports both Linux and Windows, including security updates, bug fixes, and non-security patches, and can generate compliance reports. This is the native AWS service designed specifically for patch management.

Why this answer

AWS Systems Manager Patch Manager automates the process of patching managed EC2 instances and on-premises servers. It uses patch baselines to define approved patches and can schedule patching across maintenance windows, ensuring security compliance without manual intervention.

Exam trap

The trap here is that candidates often confuse AWS Config's compliance evaluation with actual remediation actions, but Config only detects drift and can trigger automation via Systems Manager Automation documents—it does not directly patch instances.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch is a monitoring and observability service for metrics, logs, and alarms; it does not have any capability to apply patches to EC2 instances. Option C is wrong because AWS Config is a service for evaluating resource configurations against desired policies and tracking compliance, but it cannot automate the installation of patches. Option D is wrong because AWS CloudTrail records API activity for auditing and governance; it does not perform any operational actions like patching.

68
MCQeasy

A web application runs on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. To achieve high availability, what is the minimum number of Availability Zones (AZs) that must be configured for the Auto Scaling group?

A.1
B.2
C.3
D.4
AnswerB

Placing instances across two Availability Zones in an Auto Scaling group ensures that if one AZ fails, the remaining AZ's instances continue serving traffic, and the group can automatically launch new instances in that healthy AZ to maintain the desired capacity. This configuration meets the classic high-availability requirement because AWS guarantees independent failure of AZs, so two AZs provide sufficient redundancy. For most production workloads, two AZs are the minimum architectural baseline for high availability.

Why this answer

For high availability, an Auto Scaling group must span at least two Availability Zones (AZs) to ensure that if one AZ fails, the application remains available from the other AZ. A single AZ would create a single point of failure, violating the high-availability requirement. The Application Load Balancer distributes traffic across healthy instances in all configured AZs, so two AZs are the minimum to achieve fault tolerance.

Exam trap

The trap here is that candidates often think a single AZ is sufficient if the Auto Scaling group can replace failed instances, but they overlook that the AZ itself is a failure domain, and high availability requires redundancy across at least two AZs.

How to eliminate wrong answers

Option A is wrong because configuring only one AZ creates a single point of failure; if that AZ becomes unavailable, the application will be completely inaccessible, which does not meet high-availability requirements. Option C is wrong because while three AZs provide even greater resilience, the question asks for the minimum number required for high availability, and two AZs satisfy that requirement. Option D is wrong because four AZs are excessive for the minimum requirement; high availability is achieved with two AZs, and additional AZs increase cost without being necessary for the basic goal.

69
MCQhard

A company runs a production application on EC2 instances in an Auto Scaling group. The application stores data on an EBS volume. The SysOps administrator wants to ensure that the data is durable and available even if an EC2 instance fails. Which approach should the administrator take?

A.Use an instance store volume and replicate data across instances.
B.Use an EBS volume with snapshots taken every hour.
C.Move the data to an S3 bucket and access it via S3 API.
D.Migrate the data to Amazon EFS and mount it to all instances.
AnswerD

Amazon EFS provides a fully managed, elastic NFS file system that can be mounted concurrently on multiple EC2 instances across multiple Availability Zones. It is durable and highly available, with data stored redundantly across AZs, and supports standard file system semantics such as locking and concurrent access. This makes it the right choice for a shared file system that all application instances can access simultaneously.

Why this answer

Amazon EFS provides a fully managed, scalable, and shared file system that can be mounted concurrently to multiple EC2 instances across Availability Zones. By migrating the data to EFS and mounting it to all instances in the Auto Scaling group, the data remains durable and available even if an individual EC2 instance fails, because the file system persists independently of any single instance's lifecycle.

Exam trap

The trap here is that candidates often confuse EBS snapshots (which are backups, not high-availability solutions) with a truly shared, durable file system, leading them to choose Option B despite its inability to provide automatic failover and continuous availability.

How to eliminate wrong answers

Option A is wrong because instance store volumes provide only ephemeral, block-level storage that is physically attached to the host; data is lost if the instance stops, terminates, or fails, and replication across instances would require custom, complex logic without built-in durability guarantees. Option B is wrong because while EBS snapshots provide point-in-time backups, they do not ensure continuous availability or automatic failover; if the EC2 instance fails, the EBS volume is still tied to that instance and cannot be immediately attached to another instance without manual intervention and potential downtime. Option C is wrong because moving data to S3 and accessing it via the S3 API would require significant application refactoring to replace file-system semantics with object storage operations, and S3 does not support standard file locking or POSIX permissions needed by many production applications.

70
MCQhard

A company runs a stateful web application on EC2 instances in an Auto Scaling group. The application uses a sticky session (session affinity) feature of the Application Load Balancer. During a scale-in event, some users lose their session data. What should the SysOps administrator do to prevent session data loss?

A.Disable sticky sessions and use a round-robin routing algorithm.
B.Store session state in an external data store such as Amazon ElastiCache.
C.Use a lifecycle hook to back up session data before termination.
D.Increase the Auto Scaling group's cooldown period to delay termination.
AnswerB

Storing session state in an external data store such as Amazon ElastiCache decouples the session data from the lifecycle of any single EC2 instance. When an instance is terminated or replaced, other instances can immediately retrieve the same session from the shared ElastiCache cluster, so users experience no interruption. This is the correct pattern because ElastiCache, especially with Redis, provides low latency reads and writes and can be configured with replication and persistence for high availability.

Why this answer

Sticky sessions (session affinity) tie a user's session to a specific EC2 instance. When a scale-in event terminates that instance, the session data stored locally on the instance is lost. Storing session state in an external data store like Amazon ElastiCache decouples session data from individual instances, allowing any healthy instance to serve the user's request without data loss, even after a scale-in event.

Exam trap

The trap here is that candidates may think lifecycle hooks (Option C) are a valid solution, but they only delay termination and do not prevent data loss for in-flight sessions, whereas the correct approach is to externalize session state entirely.

How to eliminate wrong answers

Option A is wrong because disabling sticky sessions and using round-robin routing does not solve the problem; it would cause users to be routed to different instances on each request, which would still lose locally stored session data and could break the application entirely. Option C is wrong because a lifecycle hook can delay termination and allow a script to back up session data, but this is a complex, race-condition-prone workaround that does not guarantee zero data loss and adds significant latency to scaling events. Option D is wrong because increasing the cooldown period only delays the termination of instances, it does not prevent session data loss when the instance is eventually terminated.

71
MCQmedium

A SysOps administrator notices that an RDS instance's storage is nearly full. The instance uses General Purpose SSD (gp2) storage. The administrator needs to increase storage with minimal downtime. Which action should be taken?

A.Modify the RDS instance to increase the allocated storage size
B.Enable storage auto-scaling
C.Delete old data to free up space
D.Convert the storage type to Provisioned IOPS
AnswerA

Use the AWS Management Console, CLI (modify-db-instance), or API to increase AllocatedStorage on the existing RDS DB instance. RDS performs the storage expansion online, so the database remains available during the modification, and this immediately gives InnoDB/MyISAM or other engine files additional space. This is the correct immediate remediation because it directly addresses the current storage-full condition.

Why this answer

Modifying the RDS instance to increase the allocated storage size is the correct action because RDS for MySQL, MariaDB, PostgreSQL, Oracle, and SQL Server supports dynamic storage scaling with minimal downtime. When you modify the allocated storage for a gp2 volume, RDS performs the modification in the background, and the instance remains available during the process, though you may experience a brief performance impact. This directly addresses the near-full storage condition without requiring a full outage.

Exam trap

The trap here is that candidates often confuse storage auto-scaling (which is a preventive measure) with the immediate need to increase storage, or they mistakenly believe that deleting data will instantly free up space on an RDS instance, ignoring the filesystem and volume-level allocation behavior.

How to eliminate wrong answers

Option B is wrong because enabling storage auto-scaling only prevents future storage exhaustion by automatically increasing storage when thresholds are met; it does not resolve the current near-full condition and may take time to trigger. Option C is wrong because deleting old data from an RDS instance does not immediately free up space on the underlying EBS volume due to the way filesystems handle deletion; the storage remains allocated and the volume may still report as full until a vacuum or reorg is performed, and this approach risks data loss without guaranteeing space recovery. Option D is wrong because converting the storage type to Provisioned IOPS (io1/io2) does not increase storage capacity; it only changes the performance characteristics and incurs additional cost without solving the space shortage.

72
MCQmedium

A company runs a web application on Amazon EC2 instances in an Auto Scaling group that spans two Availability Zones. The application uses an Application Load Balancer (ALB) that is deployed across the same Availability Zones. The SysOps administrator wants to ensure the application remains available if an entire Availability Zone fails. Which configuration is essential for this high availability?

A.Configure the Auto Scaling group with at least one instance in each Availability Zone.
B.Enable cross-zone load balancing on the Application Load Balancer.
C.Use an Amazon Route 53 health check to route traffic away from a failed AZ.
D.Attach an Elastic IP address to each instance in the Auto Scaling group to ensure IP persistence.
AnswerA

Configuring the Auto Scaling group to maintain at least one instance in each Availability Zone (AZ) ensures that if an entire AZ becomes unavailable, the remaining AZs still have healthy instances to serve traffic. Auto Scaling also performs AZ rebalancing, which automatically detects when one AZ has fewer instances and launches replacements in that AZ to maintain a balanced distribution. This is the fundamental mechanism for achieving fault tolerance at the AZ level within a single region, which is exactly what the requirement demands.

Why this answer

For high availability across an Availability Zone (AZ) failure, the Auto Scaling group must have at least one healthy instance in each AZ. This ensures that if one AZ becomes unavailable, the ALB can route traffic to instances in the remaining AZ. Without this minimum distribution, a single AZ failure could leave the application with zero healthy targets if all instances were in the failed AZ.

Exam trap

The trap here is that candidates often confuse cross-zone load balancing (which balances traffic) with instance distribution across AZs (which ensures survival), leading them to select Option B instead of recognizing that without instances in each AZ, no load balancing can save the application.

How to eliminate wrong answers

Option B is wrong because cross-zone load balancing distributes traffic evenly across all registered instances in all AZs, but it does not protect against an entire AZ failure—it only balances load, not ensures instance survival. Option C is wrong because Route 53 health checks can route traffic away from a failed AZ at the DNS level, but they do not guarantee that instances exist in the surviving AZ; the Auto Scaling group must already have instances there. Option D is wrong because Elastic IP addresses are not used with Auto Scaling groups (which use dynamic scaling and replacement) and do not provide high availability; they are static IPs for individual instances, not for AZ failure resilience.

73
MCQmedium

A company runs a stateful web application on a single EC2 instance. To improve reliability, the company wants to implement a highly available architecture. What should the SysOps administrator do?

A.Refactor the application to store session state externally (e.g., ElastiCache), then deploy it across multiple AZs with an Application Load Balancer.
B.Migrate the application to a larger instance type.
C.Create a standby EC2 instance and use an Elastic IP to fail over manually.
D.Use Route 53 health checks to route traffic to a secondary instance if the primary fails.
AnswerA

Externalizing session state to a service like ElastiCache decouples user sessions from individual EC2 instances, making the web tier stateless. With an Application Load Balancer distributing traffic across multiple instances in different Availability Zones, the application survives instance or even AZ failures without manual intervention. The ALB's health checks automatically route around unhealthy targets, and the stateless design allows you to scale out or replace instances without losing any in-flight user session data.

Why this answer

It addresses the core challenge of making a stateful web application highly available. By storing session state externally in ElastiCache, the application becomes stateless from a networking perspective, allowing any EC2 instance to handle any request. Deploying these instances across multiple Availability Zones (AZs) behind an Application Load Balancer (ALB) provides fault tolerance and automatic traffic distribution, eliminating the single point of failure.

Exam trap

The trap here is that candidates often assume that simply adding a second instance or using DNS failover (Route 53) is sufficient for high availability, overlooking the critical requirement to externalize session state for stateful applications.

How to eliminate wrong answers

Option B is wrong because scaling vertically to a larger instance type does not eliminate the single point of failure; if the instance or its AZ fails, the application still goes down. Option C is wrong because manual failover using an Elastic IP is not automated and introduces significant downtime; it also does not handle session state persistence, so users would lose their sessions during failover. Option D is wrong because Route 53 health checks alone cannot provide seamless failover for a stateful application; they operate at the DNS level with TTL delays (often 60 seconds or more), and without external session storage, user sessions would be lost when traffic shifts to a secondary instance.

74
MCQhard

A company runs a critical database on an EC2 instance with an EBS volume. The administrator wants to create a disaster recovery plan that can recover the database in a different AWS Region within 4 hours. The database size is 1 TB. What is the MOST efficient approach to meet the RTO?

A.Share the AMI with the target region.
B.Copy the AMI and underlying EBS snapshots to the DR region.
C.Use EBS snapshots directly in the DR region.
D.Configure AWS Backup to copy backups to the DR region.
AnswerB

Copying the AMI to the DR region automatically copies its underlying EBS snapshots, creating fully independent regional resources you can launch immediately. This approach yields a ready-to-use, region-specific AMI with the original launch permissions, block device mappings, and tags, which directly supports the RTO for the critical database. It is the standard, most efficient way to enable cross-region instance recovery.

Why this answer

Copying the AMI and its underlying EBS snapshots to the DR region creates a fully independent, bootable image that can be launched as an EC2 instance in the target region. This approach directly supports the 4-hour RTO by allowing the administrator to pre-stage the AMI copy or initiate the copy on-demand, and then launch the instance from the copied AMI without needing to recreate the volume from individual snapshots or reconfigure instance metadata.

Exam trap

The trap here is that candidates confuse 'sharing' an AMI (which only grants cross-account access, not cross-region availability) with 'copying' an AMI (which physically replicates the image to another region), leading them to select Option A despite it not enabling DR in a different region.

How to eliminate wrong answers

Option A is wrong because sharing an AMI with the target region only grants access permissions; it does not copy the AMI or its underlying snapshots to the target region, so the AMI remains in the source region and cannot be used to launch an instance in the DR region. Option C is wrong because EBS snapshots are region-specific and cannot be used directly in another region; they must be copied to the DR region first, and even after copying, launching an instance requires creating volumes and configuring the instance manually, which adds complexity and time. Option D is wrong because AWS Backup can copy backups to the DR region, but it introduces additional overhead (backup plan, vault, restore testing) and typically has a longer restore time compared to directly copying the AMI and snapshots, making it less efficient for a 4-hour RTO.

75
MCQeasy

A company uses Amazon S3 to store critical data. They need to protect against accidental deletion of objects. Which feature should the SysOps Administrator enable?

A.Create a lifecycle policy to transition objects to Glacier.
B.Configure cross-region replication.
C.Enable versioning on the bucket.
D.Enable MFA Delete on the bucket.
AnswerC

Enabling versioning is the correct choice because it keeps every version of an object, including the original, when a delete is issued. Instead of removing the object, S3 inserts a delete marker at the top of the version stack, and you can restore the object by deleting that marker. This directly provides a recovery mechanism for both accidental overwrites and deletions.

Why this answer

Enabling S3 Versioning preserves every version of an object, including overwrites and deletes. When versioning is enabled, a DELETE request does not remove the object permanently; instead, it adds a delete marker, allowing the object to be restored by removing the marker. This directly protects against accidental deletion by providing a recoverable history of all object changes.

Exam trap

The trap here is that candidates often confuse MFA Delete (which adds a security layer but does not inherently prevent deletion) with versioning (which directly enables recovery from deletion), leading them to select D instead of C.

How to eliminate wrong answers

Option A is wrong because lifecycle policies transition objects to different storage classes (like Glacier) for cost optimization, but they do not prevent deletion; in fact, lifecycle rules can expire objects, causing permanent deletion. Option B is wrong because cross-region replication (CRR) copies objects to another bucket for disaster recovery or compliance, but it does not protect against accidental deletion in the source bucket—deletions are replicated by default, and even with delete marker replication, the source object is still deleted. Option D is wrong because MFA Delete adds an extra authentication requirement for permanent deletions and version suspension, but it does not prevent accidental deletion on its own; it must be combined with versioning to be effective, and the question asks for a feature to protect against accidental deletion, not just to add a security control.

Page 1 of 3 · 205 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Reliability and Business Continuity questions.