Courseiva

CCNA Reliability and Business Continuity Questions

55 of 205 questions · Page 3/3 · Reliability and Business Continuity · Answers revealed

151
Multi-Selectmedium

A company runs a critical application on EC2 instances in an Auto Scaling group. The group uses a dynamic scaling policy based on CPU utilization. The SysOps administrator wants to ensure that the application remains available during a planned maintenance event that will take down one of the Availability Zones. Which TWO actions should the administrator take? (Choose two.)

Select 2 answers
A.Update the Auto Scaling group to remove the affected Availability Zone from the list of enabled AZs.
B.Manually terminate all instances in the affected Availability Zone.
C.Increase the desired capacity of the Auto Scaling group to account for the lost capacity.
D.Create a new launch configuration with a different AMI.
E.Disable the dynamic scaling policy to prevent scaling.
AnswersA, C

Removing the affected Availability Zone from the Auto Scaling group's enabled AZ list stops the ASG from launching any new instances in that unhealthy AZ. The ASG will then rebalance instances across the remaining healthy AZs, and because the group is regional, it can continue to maintain capacity. This directly addresses the root cause by eliminating the faulty location from the placement logic, allowing existing instances in healthy AZs to serve traffic without additional risk.

Why this answer

Removing the affected Availability Zone from the Auto Scaling group prevents the group from launching new instances in a zone that will become unavailable during maintenance. This ensures that any new instances are launched only in the remaining healthy Availability Zones, maintaining application availability. Option C is correct because increasing the desired capacity temporarily compensates for the instances that will be terminated or become unreachable in the affected zone, ensuring the group has enough running instances to handle the load.

Exam trap

The trap here is that candidates may think manually terminating instances (Option B) is a valid proactive step, but it actually causes immediate disruption and does not prevent the Auto Scaling group from launching replacements in the same failing zone unless the zone is first removed.

152
MCQmedium

A SysOps administrator is testing the failover of an Amazon RDS for PostgreSQL Multi-AZ DB instance. The application currently writes to the primary instance in us-east-1a. Which action will manually trigger a failover to the standby instance in us-east-1b?

A.Reboot the DB instance and select 'Reboot with failover'.
B.Modify the DB instance to Single-AZ and then back to Multi-AZ.
C.Reboot the DB instance without selecting any failover option.
D.Promote the standby instance using the Amazon RDS console.
AnswerA

Selecting 'Reboot with failover' performs a forced failover by rebooting the primary instance and automatically promoting the Multi-AZ standby to primary. Because the standby is kept synchronously replicated through the same Multi-AZ architecture, this operation validates the failover path without any data loss and only incurs a brief availability interruption during the promotion. This is the officially supported method to test failover behavior in an Amazon RDS Multi-AZ deployment.

Why this answer

The 'Reboot with failover' option in the Amazon RDS console explicitly triggers a failover by rebooting the primary DB instance and forcing the Multi-AZ configuration to promote the standby instance in us-east-1b to become the new primary. This is the designed method for manually testing or initiating a failover in a Multi-AZ deployment.

Exam trap

The trap here is that candidates confuse Amazon RDS Multi-AZ failover with Amazon Aurora's reader promotion, where you can explicitly promote a read replica to primary, leading them to incorrectly select Option D.

How to eliminate wrong answers

Option B is wrong because modifying the DB instance to Single-AZ and back to Multi-AZ would delete the standby instance and then create a new one, which is not a failover but a reconfiguration that causes downtime and does not test the existing standby. Option C is wrong because rebooting the DB instance without selecting 'Reboot with failover' will simply restart the primary instance without promoting the standby, so no failover occurs. Option D is wrong because Amazon RDS does not support manually promoting a standby instance via the console; the standby is not directly accessible and failover is controlled only through the primary instance's reboot with failover option or an automatic failure.

153
MCQhard

A company runs a critical MySQL database on an Amazon RDS DB instance in a single Availability Zone. The SysOps administrator needs to implement a disaster recovery solution with a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour, while minimizing costs. Which solution meets these requirements?

A.Enable Multi-AZ deployment with a synchronous standby replica in another Availability Zone
B.Create a cross-Region read replica and promote it to a standalone DB instance during a disaster
C.Enable cross-Region automated backups to another Region
D.Take daily automated snapshots and copy them to another Region manually
AnswerC

Enabling cross-Region automated backups continuously replicates both automated snapshots and transaction logs to a chosen destination Region without running any compute resources there. This service-managed feature provides a typical RPO of about 5 minutes because transaction logs are shipped frequently, and an RTO of under 1 hour by restoring the latest snapshot and rolling forward logs. Since you only pay for storage in the destination Region until you actually perform a restore, this is the most cost-effective and operationally simple way to meet the stated DR requirements.

Why this answer

Cross-Region automated backups replicate transaction logs to another AWS Region with a typical lag of a few minutes, enabling point-in-time recovery (PITR) that can meet an RPO of 5 minutes. When a disaster occurs, you can restore the automated backup to a new DB instance in the destination Region, and the RTO depends on the restore time, which can be under 1 hour for a properly sized instance. This solution minimizes costs by avoiding the continuous compute and storage overhead of a standby replica or read replica.

Exam trap

The trap here is that candidates often confuse cross-Region read replicas (asynchronous, higher RPO) with cross-Region automated backups (log-based, lower RPO), or assume Multi-AZ provides cross-Region disaster recovery when it only covers AZ failures within a single Region.

How to eliminate wrong answers

Option A is wrong because Multi-AZ with a synchronous standby replica only protects against an Availability Zone failure within the same Region, not a cross-Region disaster, and it incurs the cost of a full standby instance. Option B is wrong because a cross-Region read replica is asynchronous and can have replication lag exceeding 5 minutes, making it unable to guarantee an RPO of 5 minutes; additionally, promoting a read replica to a standalone instance can take longer than 1 hour due to the need to stop replication and apply pending changes. Option D is wrong because daily automated snapshots provide an RPO of up to 24 hours, far exceeding the required 5-minute RPO, and manual copying adds operational overhead and delay.

154
MCQeasy

A company runs a critical web application on a single Amazon EC2 instance with a 100 GiB gp2 EBS volume. The SysOps administrator needs to ensure data durability by taking automated snapshots of the root volume every hour. The snapshots should be retained for 7 days. Which AWS service can be used to automate this task with minimal configuration?

A.Amazon Data Lifecycle Manager (DLM)
B.AWS Backup
C.Amazon CloudWatch Events
D.AWS Systems Manager
AnswerA

Amazon Data Lifecycle Manager (DLM) is a fully managed service designed specifically for automating the creation, retention, and deletion of Amazon EBS snapshots. You create a lifecycle policy that targets EBS volumes by tag or ID, schedule snapshots at specified intervals (e.g., every 12 hours), and set retention rules based on count or age. DLM natively handles snapshot cleanup without any custom code, offers optional cross-region copying, and integrates with EC2 Fast Snapshot Restore. For a single-volume critical web application, DLM provides the most direct, minimal-configuration solution.

Why this answer

Amazon Data Lifecycle Manager (DLM) is the correct choice because it is specifically designed to automate the creation, retention, and deletion of EBS snapshots with minimal configuration. It supports custom schedules (e.g., every hour) and retention policies (e.g., 7 days) directly on EBS volumes, making it ideal for this task without requiring additional scripting or infrastructure.

Exam trap

The trap here is that candidates often choose AWS Backup because it is a centralized backup service, but they overlook that DLM is the simpler, purpose-built service for EBS snapshot lifecycle automation with minimal configuration, especially for a single volume and a straightforward retention policy.

How to eliminate wrong answers

Option B (AWS Backup) is wrong because while it can automate EBS snapshots, it requires setting up a backup plan and vault, which introduces unnecessary overhead for a simple hourly snapshot retention policy that DLM handles natively. Option C (Amazon CloudWatch Events) is wrong because it can trigger a Lambda function or Systems Manager Automation to create snapshots, but it does not natively manage snapshot retention or deletion, requiring custom code and additional configuration. Option D (AWS Systems Manager) is wrong because it is primarily for operational management (e.g., patching, inventory) and does not provide a built-in, automated snapshot lifecycle management feature; any snapshot automation would require custom Automation documents or scripts.

155
MCQmedium

A company runs a critical application on a single Amazon EC2 instance. The SysOps administrator needs to ensure that if the instance fails, a new instance is automatically provisioned in a different Availability Zone. Which configuration should the administrator implement?

A.Create an Auto Scaling group with the instance in multiple Availability Zones
B.Create a placement group and launch the instance in it
C.Place the instance behind an Elastic Load Balancer
D.Configure an Amazon Route 53 health check with failover routing
AnswerA

An Auto Scaling group with a desired capacity of 1 spanning multiple Availability Zones is the correct choice because it provides automatic self-healing: if the instance becomes unhealthy or fails, the ASG terminates it and launches a replacement in another AZ to return to desired capacity. Unlike the other options, this directly restores the EC2 instance itself, not just traffic routing. The ASG also monitors health via EC2 status checks and can integrate with an Elastic Load Balancer for additional health checks.

Why this answer

An Auto Scaling group (ASG) can be configured with a minimum, desired, and maximum size of 1, spanning multiple Availability Zones (AZs). When the EC2 instance fails, the ASG health check replacement mechanism automatically terminates the unhealthy instance and launches a new one in a different AZ, ensuring the application remains available across AZ boundaries without manual intervention.

Exam trap

The trap here is that candidates often confuse the health-check and failover capabilities of Route 53 or ELB with the automatic instance provisioning provided by Auto Scaling groups, mistakenly thinking DNS or load balancer health checks alone can replace a failed instance.

How to eliminate wrong answers

Option B is wrong because a placement group is designed to influence the physical placement of instances for low-latency or high-throughput networking (e.g., cluster, spread, or partition groups), but it does not provide automatic instance replacement or cross-AZ failover. Option C is wrong because an Elastic Load Balancer (ELB) distributes traffic across healthy instances but does not automatically provision a new instance when an existing one fails; it only routes traffic away from unhealthy targets. Option D is wrong because an Amazon Route 53 health check with failover routing can redirect DNS traffic to a different endpoint (e.g., a static website or another resource) but does not automatically launch a new EC2 instance; it requires a pre-provisioned secondary resource.

156
MCQhard

A company's S3 bucket contains critical data. The bucket policy accidentally allowed public write access, and a malicious actor uploaded several objects. The company needs to recover the bucket to a known good state as quickly as possible. What should the SysOps administrator do?

A.Enable S3 Object Lock on the bucket to prevent further modifications.
B.Enable MFA Delete on the bucket to secure delete operations.
C.Use S3 Versioning to restore the bucket to a previous version.
D.Configure S3 Cross-Region Replication to replicate data to another region.
AnswerC

S3 Versioning is the correct option because it preserves every version of an object, including all overwrites and deletes, as long as versioning was enabled before the incident. With versioning active, a simple DELETE creates a delete marker instead of removing the object, and an overwrite creates a new version while the old one remains accessible. To restore, you can delete the delete marker or copy the desired previous version back to the bucket, effectively rolling the data back to its earlier state.

Why this answer

Enabling S3 Versioning allows restoring a bucket to a previous state by deleting the current versions of objects, effectively rolling back to the last good version before the malicious uploads. Option A is incorrect because S3 Object Lock prevents deletion or overwrite for a specified retention period but does not provide version rollback. Option B is incorrect because MFA Delete requires additional authentication for delete operations but does not help restore previous versions.

Option D is incorrect because S3 Cross-Region Replication replicates objects to another region but does not provide versioning or rollback capabilities.

157
MCQmedium

Regulatory requirements mandate that all RDS and EBS backups are replicated to a secondary AWS region within 24 hours of creation. The company has workloads in us-east-1 and must replicate backups to eu-west-1. Restoring from the secondary region must be possible without manual copying steps during a disaster. What service and configuration implements this requirement?

A.Create an AWS Backup plan with a cross-Region copy rule that replicates recovery points to a backup vault in eu-west-1 within 24 hours
B.Schedule a Lambda function that calls CreateDBSnapshot and CopyDBSnapshot to replicate RDS snapshots, and CreateSnapshot and CopySnapshot for EBS volumes to eu-west-1
C.Enable RDS automated backups with cross-region replication and configure EBS snapshot copy separately using Data Lifecycle Manager
D.Use S3 Cross-Region Replication to replicate the backup bucket containing RDS and EBS snapshots to eu-west-1
AnswerA

AWS Backup's cross-Region copy rule runs automatically after each successful backup job. The copy is encrypted with the destination vault's KMS key. In a disaster, operators restore directly from the eu-west-1 vault — no manual cross-region data transfer is needed. A single backup plan can cover multiple resource types (RDS and EBS), satisfying the consolidated requirement.

Why this answer

AWS Backup is the correct service because it natively supports cross-Region copy rules that automatically replicate recovery points (including RDS snapshots and EBS snapshots) to a backup vault in a secondary Region within a specified time window. This meets the 24-hour replication requirement and enables direct restores from the secondary Region without manual copying, as the backup vault in eu-west-1 contains the replicated recovery points ready for use.

Exam trap

The trap here is that candidates often assume they need to use separate services (like Lambda or DLM) for each resource type, missing that AWS Backup provides a unified, managed solution that handles both RDS and EBS snapshots with cross-Region replication and direct restore capabilities.

How to eliminate wrong answers

Option B is wrong because while a Lambda function could technically replicate snapshots, it requires custom code, error handling, and scheduling, and does not provide the native, managed cross-Region restore capability without manual steps; it also lacks the built-in compliance tracking of AWS Backup. Option C is wrong because RDS automated backups with cross-Region replication only apply to RDS, not EBS volumes, and Data Lifecycle Manager (DLM) for EBS snapshots does not support cross-Region copy natively; DLM only copies within the same Region, so EBS snapshots would not be replicated to eu-west-1. Option D is wrong because S3 Cross-Region Replication replicates objects in an S3 bucket, but RDS and EBS snapshots are not stored as S3 objects by default; they are stored in AWS-managed snapshot storage, and even if you manually copy snapshots to S3, the replication would not create usable snapshots in the secondary Region for direct restore.

158
Multi-Selecteasy

A company wants to protect its data in Amazon S3 from accidental deletion. Which TWO methods should the SysOps administrator use? (Choose TWO.)

Select 2 answers
A.Set up cross-Region replication.
B.Enable S3 Transfer Acceleration.
C.Configure S3 event notifications.
D.Enable S3 Versioning on the bucket.
E.Enable MFA Delete on the bucket.
AnswersD, E

Enabling S3 Versioning makes the bucket store every object version, including all overwrites and deletes. When a delete request is made, S3 does not physically remove the object; instead, it inserts a null version delete marker, leaving all prior versions intact and recoverable at any time. This allows you to easily undo accidental deletions by removing the delete marker and restoring an older version, making it a fundamental protection mechanism.

Why this answer

Enabling S3 Versioning preserves all versions of an object, including overwrites and deletes. When versioning is enabled, a delete operation does not permanently remove the object; instead, it adds a delete marker, allowing the object to be restored by removing the marker. This directly protects against accidental deletion.

Exam trap

The trap here is that candidates often confuse replication (CRR) or notifications as protective measures, but they do not prevent deletion; only versioning and MFA Delete directly safeguard against accidental or malicious permanent data loss.

159
MCQhard

A company runs a stateful application on EC2 instances behind a Network Load Balancer (NLB). The application uses sticky sessions (session affinity) to maintain client state. During a deployment, the SysOps administrator needs to replace instances without disrupting active sessions. Which approach should be used?

A.Stop the NLB, replace instances, and restart the NLB
B.Deregister the old instances from the target group with connection draining enabled, then register new instances
C.Update the target group health check to remove old instances faster
D.Terminate the old instances immediately and launch new ones
AnswerB

Deregistering the old instances from the target group with connection draining (the deregistration delay attribute on the NLB target group) causes the load balancer to stop sending new connections to those instances while continuing to allow existing connections to complete up to the configured timeout, which defaults to 300 seconds. This preserves active user sessions for the stateful application, and only after the drain period expires are the instances fully removed. Registering new instances after the old ones have fully drained provides a clean handoff without abrupt termination; for even higher availability, you can register the new instances before starting the drain so fleet capacity is never reduced.

Why this answer

Deregistering instances from the target group with connection draining enabled allows existing connections to complete gracefully before the instances are removed. The Network Load Balancer (NLB) continues to route new connections to the remaining healthy instances, and once draining finishes, new instances can be registered without disrupting active sessions. This approach maintains session affinity (sticky sessions) by ensuring that in-flight requests are completed before the instance is taken out of service.

Exam trap

The trap here is that candidates may think connection draining is only for Application Load Balancers (ALBs) or that stopping the NLB is required for maintenance, but NLB supports connection draining via target group deregistration delay, which is the correct method for zero-downtime deployments.

How to eliminate wrong answers

Option A is wrong because stopping the NLB would terminate all active connections and disrupt all sessions, defeating the purpose of maintaining session affinity. Option C is wrong because updating the health check interval only affects how quickly unhealthy instances are detected; it does not provide a graceful shutdown mechanism for active sessions. Option D is wrong because terminating instances immediately would abruptly drop all active connections, causing session loss and application errors.

160
MCQmedium

An organization is using AWS CloudFormation to deploy infrastructure. The SysOps administrator needs to ensure that if a stack update fails, the stack automatically rolls back to the last known good state. Which stack update option should be configured?

A.Disable rollback
B.Change sets
C.Rollback on failure
D.Stack policy
AnswerC

The 'Rollback on failure' stack setting is the direct mechanism that causes CloudFormation to revert resources to the last successfully deployed state if an update operation fails. When enabled, CloudFormation automatically initiates a rollback by undoing any resource changes that occurred during the failed update, keeping the stack consistent with its previous known-good template. This setting is the only option listed that actively restores the prior stack state after an unsuccessful update.

Why this answer

CloudFormation's 'Rollback on failure' option, enabled by default, automatically reverts a stack to its last known good state if a stack update fails. This ensures that failed updates do not leave the infrastructure in an inconsistent or partially deployed state, maintaining reliability and business continuity.

Exam trap

The trap here is that candidates may confuse 'Rollback on failure' with 'Disable rollback' or think that change sets or stack policies handle rollback behavior, when in fact only the rollback configuration directly controls automatic recovery from failed updates.

How to eliminate wrong answers

Option A is wrong because 'Disable rollback' prevents automatic rollback on failure, leaving the stack in a failed state, which is the opposite of what the requirement asks. Option B is wrong because change sets allow you to preview how changes will affect a stack before execution, but they do not provide automatic rollback behavior on failure. Option D is wrong because a stack policy defines which stack resources can be updated during a stack update, not the rollback behavior on failure.

161
MCQeasy

A SysOps administrator deploys the above CloudFormation template. The stack creation fails with an error. What is the most likely reason?

A.The EBS volume must specify a SnapshotId.
B.The template uses a deprecated AWSTemplateFormatVersion.
C.The instance and volume are in different Availability Zones.
D.The VolumeAttachment resource is missing the Device property.
AnswerD

The AWS::EC2::VolumeAttachment resource requires the Device property, which specifies the logical device name (such as /dev/sdh) presented to the instance. Without Device, CloudFormation cannot complete the attachment resource and will fail to create it, even though the volume and instance are otherwise correctly configured. This is the actual defect in the template.

Why this answer

The VolumeAttachment resource in AWS CloudFormation requires the Device property to specify the device name (e.g., /dev/sdf) for the attached EBS volume. Without this property, CloudFormation cannot determine the mount point, causing the stack creation to fail. The error occurs because the template omits this required field.

Exam trap

The trap here is that candidates often assume the Device property is optional or that CloudFormation will auto-assign a device name, but AWS requires explicit specification for EBS volume attachments.

How to eliminate wrong answers

Option A is wrong because EBS volumes can be created without a SnapshotId; they can be empty volumes or created from snapshots, but a SnapshotId is not mandatory. Option B is wrong because AWSTemplateFormatVersion '2010-09-09' is the current and only valid version, not deprecated. Option C is wrong because CloudFormation automatically places the EC2 instance and EBS volume in the same Availability Zone when they are in the same template without explicit AZ specification, so this is not the cause of failure.

162
MCQmedium

A company runs a critical application on EC2 instances in an Auto Scaling group. The application processes messages from an Amazon SQS queue. The SysOps administrator notices that during periods of high load, the SQS queue depth increases significantly, and the application takes a long time to recover. The administrator wants to improve the application's ability to handle spikes in traffic without over-provisioning resources. The application is stateless and can scale horizontally. What should the administrator do?

A.Change the SQS queue from standard to FIFO to ensure messages are processed in order.
B.Configure an auto scaling policy for the Auto Scaling group based on the SQS queue depth (ApproximateNumberOfMessagesVisible).
C.Use a larger EC2 instance type with enhanced networking to process messages faster.
D.Increase the EC2 instance size to a larger type with more CPU and memory.
AnswerB

Configuring the Auto Scaling group to scale based on the SQS queue depth (ApproximateNumberOfMessagesVisible) is the correct approach because it directly matches worker capacity to the unprocessed backlog. An Amazon CloudWatch alarm or target tracking policy on this metric can add EC2 instances as messages accumulate and terminate instances as the queue drains. For best results, the policy should use a per-instance backlog metric (queue depth divided by desired capacity) to avoid over-scaling or flapping when the queue is briefly busy.

Why this answer

Scaling on the SQS queue depth (ApproximateNumberOfMessagesVisible) directly ties Auto Scaling capacity to the backlog, so the group adds instances when messages accumulate and removes them when the queue drains. This is the canonical target-tracking or step-scaling pattern for queue-driven, stateless workloads and avoids over-provisioning during normal load.

Exam trap

SOA-C02 often tests whether candidates default to vertical scaling (bigger instances) or queue-type changes when the correct answer is elastic, metric-driven horizontal scaling based on queue depth.

How to eliminate wrong answers

Option A is wrong because switching to a FIFO queue changes ordering semantics but does not improve throughput or scaling — FIFO queues actually have lower throughput limits (300 TPS without batching) and would worsen the spike problem. Option C is wrong because a larger instance with enhanced networking speeds up a single instance but doesn't add capacity dynamically, so the queue would still back up under high load. Option D is wrong because increasing instance size is vertical scaling, which is capped and doesn't respond elastically to traffic spikes.

163
MCQeasy

A company is designing a highly available web application using an Application Load Balancer (ALB) with EC2 instances in an Auto Scaling group across two Availability Zones. Which configuration ensures that the application remains available if one Availability Zone fails?

A.Use a single large EC2 instance instead of multiple instances
B.Disable health checks on the ALB to avoid false positives
C.Configure the Auto Scaling group to launch instances in at least two Availability Zones
D.Launch all EC2 instances in a single Availability Zone
AnswerC

Configuring the Auto Scaling group to launch instances in at least two Availability Zones is the correct approach for high availability because it isolates the application from a single-AZ failure. If one AZ becomes unavailable, the ASG automatically launches new instances in another AZ and the ALB routes traffic only to healthy instances in the remaining AZs. This distribution also supports even scaling and aligns with the AWS Well-Architected Framework's reliability pillar.

Why this answer

Configuring the Auto Scaling group to launch instances in at least two Availability Zones ensures that if one AZ fails, the remaining AZ continues to serve traffic. The ALB distributes incoming requests across healthy targets in all enabled AZs, and Auto Scaling replaces failed instances in the remaining AZs automatically, maintaining application availability.

Exam trap

The trap here is that candidates may think a single large instance or disabling health checks improves availability, but in reality, these actions introduce single points of failure or prevent automatic failure detection, which is exactly what the SOA-C02 exam tests for multi-AZ resilience.

How to eliminate wrong answers

Option A is wrong because using a single large EC2 instance creates a single point of failure; if that instance or its AZ fails, the application becomes unavailable, and it does not leverage the ALB's multi-AZ load balancing. Option B is wrong because disabling health checks on the ALB prevents it from detecting unhealthy targets, causing traffic to be routed to failed instances, which degrades availability and violates the goal of high availability. Option D is wrong because launching all EC2 instances in a single Availability Zone means that a failure of that AZ will take down all instances, making the application unavailable regardless of the ALB or Auto Scaling configuration.

164
MCQmedium

A company is running a critical application on EC2 instances in an Auto Scaling group. The application experiences occasional CPU spikes. The SysOps administrator needs to configure a scaling policy that reacts quickly to increased load but avoids unnecessary scaling actions due to short bursts. Which scaling policy type should be used?

A.Manual scaling
B.Scheduled scaling policy
C.Simple scaling policy
D.Target tracking scaling policy with a CPU utilization target of 70%
AnswerD

Target tracking scaling policy automatically creates and manages CloudWatch alarms and adjusts the desired capacity proportionally to the deviation from the 70% CPU utilization target. It continuously evaluates the metric, allowing it to react to sudden spikes by adding instances in a dynamic manner, while its built-in cooldown and scale-in safeguards prevent oscillation. This policy is ideal for the scenario because it directly addresses the goal of maintaining CPU utilization at a defined level without manual or predetermined intervention.

Why this answer

Target tracking scaling policy with a CPU utilization target of 70% is correct because it dynamically adjusts the Auto Scaling group's desired capacity to maintain a specified metric (CPU utilization) at the target value. It uses a built-in algorithm that reacts quickly to sustained load increases while smoothing out short bursts by applying a cooldown and proportional control logic, preventing unnecessary scaling actions from transient spikes.

Exam trap

The trap here is that candidates often choose simple scaling (Option C) thinking it reacts quickly, but they overlook that simple scaling lacks the smoothing and proportional control needed to avoid unnecessary actions from short bursts, which is the exact requirement in the question.

How to eliminate wrong answers

Option A is wrong because manual scaling requires human intervention to change capacity, which cannot react quickly to CPU spikes and defeats the purpose of automation. Option B is wrong because scheduled scaling is based on predictable time patterns, not real-time load, so it cannot respond to occasional, unpredictable CPU spikes. Option C is wrong because simple scaling policies have a fixed cooldown period and only scale based on a single alarm breach, which can lead to either over-reaction to short bursts or slow response to sustained load, lacking the proportional and smoothing logic of target tracking.

165
MCQeasy

A company uses AWS CloudTrail to log API activity. The SysOps administrator needs to ensure that log files are protected from accidental deletion and are available for compliance audits for at least 7 years. Which service should be used to meet these requirements?

A.Enable S3 Object Lock in Compliance mode.
B.Enable S3 Versioning on the CloudTrail S3 bucket.
C.Move CloudTrail logs to Amazon S3 Glacier after 90 days.
D.Store logs in Amazon CloudWatch Logs with an expiration policy of 7 years.
AnswerA

S3 Object Lock in Compliance mode enforces a write-once-read-many (WORM) model, preventing any user — including the AWS account root user — from deleting or overwriting the log objects for the specified retention period. Because CloudTrail log files become immutable once written, this directly satisfies the need to preserve audit activity for 7 years in an unmodifiable state.

Why this answer

S3 Object Lock in Compliance mode provides a write-once-read-many (WORM) model that prevents any user, including the AWS account root user, from deleting or overwriting objects for the specified retention period. This meets both the protection from accidental deletion and the 7-year compliance audit requirement, as the retention mode cannot be shortened or removed once applied.

Exam trap

The trap here is that candidates often confuse versioning (which only preserves copies) with immutability (which prevents deletion entirely), or they assume that moving data to a cold storage tier like Glacier automatically protects it from deletion, when in fact Glacier objects are still deletable without an additional lock mechanism.

How to eliminate wrong answers

Option B is wrong because S3 Versioning alone does not prevent deletion; it only preserves previous versions of objects, but a user with s3:DeleteObject permission can still delete the current version, and versioned delete markers can be removed. Option C is wrong because moving logs to S3 Glacier after 90 days does not inherently protect them from deletion; Glacier objects can still be deleted unless additional controls like Object Lock are applied. Option D is wrong because CloudWatch Logs expiration policies only control log retention and automatic deletion, but they do not provide a WORM lock to prevent accidental or malicious deletion before the expiration date.

166
MCQeasy

A SysOps administrator is setting up a backup plan for an RDS MySQL database. The database is 500 GB in size and is used for a critical application. The company requires a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour. Which solution meets these requirements?

A.Deploy the RDS instance in a Multi-AZ configuration with automatic failover.
B.Use AWS Backup to copy snapshots to a different AWS Region.
C.Take manual snapshots every 5 minutes and store them in Amazon S3.
D.Configure a cross-region read replica and promote it during failover.
AnswerA

Multi-AZ RDS uses synchronous replication to a standby instance in a different Availability Zone; every transaction is committed on both the primary and standby before being acknowledged, giving an RPO of zero. Should the primary fail, AWS automatically flips the DNS endpoint to the standby, typically completing failover in 60–120 seconds, comfortably meeting a 1-hour RTO. This makes it the only option that satisfies both the 5-minute RPO and the 1-hour RTO without manual intervention.

Why this answer

Multi-AZ RDS with automatic failover meets the RPO of 5 minutes and RTO of 1 hour because it synchronously replicates data to a standby instance in a different Availability Zone. In the event of a failure, Amazon RDS automatically fails over to the standby, typically completing within 60–120 seconds, which satisfies the RTO. The synchronous replication ensures zero data loss (RPO of effectively 0), well within the 5-minute requirement.

Exam trap

The trap here is that candidates confuse Multi-AZ (synchronous, zero data loss, automatic failover) with read replicas (asynchronous, potential data loss, manual promotion), and assume cross-region replicas can meet tight RPO/RTO when they cannot due to replication lag and promotion time.

How to eliminate wrong answers

Option B is wrong because AWS Backup cross-region snapshot copies are asynchronous and typically run on a schedule (e.g., hourly/daily), which cannot achieve a 5-minute RPO. Option C is wrong because manual snapshots cannot be taken every 5 minutes—Amazon RDS enforces a minimum interval of 5 minutes between manual snapshots, and the process itself takes time, making it impossible to meet a 5-minute RPO consistently. Option D is wrong because a cross-region read replica uses asynchronous replication, which can introduce lag exceeding 5 minutes, and promoting it requires manual intervention or automation that often takes longer than 1 hour to complete, failing the RTO.

167
MCQmedium

A company runs a web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The application stores session state in memory on each instance. The SysOps administrator wants to make the application highly available across multiple Availability Zones without losing session data when instances are terminated or replaced. The solution must minimize application changes. Which approach should the administrator take?

A.Use sticky sessions (session affinity) on the ALB and configure the Auto Scaling group with a larger min size.
B.Store session data in a shared Amazon ElastiCache cluster and modify the application to read/write session state to ElastiCache.
C.Deploy the application in multiple AWS Regions and use Amazon Route 53 with latency-based routing.
D.Store session data in an Amazon RDS for MySQL database and configure the application to read/write session state to the database.
AnswerB

ElastiCache provides a centralized, in-memory data store (such as Redis) that can be shared by all EC2 instances in the Auto Scaling group. By moving session state to ElastiCache, the application becomes stateless at the instance level, so any instance can serve any user request without losing session data. ElastiCache supports replication and automatic failover, making session data highly available across Availability Zones. This directly satisfies the HA requirement and is the best practice for a decoupled web tier.

Why this answer

Storing session state in a shared Amazon ElastiCache cluster decouples session data from individual EC2 instances, allowing any instance in the Auto Scaling group to serve any user request without losing session data when instances are terminated or replaced. This approach requires minimal application changes (only modifying the session handler to point to ElastiCache) and supports high availability across multiple Availability Zones by using a replicated ElastiCache cluster (e.g., Redis with replication).

Exam trap

The trap here is that candidates often choose sticky sessions (Option A) because they seem to solve session affinity without code changes, but they fail to realize that sticky sessions do not persist session data across instance terminations, which is the core requirement for high availability without data loss.

How to eliminate wrong answers

Option A is wrong because sticky sessions (session affinity) bind a user's session to a specific EC2 instance; if that instance is terminated or replaced, the session data stored in memory is lost, violating the requirement to not lose session data. Option C is wrong because deploying across multiple AWS Regions with Route 53 latency-based routing does not address session state persistence within a single region; it introduces cross-region latency and complexity without solving the fundamental issue of in-memory session loss on instance termination. Option D is wrong because while storing session data in Amazon RDS for MySQL would persist session state, it introduces significant overhead (e.g., database connection management, schema design, and slower read/write compared to in-memory caching) and requires more extensive application changes than using ElastiCache, which is purpose-built for session storage.

168
MCQeasy

A company wants to ensure that its Amazon RDS database can withstand the loss of an entire Availability Zone. Which feature should the SysOps administrator enable?

A.Enable automated backups with a retention period of 35 days.
B.Enable Multi-AZ deployment.
C.Take manual snapshots and copy them to another Region.
D.Create a read replica in a different Availability Zone.
AnswerB

Multi-AZ deployment maintains a synchronous standby replica in a separate Availability Zone, enabling automatic failover if one AZ is lost. This directly satisfies the requirement to survive an entire Availability Zone failure without manual intervention.

Why this answer

Multi-AZ deployment for Amazon RDS automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary AZ fails, Amazon RDS automatically fails over to the standby, ensuring database availability without manual intervention. This is the only option that directly protects against an entire AZ loss by maintaining a hot standby in a separate AZ.

Exam trap

The trap here is that candidates often confuse a read replica with a Multi-AZ standby, assuming that a read replica in a different AZ can be promoted for failover, but read replicas are asynchronous and require manual promotion, whereas Multi-AZ provides automatic synchronous failover.

How to eliminate wrong answers

Option A is wrong because automated backups with a retention period of 35 days only provide point-in-time recovery to a specific time, not automatic failover or high availability; they do not protect against AZ loss as they are stored within the same Region but not in a separate AZ for immediate failover. Option C is wrong because manual snapshots copied to another Region provide disaster recovery across Regions, not high availability within a Region; they require manual restoration and do not offer automatic failover if an AZ fails. Option D is wrong because a read replica in a different AZ is an asynchronous copy used for offloading read traffic, not for automatic failover; it does not provide synchronous replication or automatic promotion to primary in case of AZ failure, and promoting it requires manual intervention.

169
MCQmedium

A company is running a critical web application on EC2 instances behind an Application Load Balancer. The instances are in an Auto Scaling group across two Availability Zones. The company needs to ensure that if an entire Availability Zone fails, the application remains available. Which configuration meets this requirement?

A.Use larger EC2 instance types to handle the load during a failure.
B.Use Amazon CloudFront to distribute traffic across multiple regions.
C.Launch all instances in the same Availability Zone to reduce latency.
D.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
AnswerD

Configuring the Auto Scaling group to launch instances in at least two Availability Zones is the correct approach. The Auto Scaling group automatically distributes instances across the selected AZs, and if one AZ fails, the remaining instances continue to operate and serve traffic. Additionally, Auto Scaling can replace unhealthy instances in other AZs, and when combined with an Application Load Balancer, it ensures cross‑AZ load balancing and resilient failover.

Why this answer

Configuring the Auto Scaling group to launch instances in at least two Availability Zones ensures that if one AZ fails, the remaining AZ(s) continue to serve traffic. The Application Load Balancer automatically distributes incoming requests across healthy instances in all configured AZs, and the Auto Scaling group replaces failed instances in the remaining AZs, maintaining capacity. This design provides high availability without relying on a single AZ, which would be a single point of failure.

Exam trap

The trap here is that candidates often confuse scaling (increasing instance size or count) with high availability (distributing across AZs), leading them to choose Option A or C, or they mistakenly think CloudFront provides regional failover for dynamic web applications, which it does not by default.

How to eliminate wrong answers

Option A is wrong because using larger EC2 instance types only increases compute capacity within a single AZ; it does not protect against an entire AZ failure, as all instances would still be in the same AZ and become unavailable simultaneously. Option B is wrong because Amazon CloudFront is a content delivery network (CDN) that caches content at edge locations; it does not distribute traffic across multiple regions for active-active failover of a web application behind an ALB, and it does not replace the need for multi-AZ deployment. Option C is wrong because launching all instances in the same Availability Zone creates a single point of failure; if that AZ fails, all instances become unreachable, causing complete application downtime.

170
MCQmedium

A company hosts a critical web application on Amazon EC2 instances in a single AWS Region (us-east-1). The SysOps administrator needs to implement a Disaster Recovery (DR) solution using a different AWS Region (us-west-2). The DR plan requires a Recovery Time Objective (RTO) of 1 hour and a Recovery Point Objective (RPO) of 15 minutes. The application uses an Amazon Aurora MySQL DB cluster and static assets stored in an Amazon S3 bucket. Which combination of actions should the administrator take to meet these requirements?

A.Create an Aurora cross-Region read replica in us-west-2. Configure S3 Cross-Region Replication from the source bucket to a destination bucket in us-west-2. During DR, promote the read replica to a primary cluster and update DNS.
B.Take a manual snapshot of the Aurora DB cluster every 15 minutes and copy it to us-west-2. Use S3 batch operations to copy assets to us-west-2 daily.
C.Enable Aurora Multi-AZ in us-east-1 and configure S3 transfer acceleration to us-west-2.
D.Use AWS Database Migration Service (DMS) for continuous replication to a DB instance in us-west-2. Use S3 versioning to keep previous object versions.
AnswerA

A cross-Region read replica provides continuous replication for the database, achieving RPO seconds. Promoting it can be done in minutes, meeting RTO of 1 hour. S3 CRR replicates objects asynchronously, typically within minutes, satisfying the RPO.

Why this answer

Aurora cross-Region read replicas provide asynchronous replication with an RPO typically under 1 second, easily meeting the 15-minute RPO requirement. Promoting the read replica to a primary cluster in us-west-2 can be completed within minutes, satisfying the 1-hour RTO. S3 Cross-Region Replication (CRR) automatically replicates static assets to the destination bucket in us-west-2 with near-real-time latency, ensuring the S3 data is also current within the RPO window.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which provides high availability within a single region) with cross-region disaster recovery, or they assume manual snapshots and DMS are simpler alternatives without considering the RPO/RTO constraints and operational overhead.

How to eliminate wrong answers

Option B is wrong because taking manual snapshots every 15 minutes is operationally impractical and cannot guarantee an RPO of 15 minutes due to snapshot creation and copy latency; also, copying assets daily via S3 batch operations far exceeds the 15-minute RPO. Option C is wrong because Aurora Multi-AZ in us-east-1 only provides high availability within a single region, not cross-region disaster recovery, and S3 Transfer Acceleration only improves upload speed to a single bucket, not replication to another region. Option D is wrong because AWS DMS for continuous replication to a DB instance in us-west-2 introduces additional complexity and potential lag that may not meet the 15-minute RPO as reliably as Aurora native replication; S3 versioning alone does not replicate objects to another region, so it fails to provide cross-region DR for static assets.

171
MCQmedium

A company runs a production database on Amazon RDS for MySQL with Multi-AZ enabled. During a recent Availability Zone outage, the database experienced a failover. After the failover, the application team notices that the database endpoint in the connection string no longer works. What is the most likely cause?

A.The application is using the IP address of the database instance instead of the DNS endpoint.
B.The DB instance identifier changed after the failover.
C.The security group for the RDS instance was modified during the failover.
D.The DNS CNAME record for the RDS endpoint was manually changed.
AnswerA

When a Multi-AZ failover occurs, RDS promotes the standby database instance in a different Availability Zone, which has a different private IP address. The DNS CNAME for the RDS endpoint is automatically updated to point to the new instance's IP. If the application has hardcoded the old IP address rather than using the endpoint, it will not re-resolve DNS and will fail to connect to the new active instance. Always use the RDS DNS endpoint so that connections are directed to the current primary.

Why this answer

When Multi-AZ failover occurs, RDS updates the DNS CNAME record to point to the new primary instance in a different Availability Zone. If the application uses the IP address directly instead of the DNS endpoint, it will continue to resolve to the old (now failed) instance's IP, which is no longer accessible. The DNS endpoint is the only stable reference that automatically follows the failover.

Exam trap

The trap here is that candidates assume the endpoint itself changes or that the instance identifier is modified, when in fact only the underlying IP address changes and the DNS record is updated automatically.

How to eliminate wrong answers

Option B is wrong because the DB instance identifier remains unchanged after a failover; only the underlying compute instance changes. Option C is wrong because security groups are not modified during a failover; they are attached to the RDS instance and persist unchanged. Option D is wrong because the DNS CNAME record is managed automatically by AWS RDS and is not manually changed by users; a manual change would not occur during an automated failover.

172
MCQmedium

A company uses an RDS for MySQL Multi-AZ DB instance. They want to minimize downtime during a planned maintenance update that requires a database engine version upgrade. What should the SysOps administrator do?

A.Use the AWS Console to apply the maintenance update immediately; the Multi-AZ configuration will minimize downtime.
B.Modify the DB instance to be a Single-AZ deployment, then apply the upgrade.
C.Take a snapshot, restore it as a new instance, and upgrade that instance.
D.Create a read replica in the same region, upgrade the replica, and promote it.
AnswerA

Applying the maintenance window update immediately via the AWS Console is correct for a Multi-AZ RDS MySQL instance because RDS orchestrates the engine patch as a rolling operation: it first applies the update to the standby replica, then triggers a DNS failover to promote that updated standby to primary, and finally updates the old primary as the new standby. The failover itself typically completes within 60–120 seconds, and because the standby is already patched, the primary's availability impact is limited to a brief connection drop during the automatic failover. This leverages the Multi-AZ architecture exactly as designed—minimizing downtime without manual intervention or data loss—while keeping the instance fully managed by RDS.

Why this answer

RDS for MySQL Multi-AZ deployments perform engine version upgrades with automatic failover, which minimizes downtime by updating the standby instance first, then promoting it to primary. The Multi-AZ configuration ensures that the upgrade process is applied with reduced impact, typically causing only a brief interruption during the failover rather than a full outage.

Exam trap

The trap here is that candidates may overthink the solution and choose complex workarounds like read replica promotion, not realizing that RDS Multi-AZ already provides a built-in mechanism to minimize downtime during engine version upgrades, making the simplest option (applying the update immediately) the correct one.

How to eliminate wrong answers

Option B is wrong because modifying a Multi-AZ instance to Single-AZ before an upgrade eliminates the failover benefit, increasing downtime during the upgrade and requiring a manual failover or longer maintenance window. Option C is wrong because taking a snapshot, restoring it as a new instance, and upgrading that instance does not minimize downtime for the original production instance; it creates a separate upgraded instance that requires DNS changes and data synchronization, leading to potential data loss and extended downtime. Option D is wrong because creating a read replica, upgrading it, and promoting it is a valid approach for major version upgrades with minimal downtime, but it is more complex and not the recommended method for a planned maintenance update on a Multi-AZ instance; the Multi-AZ built-in upgrade process is simpler and achieves the same goal with less operational overhead.

173
MCQeasy

A company processes orders using an Amazon SQS standard queue. The order processing application occasionally fails to process a message. The SysOps administrator wants to ensure that any message that fails to be successfully processed after three attempts is automatically moved to a separate queue for manual review. Which SQS feature should be configured?

A.Configure a Dead Letter Queue (DLQ) with a redrive policy
B.Increase the visibility timeout of the queue
C.Convert the queue to a FIFO queue
D.Enable redrive allow policy on the queue
AnswerA

A Dead Letter Queue (DLQ) with a redrive policy is the correct mechanism for handling messages that repeatedly fail to process successfully. When you set a redrive policy with a maximum receive count, SQS automatically moves a message to the configured DLQ after that many attempted receives (e.g., after the consumer has received it but failed to delete it). This isolates poison-pill messages from the main queue, allowing the main queue to continue processing healthy messages without retrying broken ones indefinitely, and enables later analysis or manual correction in the DLQ.

Why this answer

A Dead Letter Queue (DLQ) with a configured redrive policy is the correct SQS feature to automatically move messages that have failed processing after a specified number of attempts (in this case, three) to a separate queue for manual review. The redrive policy defines the source queue, the DLQ, and the maximum receive count (maxReceiveCount) threshold. When a message is received from the source queue more times than the maxReceiveCount, SQS automatically redirects it to the DLQ, isolating problematic messages without manual intervention.

Exam trap

The trap here is that candidates may confuse increasing the visibility timeout (which only delays reprocessing) with the automatic isolation provided by a Dead Letter Queue, or think that converting to FIFO or enabling a redrive allow policy alone solves the problem, when the core requirement is a DLQ with a configured redrive policy that specifies the maxReceiveCount.

How to eliminate wrong answers

Option B is wrong because increasing the visibility timeout only prevents other consumers from processing a message for a longer period, but does not move failed messages to a separate queue; it merely delays reprocessing. Option C is wrong because converting the queue to a FIFO queue changes the ordering and exactly-once processing guarantees, but does not provide automatic redirection of failed messages to a separate queue; FIFO queues also support DLQs, but the conversion itself is not the solution. Option D is wrong because enabling a redrive allow policy on the queue is not a standard SQS feature; the correct term is 'redrive policy' (or 'redrive permission policy') which is used to allow a source queue to use a specific DLQ, but the question asks for the feature that moves failed messages, which is the DLQ with a redrive policy, not just the allow policy.

174
MCQmedium

Refer to the exhibit. A SysOps administrator ran the commands shown. What is the state of the EC2 instance?

A.The instance is running and healthy.
B.The instance is terminated.
C.The instance is running but has a system impairment.
D.The instance is stopped.
AnswerC

The instance is running but has a system impairment: The EC2 instance state is 'running', meaning it has been successfully launched and is operational at the hypervisor level. However, the system status check reports 'impaired', which indicates a problem with the physical host or the AWS infrastructure that supports the instance, such as loss of network connectivity, power loss, or hardware failure. The instance may still be responsive at the guest OS level, but the underlying system is degraded, so this option correctly describes both the running state and the system impairment.

Why this answer

The `aws ec2 describe-instance-status` command with the `--instance-id` flag returns the instance's system status checks. The output shows `SystemStatus: impaired`, which indicates that the instance is running but has a system impairment (e.g., loss of network connectivity or power). This status is determined by AWS's automated status checks that run every minute, and an impaired status means the underlying host has an issue that affects the instance.

Exam trap

The trap here is that candidates confuse `SystemStatus: impaired` with instance state (e.g., stopped or terminated), when in fact the instance is still running but the underlying host has a problem that affects its reliability.

How to eliminate wrong answers

Option A is wrong because the output explicitly shows `SystemStatus: impaired`, not `ok`, so the instance is not healthy. Option B is wrong because a terminated instance would not return any instance status output; the command would return an error or an empty set. Option D is wrong because a stopped instance would show `InstanceState: stopped` and `SystemStatus: not-applicable` or no status checks, not `impaired`.

175
Drag & Dropmedium

Drag and drop the steps to set up an Amazon S3 bucket policy to grant cross-account access into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Identify the bucket and account, write the policy with correct principal and actions, save, and test.

176
MCQmedium

A SysOps administrator is designing a disaster recovery plan for a critical application hosted on AWS. The application runs on EC2 instances with data stored in an RDS MySQL database. The RPO must be less than 15 minutes, and the RTO must be less than 1 hour. Which solution meets these requirements?

A.Take daily snapshots of the RDS instance and copy them to another Region.
B.Use an RDS DB instance with a Multi-AZ standby and a cross-Region Read Replica.
C.Use an RDS Multi-AZ DB Cluster deployment.
D.Use a single-AZ RDS instance with automated backups and point-in-time recovery.
AnswerC

An RDS Multi-AZ DB Cluster deploys a writer and two readable standby nodes in three separate Availability Zones, with synchronous replication from the writer to both standbys before each transaction is committed. On failure, the cluster automatically promotes one of the standbys, typically completing failover in under a minute, which meets the RTO of less than one hour. Because every commit is synchronously duplicated, no committed transactions are lost, so the RPO is effectively zero and comfortably satisfies the 15-minute requirement.

Why this answer

An RDS Multi-AZ DB Cluster deployment provides synchronous replication across three Availability Zones and automatic failover, enabling an RTO of typically under 1 minute and an RPO of effectively zero (no data loss). This meets the strict RPO (<15 minutes) and RTO (<1 hour) requirements without relying on cross-Region replication or manual recovery steps.

Exam trap

The trap here is that candidates often confuse Multi-AZ standby (which provides high availability but not cross-Region DR) with Multi-AZ DB Cluster (which offers synchronous replication and automatic failover), or they mistakenly believe that a cross-Region Read Replica can meet strict RPO/RTO requirements due to its asynchronous nature and lack of automatic failover.

How to eliminate wrong answers

Option A is wrong because daily snapshots have an RPO of up to 24 hours, far exceeding the required <15 minutes, and copying to another Region adds latency without improving recovery speed. Option B is wrong because a cross-Region Read Replica is asynchronous (RPO can be minutes to hours depending on replication lag) and does not provide automatic failover for the primary DB instance, so RTO could exceed 1 hour. Option D is wrong because a single-AZ instance with automated backups and point-in-time recovery has an RPO of up to 5 minutes (backup window) but RTO can be much longer than 1 hour due to the need to restore from a snapshot and replay transaction logs.

177
MCQmedium

A company is using AWS CloudFormation to manage its infrastructure. The SysOps Administrator needs to ensure that updates to a critical stack do not accidentally replace the database. Which feature should be used?

A.Use drift detection to identify changes.
B.Enable termination protection on the stack.
C.Use a change set to review the updates before executing them.
D.Define a stack policy that denies updates to the database resource.
AnswerD

A CloudFormation stack policy is a JSON policy attached to the stack that explicitly denies update, delete, or replacement actions on specified resources. For example, a Deny on the database logical resource with an action such as Update:Replace prevents CloudFormation from physically replacing the resource during an update, causing the update to fail rather than destroying the database. This is the only control listed that proactively blocks the destructive update before it can happen.

Why this answer

A stack policy in AWS CloudFormation explicitly denies updates to specified resources, such as the database, preventing accidental replacement or deletion during stack updates. Unlike termination protection, which only prevents stack deletion, a stack policy controls update actions on individual resources within the stack.

Exam trap

The trap here is that candidates often confuse termination protection (which only prevents stack deletion) with the ability to protect individual resources from replacement during updates, leading them to incorrectly select option B.

How to eliminate wrong answers

Option A is wrong because drift detection identifies differences between the stack's actual state and its template, but it does not prevent updates or replacements from occurring. Option B is wrong because termination protection prevents the entire stack from being deleted, not individual resources from being updated or replaced during a stack update. Option C is wrong because a change set allows you to review the proposed changes before execution, but it does not enforce any protection; the administrator could still execute the change set and replace the database.

178
MCQhard

A company runs a critical application on Amazon EC2 instances. The application uses an NFS file system stored on an Amazon EFS file system. The SysOps administrator must ensure that the file system is highly available and can withstand an Availability Zone failure. The file system must be accessible from all Availability Zones in the region. Which configuration is required to meet these requirements?

A.Configure the EFS file system for One Zone storage class and mount it using the file system ID.
B.Configure the EFS file system for Standard storage class and mount it using the regional DNS name.
C.Configure EFS to use provisioned throughput and mount it using a mount target IP address.
D.Enable EFS lifecycle management to move files to Infrequent Access storage class.
AnswerB

The Standard storage class automatically replicates file data across multiple Availability Zones, ensuring durability even if one AZ goes down. The regional DNS name (fs-xxxx.region.efs.amazonaws.com) resolves to mount targets in every AZ where the file system is configured, allowing clients to fail over to an available mount target seamlessly. This combination provides the high availability needed for a critical application.

Why this answer

The EFS Standard storage class replicates data across multiple Availability Zones (AZs) within a region, providing high availability and resilience against an AZ failure. Mounting the file system using the regional DNS name ensures that clients in any AZ can reach the file system via the nearest mount target, as the regional DNS name resolves to the mount target IP addresses in the local AZ. This configuration meets the requirement for the file system to be accessible from all AZs in the region while withstanding an AZ outage.

Exam trap

The trap here is that candidates often confuse storage class (One Zone vs. Standard) with performance settings (provisioned throughput) or cost-saving features (lifecycle management), and overlook that the regional DNS name is essential for multi-AZ access and failover.

How to eliminate wrong answers

Option A is wrong because the One Zone storage class stores data only within a single Availability Zone, which does not provide high availability or withstand an AZ failure. Option C is wrong because provisioned throughput is a performance setting, not a high-availability or multi-AZ configuration; mounting via a mount target IP address would pin the client to a specific AZ, failing the requirement for accessibility from all AZs. Option D is wrong because lifecycle management moves files to the Infrequent Access (IA) storage class to reduce costs, but it does not affect the availability or multi-AZ resilience of the file system.

179
MCQmedium

An Auto Scaling group launches new EC2 instances when CPU exceeds 70 percent. The instances take 4 minutes to bootstrap (install software, register with a service discovery system, and warm up caches). Without a hook, the load balancer routes traffic to new instances before they are ready, causing 503 errors. What is the correct solution?

A.Add a lifecycle hook on the autoscaling:EC2_INSTANCE_LAUNCHING transition; signal CompleteLifecycleAction(CONTINUE) when bootstrap finishes
B.Increase the load balancer health check grace period to 10 minutes to give instances time to bootstrap
C.Increase the warm-up time in the Auto Scaling group's instance refresh configuration
D.Use a weighted target group with 0 weight for new instances until they are confirmed healthy
AnswerA

The hook holds the instance in Pending:Wait, outside the target group, until the signal arrives. The load balancer never routes traffic to the instance during its Pending:Wait phase. After the CONTINUE signal, the instance enters InService and the load balancer registers it normally. The heartbeat timeout (default 1 hour, configurable) should exceed the bootstrap time.

Why this answer

Lifecycle hooks allow the Auto Scaling group to pause instance launch until a custom action (e.g., bootstrap completion) is finished. By adding a hook on the autoscaling:EC2_INSTANCE_LAUNCHING transition, the instance is held in a 'pending:wait' state. Once the bootstrap script calls CompleteLifecycleAction with the CONTINUE result, the instance transitions to 'InService' and can then be registered with the load balancer, preventing premature traffic and 503 errors.

Exam trap

The trap here is that candidates often confuse the health check grace period (which only delays health checks, not registration) with lifecycle hooks (which actually control when the instance becomes available to the load balancer).

How to eliminate wrong answers

Option B is wrong because increasing the load balancer health check grace period only delays when the load balancer starts checking health; it does not prevent the load balancer from routing traffic to the instance before it is ready. The instance is still added to the target group immediately, and the grace period only affects health check status, not registration. Option C is wrong because the warm-up time in an instance refresh configuration controls how long new instances are given to become healthy during a rolling update, not the initial launch or bootstrap process for a scaling event triggered by CPU.

Option D is wrong because weighted target groups distribute traffic based on weights; setting 0 weight for new instances would prevent all traffic, but the instances would still be registered and could receive traffic if the weight is later changed manually, and this approach does not automatically signal readiness after bootstrap.

180
MCQmedium

A company runs a production Amazon RDS for PostgreSQL DB instance. The SysOps administrator needs to ensure that in the event of a database failure, there is automatic failover to a standby instance in another Availability Zone with minimal downtime. Which deployment configuration should the administrator enable?

A.Multi-AZ deployment
B.Read Replicas
C.Automated backups with point-in-time recovery
D.Amazon RDS Proxy
AnswerA

Multi-AZ maintains a synchronous standby replica in a different Availability Zone and performs automatic failover via DNS endpoint redirection when the primary fails. This satisfies the requirement for automatic cross-AZ failover with minimal downtime, unlike read replicas, which require manual promotion.

Why this answer

A Multi-AZ deployment for Amazon RDS automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of a database failure or an Availability Zone outage, Amazon RDS automatically fails over to the standby, typically within 60-120 seconds, without requiring manual intervention. This ensures high availability and minimal downtime for the production PostgreSQL DB instance.

Exam trap

The trap here is that candidates often confuse Read Replicas with Multi-AZ deployments, mistakenly believing that Read Replicas provide automatic failover, when in fact they only support manual promotion and are intended for read scaling, not high availability.

How to eliminate wrong answers

Option B is wrong because Read Replicas are designed for read traffic offloading and do not provide automatic failover; they require manual promotion to become the primary instance, which incurs significant downtime. Option C is wrong because automated backups with point-in-time recovery are for data durability and recovery from corruption or accidental deletion, not for automatic failover to a standby instance with minimal downtime. Option D is wrong because Amazon RDS Proxy is a connection pooling and management service that improves application scalability and resilience to database failures, but it does not replace the need for a standby instance or provide automatic failover itself.

181
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer in a single Availability Zone. To improve reliability, what is the MOST effective design change?

A.Place the RDS database in a Multi-AZ deployment.
B.Add a second Application Load Balancer in the same Availability Zone.
C.Increase the EC2 instance size to handle more traffic.
D.Launch EC2 instances across two or more Availability Zones.
AnswerD

Launching EC2 instances across two or more Availability Zones eliminates the AZ as a single point of failure for the compute tier. When used with an Application Load Balancer (or other load balancing) and ideally an Auto Scaling group, traffic can be distributed to healthy instances in the remaining AZs even if one entire AZ fails. This is the core architectural pattern for building a highly available web application on AWS.

Why this answer

The most effective design change to improve reliability is to launch EC2 instances across two or more Availability Zones. This eliminates the single point of failure at the Availability Zone level, ensuring that if one AZ experiences an outage, the Application Load Balancer can route traffic to healthy instances in another AZ. This directly addresses the core reliability principle of fault isolation and high availability.

Exam trap

The trap here is that candidates often focus on scaling or database redundancy (options A and C) instead of recognizing that the fundamental reliability gap is the single Availability Zone, which requires distributing compute resources across multiple AZs to achieve true high availability.

How to eliminate wrong answers

Option A is wrong because placing the RDS database in a Multi-AZ deployment improves database availability but does not address the single-AZ failure risk for the web application tier; the EC2 instances and ALB remain vulnerable to an AZ outage. Option B is wrong because adding a second Application Load Balancer in the same Availability Zone does not eliminate the single point of failure at the AZ level; both ALBs would be unavailable if that AZ fails. Option C is wrong because increasing the EC2 instance size only improves capacity within the same AZ, not fault tolerance; a larger instance still fails if the AZ goes down.

182
MCQhard

A SysOps administrator is designing a disaster recovery plan for a critical application that runs on EC2 instances with data stored on EBS volumes. The application requires an RPO of 15 minutes and an RTO of 2 hours. The current solution uses EBS snapshots taken every 6 hours. The administrator needs to improve the backup strategy to meet the RPO. What is the most cost-effective way to achieve this?

A.Enable EBS Recycle Bin with a retention rule of 15 minutes.
B.Change the EBS volume type to io2 Block Express.
C.Increase the snapshot frequency to every 15 minutes.
D.Use EBS Multi-Attach to attach volumes to multiple instances for redundancy.
AnswerC

Increasing the snapshot frequency to every 15 minutes directly reduces the recovery point objective (RPO) to 15 minutes by creating a new snapshot at that interval. Even in the worst-case failure, you can restore from a snapshot that is at most 15 minutes old, so no more than 15 minutes of changes would be lost. This is the correct approach because it actively creates more frequent, consistent recovery points and is the standard technique to meet a tight RPO for EBS-backed EC2 instances.

Why this answer

Increasing the snapshot frequency to every 15 minutes directly reduces the recovery point objective (RPO) from 6 hours to 15 minutes, meeting the requirement without incurring additional infrastructure costs. EBS snapshots are incremental and cost-effective, as only changed blocks are stored after the initial snapshot, making frequent snapshots a practical approach for achieving a low RPO.

Exam trap

The trap here is that candidates may confuse the EBS Recycle Bin (which protects against accidental deletion) with a backup frequency solution, or think that changing volume types or using Multi-Attach can improve RPO, when neither addresses the need for more frequent recovery points.

How to eliminate wrong answers

Option A is wrong because the EBS Recycle Bin is designed to recover accidentally deleted snapshots or EBS volumes, not to create frequent recovery points; its retention rule controls how long deleted resources are kept, not the frequency of backups. Option B is wrong because changing the volume type to io2 Block Express improves performance and durability but does not affect snapshot frequency or RPO; it is a performance optimization, not a backup strategy. Option D is wrong because EBS Multi-Attach allows a single volume to be attached to multiple instances for shared storage, but it does not create backups or recovery points; it provides high availability within a single Availability Zone, not disaster recovery across regions or over time.

183
MCQmedium

A company runs a stateful web application on a single Amazon EC2 instance with an Elastic IP address. The SysOps administrator needs to increase availability so that if the instance fails, a new instance can be launched quickly with the same configuration and the same IP address. The administrator also needs to ensure data is not lost. Which solution meets these requirements with the least operational overhead?

A.Use an Application Load Balancer with an Auto Scaling group and a launch configuration that includes the Elastic IP
B.Create an AMI from the instance, store data on an Amazon EFS file system, and use an Auto Scaling group with a lifecycle hook to associate the Elastic IP
C.Create a CloudFormation template that launches a new instance and associates the Elastic IP
D.Place the instance in an Auto Scaling group with a minimum of 1 and a maximum of 1, and set the health check to replace unhealthy instances
AnswerB

The AMI provides a pre-configured launch template. EFS provides durable, shared storage for application data. The Auto Scaling group automatically launches a new instance if the current one fails, and the lifecycle hook script associates the Elastic IP to the new instance, ensuring continuity with the same IP.

Why this answer

It separates the stateful data (stored on Amazon EFS) from the compute instance, ensuring data persistence even if the instance fails. Creating an AMI from the instance captures the configuration, and an Auto Scaling group with a lifecycle hook can associate the Elastic IP to the new instance automatically, providing a quick failover with minimal operational overhead.

Exam trap

The trap here is that candidates often assume an Auto Scaling group alone can handle Elastic IP association, but without a lifecycle hook or custom script, the new instance will not automatically receive the Elastic IP, leading to IP address changes and potential downtime.

How to eliminate wrong answers

Option A is wrong because an Application Load Balancer (ALB) does not support Elastic IP addresses; ALBs use DNS names and are designed for distributing traffic, not for preserving a static IP for a stateful application. Option C is wrong because a CloudFormation template requires manual or automated invocation to launch a new instance and associate the Elastic IP, which introduces additional operational overhead and does not automatically handle instance failure detection and replacement. Option D is wrong because placing the instance in an Auto Scaling group with a minimum and maximum of 1 does not automatically launch a new instance with the same configuration or data; it only replaces the instance if it becomes unhealthy, but without a lifecycle hook to associate the Elastic IP or a mechanism to preserve stateful data, the solution fails to meet the requirements.

184
Multi-Selectmedium

A company stores critical application logs in an Amazon S3 bucket. The SysOps administrator must implement a backup strategy that protects against accidental deletion of objects and allows recovery of previous versions. The solution must be cost-effective and require minimal operational overhead. (Choose two.)

Select 2 answers
A.Enable S3 Cross-Region Replication to a bucket in another region.
B.Enable S3 Object Lock in governance mode with a retention period of 30 days.
C.Enable S3 Versioning on the bucket.
D.Configure an S3 Lifecycle policy to expire noncurrent versions after 90 days.
E.Configure an S3 Lifecycle policy to transition objects to S3 Glacier Deep Archive after 30 days.
AnswersC, D

S3 Versioning preserves every version of an object, so if an object is overwritten or deleted, previous versions remain accessible. This directly addresses accidental deletion and enables recovery of prior versions without needing to copy data elsewhere. It is a native feature that requires no additional infrastructure and incurs storage costs only for the retained versions, making it cost-effective for this requirement.

Why this answer

Enabling S3 Versioning preserves all object versions, allowing recovery from accidental deletion or overwrite. Pairing it with a lifecycle policy that expires noncurrent versions after a defined period keeps storage costs under control by automatically removing older versions that are no longer needed. Together, these native features provide a cost-effective, low-maintenance backup strategy for the log bucket.

Exam trap

The trap here is assuming that replication or Object Lock alone provides version recovery, when versioning is the core feature that preserves previous versions of objects.

185
MCQmedium

An application running on Amazon ECS with Fargate launch type is experiencing intermittent failures. The tasks are spread across multiple Availability Zones. The SysOps administrator notices that failures occur only when an entire AZ becomes unavailable. What should the administrator do to improve the reliability of the application?

A.Use multiple subnets in the same AZ.
B.Use a cluster placement group.
C.Attach an Amazon EFS filesystem to all tasks.
D.Increase the desired task count to ensure sufficient capacity across AZs.
AnswerD

Increasing the desired task count for an ECS service ensures that more Fargate tasks are running, and because the service scheduler distributes tasks across the AZs represented in your VPC subnets, this gives you capacity to absorb an AZ failure. With a higher replica count, if one AZ fails, the remaining tasks in other AZs can continue servicing traffic, and the service can eventually replace tasks in healthy AZs. This is the correct way to provide compute redundancy for a Fargate-based application.

Why this answer

Increasing the desired task count ensures that ECS Fargate tasks are distributed across multiple Availability Zones, providing sufficient capacity to absorb the loss of an entire AZ. When one AZ becomes unavailable, the remaining tasks in other AZs continue to serve traffic, improving application reliability. This approach leverages the multi-AZ architecture already in place by ensuring enough tasks are running to handle the load even after an AZ failure.

Exam trap

The trap here is that candidates may think adding storage (EFS) or using placement groups (which are EC2-specific) can solve AZ-level failures, but the core issue is ensuring enough task capacity across multiple AZs to survive the loss of one.

How to eliminate wrong answers

Option A is wrong because using multiple subnets in the same AZ does not provide AZ-level redundancy; all tasks would still be in a single AZ, so an entire AZ failure would take all tasks down. Option B is wrong because a cluster placement group is used for EC2 instances to achieve low-latency network performance by placing instances close together, but it is not supported with Fargate launch type and does not improve AZ-level fault tolerance. Option C is wrong because attaching an Amazon EFS filesystem provides shared persistent storage across tasks but does not protect against AZ failures; if all tasks are in a single AZ that fails, the EFS mount point becomes unreachable, and the application still fails.

186
Multi-Selecteasy

A company uses Amazon S3 to store backup data. The SysOps administrator needs to ensure that the data is encrypted at rest and that access is limited to only authorized users. Which TWO actions should be taken? (Choose TWO.)

Select 2 answers
A.Enable default encryption on the S3 bucket using SSE-S3 or AWS KMS.
B.Block all public access to the S3 bucket.
C.Create a bucket policy that allows only specific IAM roles or users.
D.Enable S3 Versioning on the bucket.
E.Enable S3 Transfer Acceleration.
AnswersA, C

Enabling default encryption on the S3 bucket with SSE-S3 or SSE-KMS ensures that every object written to the bucket is encrypted server-side at rest automatically, regardless of how it is uploaded. SSE-S3 uses AES-256 managed by Amazon, while SSE-KMS gives you customer-managed keys and separate audit permissions. This default setting satisfies compliance requirements for encrypted backups and prevents the accidental upload of plaintext objects.

Why this answer

Enabling default encryption on the S3 bucket using SSE-S3 or AWS KMS ensures that all objects stored in the bucket are encrypted at rest automatically, meeting the encryption-at-rest requirement. This can be configured via the bucket properties, and it applies to any object uploaded without explicit encryption headers.

Exam trap

The trap here is that candidates often confuse 'blocking public access' (a network-level control) with 'encryption at rest' (a data protection control), or think that versioning or transfer acceleration somehow addresses encryption or authorization requirements.

187
Multi-Selecthard

A SysOps administrator is troubleshooting a high error rate on an Application Load Balancer (ALB). The ALB is configured with two target groups: one for EC2 instances and one for Lambda functions. The administrator notices that the EC2 target group is unhealthy. Which THREE steps should the administrator take to resolve the issue?

Select 3 answers
A.Review the ALB's DNS resolution for the target instances.
B.Verify that the EC2 instances' security groups allow traffic from the ALB.
C.Increase the size of the Auto Scaling group to distribute load.
D.Check the health check settings on the target group for correct path and interval.
E.Inspect the application logs on the EC2 instances for errors.
AnswersB, D, E

When an ALB forwards traffic to EC2 instances, it connects from the security group attached to the ALB's elastic network interfaces. If the instance security group's inbound rules do not explicitly allow TCP traffic on the listener port and health check port from that ALB security group (or the VPC CIDR), the instance silently drops both health check probes and client connections. Because security groups are stateful, outbound responses are allowed automatically, but the inbound rule must exist; otherwise the ALB sees connection timeouts, marks the targets unhealthy, and routes only to remaining instances, skyrocketing error rates.

Why this answer

The ALB communicates with EC2 instances using the private IP addresses of the instances. If the EC2 instances' security groups do not explicitly allow inbound traffic from the ALB's security group (or the ALB's VPC CIDR), the health checks and actual traffic will be blocked, causing the target group to be marked unhealthy. This is a common misconfiguration when the ALB and instances are in the same VPC.

Exam trap

The trap here is that candidates often focus on scaling or DNS issues (Options A and C) instead of recognizing that the most common cause of an unhealthy target group is either a security group misconfiguration or an incorrect health check path, both of which are directly addressed by Options B and D.

188
MCQhard

A company runs a stateful web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. Users report that their sessions are frequently lost during scaling events. What is the MOST effective solution to maintain session persistence?

A.Increase the cooldown period for the Auto Scaling group.
B.Modify the application to store session data in an external data store such as ElastiCache or DynamoDB.
C.Enable sticky sessions (session affinity) on the Application Load Balancer.
D.Use larger EC2 instance types to reduce the frequency of scaling.
AnswerB

Storing session data in an external data store such as ElastiCache (a managed in-memory cache) or DynamoDB (a managed NoSQL database) decouples session state from the compute layer, making each EC2 instance effectively stateless. Any instance in the Auto Scaling group can then serve any request by reading and writing session data to the shared store, so if an instance is terminated, the session remains intact and is immediately accessible to a replacement instance. ElastiCache is ideal for extremely low-latency session reads and writes, while DynamoDB provides persistent, highly available storage; both allow the application to survive scaling events and instance failures without losing user context. This is the architecturally correct approach for a stateful web application running on a horizontally scaled fleet.

Why this answer

Externalizing session state to a durable, shared data store like ElastiCache or DynamoDB decouples session data from individual EC2 instances. This ensures that sessions survive scaling events, including instance termination during scale-in, because any instance can retrieve the session from the external store. Sticky sessions (option C) only route traffic to the same instance, but they do not preserve session data if that instance is terminated.

Options A and D do not address session persistence at all.

Exam trap

Candidates often choose sticky sessions because they associate session persistence with keeping a user on the same server. However, sticky sessions do not protect against session loss when the serving instance is terminated during scale-in. Externalizing session state is the more robust solution for scaling events.

How to eliminate wrong answers

Option A is wrong because increasing the cooldown period only delays the next scaling activity, it does not preserve session data across instances or prevent session loss when scaling does occur. Option B is wrong because while storing session data in an external data store like ElastiCache or DynamoDB is a valid approach for session persistence, the question asks for the 'MOST effective solution' given the context of an ALB and Auto Scaling, and the correct answer (sticky sessions) directly addresses session affinity without requiring application code changes. Option D is wrong because using larger EC2 instance types reduces the frequency of scaling but does not eliminate it, and sessions will still be lost when scaling events happen.

189
Multi-Selecthard

A company uses AWS CloudFormation to deploy infrastructure. The SysOps administrator needs to ensure that if a stack update fails, the stack is automatically rolled back to the last known good state. Which TWO steps should the administrator take? (Choose two.)

Select 2 answers
A.Create a manual snapshot of the database before each update.
B.Configure the stack to use a service role with permissions to perform rollback actions.
C.Use rollback triggers to monitor CloudWatch metrics and automatically roll back the stack if a metric breache.
D.Enable termination protection on the stack.
E.Define a stack policy that prevents updates to the database resources.
AnswersB, C

CloudFormation uses a service role to make API calls when creating, updating, or deleting stack resources. On a failed update, the service must be able to reverse changes—delete newly created resources, revert modifications, and in some cases re-create previous resources—so the role needs explicit permissions for those rollback actions. Without those permissions, rollback can stall, leaving the stack in UPDATE_ROLLBACK_FAILED.

Why this answer

A service role grants CloudFormation the necessary IAM permissions to perform rollback actions on resources, such as deleting or reverting changes, even if the user who initiated the update lacks those permissions. This ensures that the stack can automatically return to its last known good state without manual intervention. Option C is correct because rollback triggers allow monitoring CloudWatch metrics; if a metric breaches the specified threshold, CloudFormation automatically rolls back the stack, providing an additional safeguard.

Option E is incorrect because a stack policy only prevents updates to protected resources but does not enable or automate rollback on failure.

Exam trap

The trap here is that candidates often confuse termination protection (which only prevents stack deletion) with rollback behavior, or they assume manual snapshots are part of CloudFormation's automatic rollback process, when in fact CloudFormation relies on resource-level reversal and service roles.

190
MCQmedium

An RDS Multi-AZ DB instance fails over to the standby. The application uses the DB instance endpoint. What should the SysOps administrator usually do in the application after failover?

A.Ensure the application retries/reconnects using the same DB endpoint.
B.Manually change the application to the standby instance IP address.
C.Restore from the latest snapshot before reconnecting.
D.Create a new read replica and promote it immediately.
AnswerA

The RDS Multi-AZ architecture abstracts the active database behind a consistent DNS endpoint. When failover occurs, AWS automatically repoints that endpoint to the newly promoted standby, so the application must simply retry/reconnect to the same hostname rather than changing any configuration. Implementing connection retry logic with backoff is essential, because the failover process typically causes existing connections to be dropped for 60–120 seconds while DNS updates propagate.

Why this answer

When an RDS Multi-AZ DB instance fails over to the standby, the DNS record for the DB instance endpoint is automatically updated to point to the new primary instance. The application should simply retry or reconnect using the same endpoint; no manual changes are needed because the endpoint remains valid. This is the standard behavior for Multi-AZ deployments, ensuring minimal disruption.

Exam trap

The trap here is that candidates may think they need to manually update the connection string or IP address, but the DNS endpoint automatically resolves to the new primary after failover, so only retry logic is required.

How to eliminate wrong answers

Option B is wrong because the application should use the DNS endpoint, not the IP address; the IP address can change after failover, and relying on it would break connectivity. Option C is wrong because restoring from a snapshot is unnecessary and would cause data loss; Multi-AZ failover preserves data without requiring a restore. Option D is wrong because creating and promoting a read replica is not the correct recovery action for a Multi-AZ failover; the standby is already promoted automatically by RDS.

191
MCQmedium

A company runs a production Amazon RDS for PostgreSQL DB instance in a single Availability Zone (AZ). The SysOps administrator needs to improve database availability so that in the event of a database failure or AZ outage, a standby instance is automatically promoted with minimal downtime. Which configuration should the administrator enable?

A.Enable automated backups with a retention period of 35 days.
B.Create a read replica in another Availability Zone.
C.Enable Multi-AZ deployment on the DB instance.
D.Schedule manual snapshots to be taken every hour and restore from the latest snapshot when needed.
AnswerC

Enabling Multi-AZ on an Amazon RDS for PostgreSQL DB instance provisions a synchronous standby replica in a different Availability Zone and automatically maintains a synchronous physical replication stream. In the event of an infrastructure failure, an availability zone outage, or a database patching event, Amazon RDS automatically performs a failover to the standby, typically completing within 60–120 seconds and preserving your data because all commits are synchronous. The DNS endpoint remains unchanged, so application connections are transparently redirected without manual intervention. This configuration meets the requirement for automatic failover and high availability.

Why this answer

Multi-AZ deployment automatically creates and maintains a synchronous standby replica in a different Availability Zone. In the event of a failure or AZ outage, Amazon RDS automatically fails over to the standby, typically within 60–120 seconds, with no manual intervention required. This meets the requirement for automatic promotion with minimal downtime.

Exam trap

The trap here is that candidates confuse read replicas (which are for read scaling and require manual promotion) with Multi-AZ (which provides automatic failover), or they overestimate the speed and automation of backups and snapshots for disaster recovery.

How to eliminate wrong answers

Option A is wrong because automated backups only provide point-in-time recovery (PITR) to restore the database to a specific time, not automatic failover with minimal downtime; restoration is a manual process that can take hours. Option B is wrong because a read replica is designed for read scaling and asynchronous replication, not automatic failover; promoting a read replica requires manual intervention and can result in data loss due to replication lag. Option D is wrong because manual snapshots require scheduling and manual restoration, which involves significant downtime and does not provide automatic failover or minimal disruption.

192
Drag & Dropmedium

Drag and drop the steps to set up an AWS Site-to-Site VPN connection into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First create and attach the virtual private gateway, then define the customer gateway, then create the VPN connection, configure the on-premises router, and verify the tunnel.

193
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer (ALB) in a single Availability Zone. The application stores session data in an RDS MySQL DB instance. To improve reliability, the company wants to deploy the application across multiple Availability Zones. Which combination of actions should the company take to achieve this? (Choose the correct course of action.)

A.Deploy EC2 instances in a single Availability Zone behind an Application Load Balancer. Enable Multi-AZ for the RDS MySQL DB instance.
B.Deploy EC2 instances in two Availability Zones and place them behind a Network Load Balancer. Keep RDS MySQL as a single-AZ deployment.
C.Deploy EC2 instances in two Availability Zones. Configure an RDS MySQL Read Replica in a second Availability Zone and route read traffic to it.
D.Deploy EC2 instances in two Availability Zones and place them behind an Application Load Balancer. Enable Multi-AZ for the RDS MySQL DB instance.
AnswerD

This architecture provides high availability for both the application and database layers. The EC2 fleet spans two Availability Zones behind an Application Load Balancer, which performs health checks and distributes request traffic to healthy targets, so an instance or AZ failure is absorbed by the remaining targets. Enabling Multi-AZ on RDS MySQL automatically provisions a synchronous standby in a second AZ and performs automatic failover to that standby if the primary fails, thereby eliminating single points of failure across the stack.

Why this answer

Deploying EC2 instances across two Availability Zones behind an Application Load Balancer (ALB) provides high availability for the web tier, while enabling Multi-AZ for RDS MySQL ensures synchronous replication to a standby instance in a different AZ, providing automatic failover and data durability. This combination addresses the requirement to improve reliability by eliminating single points of failure at both the compute and database layers.

Exam trap

The trap here is that candidates may confuse RDS Read Replicas with Multi-AZ deployments, assuming a Read Replica provides high availability, when in fact it only supports read scaling and does not offer automatic failover for the primary database.

How to eliminate wrong answers

Option A is wrong because deploying EC2 instances in a single Availability Zone behind an ALB does not provide high availability for the web tier; a failure in that AZ would still cause an outage. Option B is wrong because a Network Load Balancer (NLB) operates at Layer 4 and does not support HTTP-based routing or session stickiness required for the web application, and keeping RDS MySQL as single-AZ does not provide database redundancy. Option C is wrong because an RDS MySQL Read Replica is asynchronous and cannot be used for automatic failover; it is designed for read scaling, not high availability, and does not provide a synchronous standby for the primary database.

194
MCQhard

A company runs a critical application on Amazon EC2 instances across multiple Availability Zones. The application stores state data on a shared Amazon EFS file system. The SysOps administrator needs to ensure that the file system remains available if an entire Availability Zone fails. The file system must also provide low-latency access from all instances. Which configuration meets these requirements?

A.Create an EFS file system with the One Zone storage class and mount it from all instances.
B.Create an EFS file system with the Standard storage class, enable replication to another Region, and use DNS failover.
C.Create an EFS file system with the Standard storage class in the same Region, and mount it from all instances using the regional mount target.
D.Create an EFS file system with the Standard storage class, and enable Multi-AZ deployment.
AnswerC

The Standard storage class automatically replicates file system data redundantly across multiple Availability Zones within the Region, providing built-in resilience against an AZ failure. The regional mount target is a single DNS name that resolves to mount targets in each AZ, so instances in any AZ can mount the same file system with low-latency access. If one AZ becomes unavailable, the DNS/ELF service continues to route instances to healthy mount targets, satisfying the high availability requirement.

Why this answer

The EFS Standard storage class stores data redundantly across multiple Availability Zones (AZs) within a Region, ensuring high availability and durability even if an entire AZ fails. By mounting the file system using the regional mount target (which resolves to the EFS file system's regional DNS name), instances in any AZ can access the file system with low latency, as EFS automatically routes traffic to the most appropriate mount target in the same AZ. This configuration meets both the availability and low-latency requirements without additional replication or failover complexity.

Exam trap

The trap here is that candidates confuse EFS's Standard storage class with RDS's Multi-AZ deployment feature, or incorrectly assume that cross-Region replication is necessary for AZ-level fault tolerance, when in fact EFS's regional storage class already provides Multi-AZ redundancy within a single Region.

How to eliminate wrong answers

Option A is wrong because the One Zone storage class stores data only within a single Availability Zone, so if that AZ fails, the file system becomes unavailable, violating the requirement for continued availability during an AZ failure. Option B is wrong because enabling cross-Region replication does not provide low-latency access from all instances within the same Region; it introduces additional latency for cross-Region data access and requires DNS failover, which is not designed for intra-Region AZ failures and adds unnecessary complexity. Option D is wrong because EFS does not support a 'Multi-AZ deployment' configuration; the term 'Multi-AZ' applies to Amazon RDS, not EFS, and EFS inherently provides Multi-AZ redundancy through the Standard storage class, not through a separate deployment option.

195
MCQeasy

A company has an Auto Scaling group that launches EC2 instances in private subnets. The instances need to download software patches from the internet. Which component must be added to the VPC to allow outbound internet traffic while keeping the instances private?

A.An internet gateway attached to the VPC
B.A VPC peering connection to a VPC with internet access
C.An egress-only internet gateway
D.A NAT gateway in a public subnet
AnswerD

A NAT gateway operates in a public subnet with an Elastic IP and performs source network address translation for outbound IPv4 traffic. Private subnet instances route their default 0.0.0.0/0 traffic to the NAT gateway, which then forwards it to the internet while masking the private instance IPs. This is the standard AWS-managed method to give private instances outbound internet access without exposing them to inbound connections.

Why this answer

A NAT gateway in a public subnet allows EC2 instances in private subnets to initiate outbound traffic to the internet (e.g., to download patches) while preventing unsolicited inbound connections. The NAT gateway translates the private IPs to its own Elastic IP and routes traffic through an internet gateway attached to the VPC, keeping instances private.

Exam trap

The trap here is that candidates often confuse an internet gateway with a NAT gateway, assuming attaching an internet gateway to the VPC alone will give private instances internet access, but private subnets need a NAT device to route traffic through the internet gateway.

How to eliminate wrong answers

Option A is wrong because an internet gateway alone does not enable outbound traffic from private subnets; it only provides a target for routes in public subnets, and instances in private subnets lack a default route to it. Option B is wrong because a VPC peering connection does not provide internet access; it only enables private IP communication between two VPCs, and the peered VPC would still need its own internet gateway and NAT to reach the internet. Option C is wrong because an egress-only internet gateway is designed for IPv6 traffic only, not for IPv4 traffic, and the question implies IPv4 patches.

196
Multi-Selectmedium

A company has a production application running on Amazon ECS with Fargate. The application must be highly available across multiple Availability Zones. Which TWO configurations should be implemented?

Select 2 answers
A.Configure the ECS service to run tasks in a single Availability Zone to reduce network latency.
B.Configure the ECS service to run tasks in at least two Availability Zones.
C.Use the awsvpc network mode for the task definition.
D.Place the ECS service behind an Application Load Balancer.
E.Use Fargate Spot capacity providers to reduce costs.
AnswersB, D

Running ECS tasks in at least two Availability Zones is the fundamental high-availability pattern, because it ensures that if one AZ fails, the application continues serving from the remaining AZs. This placement strategy allows ECS to distribute tasks and maintains capacity during an AZ outage, preventing total loss of service. It directly addresses the need for fault tolerance and is the most appropriate action for a production workload.

Why this answer

Running ECS tasks across at least two Availability Zones ensures that if one AZ fails, the service continues to operate in the other AZ, meeting the high-availability requirement. Option D is correct because placing the ECS service behind an Application Load Balancer (ALB) enables health checks and automatic traffic distribution to healthy tasks, which is essential for maintaining availability during task failures or AZ disruptions.

Exam trap

The trap here is that candidates often confuse network mode (awsvpc) with high-availability configuration, but awsvpc is about networking capabilities (e.g., per-task ENI) and does not inherently provide multi-AZ resilience.

197
Multi-Selecthard

A SysOps administrator is designing a disaster recovery strategy for a production RDS MySQL database. The database must be recoverable within 15 minutes with a Recovery Point Objective (RPO) of less than 5 seconds. Which TWO actions should the administrator take? (Choose two.)

Select 2 answers
A.Create a read replica in the same AWS Region.
B.Enable Multi-AZ deployment for the RDS instance.
C.Create a cross-Region read replica in another AWS Region.
D.Enable automated backups with a retention period of 35 days.
E.Take manual snapshots every hour.
AnswersB, C

Enabling Multi-AZ deployment for the RDS instance creates a synchronous standby in a different Availability Zone, guaranteeing zero data loss (RPO=0) because transactions are committed on both the primary and standby before a write is acknowledged. Automatic failover is triggered by Amazon RDS within 60–120 seconds, providing high availability within a Region and protecting against AZ failures, database instance failures, and storage failures. This is the most appropriate choice when the specified RPO is below 5 seconds and the disaster recovery scope is limited to Availability Zone outages, as it eliminates asynchronous replication lag entirely.

Why this answer

Multi-AZ deployment (Option B) provides automatic failover to a standby replica in a different Availability Zone, enabling recovery within minutes and meeting the 15-minute RTO. Cross-Region read replicas (Option C) allow asynchronous replication to another region with an RPO typically under 5 seconds, satisfying the RPO requirement. Together, they ensure both rapid failover and minimal data loss.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read replicas, thinking Multi-AZ alone provides cross-region disaster recovery, or they assume automated backups or snapshots can meet a sub-5-second RPO, which they cannot due to their periodic nature.

198
MCQhard

A company runs a production web application on AWS using Auto Scaling groups (ASGs) behind an Application Load Balancer (ALB). The application state is stored in an Amazon RDS for MySQL Multi-AZ DB instance. The application experiences periodic traffic spikes, and the current ASG uses a simple scaling policy based on average CPU utilization. Recently, during a spike, the application became unresponsive for several minutes. The CloudWatch metrics show that the CPU utilization on the RDS instance peaked at 80%, and the DB Connections metric reached the maximum allowed. The read replica lag increased to over 10 seconds during the spike. The web servers are stateless and scale out quickly. The operations team needs to improve the reliability and performance of the application to handle future spikes. Which solution should the team implement?

A.Increase the desired capacity of the ASG and add more read replicas to distribute the database load.
B.Increase the DB instance size to a larger instance class and implement an Amazon ElastiCache cluster to cache frequent database queries.
C.Migrate the database to Amazon DynamoDB with auto scaling and rewrite the application to use a serverless architecture with AWS Lambda.
D.Reduce the maximum connections parameter on the RDS instance to prevent connection exhaustion and modify the application code to reduce the number of database queries.
AnswerB

Scaling the DB instance to a larger class directly increases available vCPU, memory, and the maximum connection limit, giving the primary database the headroom needed to absorb the current CPU spike. Implementing an ElastiCache cluster (for example, Redis or Memcached) in front of the database caches the results of frequent, repetitive queries, so those reads never reach the RDS instance, which lowers CPU usage and frees connections for writes and less frequent queries. Together these actions provide both immediate compute capacity and durable read-path relief, exactly matching the incident's requirements.

Why this answer

The RDS instance is hitting connection limits and high CPU, causing unresponsiveness. Increasing the DB instance size provides more CPU and memory, allowing it to handle more connections and process queries faster. Adding an ElastiCache cluster offloads frequent read queries from the database, reducing the load on RDS.

This combination addresses both the connection exhaustion and CPU bottleneck, improving performance during spikes.

Exam trap

SOA-C02 often tests the misconception that read replicas can solve all scaling issues, but they only help with read-heavy workloads and do not alleviate connection limits on the primary.

How to eliminate wrong answers

Option A is wrong because adding more read replicas does not help with write-heavy workloads or connection limits on the primary; the application likely uses the primary for writes, and read replicas only offload reads. Option C is wrong because migrating to DynamoDB and Lambda is a major re-architecture that may not be necessary and could introduce new complexities; it's not the most direct solution. Option D is wrong because reducing max connections would worsen the problem by causing connection failures; it doesn't address the root cause of high load.

199
MCQmedium

A SysOps administrator creates the above IAM policy for a user. The user reports that they cannot delete an object in the bucket 'my-bucket' even though they are using MFA. What is the likely cause?

A.The resource ARN is missing the bucket-level permission.
B.The condition key aws:MultiFactorAuthPresent is incorrectly spelled.
C.The user is not using MFA when making the API call.
D.The policy does not include s3:DeleteObjectVersion.
AnswerC

The condition likely sets `aws:MultiFactorAuthPresent` to `false` or uses the `Bool` operator to deny access when MFA is absent. Because the user made the API call without an MFA token, the condition evaluates to `false`, triggering the `Deny` statement. This is the explicit reason why the delete request fails, as the policy mandates MFA for all actions by this user.

Why this answer

The policy requires MFA for all s3:DeleteObject actions, as indicated by the condition key aws:MultiFactorAuthPresent set to 'true'. If the user reports they cannot delete an object despite using MFA, the most likely cause is that they are not actually using MFA when making the API call — for example, they may have authenticated with long-term credentials (access key/secret key) without a multi-factor authentication session. The condition key checks the presence of an MFA-authenticated session token, not just whether the user has MFA enabled on their account.

Exam trap

The trap here is that candidates confuse 'having MFA enabled on the user account' with 'using MFA in the API call session' — the condition key aws:MultiFactorAuthPresent checks the latter, not the former.

How to eliminate wrong answers

Option A is wrong because the resource ARN 'arn:aws:s3:::my-bucket/*' correctly specifies object-level permissions for all objects in the bucket, and bucket-level permissions (e.g., s3:ListBucket) are not required for the s3:DeleteObject action. Option B is wrong because the condition key 'aws:MultiFactorAuthPresent' is correctly spelled — it is case-sensitive and matches the official AWS documentation. Option D is wrong because s3:DeleteObjectVersion is a separate action for deleting a specific version of an object, and the policy already includes s3:DeleteObject, which covers deleting the current version of an object (the most common operation).

200
MCQeasy

A company runs a web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application stores session data in an in-memory cache on the EC2 instances. During an instance refresh, users lose their session data. Which action should be taken to improve reliability without major application changes?

A.Use ElastiCache for Memcached with auto-discovery.
B.Move session state to Amazon ElastiCache for Redis.
C.Increase the minimum size of the Auto Scaling group.
D.Enable sticky sessions (session affinity) on the ALB target group.
AnswerB

Moving session state to Amazon ElastiCache for Redis is the correct solution because Redis supports replication across multiple Availability Zones, persistence through snapshots and AOF log, and atomic operations for managing session TTLs. By externalizing sessions, the EC2 instances become stateless: any instance in the Auto Scaling group can serve any user request, and session data survives instance termination, replacement, or scaling events. ElastiCache for Redis can also be configured with cluster mode to scale horizontally, ensuring session capacity grows with application demand while maintaining high availability.

Why this answer

Amazon ElastiCache for Redis provides a fully managed, external, and highly available in-memory data store that can be used to persist session state outside of the EC2 instances. By moving session data to Redis, the session state survives instance refreshes, terminations, or scaling events without requiring any changes to the application's session management logic beyond pointing to the Redis endpoint. This decouples session state from the compute layer, ensuring reliability and data durability during Auto Scaling lifecycle events.

Exam trap

The trap here is that candidates often confuse sticky sessions (session affinity) with session persistence, mistakenly believing that routing requests to the same instance prevents data loss, when in fact sticky sessions do not protect against instance termination or replacement during scaling events.

How to eliminate wrong answers

Option A is wrong because ElastiCache for Memcached is a pure caching solution that does not offer built-in persistence, replication, or failover capabilities; if the Memcached node fails, all session data is lost, and auto-discovery only helps with client connection management, not data durability. Option C is wrong because increasing the minimum size of the Auto Scaling group does not prevent session data loss during an instance refresh; it only ensures more instances are running, but the in-memory cache on each instance is still ephemeral and lost when instances are replaced. Option D is wrong because enabling sticky sessions (session affinity) on the ALB target group only ensures that a user's requests are routed to the same instance during a session, but it does not preserve session data when that instance is terminated or replaced during an instance refresh; the data is still stored in the instance's local memory and is lost upon instance termination.

201
MCQeasy

A company has an RDS PostgreSQL database with a Multi-AZ deployment. The primary instance fails. What happens to the application connections?

A.The application must reconnect to the same endpoint; it will be redirected to the standby instance.
B.The application will be automatically redirected to a read replica.
C.The administrator must manually change the CNAME to point to the standby.
D.The application must use a new endpoint in a different AWS Region.
AnswerA

In a Multi-AZ RDS deployment, a standby DB instance is provisioned in a different Availability Zone and is kept in sync with the primary via synchronous replication. When a failure occurs, Amazon RDS automatically promotes the standby and updates the DNS record for the same endpoint to point to the new primary. Existing application connections are dropped, so the application must reconnect, but it must reconnect to the same endpoint; the DNS update happens automatically and transparently to the client.

Why this answer

When the primary RDS PostgreSQL instance fails in a Multi-AZ deployment, AWS automatically fails over to the standby instance in a different Availability Zone. The DNS endpoint (CNAME) remains the same, but its resolution is updated to point to the standby instance's IP address. The application must reconnect to the same endpoint, as existing connections to the failed primary are dropped; after reconnection, traffic is seamlessly routed to the new primary.

Exam trap

The trap here is that candidates often confuse Multi-AZ failover with read replicas, assuming the application is automatically redirected to a read replica for writes, but read replicas are asynchronous and cannot accept write traffic.

How to eliminate wrong answers

Option B is wrong because read replicas are used for read scaling, not automatic failover; Multi-AZ failover uses a standby in a different AZ, not a read replica. Option C is wrong because AWS RDS Multi-AZ automatically updates the DNS CNAME to point to the standby instance; no manual administrator intervention is required. Option D is wrong because the endpoint remains the same across the failover; the application does not need a new endpoint in a different AWS Region, as Multi-AZ operates within a single region.

202
MCQmedium

A SysOps administrator is designing a disaster recovery plan for a critical application that runs on EC2 instances in a single region. The RTO is 1 hour, and the RPO is 15 minutes. The application data is stored on an Amazon EBS volume. Which approach meets these requirements at the lowest cost?

A.Deploy the application across multiple Availability Zones using an Auto Scaling group and an Application Load Balancer.
B.Use AWS Database Migration Service (DMS) for continuous replication to an EC2 instance in the DR region.
C.Take automated EBS snapshots every 15 minutes and copy them to the DR region. Use a pre-configured Amazon Machine Image (AMI) to launch EC2 instances from the latest snapshot.
D.Use S3 Cross-Region Replication to replicate the EBS volume data to an S3 bucket in the DR region.
AnswerC

Automated EBS snapshots taken every 15 minutes and copied to the DR region deliver a consistent, block-level backup of the instance's data with a recoverable point objective of at most 15 minutes. Because EBS snapshots are stored in Amazon S3 and support cross-region copying, you can recreate the volume in the DR region. Launching EC2 instances from a pre-configured AMI that uses the latest restored snapshot as its root volume minimizes the recovery time objective by avoiding manual OS and application provisioning, making this a cost-effective and reliable DR strategy.

Why this answer

Automated EBS snapshots taken every 15 minutes meet the RPO of 15 minutes, and copying them to the DR region allows launching EC2 instances from the latest snapshot using a pre-configured AMI, which can achieve an RTO of 1 hour. This approach is the lowest cost as it only incurs snapshot storage and cross-region data transfer costs, without requiring continuous replication infrastructure or additional compute resources.

Exam trap

The trap here is that candidates may confuse cross-region replication for EBS volumes with S3 Cross-Region Replication, assuming EBS data can be directly replicated to S3, but EBS volumes are block-level storage and cannot be replicated via S3 CRR without an intermediary like AWS Backup or snapshot copy.

How to eliminate wrong answers

Option A is wrong because deploying across multiple Availability Zones within the same region does not provide disaster recovery for a regional failure; it only provides high availability within the region, failing to meet the DR requirement for a separate region. Option B is wrong because AWS DMS is designed for database replication, not for replicating arbitrary EBS volume data; it would require a database engine and continuous replication infrastructure, increasing cost and complexity unnecessarily. Option D is wrong because S3 Cross-Region Replication cannot directly replicate EBS volume data; EBS volumes are block storage, not object storage, and S3 CRR only works with S3 objects, so this approach is technically infeasible.

203
MCQeasy

A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which configuration should be used?

A.Use AWS Lambda to periodically check instance health and reboot if necessary.
B.Enable termination protection on the instance.
C.Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
D.Place the instance in an Auto Scaling group with a minimum size of 1.
AnswerC

A CloudWatch alarm on the StatusCheckFailed metric can be configured with the EC2 recovery action, which automatically stops and starts the instance on a different physical host when the underlying hardware or network is impaired. This recovery process preserves the instance ID, private IP address, Elastic IP, and instance metadata, and the EBS root volume remains attached with its existing data. Because the action moves the instance to healthy hardware while maintaining its identity and configuration, it is the correct mechanism to recover from a physical host failure without requiring manual intervention.

Why this answer

A CloudWatch alarm on the StatusCheckFailed metric can be configured with an EC2 recovery action. When the alarm triggers (e.g., due to an underlying hardware failure), the recovery action automatically stops and starts the instance on healthy hardware, preserving the instance ID, private IP, Elastic IP, and EBS attachments. This is the native AWS mechanism for automatic instance recovery from hardware failures.

Exam trap

The trap here is that candidates often confuse termination protection (which only prevents deletion) or Auto Scaling replacement (which creates a new instance) with the native recovery action that preserves the instance's identity and state.

How to eliminate wrong answers

Option A is wrong because using AWS Lambda to periodically check instance health and reboot is an unnecessary, custom workaround that adds complexity and latency; AWS already provides the built-in CloudWatch alarm recovery action for this purpose. Option B is wrong because termination protection only prevents accidental deletion of an instance via the console or API; it does not detect or recover from hardware failures. Option D is wrong because placing the instance in an Auto Scaling group with a minimum size of 1 will replace a failed instance with a new one (different instance ID, IP, and metadata), which does not preserve the original instance's identity and attached resources like the recovery action does.

204
Matchingmedium

Match each AWS backup and disaster recovery service to its feature.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Centralized backup management

Automatic object replication across regions

High availability with standby replica

Read scaling and cross-region disaster recovery

Continuous replication for DR

Why these pairings

The correct matches are: AWS Backup for centralized automation, AWS Elastic Disaster Recovery for continuous replication, AWS Storage Gateway for hybrid cloud storage, and AWS Snowball for offline transfer. Common confusions involve swapping the centralized backup and replication functions.

205
MCQeasy

A SysOps administrator needs to ensure that an Amazon S3 bucket can withstand the loss of an entire AWS Availability Zone. What is the SIMPLEST configuration to meet this requirement?

A.Enable cross-region replication to a bucket in another Region.
B.Use S3 Standard storage class.
C.Use S3 One Zone-IA storage class.
D.Enable MFA Delete on the bucket.
AnswerB

The S3 Standard storage class is the correct choice because it automatically stores each object redundantly across a minimum of three Availability Zones in the same AWS Region. S3 Standard is engineered for 99.999999999% (11 nines) object durability and 99.99% availability, so if one Availability Zone becomes unavailable, S3 can continue serving reads and writes from the remaining AZs without any administrative action. This built-in synchronous replication across AZs is precisely what satisfies the requirement for AZ failure resilience.

Why this answer

S3 Standard automatically stores objects across a minimum of three Availability Zones (AZs) within the same AWS Region. This design ensures that the bucket can withstand the loss of an entire AZ without any additional configuration, making it the simplest option to meet the requirement.

Exam trap

The trap here is that candidates often overthink and choose cross-region replication (Option A) for high availability, but the question specifically asks for resilience against an AZ loss, which S3 Standard already provides within a single Region without the complexity and cost of CRR.

How to eliminate wrong answers

Option A is wrong because cross-region replication (CRR) adds complexity and cost, and is not the simplest way to withstand an AZ loss—S3 Standard already provides AZ resilience within a single Region. Option C is wrong because S3 One Zone-IA stores data in only a single AZ, so the loss of that AZ would result in permanent data loss, failing the requirement. Option D is wrong because MFA Delete is a security feature that adds extra authentication for delete operations; it does not provide any data durability or AZ resilience.

← PreviousPage 3 of 3 · 205 questions total

Ready to test yourself?

Try a timed practice session using only Reliability and Business Continuity questions.