Courseiva

CCNA Soa Reliability Bc Questions

75 of 205 questions · Page 2/3 · Soa Reliability Bc topic · Answers revealed

76
Multi-Selectmedium

Which TWO actions improve the availability of an application running on EC2 instances in a single Auto Scaling group? (Choose 2)

Select 2 answers
A.Use an Elastic Load Balancer with health checks to route traffic only to healthy instances.
B.Increase the instance size to handle higher load.
C.Create a CloudWatch alarm to notify when the CPU utilization exceeds 80%.
D.Enable EBS optimization on the instances.
E.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
AnswersA, E

An Elastic Load Balancer continuously performs health checks against targets in its target group, using protocols like HTTP or TCP. If an instance fails a configured number of consecutive checks, the ELB automatically marks it unhealthy and stops routing traffic to it, while continuing to send requests only to healthy instances. This provides built-in fault tolerance for instance failures and is essential for maintaining application availability.

Why this answer

An Elastic Load Balancer (ELB) performs health checks against EC2 instances and automatically routes traffic only to instances that pass those checks. If an instance becomes unhealthy, the ELB stops sending traffic to it, preventing user requests from reaching a failing instance and thereby improving application availability.

Exam trap

The trap here is that candidates often confuse scaling actions (like increasing instance size) or monitoring (like CloudWatch alarms) with direct availability improvements, failing to recognize that only redundancy (multiple AZs) and health-based traffic routing (ELB health checks) actively mitigate failures.

77
MCQeasy

A SysOps administrator is tasked with ensuring that an Amazon S3 bucket can withstand the loss of an entire AWS Region. The bucket stores critical data that must be accessible with minimal latency from multiple regions. Which solution meets these requirements?

A.Enable S3 Versioning and configure a lifecycle policy to transition objects to S3 Glacier Deep Archive.
B.Configure S3 Cross-Region Replication to a bucket in another AWS Region. Use Amazon CloudFront with multiple origins pointing to both buckets.
C.Enable S3 Versioning and MFA Delete on the bucket. Use S3 Object Lock to prevent object deletion.
D.Enable S3 Transfer Acceleration on the bucket and use a CloudFront distribution with the bucket as the origin.
AnswerB

S3 Cross-Region Replication asynchronously copies every uploaded object to a destination bucket in a different Region, giving the bucket contents regional redundancy; versioning must be enabled on both source and destination for CRR to function. Amazon CloudFront can be configured with an origin group containing both buckets—one as primary and one as secondary—so that if the primary origin returns an error or is unreachable, CloudFront automatically fails over to the replicated bucket in the other Region. This combination provides both replicated durability and low-latency edge delivery, supporting failover without manual intervention.

Why this answer

S3 Cross-Region Replication (CRR) automatically replicates objects to a bucket in another AWS Region, ensuring data survives a regional failure. Using Amazon CloudFront with multiple origins pointing to both buckets provides low-latency access from any region by routing requests to the nearest healthy origin, meeting both durability and performance requirements.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration or Versioning with multi-region replication, failing to recognize that only CRR provides cross-region data redundancy and that CloudFront with multiple origins is needed for low-latency access and automatic failover.

How to eliminate wrong answers

Option A is wrong because S3 Glacier Deep Archive is designed for long-term archival with retrieval times of hours, not for low-latency access, and versioning alone does not provide multi-region redundancy. Option C is wrong because S3 Versioning, MFA Delete, and Object Lock protect against accidental deletion but do not replicate data across regions or enable low-latency access from multiple regions. Option D is wrong because S3 Transfer Acceleration speeds up uploads over long distances but does not replicate data to another region or provide failover for regional outages; CloudFront with a single bucket origin does not offer multi-region redundancy.

78
MCQhard

A company runs a stateful application on EC2 instances in an Auto Scaling group. The application maintains state in memory. The SysOps administrator wants to ensure that when an instance is terminated, the state is preserved and a new instance can resume operation. Which approach should the administrator use?

A.Use an Auto Scaling lifecycle hook to offload state before termination.
B.Use a warm standby instance that takes over when the primary fails.
C.Configure the EC2 instance to run a script on shutdown to save state locally.
D.Enable connection draining on the ALB.
AnswerA

Auto Scaling lifecycle hooks pause the instance in the 'terminating:wait' state, providing a window for a script or Lambda function to copy application state to durable external storage such as Amazon S3 or an EFS mount. Only after the state is successfully stored does the hook signal the Auto Scaling group to continue the termination, ensuring no data is lost during a scale-in event. This is the only option that explicitly performs an offload operation before the instance is destroyed.

Why this answer

An Auto Scaling lifecycle hook can be configured to execute a custom action (e.g., offloading in-memory state to Amazon S3 or ElastiCache) before the instance is terminated. The lifecycle hook places the instance in a 'terminating:wait' state, allowing the administrator to run a script that preserves state, and then completes the termination via CompleteLifecycleAction. This ensures state is saved before the instance is fully terminated, enabling a new instance to resume operation.

Exam trap

The trap here is that candidates assume a shutdown script (Option C) is sufficient, but they overlook that local instance store is ephemeral and that the OS shutdown sequence may not complete before the instance is forcefully terminated, making lifecycle hooks the only reliable mechanism for state preservation.

How to eliminate wrong answers

Option B is wrong because a warm standby instance does not address the need to preserve state from a terminating instance; it simply provides a pre-provisioned replacement that would still lack the in-memory state of the failed instance. Option C is wrong because a shutdown script runs after the termination signal is sent, but the instance may be forcibly terminated before the script completes, and local storage (instance store or ephemeral volumes) is lost on termination, making local save unreliable. Option D is wrong because connection draining on an ALB only stops new connections and allows in-flight requests to complete before deregistering the target; it does not preserve or offload application state stored in memory.

79
MCQhard

A company runs a microservices architecture on Amazon ECS with Fargate. They need to ensure that if a task fails, it is automatically restarted. Which configuration is required?

A.Configure a CloudWatch alarm to restart the task
B.Define tasks using the ECS RunTask API
C.Use an Auto Scaling group for the Fargate tasks
D.Create an ECS service with a desired count and task definition
AnswerD

An ECS service is the correct mechanism for running a long-lived microservice: you specify a task definition, a desired count, and a cluster, and the service scheduler ensures that the desired number of tasks is always running. When a task crashes, is killed, or fails a health check, the service automatically replaces it by starting a new task from the same task definition. This provides self-healing and integrates with Application Load Balancers, target groups, and service discovery for production traffic.

Why this answer

An ECS service with a desired count and task definition ensures that Fargate tasks are automatically restarted if they fail. The ECS service scheduler monitors the desired count and replaces any stopped or failed tasks to maintain the specified number of running instances, providing built-in resilience without additional infrastructure.

Exam trap

The trap here is that candidates confuse the RunTask API (for one-off tasks) with ECS services (for long-running, self-healing tasks), or mistakenly think CloudWatch alarms or Auto Scaling groups are needed for task restart logic when the ECS service itself provides this capability.

How to eliminate wrong answers

Option A is wrong because CloudWatch alarms can trigger actions like sending notifications or scaling, but they cannot directly restart an ECS task; restarting requires a service or custom automation. Option B is wrong because the RunTask API launches tasks on-demand for batch or one-off jobs, not for continuous availability or automatic restart on failure. Option C is wrong because Auto Scaling groups are used for EC2 instances, not for Fargate tasks; Fargate tasks are managed by ECS services or standalone RunTask calls, and scaling is handled via ECS Service Auto Scaling, not an Auto Scaling group.

80
MCQhard

A company runs a stateful web application on a single EC2 instance. The SysOps Administrator wants to improve fault tolerance. Which design should they implement?

A.Create a Multi-AZ RDS instance and attach it to the existing EC2 instance.
B.Add a second EC2 instance in the same Availability Zone and use a Network Load Balancer.
C.Use an Auto Scaling group with a launch configuration that stores session data on instance store.
D.Place instances in an Auto Scaling group across two Availability Zones, use an Application Load Balancer, and store session state in ElastiCache.
AnswerD

Placing the application instances in an Auto Scaling group across two Availability Zones ensures that both an instance failure and a full AZ failure can be absorbed automatically, as the ASG launches a replacement instance in the remaining healthy AZ. An Application Load Balancer distributes traffic and performs health checks, routing requests only to healthy instances. Storing session state in ElastiCache (or a similar external store) decouples sessions from compute instances, so any instance can handle any request without losing user data; if an instance terminates, a new instance can access the same session state from ElastiCache. This combination provides both high availability and fault tolerance for a stateful web application.

Why this answer

It distributes the web application across multiple Availability Zones for high availability, uses an Application Load Balancer to route traffic to healthy instances, and stores session state externally in ElastiCache. This decouples the stateful session data from the EC2 instances, allowing any instance to handle any request without losing session context, which is essential for fault tolerance in a stateful application.

Exam trap

The trap here is that candidates often confuse high availability with fault tolerance, choosing a single-AZ solution (Option B) or a database-only fix (Option A), failing to recognize that stateful applications require externalized session state to survive instance or AZ failures.

How to eliminate wrong answers

Option A is wrong because Multi-AZ RDS provides database high availability, but the web application itself remains a single point of failure on one EC2 instance; it does not address the fault tolerance of the application tier. Option B is wrong because adding a second EC2 instance in the same Availability Zone does not protect against an Availability Zone failure; the entire zone could go down, taking both instances with it. Option C is wrong because instance store is ephemeral and data is lost if the instance stops, terminates, or fails; storing session data on instance store would cause session loss during any fault event, defeating the purpose of improving fault tolerance.

81
MCQeasy

A company has an application that runs on EC2 instances behind an Application Load Balancer. The application uses an RDS Multi-AZ database. The company wants to ensure that the application remains available during a database failover. What should the SysOps administrator do?

A.Ensure the application retries database connections during failover.
B.Create a read replica of the database to offload read traffic.
C.Increase the EC2 instance size to handle the load.
D.Enable termination protection on the EC2 instances.
AnswerA

During a Multi-AZ RDS failover, the RDS DNS CNAME is repointed to the standby database, a process that typically takes 60–120 seconds, during which existing database connections are forcibly terminated. A resilient application must catch connection errors and retry with exponential backoff or a fresh connection attempt so it can ride through the interruption; without this retry logic, requests that arrive during the failover window will return errors even after the new primary is ready.

Why this answer

During an RDS Multi-AZ failover, the database DNS record is updated to point to the standby instance, which can take up to 60 seconds. Existing connections are dropped, so the application must implement connection retry logic with exponential backoff to re-establish connections to the new primary. Without retries, the application will fail to serve requests until the database is reachable again, breaking availability.

Exam trap

The trap here is that candidates confuse high-availability infrastructure (Multi-AZ, termination protection) with application-level resilience, assuming the infrastructure alone guarantees uptime without requiring the application to handle transient connection failures.

How to eliminate wrong answers

Option B is wrong because creating a read replica offloads read traffic but does not help the application survive a Multi-AZ failover; the read replica is a separate instance and does not become the new primary during failover. Option C is wrong because increasing EC2 instance size addresses compute capacity, not database connectivity issues caused by a failover. Option D is wrong because termination protection prevents accidental EC2 instance termination, but has no effect on database failover behavior or connection handling.

82
Multi-Selectmedium

A company is designing a highly available architecture for a web application using AWS services. Which TWO actions should the SysOps administrator take to improve reliability? (Choose TWO.)

Select 2 answers
A.Use a single large EC2 instance to eliminate complexity.
B.Use an Auto Scaling group with an Elastic Load Balancer.
C.Use a single NAT gateway for outbound traffic.
D.Deploy EC2 instances in multiple Availability Zones.
E.Store application data on a single EBS volume.
AnswersB, D

An Auto Scaling group combined with an Elastic Load Balancer delivers high availability by maintaining a desired instance count across healthy capacity and automatically replacing failed instances based on health checks. The ELB distributes incoming traffic across all healthy instances, preventing any single node from being overwhelmed and allowing the fleet to scale in and out with demand. When integrated across multiple Availability Zones, this pattern provides both elasticity and fault tolerance, making it a core AWS high-availability design.

Why this answer

An Auto Scaling group combined with an Elastic Load Balancer automatically distributes incoming traffic across healthy EC2 instances and replaces any failed instances, ensuring the application remains available even during instance failures or traffic spikes. This architecture is a core AWS best practice for building highly available and resilient web applications.

Exam trap

The trap here is that candidates often think a single large instance is simpler and more reliable, but AWS explicitly recommends horizontal scaling with multiple instances across Availability Zones to eliminate single points of failure.

83
MCQhard

Refer to the exhibit. A SysOps administrator deployed the CloudFormation template. Which statement is true about data protection?

A.Deleted objects are immediately and permanently removed.
B.The bucket cannot be deleted by anyone.
C.Objects cannot be deleted from the bucket.
D.Deleted objects become noncurrent versions and are retained for 30 days.
AnswerD

Because versioning is enabled on the bucket, deleting an object creates a delete marker and the object's current version transitions to noncurrent status, rather than being destroyed. The lifecycle configuration specifies that noncurrent versions expire 30 days after they become noncurrent, at which point S3 permanently removes them. Consequently, deleted objects are retained for exactly 30 days before being purged, matching the behavior described in this option.

Why this answer

The CloudFormation template likely configures an S3 bucket with versioning enabled and a lifecycle policy that transitions noncurrent versions to the 'ExpiredObjectDeleteMarker' or retains them for a specified period. By default, when versioning is enabled, deleted objects are not permanently removed but become noncurrent versions, and the lifecycle rule can be set to permanently delete these noncurrent versions after 30 days, as indicated in the template.

Exam trap

The trap here is that candidates often confuse S3 versioning's behavior with standard (non-versioned) buckets, assuming that 'deleted' means 'permanently removed' immediately, or they overlook that lifecycle policies only apply to noncurrent versions and do not prevent direct deletion of objects.

How to eliminate wrong answers

Option A is wrong because with S3 versioning enabled, deleted objects are not immediately and permanently removed; they become noncurrent versions, and permanent removal only occurs after the lifecycle policy's expiration period or manual deletion of all versions. Option B is wrong because the bucket can be deleted by anyone with the appropriate IAM permissions (e.g., s3:DeleteBucket), and the template does not include a bucket policy or resource-based policy that explicitly denies deletion; the 'DeletionPolicy' attribute in CloudFormation (e.g., 'Retain') only protects the bucket from being deleted during stack deletion, not from direct API calls. Option C is wrong because objects can be deleted from the bucket when versioning is enabled; deletion creates a delete marker (for the current version) or permanently deletes a specific version ID, and the lifecycle policy does not prevent deletion—it only governs the expiration of noncurrent versions.

84
MCQmedium

An application running on EC2 instances stores session data in an attached EBS volume. The company wants to ensure session data is not lost if an instance fails. Which solution should the administrator implement?

A.Move session storage to Amazon ElastiCache for Redis with replication.
B.Use EBS Multi-Attach to attach the volume to multiple instances.
C.Take frequent EBS snapshots of the volume.
D.Use a larger EC2 instance type with more memory.
AnswerA

ElastiCache for Redis with replication provides a highly available, in-memory session store that is accessible from all EC2 instances. By using Multi-AZ replication, Redis automatically fails over to a replica if the primary node fails, so session data remains available. You can also enable append-only file persistence to guard against data loss, making it a robust and widely used pattern for storing session state in distributed applications.

Why this answer

Amazon ElastiCache for Redis with replication provides a highly available, in-memory data store that persists session data independently of EC2 instances. If an instance fails, the session data remains intact in the replicated Redis cluster, ensuring zero data loss and seamless failover. This decouples session state from ephemeral compute resources, aligning with the reliability and business continuity requirements.

Exam trap

The trap here is that candidates often assume EBS snapshots or Multi-Attach provide real-time session durability, but they fail to recognize that session data is ephemeral and requires a separate, highly available data store like ElastiCache to survive instance failures without data loss.

How to eliminate wrong answers

Option B is wrong because EBS Multi-Attach only supports io1/io2 volumes and is limited to a single Availability Zone; it does not provide cross-instance failover or protect against instance failure, as all attached instances share the same underlying storage and would all lose access if the volume fails. Option C is wrong because frequent EBS snapshots are point-in-time backups stored in Amazon S3, not real-time session storage; restoring from a snapshot would lose any session data written after the last snapshot and introduces significant downtime. Option D is wrong because using a larger EC2 instance type with more memory only increases the capacity for in-memory session data on that single instance; if the instance fails, all session data in memory is lost, regardless of instance size.

85
MCQmedium

A company runs a web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application is deployed in a single Availability Zone. The SysOps administrator notices that during peak hours, the application becomes slow and some requests fail. CloudWatch metrics show that CPU utilization on the instances reaches 90%, but the Auto Scaling group does not scale out. The administrator has configured a target tracking scaling policy based on average CPU utilization with a target value of 75%. The Auto Scaling group has a minimum of 2, maximum of 10, and desired capacity of 2. What is the MOST likely reason the Auto Scaling group is not scaling out?

A.The Auto Scaling group is configured with a single Availability Zone, and the target tracking policy cannot scale out beyond the capacity of that single AZ.
B.The target tracking policy uses a target value of 75%, but the average CPU is above that, so it should scale out.
C.The target tracking policy requires detailed monitoring to be enabled on the instances.
D.The Auto Scaling group has reached its maximum capacity of 10 instances.
AnswerA

The Auto Scaling group is constrained to a single Availability Zone, and the target tracking policy cannot add instances if that specific AZ does not have sufficient capacity for the instance type defined in the launch template. Auto Scaling attempts to launch new instances, but because the group does not span multiple AZs, it cannot shift to another AZ when the sole AZ reports InsufficientInstanceCapacity. Quotas such as vCPUs typically apply per-region, but the AZ-level capacity limitation is the specific reason here—even though the group's maximum is 10, scale-out fails because there is no available capacity in that one AZ.

Why this answer

The most likely reason is that the Auto Scaling group is configured with a single Availability Zone. Target tracking scaling policies operate within the constraints of the configured subnets. If the group is only in one AZ, the subnet may have insufficient IP addresses or the AZ may have reached its instance limit, preventing the group from scaling out beyond that AZ's capacity.

Since the group has a maximum of 10, but the AZ capacity is limited, the scaling policy cannot add more instances. This is a common issue: AWS recommends using multiple Availability Zones for Auto Scaling groups to allow scaling across AZs and avoid single points of failure.

86
MCQeasy

An organization needs to back up an Amazon EFS file system daily and retain backups for 30 days. Which AWS service provides a managed backup solution for EFS?

A.AWS Backup
B.Amazon Data Lifecycle Manager (DLM)
C.EFS replication to another region
D.S3 Lifecycle policies
AnswerA

AWS Backup is the correct service because it is the native, fully managed backup service designed specifically for Amazon EFS file systems. A backup plan schedules recurring EFS backups, stores them as recovery points, and applies configurable retention policies, including cross-region and cross-account copies. It also supports point-in-time restore, lifecycle rules that transition recovery points to cold storage, and integration with AWS Organizations for compliance. Thus AWS Backup provides the backup-with-retention capability the organization requires.

Why this answer

AWS Backup is the correct answer because it is a fully managed, policy-based backup service that supports Amazon EFS natively. It allows you to define backup plans with daily schedules and retention rules (e.g., 30 days) without needing to manage any backup infrastructure or scripts. AWS Backup handles the lifecycle of backups, including incremental backups and automatic deletion of expired recovery points.

Exam trap

The trap here is that candidates often confuse Amazon Data Lifecycle Manager (DLM) as a general-purpose backup tool, but DLM is strictly limited to EBS snapshots and AMIs, not EFS file systems.

How to eliminate wrong answers

Option B (Amazon Data Lifecycle Manager) is wrong because DLM is designed for managing the lifecycle of Amazon EBS snapshots and EBS-backed AMIs, not for backing up Amazon EFS file systems. Option C (EFS replication to another region) is wrong because replication provides a cross-region copy of the file system for disaster recovery, but it does not offer point-in-time backup retention or automated deletion after a specific period like 30 days. Option D (S3 Lifecycle policies) is wrong because S3 Lifecycle policies manage the transition and expiration of objects within S3 buckets, and they cannot be applied directly to an EFS file system, which is a separate storage service.

87
Multi-Selecteasy

A SysOps administrator is planning for disaster recovery of an RDS MySQL database. The database is currently in a single AZ. Which TWO actions will improve recovery time and reduce data loss? (Select TWO.)

Select 2 answers
A.Create a read replica in a different AWS Region.
B.Enable automated backups with a retention period of 7 days.
C.Enable Multi-AZ deployment for automatic failover.
D.Enable deletion protection on the RDS instance.
E.Increase the allocated storage to improve performance.
AnswersB, C

Automated backups with a retention period of 7 days enable point-in-time recovery (PITR), allowing you to restore the database to any second within that retention window. This protects against logical errors, such as accidental table drops or erroneous UPDATE statements, by letting you recover a clean copy of the data as of a specific timestamp. The backup is stored in S3 and is essential for meeting recovery point objectives (RPO) when data corruption occurs, making it a core disaster recovery tool.

Why this answer

Enabling automated backups with a 7-day retention period allows point-in-time recovery (PITR) to any second within the retention window, minimizing data loss (RPO) by restoring to the most recent backup. Option C is correct because Multi-AZ deployment provides automatic synchronous standby replication to a different Availability Zone, enabling automatic failover with minimal downtime (RTO) in case of an AZ failure or instance issue.

Exam trap

The trap here is that candidates often confuse read replicas with Multi-AZ failover, thinking a read replica can serve as a quick disaster recovery option, but read replicas are asynchronous and do not provide automatic failover or synchronous data protection.

88
MCQeasy

A company is designing a highly available web application on AWS. The application runs on EC2 instances behind an Application Load Balancer. Which configuration ensures that the application remains available if an entire AWS Availability Zone fails?

A.Deploy EC2 instances in multiple subnets of the same Availability Zone.
B.Launch EC2 instances in at least two different Availability Zones.
C.Use a larger EC2 instance type to handle the load.
D.Use EC2 instances in multiple AWS Regions.
AnswerB

Launching EC2 instances in at least two different Availability Zones is the correct approach because each AZ is an isolated, independent failure domain with separate power, cooling, and physical infrastructure. A load balancer can then distribute traffic across these AZs, so if one AZ becomes unavailable, the remaining AZs continue to serve traffic, maintaining high availability within the same AWS Region without cross-region latency.

Why this answer

Deploying EC2 instances in at least two different Availability Zones (AZs) ensures that if one AZ fails, the Application Load Balancer (ALB) can route traffic to healthy instances in the remaining AZ(s). ALBs are regional constructs that automatically distribute traffic across registered targets in multiple AZs, and they perform health checks to detect and route away from failed AZs. This design meets the high availability requirement by eliminating the AZ as a single point of failure.

Exam trap

The trap here is that candidates often confuse high availability with scalability or performance, mistakenly thinking that larger instances (Option C) or multi-Region deployment (Option D) are required, when the core requirement is simply eliminating a single AZ as a point of failure by using multiple AZs within the same region.

How to eliminate wrong answers

Option A is wrong because deploying instances in multiple subnets within the same Availability Zone does not protect against an AZ failure; if that single AZ goes down, all instances become unavailable. Option C is wrong because using a larger EC2 instance type only increases the compute capacity of a single instance, but does not provide redundancy or fault tolerance; a failure of that instance or its AZ still causes downtime. Option D is wrong because using multiple AWS Regions provides disaster recovery across geographic regions, but it is overkill and not necessary for availability within a single region; it also introduces higher latency and complexity not required for the stated goal of surviving an AZ failure.

89
MCQmedium

A company runs a global e-commerce application that uses Amazon DynamoDB as its primary database. The application requires single-digit millisecond read and write latency from any region and must continue to operate during a regional outage with minimal data loss. Which DynamoDB feature should the SysOps administrator enable to meet these requirements?

A.DynamoDB Accelerator (DAX)
B.DynamoDB global tables
C.DynamoDB Point-in-Time Recovery (PITR)
D.DynamoDB Auto Scaling
AnswerB

DynamoDB global tables automatically replicate each item write to all selected AWS Regions, creating active-active replica tables with multi-region read and write capability. This cross-region replication gives users low-latency access because they can be served by a nearby replica, and it provides business continuity by allowing another Region to continue serving traffic during a Regional outage without manual data restore. Because every replica holds a full copy of the data, a Region failure is effectively transparent at the table level, assuming your application can reroute traffic.

Why this answer

DynamoDB global tables provide multi-Region, multi-active replication, enabling single-digit millisecond reads and writes from any Region while offering automatic failover and recovery during a regional outage. This feature uses DynamoDB Streams to replicate data across Regions with eventual consistency, meeting the requirement for continued operation with minimal data loss.

Exam trap

The trap here is that candidates often confuse DynamoDB Accelerator (DAX) with global tables, assuming a caching layer can provide multi-Region availability, but DAX is Region-specific and does not replicate data across Regions.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that reduces read latency but does not provide multi-Region replication or write availability during a regional outage. Option C is wrong because Point-in-Time Recovery (PITR) enables backup restoration to any point within the last 35 days but does not provide real-time failover or cross-Region read/write capability. Option D is wrong because Auto Scaling adjusts provisioned throughput based on traffic but does not replicate data across Regions or ensure availability during a regional outage.

90
MCQhard

A company runs a production application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application uses an RDS for PostgreSQL database. The SysOps administrator has configured a read replica in a different AWS Region for disaster recovery. During a disaster, the primary region becomes unavailable. The administrator promotes the read replica to a standalone instance. After promoting, the application fails to connect to the new database because the endpoint changed. The administrator needs to minimize downtime. What should the administrator do to handle the endpoint change automatically?

A.Assign an Elastic IP address to the RDS instance.
B.Use Amazon Route 53 with a weighted alias record that points to the primary database endpoint, and configure a health check to fail over to the secondary endpoint.
C.Update the application configuration files to point to the new endpoint.
D.Use an RDS proxy and configure it to automatically failover to the promoted replica.
AnswerB

Route 53 weighted alias records can point to the RDS primary and secondary endpoints, allowing you to control traffic distribution. By associating a health check with the primary record, Route 53 automatically removes the primary from DNS resolution when the health check fails, causing traffic to be sent to the secondary endpoint. This achieves automated failover with minimal downtime, and the application continues using the same DNS name throughout.

Why this answer

By using Amazon Route 53 with a weighted alias record that points to the primary database endpoint and configuring a health check, the administrator can automate DNS failover. When the primary region becomes unavailable, the health check fails, and Route 53 automatically routes traffic to the secondary record that points to the promoted read replica's endpoint. This minimizes downtime.

Option A is wrong because Elastic IP addresses cannot be assigned to RDS instances; they are used for EC2 instances. Option C is wrong because manually updating application configuration files would increase downtime and is not automatic. Option D is wrong because an RDS proxy does not automatically update the endpoint after a disaster recovery promotion; it still requires the endpoint to be changed in the application configuration.

91
MCQmedium

A company stores critical data in an Amazon S3 bucket in the us-west-2 Region. The SysOps administrator needs to ensure that all objects are automatically replicated to another AWS Region for disaster recovery. The Recovery Point Objective (RPO) must be less than 15 minutes, and existing objects must also be replicated. Which S3 feature should the administrator use?

A.S3 Cross-Region Replication (CRR) with Replication Time Control (RTC)
B.S3 Same-Region Replication (SRR)
C.S3 Event Notifications with an AWS Lambda function to copy objects to another region
D.S3 Transfer Acceleration
AnswerA

S3 Cross-Region Replication (CRR) with Replication Time Control (RTC) is the correct choice because CRR asynchronously copies objects from the source bucket in the US to a destination bucket in another AWS Region, satisfying the geographic disaster recovery requirement. RTC adds a guaranteed 15-minute replication SLA for 99.99% of objects, meeting the stated RPO. Additionally, CRR can be configured to replicate existing objects using S3 Batch Replication, while RTC provides monitoring via CloudWatch metrics and events, ensuring timely, trackable replication.

Why this answer

S3 Cross-Region Replication (CRR) with Replication Time Control (RTC) is the correct choice because it provides automatic, asynchronous replication of objects to a different AWS Region, meeting the RPO of less than 15 minutes by guaranteeing replication within 15 minutes for most objects (99.99% of objects are replicated within 15 minutes). Additionally, CRR can replicate existing objects when configured with the appropriate replication rule and batch operations, satisfying the requirement to replicate all objects.

Exam trap

The trap here is that candidates may choose S3 Event Notifications with Lambda (Option C) because it seems like a flexible custom solution, but they overlook the lack of a guaranteed RPO, the inability to replicate existing objects without additional effort, and the operational overhead compared to the managed, SLA-backed CRR with RTC.

How to eliminate wrong answers

Option B (S3 Same-Region Replication) is wrong because it replicates objects within the same AWS Region, not across regions, so it does not meet the disaster recovery requirement for cross-region replication. Option C (S3 Event Notifications with Lambda) is wrong because it is a custom, event-driven approach that introduces latency, complexity, and potential failure points, and it cannot guarantee the 15-minute RPO or reliably replicate existing objects without additional scripting. Option D (S3 Transfer Acceleration) is wrong because it is designed to speed up uploads over long distances using edge locations, not to replicate objects between buckets or regions.

92
MCQeasy

A production RDS MySQL database stores financial records. The team needs the ability to restore the database to any point within the last 7 days in case of accidental data deletion. Automated backups are currently disabled. What must be configured?

A.Enable automated backups and set the backup retention period to 7 days
B.Create a manual DB snapshot every night using the AWS CLI on a schedule
C.Enable Multi-AZ to maintain a synchronous standby replica in a second Availability Zone
D.Enable RDS read replicas and promote one if data deletion occurs
AnswerA

Automated backups with a 7-day retention period keep daily snapshots and transaction logs for 7 days. Any point within the retention window is recoverable. Transaction logs allow recovery to any 5-minute interval within that window. Setting the period to 0 disables automated backups and PITR entirely.

Why this answer

To restore an RDS MySQL database to any point within the last 7 days, you must enable automated backups and set the backup retention period to 7 days. Automated backups enable point-in-time recovery (PITR), which allows restoration to any second within the retention window using binary logs. Without automated backups, RDS cannot perform PITR, even if manual snapshots exist.

Exam trap

The trap here is that candidates often confuse manual snapshots with automated backups, not realizing that only automated backups enable point-in-time recovery, while manual snapshots are static and cannot be used for granular restoration.

How to eliminate wrong answers

Option B is wrong because manual DB snapshots capture only a single point in time and do not provide the continuous binary log data needed for point-in-time recovery to any arbitrary moment within 7 days. Option C is wrong because Multi-AZ provides high availability and automatic failover, but it does not create backups or enable point-in-time recovery; it only maintains a synchronous standby replica. Option D is wrong because RDS read replicas are designed for read scaling and, while they can be promoted to a standalone instance, they do not provide point-in-time recovery capabilities and rely on the same backup configuration as the source instance.

93
MCQmedium

A company runs a production Amazon RDS for MySQL DB instance in a single Availability Zone. The SysOps administrator needs to improve database availability to ensure automatic failover if the primary instance fails. Which configuration should the administrator enable?

A.Create a Read Replica in another Availability Zone and promote it on failure.
B.Enable Multi-AZ deployment on the DB instance.
C.Take hourly snapshots and automate restoration in another AZ.
D.Use Amazon RDS Proxy to manage connection failover.
AnswerB

Enabling Multi-AZ deployment creates a synchronous standby replica in a different Availability Zone with automatic failover, ensuring that if the primary instance fails, Amazon RDS automatically switches to the standby with minimal downtime. The synchronous replication means the standby is always up to date with the primary, so there is no data loss. This configuration provides high availability with a single database endpoint, so applications can maintain connectivity during failover.

Why this answer

Enabling Multi-AZ deployment on the DB instance automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary instance fails, Amazon RDS automatically fails over to the standby, providing high availability without manual intervention. This is the native AWS solution for automatic failover for RDS MySQL.

Exam trap

The trap here is that candidates often confuse Read Replicas (asynchronous, for read scaling) with Multi-AZ (synchronous, for high availability), assuming promoting a Read Replica provides the same automatic failover guarantee.

How to eliminate wrong answers

Option A is wrong because creating a Read Replica and promoting it on failure is a manual process that introduces downtime and does not provide automatic failover; Read Replicas are designed for read scaling, not synchronous high availability. Option C is wrong because taking hourly snapshots and automating restoration in another AZ would result in significant data loss (up to one hour) and long recovery times, not automatic failover. Option D is wrong because Amazon RDS Proxy manages database connections and connection pooling, but it does not provide automatic failover of the database instance itself; it can work with Multi-AZ but is not a substitute for it.

94
MCQeasy

A company runs a web application on Amazon EC2 instances in a single Availability Zone. The SysOps administrator wants to increase the availability of the application so that it can survive an Availability Zone failure. Which action is the most effective?

A.Deploy an additional EC2 instance in the same Availability Zone.
B.Launch EC2 instances in two different Availability Zones and place them behind an Application Load Balancer.
C.Enable termination protection on all EC2 instances.
D.Use an Amazon RDS Multi-AZ deployment for the database tier.
AnswerB

Launching EC2 instances in two different Availability Zones and placing them behind an Application Load Balancer is the canonical web-tier high availability pattern. The ALB performs continuous health checks and distributes traffic to healthy targets across both AZs. If an entire AZ becomes unhealthy, the ALB automatically reroutes requests to instances in the remaining AZ, preserving the application's availability. This design eliminates the AZ as a single point of failure for the web tier.

Why this answer

Deploying EC2 instances across two different Availability Zones and placing them behind an Application Load Balancer (ALB) provides fault isolation. If one AZ fails, the ALB automatically routes traffic to the healthy instances in the other AZ, ensuring the application remains available. This architecture directly addresses the goal of surviving an AZ failure by eliminating the single point of failure at the AZ level.

Exam trap

The trap here is that candidates often confuse high availability with fault tolerance at a single component level, mistakenly thinking that adding more instances in the same AZ or enabling termination protection improves availability, when in fact only distributing resources across multiple isolated Availability Zones can survive an AZ failure.

How to eliminate wrong answers

Option A is wrong because adding more instances in the same Availability Zone does not protect against an AZ failure; all instances would still be affected if that single AZ goes down. Option C is wrong because termination protection only prevents accidental deletion of instances, it does not provide any redundancy or fault tolerance for an AZ outage. Option D is wrong because while an Amazon RDS Multi-AZ deployment improves database availability, it does not address the availability of the web application tier running on EC2; the question asks for the most effective action to increase application availability, which requires a multi-AZ architecture for the compute layer.

95
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer. The instances are in an Auto Scaling group across three Availability Zones. To improve reliability, the company wants to ensure that if an entire Availability Zone fails, the application remains available. Which configuration should be implemented?

A.Remove the load balancer and use Route 53 weighted routing to distribute traffic.
B.Launch all instances in a single Availability Zone to reduce latency.
C.Configure the Auto Scaling group to launch instances in three Availability Zones.
D.Use a Network Load Balancer instead of an Application Load Balancer.
AnswerC

Configuring the Auto Scaling group to launch instances in three Availability Zones is the correct approach because it distributes the web application across separate physical data centers within the Region. The Auto Scaling group manages instance health and automatically replaces unhealthy or failed instances in the remaining AZs, while the Application Load Balancer routes traffic only to healthy instances across all three zones. This design ensures that if an entire AZ experiences an outage, the load balancer shifts traffic to instances in the other AZs, keeping the application available and meeting the reliability requirement.

Why this answer

Configuring the Auto Scaling group to launch instances across three Availability Zones ensures that if one AZ fails, the remaining AZs continue to serve traffic. The Application Load Balancer automatically distributes incoming requests to healthy instances in the surviving AZs, maintaining application availability without manual intervention.

Exam trap

The trap here is that candidates may think changing the load balancer type (e.g., to NLB) is the solution, but the core requirement is ensuring the Auto Scaling group spans multiple Availability Zones, not the load balancer protocol or layer.

How to eliminate wrong answers

Option A is wrong because removing the load balancer and using Route 53 weighted routing would not provide automatic health checking and failover at the instance level; Route 53 weighted routing distributes traffic based on weights but does not monitor instance health or reroute traffic if an AZ fails, leading to potential downtime. Option B is wrong because launching all instances in a single Availability Zone creates a single point of failure; if that AZ fails, the entire application becomes unavailable, directly contradicting the goal of improving reliability. Option D is wrong because replacing the Application Load Balancer with a Network Load Balancer does not inherently improve AZ-level resilience; both ALB and NLB support cross-zone load balancing and can distribute traffic across multiple AZs, but the key requirement is that the Auto Scaling group spans multiple AZs, not the type of load balancer.

96
MCQmedium

A company runs a production database on an Amazon RDS for PostgreSQL DB instance in a single Availability Zone. The SysOps administrator needs to improve the database's availability to meet an SLA of 99.99% and ensure automatic failover in case of a database failure. Which configuration change should be made?

A.Enable a Multi-AZ deployment
B.Create a read replica in a different AWS Region
C.Configure automated backups with cross-region copy
D.Enable deletion protection on the DB instance
AnswerA

Multi-AZ deployment provisions a standby replica in a different Availability Zone with synchronous replication. If the primary instance fails, AWS automatically fails over to the standby, typically within 60 seconds, preserving your DNS endpoint so applications continue without manual intervention. This is the only option that directly provides automatic high availability for the primary RDS instance.

Why this answer

Enabling a Multi-AZ deployment for an Amazon RDS PostgreSQL DB instance automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of a database failure or an Availability Zone outage, Amazon RDS automatically fails over to the standby replica, typically within 60-120 seconds, meeting the 99.99% SLA requirement without manual intervention.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and disaster recovery) with Multi-AZ deployments (which are for high availability and automatic failover), leading them to incorrectly select the cross-region read replica option.

How to eliminate wrong answers

Option B is wrong because creating a read replica in a different AWS Region provides read scalability and disaster recovery, but it does not support automatic failover for the primary DB instance; failover requires manual promotion of the read replica, which cannot meet a 99.99% SLA. Option C is wrong because configuring automated backups with cross-region copy protects against data loss by storing backups in another region, but it does not provide automatic failover or high availability for the database instance itself. Option D is wrong because enabling deletion protection on the DB instance only prevents accidental deletion of the database; it has no effect on availability, failover, or resilience against failures.

97
MCQeasy

Refer to the exhibit. A SysOps administrator creates a CloudFormation stack with the template shown. After 30 days, what happens to noncurrent versions of objects in the bucket?

A.They are permanently deleted.
B.They are moved to Amazon S3 Glacier.
C.They become the current version.
D.They are moved to Amazon S3 Standard-Infrequent Access.
AnswerA

An S3 Lifecycle expiration rule for noncurrent versions explicitly deletes those versions after the specified number of days, regardless of the storage class. Because the rule only includes an Expiration action and no Transition action, old versions are permanently removed from the bucket rather than being archived or moved. Once permanently deleted, these objects cannot be recovered, so it is critical to have a backup or a separate replication rule if you need to preserve them.

Why this answer

The CloudFormation template configures an S3 bucket with a lifecycle rule that sets 'NoncurrentVersionExpirationInDays' to 30. This rule permanently deletes noncurrent versions of objects after 30 days, as S3 lifecycle policies for noncurrent versions do not transition to other storage classes unless explicitly specified with a separate transition action. After 30 days, the noncurrent versions are removed from the bucket entirely.

Exam trap

The trap here is that candidates may assume noncurrent versions are automatically transitioned to cheaper storage classes like Glacier or S3 Standard-IA, but without explicit transition actions in the lifecycle rule, only expiration (deletion) occurs.

How to eliminate wrong answers

Option B is wrong because the lifecycle rule only specifies 'NoncurrentVersionExpirationInDays' with no 'Transitions' action for noncurrent versions, so objects are not moved to S3 Glacier. Option C is wrong because noncurrent versions cannot become the current version; S3 versioning maintains a distinct current version, and noncurrent versions are older versions that are immutable. Option D is wrong because there is no 'NoncurrentVersionTransition' action defined in the lifecycle rule to move objects to S3 Standard-Infrequent Access; the rule only sets expiration.

98
MCQmedium

A SysOps administrator needs to ensure that an S3 bucket can recover from accidental deletions by users. The bucket stores versioned objects. What additional configuration should be enabled to prevent permanent deletion?

A.Enable S3 Server-Side Encryption.
B.Enable S3 Lifecycle rules to expire objects.
C.Enable MFA Delete on the bucket.
D.Configure a bucket policy to deny s3:DeleteObject.
AnswerC

MFA Delete requires the principal making a destructive request to supply a valid one-time code from a hardware or virtual MFA device, in addition to normal AWS authentication. When applied to a versioned bucket, it protects against permanently deleting an object version and against changing the bucket's versioning state. This means an accidental delete creates a recoverable delete marker, and even compromised AWS credentials cannot irreversibly purge data without the MFA code.

Why this answer

Enabling MFA Delete on the S3 bucket adds an extra layer of protection by requiring multi-factor authentication for any DeleteObject or DeleteBucket operations. Even if a user has s3:DeleteObject permission, they cannot permanently delete versioned objects unless they present a valid MFA code. This prevents accidental or unauthorized permanent deletions while still allowing versioned objects to be recovered.

Exam trap

The trap here is that candidates assume a bucket policy denying s3:DeleteObject is sufficient, but it does not prevent accidental deletion by authorized users who have delete permissions and can simply remove the policy; MFA Delete is the only way to enforce an additional authentication factor for permanent deletions in versioned buckets.

How to eliminate wrong answers

Option A is wrong because S3 Server-Side Encryption protects data at rest from unauthorized access, not from accidental deletion. Option B is wrong because S3 Lifecycle rules to expire objects actually automate the deletion of objects, which increases the risk of permanent deletion rather than preventing it. Option D is wrong because a bucket policy denying s3:DeleteObject would block all delete operations, including the ability to delete non-current versions or markers, which is overly restrictive and does not leverage versioning recovery; it also does not prevent accidental deletion by authorized users who could simply remove the policy.

99
MCQeasy

A company runs a web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application stores session data on local instance storage. Users report that they are unexpectedly logged out during peak traffic. Which action should the SysOps Administrator take to improve reliability?

A.Move the session storage to an instance store volume.
B.Enable sticky sessions on the Application Load Balancer.
C.Increase the size of the Auto Scaling group to handle peak traffic.
D.Configure an ElastiCache Redis cluster to store session state externally.
AnswerD

Configuring an ElastiCache Redis cluster to store session state externally solves the persistence problem by creating a centralized, in-memory data store that every EC2 instance can read from and write to. Because sessions live outside any single instance, any instance in the Auto Scaling group can service any request, making the application effectively stateless from the instance perspective. Redis offers sub-millisecond latency, supports replication and backup/restore for durability, and integrates cleanly with common session-management libraries, so instance terminations or scaling events no longer disrupt active user sessions.

Why this answer

Storing session data on local instance storage is ephemeral; if an instance is terminated or replaced during scaling events, session data is lost, causing users to be logged out. Moving session state to an external, highly available service like ElastiCache Redis ensures persistence across instance lifecycles and improves reliability under peak traffic.

Exam trap

The trap here is that candidates often confuse sticky sessions (session affinity) with session persistence, not realizing that sticky sessions only maintain routing to a specific instance but do not protect against data loss when that instance is terminated.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral and data is lost on instance stop/termination, which would not solve the logout issue. Option B is wrong because sticky sessions (session affinity) only route a user to the same instance, but if that instance is terminated during scaling, the session data is still lost. Option C is wrong because increasing the Auto Scaling group size does not address the root cause—session data loss on instance replacement—and may even increase the frequency of scaling events.

100
MCQmedium

A company uses AWS CloudFormation to deploy its infrastructure. The SysOps administrator needs to ensure that the application stack can be recreated in another AWS Region in the event of a disaster. The stack includes an RDS MySQL database and an EC2 instance running a web server. The administrator wants to automate the backup of the RDS database and the EC2 instance configuration. What is the MOST efficient way to achieve this?

A.Use S3 to store database dump files and instance configuration scripts.
B.Create manual snapshots of the RDS database and EC2 instance every day and copy them to the secondary region.
C.Store the CloudFormation template in S3 and use it to recreate the stack in the secondary region.
D.Use AWS Backup to create backup plans that include the RDS instance and EC2 instance, and copy backups to the secondary region.
AnswerD

AWS Backup provides a fully managed, policy-based backup solution that can target both RDS instances and EC2 instances (via Amazon Machine Images) within a single backup plan. You can schedule automated backups, apply retention and lifecycle policies, and configure cross-region replication to the secondary region, ensuring consistent disaster recovery without custom scripting or manual snapshots. This is the most efficient and reliable approach because it centralizes backup management and automates the entire DR copy process.

Why this answer

AWS Backup provides a centralized, automated way to create backup plans that include both RDS and EC2 instances, and it supports cross-region backup copies. This meets the requirement to automate backups and ensure they are available in a secondary region for disaster recovery. It is the most efficient because it eliminates manual scripting and snapshot management.

Exam trap

SOA-C02 often tests the difference between infrastructure-as-code recreation and data backup; candidates may choose CloudFormation template storage thinking it covers backups, but it does not protect data.

How to eliminate wrong answers

Option A is wrong because using S3 to store database dump files and configuration scripts requires manual or scripted processes, which is not automated and may not ensure consistency. Option B is wrong because creating manual snapshots daily is not automated and is error-prone; it also does not cover EC2 instance configuration in a unified way. Option C is wrong because storing the CloudFormation template in S3 only allows recreating the stack, but it does not automate backups of the RDS database or EC2 instance configuration; it addresses infrastructure recreation, not data backup.

101
MCQmedium

A company has an Amazon DynamoDB table with on-demand capacity mode. The SysOps administrator needs to ensure that the table can survive a regional outage. The table is currently in us-east-1. Which feature should be configured to achieve regional resilience with minimal data loss?

A.DynamoDB Accelerator (DAX)
B.DynamoDB global tables
C.DynamoDB point-in-time recovery
D.DynamoDB auto scaling
AnswerB

DynamoDB global tables replicate your table automatically across multiple AWS Regions using DynamoDB Streams, creating a multi-active, fully managed solution. In the event of a regional outage, applications can read and write to the table in another Region with minimal downtime because each replica is independently accessible. Global tables use last-writer-wins conflict resolution to reconcile concurrent updates, providing eventual consistency and effectively meeting disaster recovery needs with a low RTO and RPO.

Why this answer

DynamoDB global tables provide multi-Region, fully replicated tables that automatically propagate writes to all configured Regions, enabling the table to survive a regional outage with minimal data loss. This feature uses DynamoDB Streams to replicate data asynchronously across Regions, offering recovery point objectives (RPO) of typically under one second. For the requirement of regional resilience, global tables are the correct choice because they maintain active copies in multiple AWS Regions.

Exam trap

The trap here is that candidates often confuse point-in-time recovery (PITR) with cross-Region disaster recovery, not realizing that PITR only protects against accidental deletes or corruption within a single Region, not a full regional outage.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read performance but does not provide any cross-Region replication or regional resilience. Option C is wrong because point-in-time recovery (PITR) enables restoring a table to any point within the last 35 days within the same Region, but it does not protect against a regional outage since the backups are stored in the same Region. Option D is wrong because DynamoDB auto scaling adjusts read/write capacity based on traffic but does not replicate data across Regions or provide any disaster recovery capability.

102
MCQhard

A company uses AWS CloudFormation to deploy a multi-tier application. The stack includes an RDS DB instance with automated backups enabled. The SysOps administrator needs to ensure that the database can be recovered to any point within the last 35 days with minimal data loss. What should the administrator do?

A.Create a manual snapshot daily and retain 35 snapshots.
B.Set the backup retention period to 35 days.
C.Enable Multi-AZ on the RDS instance.
D.Configure AWS Backup with a 35-day backup plan.
AnswerB

Setting the backup retention period to 35 days on the RDS instance enables automated backups that include daily snapshots and transaction logs. These logs allow point-in-time recovery to any second within the retention window, which is exactly what the application needs. RDS supports a maximum retention period of 35 days, so this directly meets the stated requirement without additional tooling.

Why this answer

Automated backups in RDS allow point-in-time recovery (PITR) to any second within the backup retention period. By setting the retention period to 35 days, the administrator enables recovery to any point within that window, minimizing data loss to the last committed transaction before the restore time.

Exam trap

The trap here is that candidates confuse manual snapshots or AWS Backup with the point-in-time recovery capability that only automated backups provide, leading them to choose options that offer only full snapshot recovery rather than granular log-based restore.

How to eliminate wrong answers

Option A is wrong because manual snapshots are not point-in-time recoverable; they capture only the state at the time of creation, so you cannot recover to an arbitrary point between snapshots, leading to potential data loss of up to 24 hours. Option C is wrong because Multi-AZ provides high availability and automatic failover, not point-in-time recovery; it does not extend the backup retention or enable granular restore capabilities. Option D is wrong because AWS Backup can manage RDS snapshots but does not support point-in-time recovery for RDS; it only creates full snapshots, which lack the granularity needed for recovery to any second within 35 days.

103
MCQhard

A company runs a stateful web application on EC2 instances in an Auto Scaling group across two Availability Zones. The application uses an Application Load Balancer for traffic distribution. Users report that their sessions are frequently lost during scale-in events. The SysOps administrator needs to minimize session loss without introducing significant latency. What should the administrator do?

A.Replace the Application Load Balancer with a Network Load Balancer. Enable proxy protocol v2 to pass client IP addresses.
B.Enable sticky sessions (session affinity) on the ALB. Configure a lifecycle hook on the Auto Scaling group with a wait time equal to the ALB's connection draining timeout.
C.Increase the Auto Scaling group's cooldown period to 600 seconds. Configure the ALB to have a deregistration delay of 600 seconds.
D.Configure the Auto Scaling group to scale based on memory utilization instead of CPU. Set the cooldown period to 300 seconds.
AnswerB

Enabling sticky sessions on the ALB uses the AWSALB cookie to pin a client's requests to the same target instance, preserving session state across the fleet. A lifecycle hook on the Auto Scaling group pauses the termination process during scale-in, and setting its wait time to match the ALB's connection draining timeout (deregistration delay) allows in-flight requests to finish before the instance is removed. This coordinated approach ensures that a terminating instance drains gracefully while session continuity is maintained.

Why this answer

Enabling sticky sessions (session affinity) on the ALB ensures that a client's requests are consistently routed to the same EC2 instance, preventing session loss during scale-in. Configuring a lifecycle hook on the Auto Scaling group with a wait time equal to the ALB's connection draining timeout (deregistration delay) allows in-flight requests to complete before the instance is terminated, minimizing session disruption without adding significant latency.

Exam trap

The trap here is that candidates often assume increasing timeouts (cooldown or deregistration delay) alone is sufficient, but without a lifecycle hook to coordinate the Auto Scaling group with the ALB's draining process, instances can be terminated prematurely, causing session loss.

How to eliminate wrong answers

Option A is wrong because replacing the ALB with a Network Load Balancer (NLB) does not provide session affinity (sticky sessions) natively; NLB operates at Layer 4 and cannot inspect HTTP session cookies, so it would not prevent session loss during scale-in. Option C is wrong because increasing the cooldown period to 600 seconds only delays subsequent scaling activities, not the termination of instances during scale-in, and the deregistration delay on the ALB alone does not coordinate with the Auto Scaling group to hold instances; without a lifecycle hook, instances can be terminated while still handling active sessions. Option D is wrong because changing the scaling metric to memory utilization and setting a cooldown period does not address session persistence or graceful instance termination; it only alters when scaling occurs, not how sessions are maintained during scale-in.

104
MCQhard

A company has a production application running on Amazon ECS with Fargate launch type. The application uses an Application Load Balancer. The SysOps administrator notices that during deployments, the application experiences a brief period of downtime. Which combination of actions should the administrator take to achieve zero-downtime deployments?

A.Configure the ECS service to use a rolling update with a minimum healthy percent of 0 and a maximum percent of 100.
B.Increase the deregistration delay on the ALB target group to 300 seconds.
C.Use a blue/green deployment with CodeDeploy and set the 'Minimum healthy percent' to 50.
D.Configure the ECS service to use a rolling update with a minimum healthy percent of 100 and a maximum percent of 200.
AnswerD

This rolling update configuration ensures that ECS starts new tasks up to twice the desired count (maximum percent 200) before terminating any existing tasks, while the minimum healthy percent of 100 guarantees that the service never drops below the desired number of running tasks. Because new tasks are registered with the load balancer and pass health checks before old tasks are stopped, traffic is served continuously without interruption. This is the classic zero-downtime deployment strategy for ECS services behind an ALB.

Why this answer

Setting the minimum healthy percent to 100 and maximum percent to 200 ensures that during a rolling update, the ECS service first launches new tasks (up to 200% of the desired count) before terminating any old tasks. This guarantees that the ALB always has a sufficient number of healthy targets to serve traffic, eliminating downtime. The Application Load Balancer distributes traffic between old and new tasks during the transition, achieving zero-downtime deployments.

Exam trap

The trap here is that candidates often confuse the minimum healthy percent and maximum percent values, mistakenly thinking that allowing all tasks to be replaced at once (0/100) is acceptable, or that blue/green deployments inherently guarantee zero downtime without proper configuration.

How to eliminate wrong answers

Option A is wrong because setting minimum healthy percent to 0 and maximum percent to 100 allows all existing tasks to be terminated before new ones are started, causing a period with zero healthy targets and thus downtime. Option B is wrong because increasing the deregistration delay to 300 seconds only affects how long the ALB waits before removing a target that is deregistering; it does not prevent the underlying issue of insufficient healthy targets during the update. Option C is wrong because blue/green deployments with CodeDeploy and a minimum healthy percent of 50 still allow up to half of the targets to be unhealthy during the transition, which can cause downtime if the ALB’s health checks fail; moreover, blue/green deployments typically require a full set of new targets before switching, but the 50% setting contradicts that goal.

105
Multi-Selectmedium

A company is designing a backup strategy for its on-premises file servers to AWS. Which TWO services can be used to back up data to AWS? (Choose TWO.)

Select 2 answers
A.AWS Backup
B.AWS Snowball
C.AWS Storage Gateway (File Gateway)
D.Amazon EFS
E.S3 Transfer Acceleration
AnswersA, C

AWS Backup is the correct answer because it natively supports backing up on-premises workloads via the AWS Backup Gateway, which connects your on-premises virtual machines to AWS Backup. This service allows you to define backup policies, retention rules, and lifecycle management in a single place, covering both cloud and on-premises resources. Unlike simple data replication or file syncing tools, AWS Backup provides a centralized, scheduled, and auditable backup solution that ensures recoverability of on-premises VMs.

Why this answer

AWS Backup is correct because it provides a fully managed, policy-based backup service that can centrally automate and manage backups for on-premises file servers via the AWS Backup Gateway (formerly Storage Gateway Virtual Tape Library). It integrates with AWS Storage Gateway to back up on-premises data to S3 and Glacier, supporting file-level recovery without needing custom scripts.

Exam trap

The trap here is that candidates confuse data transport services (Snowball) or storage targets (EFS) with backup services, or mistake a performance feature (S3 Transfer Acceleration) for a backup solution, when the question specifically asks for services that can be used to back up data to AWS.

106
MCQeasy

A SysOps administrator is configuring an Amazon RDS for MySQL Multi-AZ deployment. What is the primary benefit of using Multi-AZ?

A.Improved read performance by distributing queries across multiple instances.
B.Automatic failover to a standby instance in a different Availability Zone.
C.Synchronous replication across AWS Regions.
D.Automatic creation of read replicas for disaster recovery.
AnswerB

This is the core benefit of a Multi-AZ Amazon RDS deployment: the primary DB instance synchronously replicates data to a standby instance in a different Availability Zone within the same Region. If the primary fails, RDS automatically detects the issue and promotes the standby to primary, updating the DNS endpoint so applications can reconnect with minimal downtime. This provides high availability without requiring manual intervention.

Why this answer

In Amazon RDS for MySQL Multi-AZ deployments, the primary benefit is automatic failover to a standby instance in a different Availability Zone. This is achieved through synchronous replication to a standby in a separate AZ, ensuring that if the primary instance fails, RDS automatically promotes the standby to become the new primary, minimizing downtime and maintaining data durability.

Exam trap

The trap here is that candidates confuse Multi-AZ with read replicas, assuming Multi-AZ provides read scaling, when in fact it is solely for high availability and automatic failover.

How to eliminate wrong answers

Option A is wrong because Multi-AZ does not improve read performance; read traffic is always directed to the primary instance, and the standby is not used for serving reads. Option C is wrong because Multi-AZ replication is within a single AWS Region, not across Regions; cross-Region replication is handled by Aurora Global Database or manual read replicas. Option D is wrong because Multi-AZ does not automatically create read replicas; read replicas are a separate feature for offloading read traffic and are not part of the Multi-AZ failover mechanism.

107
MCQmedium

A company has an S3 bucket that stores critical data. The bucket has versioning enabled. A SysOps administrator accidentally deletes a version of an object. What is the quickest way to recover the deleted version?

A.Use the S3 bucket's 'Undelete' feature.
B.Enable MFA Delete and then restore the object.
C.Contact AWS Support to restore the object.
D.Copy the deleted version from the bucket's version history.
AnswerD

With S3 Versioning enabled, deleting an object does not erase the data; it simply creates a delete marker that becomes the current version while the original object remains as a noncurrent version. To restore it, copy the most recent noncurrent (deleted) version back into the bucket, which removes the delete marker and makes that version current again. This self-service method is the standard and reliable way to recover a deleted object.

Why this answer

S3 object versioning maintains a version history for each object, including deleted versions. When a version is deleted, it is not permanently removed; instead, a delete marker is created, and the deleted version remains in the version history. The quickest way to recover it is to copy the deleted version ID from the bucket's version history (e.g., using the AWS CLI `aws s3api copy-object` with the `--version-id` parameter) to restore the object without contacting support or enabling additional features.

Exam trap

The trap here is that candidates may think S3 has an 'Undelete' feature (Option A) or that MFA Delete (Option B) can reverse deletions, but S3's versioning model requires manual recovery via version history, not a built-in undo function.

How to eliminate wrong answers

Option A is wrong because S3 does not have an 'Undelete' feature; deleted versions are recovered by removing the delete marker or copying from version history. Option B is wrong because MFA Delete is a security feature that requires multi-factor authentication for permanent deletions, but it does not provide a recovery mechanism for already-deleted versions; enabling it after deletion does not restore the object. Option C is wrong because AWS Support cannot restore deleted S3 object versions; recovery is entirely self-service through the S3 versioning feature.

108
MCQeasy

A company runs a stateless web application on EC2 instances in an Auto Scaling group. The instances are behind an Application Load Balancer. The Auto Scaling group uses a dynamic scaling policy based on average CPU utilization. During a traffic spike, new instances are launched but take 5 minutes to become healthy. Users experience errors during this time. Which solution would reduce the time to serve traffic from new instances?

A.Use a launch template with a pre-provisioned AMI.
B.Add a lifecycle hook to delay instance termination.
C.Increase the cooldown period for the scaling policy.
D.Use a larger instance type.
AnswerA

A pre-provisioned (pre-warmed) AMI has the application binary, runtime, and all dependencies already installed and initialized during image creation. When the Auto Scaling group launches a new instance from a launch template using this AMI, the app is immediately ready to serve traffic, avoiding the need for first-boot user-data scripts to install and configure software. This directly reduces the instance's "time to healthy" and is the correct way to speed up scaling in a stateless web tier.

Why this answer

A pre-provisioned AMI eliminates the need for software installation and configuration at launch time, reducing the time for new EC2 instances to become healthy. By baking the application and dependencies into the AMI, instances can start serving traffic almost immediately after booting, rather than waiting for user data scripts or configuration management tools to complete.

Exam trap

The trap here is that candidates may think increasing the cooldown period (Option C) helps stabilize scaling, but it actually delays the launch of new instances, making the problem worse during traffic spikes.

How to eliminate wrong answers

Option B is wrong because a lifecycle hook to delay instance termination would only affect instances being terminated, not newly launched instances, and would actually increase the time before instances become healthy. Option C is wrong because increasing the cooldown period would prevent the Auto Scaling group from launching new instances quickly during a traffic spike, worsening the problem. Option D is wrong because using a larger instance type does not reduce the time to serve traffic from new instances; it only provides more resources per instance, but the boot and configuration time remains the same.

109
MCQhard

A company runs a critical application on a single Amazon EC2 instance with an attached Amazon EBS volume. The SysOps administrator needs to implement a disaster recovery solution that meets a Recovery Point Objective (RPO) of 15 minutes and a Recovery Time Objective (RTO) of 30 minutes. The application runs continuously and data changes frequently. Which solution should the administrator implement?

A.Use Amazon Data Lifecycle Manager (DLM) to take EBS snapshots every 15 minutes and automate the creation of a new AMI.
B.Use AWS Backup to schedule backups every 15 minutes and restore from the latest backup when needed.
C.Use AWS Elastic Disaster Recovery (AWS DRS) to continuously replicate the instance to a staging area in another region.
D.Use an Auto Scaling group with a custom AMI that is updated every 15 minutes by a Lambda function.
AnswerC

AWS Elastic Disaster Recovery (AWS DRS) continuously replicates block-level changes from the source EC2 instance to a staging area in a second region using a lightweight agent, with typical RPO in the seconds and RTO in the single-digit minutes. DRS maintains a converted, continuously updated copy of the source volumes on low-cost staging instances, so at failover time it simply powers on the target instance using the latest replicated state rather than constructing it from snapshots. This always-ready, continuous replication approach is precisely what is required to meet a 30-minute recovery window while ensuring data loss is limited to a few seconds.

Why this answer

AWS Elastic Disaster Recovery (AWS DRS) continuously replicates the entire EC2 instance, including the EBS volume, to a staging area in another AWS Region with sub-second data changes. This meets the RPO of 15 minutes and RTO of 30 minutes because you can launch a fully recovered instance in the target region within minutes from the latest consistent point, without relying on periodic snapshots or backups that would miss frequent data changes.

Exam trap

The trap here is that candidates often choose periodic snapshot or backup solutions (like DLM or AWS Backup) because they think 15-minute intervals satisfy the RPO, but they overlook the RTO constraint and the fact that frequent data changes require continuous replication, not periodic snapshots, to avoid data loss between intervals.

How to eliminate wrong answers

Option A is wrong because Amazon Data Lifecycle Manager (DLM) can take EBS snapshots every 15 minutes, but creating a new AMI from those snapshots is not automated by DLM and the process would take longer than 30 minutes to build and register an AMI, failing the RTO. Option B is wrong because AWS Backup scheduled backups every 15 minutes still rely on periodic snapshots, which cannot capture every data change between intervals, and restoring from the latest backup can take longer than 30 minutes due to volume creation and attachment time. Option D is wrong because an Auto Scaling group with a custom AMI updated every 15 minutes by a Lambda function does not provide continuous replication; the AMI creation process itself takes time and the instance launched from an older AMI would miss data changes made in the interim, failing the RPO.

110
MCQhard

A company runs a critical application on EC2 instances in an Auto Scaling group with a minimum of 2 instances. The instances are in a single Availability Zone. The company wants to achieve 99.99% availability. Which change should they make?

A.Modify the Auto Scaling group to launch instances in multiple Availability Zones and place an Application Load Balancer in front.
B.Increase the minimum size of the Auto Scaling group to 4 instances.
C.Use a larger EC2 instance type to handle more traffic.
D.Configure the Auto Scaling group to scale based on memory utilization.
AnswerA

By distributing instances across multiple Availability Zones and using an Application Load Balancer to route traffic, the application can withstand an entire AZ failure because the remaining AZs continue serving. The ALB performs health checks and only forwards requests to healthy registered targets, automatically shifting load away from failed AZs. This is the core high-availability design for EC2-based architectures.

Why this answer

To achieve 99.99% availability, the application must survive an Availability Zone (AZ) failure. Running instances in a single AZ creates a single point of failure. By modifying the Auto Scaling group to launch instances in multiple AZs and placing an Application Load Balancer (ALB) in front, traffic is automatically distributed across healthy instances in different AZs, ensuring fault tolerance even if an entire AZ becomes unavailable.

Exam trap

The trap here is that candidates often focus on increasing instance count or scaling metrics, overlooking that 99.99% availability requires geographic redundancy across Availability Zones, not just more instances in a single zone.

How to eliminate wrong answers

Option B is wrong because simply increasing the minimum size to 4 instances within the same single AZ does not protect against an AZ outage; all instances would still fail if that AZ goes down. Option C is wrong because using a larger EC2 instance type only improves performance and capacity, not availability; it does not address the risk of an AZ failure. Option D is wrong because scaling based on memory utilization helps with performance and cost optimization but does not provide redundancy across AZs; it cannot mitigate an AZ-level failure.

111
Multi-Selecthard

A company uses Amazon S3 to store backup data. The SysOps administrator needs to ensure that the data is protected against accidental deletion by users with administrative privileges. Which combination of actions should the administrator take? (Choose TWO.)

Select 2 answers
A.Enable MFA Delete on the S3 bucket.
B.Apply an S3 bucket policy that denies s3:DeleteObject for all users.
C.Enable versioning on the S3 bucket.
D.Configure a lifecycle policy to transition objects to S3 Glacier.
E.Enable AWS CloudTrail to log all S3 API calls.
AnswersA, C

MFA Delete is the correct safeguard for this scenario because it forces any request that permanently deletes an object version, or that suspends versioning on the bucket, to include a valid code from a hardware or virtual MFA device. This prevents an attacker who has stolen console credentials or an IAM access key from irrevocably erasing backup data, since they would also need possession of the MFA token. It is important to note that MFA Delete can only be enabled when versioning is turned on, and it must be set via the AWS CLI or API rather than the console.

Why this answer

Option A is correct because MFA Delete adds an additional authentication factor requirement for permanently deleting object versions or changing the versioning state of the bucket, which specifically protects against accidental deletion even by users with administrative privileges. Option C is correct because enabling versioning ensures that overwritten or deleted objects are retained as noncurrent versions, allowing recovery of data that would otherwise be lost. Together, versioning preserves deleted objects and MFA Delete prevents an administrator from permanently removing them or disabling versioning without an MFA token.

Option B is not appropriate because a blanket deny of s3:DeleteObject would break legitimate deletion workflows and can be bypassed or modified by users with administrative privileges who can edit bucket policies. Option D is incorrect because lifecycle transitions to S3 Glacier only change storage class and do not protect against deletion. Option E is incorrect because CloudTrail only records API activity for auditing; it does not prevent or recover from accidental deletion.

Exam trap

SOA-C02 often tests the misconception that bucket policies or CloudTrail can prevent deletion, when in fact only MFA Delete and versioning provide protection against accidental deletion by privileged users.

112
MCQmedium

A company runs a stateless web application on EC2 instances in an Auto Scaling group across multiple Availability Zones. The application experiences increased latency during peak hours. The SysOps administrator needs to improve the application's performance and reliability. Which action should be taken?

A.Use larger EC2 instance types instead of smaller ones.
B.Reduce the Auto Scaling group's cooldown period to scale out faster.
C.Change the scaling metric from CPU utilization to memory utilization.
D.Increase the maximum instance count in the Auto Scaling group.
AnswerD

Increasing the maximum instance count in the Auto Scaling group directly expands the group's capacity ceiling, allowing it to launch additional EC2 instances when demand spikes. This is the appropriate solution for a stateless web application because horizontal scaling distributes traffic across multiple instances and provides both elastic capacity and fault tolerance. With a properly configured scaling policy, the group can now add instances beyond its previous limit, accommodating higher traffic volumes without modifying the application.

Why this answer

Increasing the maximum instance count allows the Auto Scaling group to launch more EC2 instances during peak hours, distributing the load across more resources and reducing latency. This directly improves both performance (by handling more requests) and reliability (by maintaining capacity under increased demand).

Exam trap

The trap here is that candidates often confuse scaling metrics or instance sizing with the fundamental need to increase the maximum capacity limit when the group is already hitting its cap during peak load.

How to eliminate wrong answers

Option A is wrong because using larger instance types may improve per-instance performance but does not inherently increase the total capacity to handle peak load; it also reduces granularity for scaling and can increase costs without addressing the need for more instances. Option B is wrong because reducing the cooldown period can cause rapid, unstable scaling (thrashing) and does not solve the underlying capacity shortage; it may lead to unnecessary scaling actions and increased costs. Option C is wrong because memory utilization is not a suitable scaling metric for a stateless web application where CPU utilization directly reflects request processing load; memory usage remains relatively stable and would not trigger timely scaling during CPU-bound latency spikes.

113
Multi-Selectmedium

A SysOps administrator is responsible for an Auto Scaling group that runs a critical application. The administrator wants to ensure that the application can recover from an AZ failure. Which THREE steps should the administrator take? (Choose three.)

Select 3 answers
A.Use EC2 instances in a single Availability Zone to reduce latency.
B.Place subnets in each Availability Zone used by the Auto Scaling group.
C.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
D.Attach an Application Load Balancer that is enabled for multiple Availability Zones.
E.Use a single subnet in one Availability Zone to simplify network design.
AnswersB, C, D

An Auto Scaling group defines which subnets it can launch instances into, and each subnet maps to exactly one Availability Zone. To spread instances across multiple Availability Zones, you must explicitly configure the ASG to include subnets from each of those zones. Without a corresponding subnet in a given AZ, the group cannot place instances there, so omitting subnets limits the group's ability to recover from zone-level failures.

Why this answer

Placing subnets in each Availability Zone (AZ) used by the Auto Scaling group allows the group to launch EC2 instances in multiple AZs, which is essential for fault isolation. If one AZ fails, the Auto Scaling group can still launch and maintain instances in the other AZs, ensuring application availability. Without subnets in each AZ, the Auto Scaling group cannot distribute instances across AZs, defeating the purpose of high availability.

Exam trap

The trap here is that candidates often think using a single AZ simplifies management and reduces latency, but they overlook that this creates a critical single point of failure, which is unacceptable for a critical application requiring AZ failure recovery.

114
MCQhard

Refer to the exhibit. A SysOps administrator needs to restore the database 'mydb' to the most recent restorable time shown. However, the administrator cannot restore to that time. What is the MOST likely reason?

A.The database engine does not support point-in-time recovery.
B.Automated backups are disabled (BackupRetentionPeriod is 0).
C.The backup window has already passed.
D.The database is not Multi-AZ.
AnswerB

With BackupRetentionPeriod set to 0, automated backups are completely disabled, meaning RDS never takes daily snapshots and does not retain transaction logs for PITR. Without any automated backup or retained log, there is no recovery point available to restore from, which directly explains why the restore fails. Manual snapshots could still be used if they were created separately, but the scenario indicates no usable backup exists for restoration.

Why this answer

Automated backups must be enabled with a BackupRetentionPeriod greater than 0 for point-in-time recovery (PITR) to be available. When BackupRetentionPeriod is set to 0, automated backups are disabled, and the database cannot be restored to any point in time within the retention window. The exhibit shows that the most recent restorable time is not available because no automated backups exist to support PITR.

Exam trap

The trap here is that candidates may assume the most recent restorable time is always available or confuse the backup window with the ability to perform PITR, when in fact the root cause is that automated backups are disabled entirely.

How to eliminate wrong answers

Option A is wrong because all supported RDS database engines (MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, and Aurora) support point-in-time recovery when automated backups are enabled. Option C is wrong because the backup window defines when automated backups are taken, but it does not prevent restoring to the most recent restorable time; the most recent restorable time is determined by the last successful backup and transaction logs, not by whether the backup window has passed. Option D is wrong because Multi-AZ deployment is not a prerequisite for point-in-time recovery; PITR works on single-AZ instances as long as automated backups are enabled.

115
MCQeasy

A company has an S3 bucket that stores critical financial data. The bucket versioning is enabled. A SysOps administrator needs to ensure that data can be recovered after accidental deletion by users. What is the MOST effective way to protect against accidental deletion?

A.Configure a lifecycle policy to transition objects to Glacier after 30 days.
B.Apply a bucket policy that denies s3:DeleteObject for all users.
C.Replicate objects to another S3 bucket in a different AWS Region.
D.Enable MFA Delete on the S3 bucket.
AnswerD

MFA Delete, when enabled on a versioned S3 bucket, requires the bucket owner to present the root credential plus a valid MFA code for two operations: changing the bucket's versioning state and permanently deleting object versions. An ordinary DeleteObject request without a version ID only creates a delete marker, while any attempt to hard-delete a specific version without the MFA token fails. This prevents an authorized but careless user from irretrievably removing critical financial data.

Why this answer

MFA Delete adds an additional authentication factor (a hardware or virtual MFA device) that must be provided to permanently delete an object version or to change the versioning state of the bucket. This makes accidental or malicious deletion significantly harder because a compromised access key alone is insufficient. It is the most effective control specifically designed to protect versioned S3 data from deletion.

Exam trap

SOA-C02 often tests the misconception that bucket policies or cross-region replication prevent accidental deletion, when in fact only MFA Delete (or Object Lock) provides a strong, version-level deletion protection mechanism for versioned buckets.

How to eliminate wrong answers

Option A is wrong because a lifecycle policy transitions objects to Glacier for cost optimization; it does not prevent deletion and may even delete objects if configured with expiration actions. Option B is wrong because a bucket policy denying s3:DeleteObject can be modified or removed by an administrator with sufficient permissions, and it does not protect against accidental deletion by users who have policy-editing rights; it also does not address version deletion. Option C is wrong because cross-region replication provides durability and disaster recovery but does not prevent deletion in the source bucket — if an object is deleted in the source, replication can propagate the deletion to the destination depending on configuration.

116
MCQmedium

Refer to the exhibit. A SysOps administrator creates an IAM policy to allow an EC2 instance to upload objects to an S3 bucket. However, the instance is unable to upload objects. What is the MOST likely reason?

A.The S3 bucket has server-side encryption enabled.
B.The policy does not include s3:GetObject permission.
C.The bucket policy denies all access.
D.The IAM role is not attached to the EC2 instance.
AnswerD

An EC2 instance can only use IAM permissions if an instance profile containing a role is attached at launch time or later. Without an attached role, the instance has no AWS credentials to sign API requests, so any S3 operation (including upload) fails with an error such as 'Unable to locate credentials' or an access denied error because the request is not authenticated with an authorized IAM identity. The role must be attached to the EC2 instance and the instance must have the necessary permissions in its trust and permissions policies. This is the root cause of the upload failure.

Why this answer

The IAM role must be attached to the EC2 instance as an instance profile for the instance to assume the role and obtain temporary credentials. Without this attachment, the instance has no valid AWS credentials to sign API requests, so the s3:PutObject action will fail regardless of the permissions defined in the role's policy.

Exam trap

The trap here is that candidates often assume the IAM policy alone is sufficient, forgetting that the EC2 instance must have a mechanism (the instance profile) to assume the role and obtain credentials.

How to eliminate wrong answers

Option A is wrong because server-side encryption (SSE) does not inherently block uploads; the instance can still upload objects if it has the correct permissions and the bucket policy does not explicitly deny access. Option B is wrong because the policy only needs s3:PutObject to upload objects; s3:GetObject is required for reading, not writing. Option C is wrong because the question states the policy allows uploads, and a bucket policy that denies all access would be an explicit deny, but the most likely reason given the scenario is the missing attachment of the IAM role to the instance.

117
Multi-Selectmedium

A company runs a stateless web application on EC2 instances in an Auto Scaling group. To improve reliability during a traffic spike, which THREE actions should the SysOps administrator take? (Choose three.)

Select 3 answers
A.Configure a target tracking scaling policy based on average CPU utilization.
B.Enable detailed monitoring for EC2 instances.
C.Use a larger instance type to handle more traffic per instance.
D.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
E.Place the instances behind an Application Load Balancer with health checks.
AnswersA, D, E

A target tracking scaling policy automatically adjusts the Auto Scaling group's desired capacity based on a CloudWatch metric, such as average CPU utilization, to keep the metric close to a specified target value. For a stateless web app with varying traffic, this provides true elasticity, scaling out to handle spikes and scaling in when demand falls, without manual intervention. Because it directly maps load to capacity, it is the most effective single choice among the listed options.

Why this answer

A target tracking scaling policy based on average CPU utilization allows the Auto Scaling group to automatically adjust the number of EC2 instances in response to traffic spikes. This policy maintains the target metric (e.g., 50% CPU) by adding or removing instances, ensuring the application remains responsive without manual intervention. It is a key mechanism for improving reliability under variable load.

Exam trap

The trap here is that candidates often confuse detailed monitoring (which only improves metric granularity) with a direct reliability improvement, or they mistakenly believe that scaling up (larger instances) is equivalent to scaling out (more instances) for fault tolerance.

118
MCQhard

A company runs a critical stateful web application on Amazon EC2 instances in a single AWS region. The application stores user session data in an Amazon ElastiCache for Redis cluster. The SysOps administrator must design a disaster recovery (DR) strategy that can survive a complete regional outage with a Recovery Point Objective (RPO) of 15 minutes and a Recovery Time Objective (RTO) of 1 hour. The application must be able to redirect users to the DR region with minimal manual effort. Which combination of actions meets these requirements?

A.Use Amazon Route 53 with weighted routing to distribute traffic between the two regions. Use a global DynamoDB table for session data, and launch EC2 instances in the DR region only when a failure is detected using AWS CloudFormation StackSets.
B.Create a read replica of the ElastiCache Redis cluster in the DR region using the native cross-region replication feature. Use Route 53 with failover routing to point to the DR region ALB when the primary health check fails. Pre-configure EC2 instances in an Auto Scaling group in the DR region.
C.Use an Amazon CloudFront distribution with multiple origins (primary and DR). Enable session stickiness at the CloudFront level. Use EC2 instances in both regions behind separate ALBs. No special data replication is needed because sessions are stored in Redis.
D.Use EC2 instances with an Auto Scaling group in both regions. Schedule a Lambda function to take snapshots of the Redis cluster every 15 minutes and copy them to the DR region. Use Route 53 latency routing to direct users to the nearest region.
AnswerB

Global Datastore for Redis provides cross-Region replication with low RPO. Pre-configured Auto Scaling groups in the DR region ensure that compute capacity is ready. Route 53 failover routing automatically redirects traffic when the primary ALB health check fails. This combination meets the RPO and RTO requirements with minimal manual effort.

Why this answer

ElastiCache for Redis supports cross-region replication via a read replica in the DR region, which can keep session data synchronized with minimal lag, meeting the 15-minute RPO. Route 53 failover routing with health checks on the primary region's ALB automatically redirects traffic to the pre-configured DR region EC2 instances and ALB, achieving the 1-hour RTO with minimal manual effort. Pre-configuring the DR region with an Auto Scaling group ensures compute capacity is ready, while the read replica provides the required data availability.

Exam trap

The trap here is that candidates may assume snapshot-based replication (Option D) is sufficient for a 15-minute RPO, but they overlook the inherent latency and potential data loss from periodic snapshots, and that latency routing (Option D) does not provide health-based failover, while weighted routing (Option A) lacks automatic failover capability.

How to eliminate wrong answers

Option A is wrong because weighted routing does not automatically fail over during a regional outage; it distributes traffic based on weights, not health, and using a global DynamoDB table for session data is unnecessary since the application uses ElastiCache for Redis, not DynamoDB. Option C is wrong because CloudFront does not natively support session stickiness based on ElastiCache session data, and without cross-region replication of Redis, the DR region would have no session data, violating the RPO. Option D is wrong because scheduling snapshots every 15 minutes and copying them to the DR region cannot guarantee an RPO of 15 minutes due to snapshot timing and transfer delays, and latency routing does not provide automatic failover during a regional outage; it routes based on latency, not health.

119
MCQmedium

A company runs a stateless web application on Amazon EC2 instances in an Auto Scaling group across two Availability Zones. The SysOps administrator needs to ensure that the application can tolerate a failure of an entire Availability Zone. Which configuration is required?

A.Use an Application Load Balancer (ALB) that spans both Availability Zones with health checks enabled.
B.Enable termination protection on all Amazon EC2 instances.
C.Place the Amazon EC2 instances in a cluster placement group.
D.Associate an Elastic IP address with the primary instance.
AnswerA

An Application Load Balancer (ALB) is a regional service that spans all Availability Zones (AZs) in its subnet configuration and actively sends health-check requests to each registered target. When an EC2 instance or an entire AZ fails health checks, the ALB automatically stops routing new traffic to that target and continues serving requests from healthy instances in other AZs. Coupled with an Auto Scaling group that spans multiple AZs, this design provides both elasticity and zone-failure tolerance, because the ALB constantly updates its target membership based on instance health and scaling events.

Why this answer

An Application Load Balancer (ALB) that spans both Availability Zones with health checks enabled distributes incoming traffic across EC2 instances in multiple AZs. If an entire AZ fails, the ALB automatically routes traffic only to healthy instances in the remaining AZ, ensuring the stateless web application remains available. Health checks detect instance or AZ failure and remove unhealthy targets from the load balancer's target group, which is essential for fault tolerance.

Exam trap

The trap here is that candidates often confuse high availability with data durability or instance protection, leading them to choose termination protection or Elastic IPs, when the core requirement is automatic traffic rerouting across AZs, which only a load balancer with health checks can provide.

How to eliminate wrong answers

Option B is wrong because termination protection prevents accidental deletion of an instance but does not provide any resilience against an Availability Zone failure; it does not reroute traffic or maintain application availability. Option C is wrong because a cluster placement group is designed for low-latency, high-throughput networking within a single AZ; it actually increases the risk of simultaneous failure if that AZ goes down, as all instances are in the same AZ. Option D is wrong because associating an Elastic IP with the primary instance only provides a static public IP, which does not survive an AZ failure and does not offer automatic failover or load balancing across AZs.

120
MCQmedium

A company is running a web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application stores session data on the local instance storage. Users are experiencing session loss during scaling events. What should a SysOps administrator do to maintain session persistence?

A.Move session data to an ElastiCache for Redis cluster.
B.Increase the EC2 instance size to reduce the frequency of scaling events.
C.Attach an Amazon EBS volume to each instance and store session data there.
D.Enable sticky sessions on the Application Load Balancer.
AnswerA

Moving session state to ElastiCache for Redis decouples user session data from the ephemeral lifecycle of EC2 instances. Because Redis is a centralized, network-accessible store, any instance in the Auto Scaling group can read and write the same session data, and the data survives instance termination or replacement (especially with Redis persistence/AOF). This enables the Application Load Balancer to route requests to any healthy instance, eliminating the single point of failure tied to local storage.

Why this answer

Session data stored on local instance storage is ephemeral and lost when instances are terminated or replaced during scaling events. Moving session data to ElastiCache for Redis provides a centralized, durable, and low-latency session store that persists independently of EC2 instance lifecycles, ensuring session continuity across scaling operations.

Exam trap

The trap here is that candidates often confuse sticky sessions (which maintain request routing) with session persistence (which maintains session data), leading them to choose option D without realizing that sticky sessions do not protect against data loss when the target instance is terminated.

How to eliminate wrong answers

Option B is wrong because increasing instance size does not eliminate session loss; it only reduces the frequency of scaling events but does not address the fundamental issue of ephemeral storage being lost on instance termination. Option C is wrong because attaching an EBS volume to each instance still ties session data to individual instances; if an instance is terminated during scale-in, the EBS volume is detached and the session data is lost unless the volume is manually reattached, which is not automated and defeats the purpose of Auto Scaling. Option D is wrong because sticky sessions (session affinity) only route subsequent requests from the same user to the same instance, but they do not preserve session data if that instance is terminated; the session data on local storage is still lost when the instance is replaced.

121
MCQhard

A company runs a critical application on AWS Lambda functions. The functions are invoked by an API Gateway endpoint. The SysOps administrator needs to ensure that the application continues to work if an entire AWS Region becomes unavailable. What should the administrator do?

A.Use AWS Global Accelerator to route traffic to the closest Region.
B.Configure Lambda functions with provisioned concurrency in multiple Regions.
C.Use Lambda@Edge to run the functions at edge locations.
D.Deploy the same API Gateway and Lambda setup in a second Region and use Route 53 with failover routing.
AnswerD

Deploying an identical API Gateway and Lambda setup in a second Region and using Route 53 failover routing with health checks on the primary endpoint provides an active-passive disaster recovery pattern. When the health check fails, Route 53 automatically returns the secondary Region's IP address, directing clients to the standby stack. This approach ensures cross-Region availability and is the standard solution for making a Lambda-based critical application resilient to an entire Region outage.

Why this answer

Deploying the same API Gateway and Lambda setup in a second AWS Region and using Route 53 with failover routing creates an active-passive disaster recovery architecture. Route 53 health checks monitor the primary Region's endpoint, and if it becomes unhealthy (e.g., due to a regional outage), DNS failover automatically routes traffic to the secondary Region, ensuring continuous operation of the application.

Exam trap

The trap here is that candidates often confuse high-availability features like Global Accelerator or provisioned concurrency with true disaster recovery across Regions, failing to recognize that only Route 53 failover routing provides the DNS-level traffic redirection needed when an entire Region becomes unavailable.

How to eliminate wrong answers

Option A is wrong because AWS Global Accelerator improves performance and availability by routing traffic to the closest healthy endpoint within a single Region or across multiple Regions, but it does not provide automatic failover to a completely separate Region if the entire primary Region becomes unavailable; it relies on existing endpoints in the same Region. Option B is wrong because configuring Lambda functions with provisioned concurrency in multiple Regions initializes the functions to reduce cold starts, but it does not include the necessary API Gateway endpoints or DNS-based routing to redirect traffic if the primary Region fails. Option C is wrong because Lambda@Edge runs functions at CloudFront edge locations, which are designed for lightweight request/response modifications and cannot host the full application logic or API Gateway integration required for this critical application; it also does not provide regional failover.

122
MCQhard

An EC2 instance runs a database on a 2 TB EBS gp3 volume. After a corruption event, the team must restore from a snapshot. When they detach the corrupted volume, attach a new volume restored from the snapshot, and start the database, performance is 10 to 20 times lower than normal for the first two hours. What causes this behavior, and what feature eliminates it?

A.Enable Fast Snapshot Restore (FSR) on the snapshot in the target Availability Zone before creating the replacement volume
B.Use a Provisioned IOPS (io2) volume type instead of gp3 to get higher IOPS during initialization
C.Run a full dd or fio pre-warm pass over the volume after attaching it but before starting the database
D.Increase the EBS volume size to 4 TB when restoring from the snapshot to get double the throughput baseline
AnswerA

FSR fully initializes the volume's block index immediately upon creation. The first I/O to any block is served from EBS at full throughput rather than waiting for lazy initialization from S3. For a 2 TB database volume where I/O latency determines restore time, FSR eliminates the 2-hour performance degradation period entirely.

Why this answer

When you create an EBS volume from a snapshot, the volume's data blocks are lazily loaded from Amazon S3 on first access. This causes high latency and low IOPS until all blocks are fetched. Fast Snapshot Restore (FSR) pre-initializes the volume in a specific Availability Zone, eliminating the need for lazy loading and providing full performance immediately.

Exam trap

The trap here is that candidates assume performance issues are due to volume type (gp3 vs io2) or size, rather than recognizing the fundamental lazy-load initialization behavior of EBS snapshots and the specific feature (FSR) designed to mitigate it.

How to eliminate wrong answers

Option B is wrong because Provisioned IOPS (io2) volumes do not eliminate the lazy-load initialization penalty; they only provide consistent IOPS after the volume is fully initialized, but the initial access still suffers from the same on-demand fetch from S3. Option C is wrong because running dd or fio pre-warms the volume manually, but this is a workaround, not a feature that eliminates the behavior, and it still requires the same time-consuming initialization process. Option D is wrong because increasing the volume size to 4 TB does not change the lazy-load behavior; it only increases the baseline throughput for the volume after initialization, but the initial performance degradation remains until all blocks are loaded.

123
MCQmedium

A company runs a production application on Amazon EC2 instances in an Auto Scaling group across two Availability Zones. The application uses an Amazon RDS Multi-AZ DB instance. The SysOps administrator wants to test the application's behavior during an Availability Zone failure of the database. Which action should the administrator take to simulate a failure with minimal impact on production?

A.Reboot the DB instance with the 'Reboot with failover' option
B.Modify the DB instance to be a single-AZ deployment
C.Delete the standby replica in the other Availability Zone
D.Stop the DB instance
AnswerA

Choosing 'Reboot with failover' instructs Amazon RDS to perform a graceful, forced failover from the primary to the standby replica in the other Availability Zone. This simulates an AZ outage or primary instance failure, letting you validate that your application reconnects and continues operating after RDS promotes the standby and updates the DNS endpoint. The reboot causes only a short interruption while the failover completes, making it the correct way to test resilience without manual infrastructure changes.

Why this answer

Rebooting the RDS Multi-AZ DB instance with the 'Reboot with failover' option forces a synchronous failover to the standby replica in the other Availability Zone. This simulates an AZ failure of the primary database with minimal impact because the application's Auto Scaling group spans two AZs and the RDS Multi-AZ deployment provides automatic failover, so the application should experience only a brief interruption during the DNS change to the new primary.

Exam trap

The trap here is that candidates may think stopping or deleting the standby replica simulates an AZ failure, but those actions either cause a full outage or permanently remove redundancy, whereas 'Reboot with failover' is the only option that triggers a controlled failover with minimal production impact.

How to eliminate wrong answers

Option B is wrong because modifying the DB instance to be a single-AZ deployment permanently removes the standby replica and changes the architecture, which does not simulate a transient AZ failure and has a greater impact on production. Option C is wrong because deleting the standby replica in the other AZ is a destructive action that removes high availability entirely, and it does not simulate a failover event; it also requires manual intervention to recreate the standby. Option D is wrong because stopping the DB instance halts the database completely, causing a full outage rather than a controlled failover, and it does not test the application's behavior during an AZ failure of the database.

124
Multi-Selectmedium

A company is designing a disaster recovery strategy for its production database hosted on Amazon RDS for MySQL. The primary database is in us-east-1. The company requires an RPO of less than 5 minutes and an RTO of less than 1 hour in the event of a Regional failure. Which TWO actions should the company take to meet these requirements?

Select 2 answers
A.Take daily manual snapshots and copy them to us-west-2.
B.Create a cross-Region Read Replica in us-west-2 and promote it during a disaster.
C.Enable cross-Region automated backups.
D.Enable Multi-AZ deployment for the RDS instance.
E.Use a single-AZ RDS instance with automated backups.
AnswersB, C

A cross-Region Read Replica uses Amazon RDS's asynchronous replication to continuously stream transactions from the primary DB instance in the source Region to a read replica in us-west-2. In a disaster, you can promote the replica to a standalone primary instance within minutes, and the replication lag is typically well under 5 minutes when the network is healthy, meeting the RPO requirement. This gives you a warm standby database in the target Region, enabling fast failover without needing to restore from backups.

Why this answer

A cross-Region Read Replica in us-west-2 can be promoted to a standalone primary database during a disaster, enabling failover with an RTO typically under 1 hour. The replication lag is usually seconds to a few minutes, meeting the RPO of less than 5 minutes. This approach provides a warm standby in another Region without requiring manual snapshot restores.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which only protects against AZ failures within the same Region) with cross-Region disaster recovery, leading them to incorrectly select Option D as a solution for Regional failures.

125
MCQeasy

A company runs a web application on Amazon EC2 instances in an Auto Scaling group. The application stores session state locally on each instance, so users lose their sessions when an instance is replaced. The SysOps administrator needs to make the application stateless so that instances can be replaced without disrupting users. Which action should the administrator take?

A.Configure the Auto Scaling group to use a termination policy that terminates the oldest instance.
B.Increase the instance size to reduce the likelihood of replacement.
C.Store session state in an Amazon ElastiCache for Redis cluster.
D.Enable sticky sessions on the Application Load Balancer.
AnswerC

Externalizing session state to ElastiCache for Redis removes the dependency on local instance storage, allowing any instance in the Auto Scaling group to handle any user's request. When an instance is replaced, the session data remains available in the Redis cluster. This makes the application stateless and supports seamless scaling and recovery.

Why this answer

Moving session state to a centralized, highly available store such as ElastiCache for Redis decouples the application from individual instances. Any instance can then serve any user, so instance replacement or scaling does not cause session loss. This is the standard approach to making a stateful web application stateless in a dynamic Auto Scaling environment.

Exam trap

The trap here is thinking that sticky sessions solve session loss, when they only delay the problem until the bound instance is replaced or fails.

126
Multi-Selecthard

A company is designing a disaster recovery strategy for its AWS environment. The primary Region is us-east-1, and the secondary Region is us-west-2. The application uses Amazon RDS for MySQL, Amazon S3 for static assets, and EC2 instances in an Auto Scaling group. The RTO is 30 minutes, and the RPO is 15 minutes. Which TWO actions should the SysOps administrator take to meet these requirements? (Choose two.)

Select 2 answers
A.Configure a cross-Region read replica for the RDS instance in us-west-2.
B.Deploy a single RDS instance in us-west-2 as a standby.
C.Create AMI backups of EC2 instances every hour and copy to us-west-2.
D.Take daily EBS snapshots and copy them to us-west-2.
E.Enable S3 cross-Region replication for the static assets bucket.
AnswersA, E

A cross-Region read replica in us-west-2 maintains an asynchronous copy of the primary RDS database using the engine's native replication. In a disaster, you can promote the replica to a standalone primary DB instance within minutes, which meets a 15-minute RPO because replication lag is typically well below that threshold and the RTO is minimal. The replica also supports read traffic in the DR Region before promotion, making it a cost-effective and operationally proven pattern for RDS disaster recovery.

Why this answer

A cross-Region read replica in us-west-2 can be promoted to a primary instance in under 30 minutes, and with binary log (binlog) replication it can achieve an RPO of 15 minutes or less. This allows the RDS database to be recovered with minimal data loss in the secondary Region, meeting both the RTO and RPO targets.

Exam trap

The trap here is that candidates often confuse cross-Region read replicas with Multi-AZ deployments, assuming a standby in another Region is sufficient, but Multi-AZ only provides high availability within a single Region and does not support cross-Region disaster recovery with the required RPO.

127
MCQeasy

A company runs a static website on Amazon S3 with a custom domain name (www.example.com). The website is accessed via Amazon CloudFront. The company's marketing team recently updated the website content, but users are reporting that they still see the old content. The SysOps administrator checks the S3 bucket and confirms that the new files are present. The administrator also checks CloudFront and finds that the default TTL for the cache behavior is 24 hours. The marketing team needs the new content to be visible immediately. What should the administrator do to make the new content available to users as quickly as possible?

A.Disable the CloudFront distribution and re-enable it after 5 minutes.
B.Change the default TTL for the CloudFront cache behavior to 0 seconds.
C.Create a CloudFront invalidation for the path '/*' to remove all cached files.
D.Change the S3 bucket's lifecycle policy to expire objects after 1 day.
AnswerC

Creating an invalidation for the path '/*' issues a request that removes all cached objects from every edge location in the CloudFront distribution. When the invalidation completes, the next request for any of those objects causes CloudFront to go back to the S3 origin and fetch the latest version, thereby ensuring users see updated content. This is the immediate, targeted mechanism designed for exactly this scenario.

Why this answer

CloudFront caches objects at edge locations according to the cache behavior's TTL settings; changing the TTL only affects future cache fills, not objects already cached. To immediately remove stale content from all edge locations, the administrator must create an invalidation for the affected paths — here '/*' invalidates everything. This forces CloudFront to fetch fresh copies from the S3 origin on the next request, making the new content visible right away.

Exam trap

SOA-C02 often tests the misconception that lowering the TTL or restarting the distribution refreshes already-cached content — candidates must remember that only an invalidation (or versioned object keys) evicts objects already stored at edge locations.

How to eliminate wrong answers

Option A is wrong because disabling and re-enabling a distribution does not purge cached objects — the edge caches retain their content, and the distribution also takes time to redeploy, so users would still see stale files. Option B is wrong because setting the default TTL to 0 only affects objects cached after the change; it does not evict objects already stored at edge locations, so existing users continue to receive the old content. Option D is wrong because an S3 lifecycle expiration policy deletes objects after a retention period and has no effect on CloudFront's edge cache or on how quickly updated content propagates.

128
MCQmedium

A company runs a stateful web application on EC2 instances behind a Network Load Balancer. The application requires that client requests from a particular session are always sent to the same target instance. Which feature should the SysOps administrator configure on the NLB to meet this requirement?

A.Configure path-based routing rules
B.Enable health checks on the target group
C.Configure sticky sessions using flow-based routing (client IP affinity)
D.Enable cross-zone load balancing
AnswerC

Sticky sessions on an NLB are implemented through flow-based routing, also known as client IP affinity, which uses a consistent hash of the connection's protocol, source IP, source port, destination IP, and destination port to route all flows from a client to the same target. Unlike ALBs, NLBs cannot rely on cookie-based stickiness because they operate at Layer 4, so this affinity mechanism is the correct way to ensure a stateful web application's user requests consistently hit the same EC2 instance, preserving session state.

Why this answer

Network Load Balancers (NLBs) support session stickiness through flow-based routing based on the client's IP address and port (Layer 4). This ensures that requests from the same client session are consistently routed to the same target instance, meeting the requirement without using application-layer cookies.

Exam trap

The trap here is that candidates often confuse NLB features with ALB features, assuming path-based routing or HTTP-level cookies apply to NLBs, when in fact NLBs operate at Layer 4 and use flow-based stickiness rather than application-layer session persistence.

How to eliminate wrong answers

Option A is wrong because path-based routing rules are a feature of Application Load Balancers (ALBs), not Network Load Balancers (NLBs); NLBs route traffic based on TCP/UDP/TLS protocols and do not inspect HTTP paths. Option B is wrong because health checks determine target availability but do not influence session persistence; they only mark unhealthy targets as out of service. Option D is wrong because cross-zone load balancing distributes traffic evenly across targets in all Availability Zones, which can actually break session affinity by sending requests to different targets; it does not provide session stickiness.

129
MCQeasy

A company wants to ensure that its S3 bucket is accessible only from a VPC. Which configuration should the SysOps Administrator implement?

A.Create an S3 VPC endpoint and attach a bucket policy that restricts access to that endpoint.
B.Configure a bucket policy that allows access from the public internet.
C.Make the bucket public and rely on IAM roles.
D.Attach a security group to the S3 bucket.
AnswerA

Creating an S3 VPC gateway endpoint gives your VPC private, routable connectivity to S3 without traversing the public internet. To actually enforce that restriction, the bucket policy must include a condition such as "aws:SourceVpce" or "aws:SourceVpc" so that only requests originating from that endpoint or VPC are allowed; otherwise, the endpoint alone does not block other network paths. This is the only option that both enables private access and explicitly limits the source network to your VPC.

Why this answer

An S3 VPC endpoint (either gateway or interface type) allows private connectivity between a VPC and S3 without traversing the public internet. By attaching a bucket policy that includes a condition like `aws:SourceVpce` or `aws:SourceVpc`, access is explicitly restricted to traffic originating from that specific VPC endpoint, ensuring the bucket is not accessible from any other network.

Exam trap

The trap here is that candidates may think security groups can be applied to S3 buckets (Option D) because they are familiar with security groups for EC2, but S3 operates at the service level and uses bucket policies and IAM for access control, not security groups.

How to eliminate wrong answers

Option B is wrong because allowing access from the public internet would make the bucket accessible from anywhere, defeating the requirement to restrict access to a VPC. Option C is wrong because making the bucket public and relying on IAM roles still exposes the bucket to the public internet; IAM roles control who can make requests but do not restrict network-level access. Option D is wrong because security groups are a network access control mechanism for EC2 instances and other resources within a VPC, but S3 buckets are not VPC resources and cannot have security groups attached to them.

130
MCQhard

A company runs a critical microservices application on Amazon ECS with Fargate launch type. The application consists of several services that communicate via internal HTTP calls. The SysOps Administrator notices that during periods of increased load, some services become unresponsive and the health checks fail. The ECS service auto scaling is configured based on CPU utilization, but it does not scale quickly enough. The administrator needs to improve the reliability and responsiveness of the application. The services are stateless and can be scaled horizontally. The current architecture uses a single Application Load Balancer for each service. The ALB health checks are set to a 30-second interval with a 5-second timeout and 2 unhealthy thresholds. The administrator has observed that when a service instance becomes unhealthy, it takes too long for the ALB to stop sending traffic to it, causing errors. What should the SysOps Administrator do to improve the reliability and responsiveness of the application?

A.Increase the ALB health check interval to 60 seconds and unhealthy threshold to 5.
B.Configure the ALB health check to have a 5-second interval, 2-second timeout, and 2 unhealthy threshold.
C.Increase the ECS service auto scaling target CPU utilization to 90%.
D.Replace the ALB with a Network Load Balancer and use TCP health checks.
AnswerB

Setting the ALB health check interval to 5 seconds, timeout to 2 seconds, and unhealthy threshold to 2 is the fastest configuration AWS allows; a target that fails two consecutive checks is deregistered within roughly 10 seconds, so ECS can quickly stop the unhealthy task and start a replacement. The 2-second timeout ensures a hung or unresponsive process is detected promptly, and because the unhealthy threshold is 2, transient blips are still somewhat filtered while maintaining minimal delay.

Why this answer

Reducing the health check interval to 5 seconds and timeout to 2 seconds with an unhealthy threshold of 2 allows the ALB to detect service failures much faster. With the original settings (30s interval, 5s timeout, 2 threshold), detection could take up to 65 seconds. With the new settings, detection time is reduced to about 12 seconds, so traffic is stopped sooner and errors are minimized.

Option A is wrong because increasing the interval and threshold would make detection even slower. Option C is wrong because increasing the CPU target delays scaling, which does not address health check responsiveness. Option D is wrong because a Network Load Balancer uses TCP health checks which are less application-aware and may not detect HTTP-level failures, and it does not inherently speed up health check detection.

131
MCQhard

A company runs a production workload on a fleet of EC2 instances in an Auto Scaling group (ASG). The ASG spans three Availability Zones. To avoid regional failure, the company wants to replicate the infrastructure in a second AWS Region and be able to fail over within 30 minutes. The application state is stored in an RDS MySQL database. What is the MOST cost-effective and reliable solution?

A.Create a cross-Region read replica of the RDS database. In the secondary Region, deploy a duplicate ASG and ALB. In a disaster, promote the read replica to a standalone instance and update Route 53 DNS.
B.Use an Application Load Balancer with cross-Region load balancing to distribute traffic to both Regions.
C.Use RDS Multi-AZ in both Regions and configure synchronous replication between them.
D.Take daily snapshots of the RDS database and copy them to the secondary Region. In the event of a failure, restore the latest snapshot.
AnswerA

A cross-Region read replica uses asynchronous replication from the primary RDS instance to a secondary Region, keeping the replica typically within seconds of the source. In a disaster, you can promote the replica to a standalone primary database in minutes, minimizing RPO to near zero, and then repoint Route 53 to the new region's ALB. Having a duplicate ASG and ALB pre-provisioned ensures your application layer is ready to accept traffic immediately after the DNS switch.

Why this answer

It provides a cost-effective and reliable disaster recovery solution that meets the 30-minute failover requirement. A cross-Region read replica of RDS MySQL allows you to maintain an up-to-date copy of the database in the secondary Region with minimal cost (only paying for the replica instance and data transfer). In a disaster, you can promote the read replica to a standalone instance in minutes, and with a pre-deployed ASG and ALG in the secondary Region, you can update Route 53 DNS to redirect traffic, achieving failover within the required timeframe.

Exam trap

The trap here is that candidates often confuse RDS Multi-AZ with cross-Region replication, not realizing that Multi-AZ is a single-Region feature for high availability, while cross-Region read replicas are the correct solution for disaster recovery across Regions.

How to eliminate wrong answers

Option B is wrong because Application Load Balancers do not support cross-Region load balancing; ALBs are regional services and cannot distribute traffic across multiple AWS Regions. Option C is wrong because RDS Multi-AZ is designed for high availability within a single Region, not for cross-Region replication; synchronous replication across Regions would introduce unacceptable latency and is not supported by RDS Multi-AZ. Option D is wrong because taking daily snapshots and restoring them in a disaster would result in up to 24 hours of data loss (RPO) and a restore time that likely exceeds the 30-minute RTO, making it neither reliable nor fast enough for the stated requirements.

132
MCQhard

A SysOps administrator runs the above command for an EC2 instance. The instance is running but the system status check is impaired. What does this indicate?

A.The instance is still running but the application is not responding.
B.The instance is unreachable due to a misconfigured security group.
C.There is a problem with the underlying physical host that requires stopping and starting the instance.
D.The operating system on the instance has crashed.
AnswerC

The 'system status check failed' event indicates that AWS has detected a problem with the physical host that cannot be resolved by the instance or the guest OS. Common causes include loss of system power, loss of network connectivity, or hardware degradation on the host. Because the instance is tied to that host, a stop and start operation forces AWS to provision the instance on a fresh, healthy host, which is the documented remediation for a failed system status check.

Why this answer

The system status check in AWS EC2 monitors the underlying physical host for issues such as loss of network connectivity, power loss, or hardware failure. When this check is impaired, it indicates a problem with the host that requires stopping and starting the instance to migrate it to a new healthy host. Option C is correct because stopping and starting the instance forces a migration to a different physical host, resolving the underlying hardware issue.

Exam trap

The trap here is that candidates confuse system status checks (host-level) with instance status checks (guest-level), leading them to incorrectly attribute the failure to OS or application issues rather than the underlying physical host.

How to eliminate wrong answers

Option A is wrong because a system status check failure does not indicate application-level unresponsiveness; that would be detected by an instance status check, which monitors the guest OS and application. Option B is wrong because a misconfigured security group would cause network connectivity issues but would not affect the system status check, which tests the health of the physical host. Option D is wrong because an OS crash would be detected by the instance status check (e.g., failed system log or impaired guest OS), not by the system status check, which focuses on the underlying host.

133
MCQmedium

An application uses an Amazon DynamoDB table with on-demand capacity. The SysOps administrator needs to ensure the table remains available during an AWS regional outage. Which strategy should be used?

A.Enable DynamoDB Accelerator (DAX).
B.Create a read replica in another region.
C.Use DynamoDB global tables.
D.Increase read and write capacity units.
AnswerC

Global tables are DynamoDB's multi-Region replication feature, providing active-active copies of a table in up to six AWS Regions. When enabled, all data mutations are automatically propagated to every replica, and each replica supports both reads and writes with conflict resolution, so application traffic can be failed over to a healthy Region. This availability and automatic synchronization directly address the requirement to survive a regional outage.

Why this answer

DynamoDB global tables provide multi-region, multi-active replication, ensuring the table remains available during an AWS regional outage by automatically replicating data across selected AWS Regions. This is the only option that addresses regional fault tolerance by design, as it uses DynamoDB's built-in replication to maintain availability and data durability across regions.

Exam trap

The trap here is that candidates often confuse read replicas (an RDS concept) with DynamoDB's global tables, or assume that DAX or scaling capacity can provide regional resilience, when in fact only global tables offer multi-region active-active replication for DynamoDB.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read performance but operates within a single region and does not provide any cross-region availability or disaster recovery. Option B is wrong because DynamoDB does not support read replicas in the traditional RDS sense; the correct multi-region replication feature is global tables, not read replicas. Option D is wrong because increasing read and write capacity units (even with on-demand scaling) only affects performance within a single region and cannot protect against a regional outage.

134
MCQeasy

A company runs a stateless web application on EC2 instances in an Auto Scaling group. The application is deployed across multiple Availability Zones. The SysOps administrator wants to ensure that the application remains available even if an entire Availability Zone fails. What is the MOST effective way to achieve this?

A.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
B.Create a CloudWatch alarm to reboot instances when they become unhealthy.
C.Use a single Availability Zone to reduce complexity.
D.Use a larger instance type to handle more traffic.
AnswerA

Configuring the Auto Scaling group to span at least two Availability Zones distributes your stateless web application's instances across independent failure domains. If an entire AZ becomes unavailable—due to power loss, cooling failure, or networking issues—the remaining healthy AZs continue to serve traffic, and the Auto Scaling group automatically replaces failed instances in the healthy zones. This is the core mechanism for high availability, because it removes any single data-center-level dependency and lets a load balancer route only to surviving instances.

Why this answer

By configuring the Auto Scaling group to launch instances in at least two Availability Zones, the application can survive the failure of an entire AZ because the Auto Scaling group will automatically replace failed instances in the remaining healthy AZs. This design ensures that the stateless web application remains available as long as at least one AZ is operational, leveraging the fault isolation that AWS Availability Zones provide. The Auto Scaling group distributes instances across the specified AZs and will maintain the desired capacity even if one AZ becomes completely unavailable.

Exam trap

The trap here is that candidates often confuse instance-level recovery mechanisms (like rebooting or replacing a single failed instance) with AZ-level fault tolerance, leading them to choose options that address individual instance health rather than the architectural redundancy required for AZ failure scenarios.

How to eliminate wrong answers

Option B is wrong because a CloudWatch alarm to reboot instances only addresses individual instance failures, not the failure of an entire Availability Zone; if the AZ itself fails, all instances in that AZ become unreachable and rebooting them does not restore availability. Option C is wrong because using a single Availability Zone creates a single point of failure; if that AZ fails, the entire application becomes unavailable, which directly contradicts the goal of remaining available during an AZ failure. Option D is wrong because using a larger instance type to handle more traffic does not provide any fault tolerance or redundancy; it only increases capacity within a single AZ and does not protect against an AZ-wide outage.

135
Multi-Selectmedium

A company runs a web application on EC2 instances in an Auto Scaling group behind an ALB. The application uses an RDS MySQL database. The SysOps administrator needs to improve the reliability of the database layer. Which TWO actions should the administrator take? (Choose two.)

Select 2 answers
A.Enable Multi-AZ on the RDS instance.
B.Take a manual snapshot every hour.
C.Create a read replica in a different Region.
D.Configure automated backups with a retention period of 30 days.
E.Increase the DB instance class to the largest available.
AnswersA, D

Enable Multi-AZ on the RDS instance: Multi-AZ creates a synchronous standby replica in a different Availability Zone. When the primary DB instance fails or its Availability Zone becomes unavailable, Amazon RDS automatically flips the DNS record to the standby, providing rapid, automated failover. This directly addresses the high availability requirement for the web application's database layer by eliminating a single point of failure.

Why this answer

Enabling Multi-AZ on an RDS MySQL instance automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary instance fails, Amazon RDS automatically fails over to the standby, providing high availability and improving database reliability without manual intervention. This directly addresses the need for a resilient database layer.

Exam trap

The trap here is confusing backup strategies (automated backups, manual snapshots) or read replicas with high-availability features like Multi-AZ, leading candidates to select options that improve data durability or read performance instead of database layer reliability.

136
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer (ALB) across three Availability Zones. The application stores session data in memory on the EC2 instances. During a deployment, a new version of the application is released by terminating and replacing instances. Users report that they are unexpectedly logged out during the deployment. What should a SysOps administrator do to improve the reliability of the application during deployments?

A.Reduce the deployment to a single Availability Zone to minimize instance churn.
B.Store session data in an RDS Multi-AZ database.
C.Enable sticky sessions (session affinity) on the ALB.
D.Use an ElastiCache cluster to store session state externally.
AnswerD

An ElastiCache cluster, particularly using Redis, provides a dedicated in-memory data store that can serve as a centralized session repository external to the EC2 instances. Because the session state is decoupled from any individual compute resource, a replacement instance launched by Auto Scaling or a new deployment can immediately read the same session data from the cache, ensuring seamless user continuity. ElastiCache supports native TTL expiration for sessions and can be configured with Multi-AZ replication to provide high availability for the session data itself.

Why this answer

Storing session state externally in an ElastiCache cluster decouples session data from individual EC2 instances. When instances are terminated and replaced during deployment, the new instances can retrieve session state from the shared ElastiCache cluster, preventing users from being logged out. This approach ensures that session data persists independently of the EC2 instance lifecycle, maintaining application reliability during rolling updates.

Exam trap

The trap here is that candidates often confuse sticky sessions (which only route traffic to the same instance) with external session storage, failing to realize that sticky sessions do not preserve session data when the instance itself is terminated.

How to eliminate wrong answers

Option A is wrong because reducing to a single Availability Zone eliminates fault tolerance and increases the risk of downtime, which contradicts reliability goals. Option B is wrong because RDS Multi-AZ is designed for relational database high availability, not for low-latency session state storage; using a database for session data introduces unnecessary overhead and latency compared to an in-memory cache. Option C is wrong because sticky sessions (session affinity) tie a user's session to a specific EC2 instance; when that instance is terminated during deployment, the session data is lost, causing users to be logged out.

137
Multi-Selecthard

A SysOps administrator is designing a highly available architecture for a web application using an Application Load Balancer and an Auto Scaling group across three Availability Zones. The application must be able to withstand the loss of an entire AZ. Which THREE components are necessary to meet this requirement? (Choose THREE.)

Select 3 answers
A.Use a single NAT gateway to provide internet access.
B.Launch EC2 instances in at least two Availability Zones.
C.Configure health checks on the ALB target group.
D.Use a cluster placement group for EC2 instances.
E.Enable cross-zone load balancing on the ALB.
AnswersB, C, E

Launching EC2 instances in at least two Availability Zones satisfies the requirement to withstand the loss of an entire AZ because the Auto Scaling group distributes instances across those zones. If one AZ fails, the load balancer routes traffic only to healthy instances in the remaining zones, ensuring continued application availability. This directly meets the stem’s constraint of surviving a full AZ outage.

Why this answer

Launching EC2 instances in at least two Availability Zones (AZs) ensures that if one AZ fails, the Auto Scaling group can still serve traffic from instances in the remaining AZs. This is a fundamental requirement for high availability, as an Auto Scaling group spanning multiple AZs can automatically replace failed instances in other zones. Without multi-AZ deployment, a single AZ failure would cause complete application downtime.

Exam trap

The trap here is that candidates often confuse a single NAT gateway with high availability, not realizing that a NAT gateway is AZ-specific and requires one per AZ for fault tolerance, or they mistakenly think a cluster placement group improves availability when it actually concentrates instances into a single failure domain.

138
Multi-Selectmedium

A company runs a stateless web application on EC2 instances behind an Application Load Balancer. The company wants to improve the application's availability and fault tolerance. Which TWO actions should the SysOps administrator take?

Select 2 answers
A.Configure Auto Scaling to maintain a minimum number of instances.
B.Use Amazon CloudFront as an origin for the ALB.
C.Deploy EC2 instances across multiple Availability Zones.
D.Disable termination protection on EC2 instances.
E.Use larger EC2 instance types.
AnswersA, C

Auto Scaling with a minimum instance count is the core corrective mechanism for this scenario: it continuously monitors instance health and replaces any failed instance by launching a new one in its place, while also maintaining the desired capacity across healthy instances. Without a minimum, a single instance failure or an Availability Zone disturbance can drop the fleet to zero and take the stateless web application offline. The auto scaling group also integrates with Elastic Load Balancing to stop sending traffic to unhealthy instances, reinforcing both self-healing and capacity preservation.

Why this answer

Auto Scaling can maintain a minimum number of EC2 instances, ensuring that if an instance fails, a replacement is automatically launched to keep the application running. This directly improves availability by providing automatic recovery from instance failures without manual intervention.

Exam trap

The trap here is that candidates often confuse performance optimization (CloudFront) or capacity scaling (larger instances) with fault tolerance, when the correct approach is to distribute workloads across multiple failure domains and enable automatic recovery.

139
MCQmedium

A company runs a web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application stores session data locally on each instance. During a traffic spike, the Auto Scaling group launches new instances, but users report that they are logged out and lose session data. Which solution addresses this issue without modifying the application?

A.Modify the application to use ElastiCache for session storage.
B.Enable sticky sessions (session affinity) on the Application Load Balancer.
C.Increase the cooldown period for the Auto Scaling group.
D.Use larger EC2 instance types to handle the traffic spike.
AnswerB

Enabling sticky sessions (session affinity) on the Application Load Balancer routes requests from a particular client to the same EC2 instance for the duration of the session. This preserves the session data as long as that instance remains healthy and within the Auto Scaling group, even while the group scales out. It is a configuration-only change to the load balancer, requiring no application code modifications, making it the correct solution.

Why this answer

The application stores session data locally on each EC2 instance, so when new instances are launched during a traffic spike, the load balancer may route a user's subsequent request to a different instance that does not have their session data, causing them to be logged out. Enabling sticky sessions (session affinity) on the Application Load Balancer ensures that all requests from a user during a session are sent to the same instance, preserving the locally stored session data without requiring any application modifications.

Exam trap

The trap here is that candidates often assume that scaling out (adding more instances) or scaling up (using larger instances) will solve session persistence issues, but they overlook that the real problem is the lack of a shared session store or a mechanism to pin users to the same instance, which sticky sessions directly address without code changes.

How to eliminate wrong answers

Option A is wrong because modifying the application to use ElastiCache for session storage would require code changes, which contradicts the requirement to not modify the application. Option C is wrong because increasing the cooldown period for the Auto Scaling group only delays the launch of new instances, but does not solve the session data loss when requests are routed to different instances. Option D is wrong because using larger EC2 instance types may help handle the traffic spike but does not address the fundamental issue of session data being stored locally and lost when requests are distributed across multiple instances.

140
MCQhard

A company has an Amazon RDS for PostgreSQL DB instance with Multi-AZ deployment in us-east-1. The SysOps administrator must design a disaster recovery strategy to recover from a regional outage. The Recovery Time Objective (RTO) is 1 hour and the Recovery Point Objective (RPO) is 5 minutes. Which solution meets these requirements at the lowest cost?

A.Create a Read Replica in a different region and promote it during a disaster.
B.Take daily snapshots and copy them to another region.
C.Use cross-region automated backups.
D.Deploy a second Multi-AZ DB instance in another region.
AnswerA

A cross-region Read Replica for Amazon RDS for PostgreSQL uses asynchronous streaming replication, with typical lag anywhere from a few seconds to several minutes — well within the required 5-minute RPO. During a regional disaster, promoting the replica transitions it to a standalone writable instance in minutes to tens of minutes, comfortably meeting the 1-hour RTO. It is also the most cost-effective option because you pay only for a single replica instance and its storage, and you can use it for read traffic before disaster, unlike a dedicated standby.

Why this answer

A cross-region Read Replica meets the RPO of 5 minutes because replication is continuous (asynchronous) with minimal lag, and the RTO of 1 hour is achievable by promoting the replica during a disaster. This is the lowest-cost option because it uses a single standby instance in another region without the overhead of a full Multi-AZ deployment or frequent snapshot transfers.

Exam trap

The trap here is that candidates confuse 'cross-region automated backups' (which do not exist as a native feature) with automated snapshot copying, or assume that daily snapshots can meet a 5-minute RPO by increasing snapshot frequency, ignoring the fundamental limitation of snapshot scheduling and transfer time.

How to eliminate wrong answers

Option B is wrong because daily snapshots cannot achieve an RPO of 5 minutes (snapshots are taken at most every 24 hours, and copying to another region adds latency). Option C is wrong because cross-region automated backups are not a native RDS feature; automated backups are region-specific and cannot be automatically copied to another region without manual or scripted snapshot copy operations. Option D is wrong because deploying a second Multi-AZ DB instance in another region incurs the cost of a full primary and standby pair, which is significantly more expensive than a single Read Replica, and does not provide a faster RTO/RPO than a promoted Read Replica.

141
Multi-Selecthard

A company runs a web application on EC2 instances behind an Application Load Balancer. The instances are in an Auto Scaling group. The SysOps administrator wants to ensure that the application can handle a sudden increase in traffic without downtime. Which THREE actions should be taken?

Select 3 answers
A.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
B.Configure a target tracking scaling policy based on the ALB's RequestCountPerTarget metric.
C.Configure the Auto Scaling group with a dynamic scaling policy.
D.Configure a scheduled scaling policy to add instances during known peak hours.
E.Use Spot Instances to reduce costs.
AnswersA, B, C

Placing the Auto Scaling group across multiple Availability Zones distributes the EC2 instances between isolated data centers, so the application remains available even if one entire AZ becomes unavailable. The ALB also performs cross-zone load balancing, sending traffic only to healthy instances in the remaining AZs. This configuration is primarily a high-availability measure, not a scaling action, but it underpins the group's ability to replace lost capacity without redirecting all traffic to a single location.

Why this answer

Launching instances in multiple Availability Zones (AZs) ensures high availability and fault tolerance. If one AZ experiences a failure, the Auto Scaling group can still serve traffic from instances in other AZs, preventing downtime during sudden traffic spikes or infrastructure issues.

Exam trap

The trap here is that candidates often confuse scheduled scaling (for predictable patterns) with dynamic scaling (for unpredictable spikes), and they may overlook that Spot Instances are unsuitable for workloads requiring high availability and no downtime.

142
MCQmedium

A SysOps administrator is designing a disaster recovery strategy for a critical application that runs on EC2 instances. The application data is stored on EBS volumes. The recovery point objective (RPO) is 15 minutes, and the recovery time objective (RTO) is 1 hour. Which solution meets these requirements MOST cost-effectively?

A.Use AWS CloudEndure to continuously replicate the EC2 instances to another Region.
B.Use AWS Backup to back up the application data to Amazon S3 every 15 minutes.
C.Take hourly AMIs of the instances and copy them to another Region.
D.Take EBS snapshots every 15 minutes and copy them to another Region using cross-region snapshot copy.
AnswerD

EBS snapshots are block-level, incremental backups of your EC2 volumes, so taking them every 15 minutes captures only the changed blocks since the previous snapshot, keeping storage costs low while meeting the 15-minute RPO. Using cross-region snapshot copy asynchronously replicates those snapshots to the DR Region, ensuring you have recent recoverable point-in-time data there without needing to keep duplicate running instances. On failover, you can create new EBS volumes from the latest snapshot, attach them to pre-provisioned or restored EC2 instances, and start the application—comfortably within a 1-hour RTO because snapshot-to-volume creation is fast and the infrastructure can be pre-staged as a stopped AMI or launch template.

Why this answer

Taking EBS snapshots every 15 minutes and using cross-region snapshot copy meets the 15-minute RPO and 1-hour RTO while being the most cost-effective. EBS snapshots are incremental, storing only changed blocks, which minimizes storage costs compared to full AMIs. Cross-region snapshot copy ensures data is available in another Region for recovery within the RTO.

Exam trap

The trap here is that candidates often choose hourly AMIs (Option C) thinking they are faster to restore, but they fail to recognize that AMIs include full volume data and are taken less frequently, missing the 15-minute RPO and costing more due to full copies rather than incremental snapshots.

How to eliminate wrong answers

Option A is wrong because AWS CloudEndure (now AWS Application Migration Service) continuously replicates entire servers, which incurs high licensing and infrastructure costs, making it overkill for an RPO of 15 minutes and RTO of 1 hour. Option B is wrong because AWS Backup to Amazon S3 every 15 minutes does not natively support EBS volume restoration for EC2 instances; it backs up data to S3, not as EBS snapshots, and restoring to a bootable volume would exceed the 1-hour RTO. Option C is wrong because hourly AMIs are too infrequent to meet the 15-minute RPO, and copying full AMIs to another Region incurs higher storage and transfer costs compared to incremental snapshots.

143
MCQmedium

A SysOps administrator is troubleshooting an issue where an Auto Scaling group is not launching EC2 instances despite having a scaling policy that should trigger when CPU utilization exceeds 80%. The CloudWatch alarm shows that the metric is breaching the threshold, but no instances are launched. What is the most likely cause?

A.The scaling policy is incorrectly configured to use a simple scaling policy instead of a step scaling policy.
B.The health check grace period is too long.
C.The Auto Scaling group has reached its maximum size.
D.The CloudWatch alarm is in insufficient data state.
AnswerC

The Auto Scaling group has reached its maximum size, which is a hard limit that prevents any scale-out activity from adding new instances. When a scale-out policy is triggered while the group is already at MaxSize, the scaling process simply skips or fails the launch because it would exceed the group's maximum capacity. This is the most direct explanation for the alarm breaching but no new instances appearing.

Why this answer

When an Auto Scaling group reaches its maximum size, it cannot launch new instances even if a scaling policy is triggered. The CloudWatch alarm breaching the threshold indicates the scaling condition is met, but the group's capacity limit prevents any new instance launches. This is a common misconfiguration where the max size is set too low relative to the desired or current capacity.

Exam trap

The trap here is that candidates often focus on the scaling policy type or alarm state, overlooking the fundamental capacity constraint of the Auto Scaling group's maximum size.

How to eliminate wrong answers

Option A is wrong because both simple and step scaling policies can trigger instance launches; the policy type affects how the scaling adjustment is applied (e.g., step scaling allows more granular adjustments based on metric breach size), but neither prevents launch execution. Option B is wrong because the health check grace period only delays the Auto Scaling group from replacing an instance that fails health checks after launch; it does not block new launches triggered by a scaling policy. Option D is wrong because the question explicitly states the CloudWatch alarm is breaching the threshold, meaning it is in ALARM state, not insufficient data state.

144
MCQmedium

A company has an AWS Lambda function that processes S3 events. The function is critical and must be available even if one Availability Zone fails. How can a SysOps administrator ensure high availability for the Lambda function?

A.Use an Application Load Balancer to distribute events to multiple Lambda functions.
B.No action is required; Lambda functions are inherently highly available within a region.
C.Configure the Lambda function to run in two subnets in different Availability Zones.
D.Deploy the Lambda function in two separate regions and use Route 53 failover.
AnswerB

No configuration is required because AWS Lambda is a regional, highly available service that executes function code in multiple Availability Zones automatically. When an S3 event triggers a Lambda function asynchronously, the event is queued and the Lambda service manages capacity, scaling, and redundant infrastructure to ensure the invocation can be processed successfully. The function is also covered by the Lambda service SLA, and the S3 event notification mechanism is designed to be reliable within the same region. Therefore, taking no action is the correct and sufficient approach.

Why this answer

AWS Lambda functions are inherently highly available within an AWS Region. The Lambda service automatically runs your function across multiple Availability Zones (AZs) to handle failures of individual AZs. No additional configuration is required to achieve AZ-level resilience; the service manages the underlying compute fleet and network infrastructure to ensure continued operation even if one AZ fails.

Exam trap

The trap here is that candidates often confuse the need to configure VPC subnets for Lambda functions with the Lambda service's inherent AZ resilience, leading them to incorrectly select Option C, which is only relevant for VPC-attached functions accessing private resources, not for the function's own availability.

How to eliminate wrong answers

Option A is wrong because an Application Load Balancer (ALB) is used to distribute HTTP/S traffic to targets like Lambda functions via function URLs or as a target group, but it does not provide high availability for the Lambda service itself—Lambda already runs across AZs natively. Option C is wrong because Lambda functions do not run in subnets unless they are attached to a VPC; even then, configuring the function in two subnets in different AZs only provides high availability for VPC resources (e.g., RDS, ElastiCache) accessed by the function, not for the Lambda service itself. Option D is wrong because deploying the Lambda function in two separate regions with Route 53 failover is unnecessary for AZ-level resilience and introduces cross-region latency and complexity; Lambda is already regionally resilient across AZs without multi-region setup.

145
MCQmedium

A company runs a critical database on an Amazon RDS for MySQL DB instance. The SysOps administrator needs to ensure that the database can survive a single Availability Zone failure with minimal downtime. Which configuration should the administrator implement?

A.Enable automatic backups.
B.Deploy a read replica in a different AZ.
C.Enable Multi-AZ deployment.
D.Take a manual snapshot and copy it to another AZ.
AnswerC

Enabling Multi-AZ deployment automatically provisions and maintains a synchronous standby replica in a different Availability Zone. RDS continuously monitors the health of the primary database and, when it detects an AZ failure or instance failure, automatically flips the DNS record to the standby, typically completing failover within 60–120 seconds. Because replication to the standby is synchronous, committed transactions are preserved, providing both high availability and zero data loss for committed data.

Why this answer

Multi-AZ deployment for Amazon RDS automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary AZ fails, Amazon RDS automatically fails over to the standby, typically within 60–120 seconds, minimizing downtime without manual intervention. This meets the requirement for surviving a single AZ failure with minimal downtime.

Exam trap

The trap here is that candidates often confuse read replicas (which require manual promotion and are for read scaling) with Multi-AZ (which provides automatic failover for high availability), leading them to select Option B incorrectly.

How to eliminate wrong answers

Option A is wrong because automatic backups only provide point-in-time recovery to restore data to a new instance, not automatic failover or high availability during an AZ failure. Option B is wrong because a read replica in a different AZ is designed for read scaling and can be promoted to a primary, but promotion is a manual process that takes time and does not provide automatic failover with minimal downtime. Option D is wrong because taking a manual snapshot and copying it to another AZ requires manual restoration steps, which results in significant downtime and does not provide automatic failover.

146
MCQeasy

A SysOps administrator is designing a backup strategy for an Amazon EFS file system. The file system stores critical data that must be recoverable within 15 minutes of a failure. Which solution meets these requirements?

A.Use AWS Backup to create automated backups of the EFS file system and restore to a new file system if needed.
B.Configure EFS lifecycle management to move files to Infrequent Access storage class.
C.Take periodic EBS snapshots of the EC2 instance that mounts the EFS volume.
D.Create a Lambda function that copies files to an S3 bucket every hour.
AnswerA

AWS Backup natively integrates with Amazon EFS, enabling scheduled, automated backups via backup plans and lifecycle policies. Restores create a new EFS file system from the backup vault, preserving file metadata and permissions. This is the standard managed service for centralized EFS backup, supporting cross-region/cross-account copies and retention management.

Why this answer

AWS Backup provides a fully managed, policy-driven backup solution for Amazon EFS that supports automated, scheduled backups with point-in-time recovery. Restoring an EFS file system from a backup typically completes within minutes, meeting the 15-minute recovery time objective (RTO). This is the only option that directly addresses the requirement for recoverable backups with a defined RTO.

Exam trap

The trap here is that candidates may confuse EFS lifecycle management or EBS snapshots with backup solutions, not realizing that EFS requires a dedicated backup service (AWS Backup) to meet specific RTO/RPO requirements, and that EBS snapshots do not capture data stored on a separate network file system.

How to eliminate wrong answers

Option B is wrong because EFS lifecycle management only moves files between storage classes (e.g., Standard to Infrequent Access) to reduce costs; it does not create backups or enable recovery after failure. Option C is wrong because EBS snapshots capture the state of an EC2 instance's block storage, not the data stored in a separate EFS file system mounted over NFS; EFS data is not included in EBS snapshots. Option D is wrong because a Lambda function that copies files to S3 every hour introduces at least a 1-hour recovery point objective (RPO) and does not guarantee recovery within 15 minutes; additionally, it requires custom scripting and does not provide the consistency guarantees of a native backup service.

147
MCQmedium

A SysOps administrator is designing a disaster recovery plan for a critical RDS MySQL database. The database must be available with a Recovery Point Objective (RPO) of less than 1 hour and a Recovery Time Objective (RTO) of less than 2 hours. The primary region is us-east-1. Which solution meets these requirements?

A.Enable Multi-AZ deployment in the primary region
B.Enable automated backups and restore to a new region when needed
C.Take daily manual snapshots and copy them to another region
D.Create a cross-Region read replica in us-west-2 and promote it during a disaster
AnswerD

A cross-Region read replica in us-west-2 uses RDS's asynchronous replication from the primary DB instance, which typically incurs only seconds of lag. Because the replica is already hydrated with data, promoting it via the 'Promote Read Replica' operation is a single metadata change that transitions it to a standalone writable instance, avoiding the time-consuming restore of a backup. This makes it the only option that can satisfy a tight RTO while also offering a very low RPO (near-real-time) under normal operation.

Why this answer

A cross-Region read replica in us-west-2 can be promoted to a standalone primary database in under 2 hours, and the replication lag is typically less than 1 hour, meeting both RPO and RTO requirements. This solution provides continuous asynchronous replication from the primary region, ensuring minimal data loss and fast failover without needing to restore from backups.

Exam trap

The trap here is that candidates often confuse Multi-AZ (high availability within a region) with cross-region disaster recovery, or assume that automated backups or manual snapshots can be restored quickly enough to meet aggressive RPO/RTO targets, ignoring the time required for cross-region data transfer and restore operations.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment only provides high availability within a single region, not cross-region disaster recovery, so it cannot meet the RPO/RTO if the entire region fails. Option B is wrong because automated backups are stored in the same region and restoring to a new region requires copying the backup across regions, which can take longer than 2 hours and may not achieve an RPO of less than 1 hour due to backup frequency (typically daily). Option C is wrong because daily manual snapshots have an RPO of up to 24 hours, which exceeds the required 1-hour RPO, and copying them to another region adds additional time, likely exceeding the 2-hour RTO.

148
MCQmedium

A company has a production AWS account with a single VPC and multiple subnets across two Availability Zones. The company hosts a web application on EC2 instances in an Auto Scaling group. The application uses an Amazon Aurora MySQL database cluster with one writer and two reader instances in the same VPC. The SysOps administrator configured AWS CloudTrail to log API calls and Amazon CloudWatch alarms for operational monitoring. After a recent network partition event in one Availability Zone, the application became unavailable for several minutes. The administrator wants to improve the application's resilience to such events without changing the database cluster configuration. The administrator has budget for additional resources but wants to minimize costs. What should the administrator do?

A.Enable Multi-AZ for the Aurora cluster, which is already enabled by default. Additionally, increase the instance size of the writer instance.
B.Configure a cross-region read replica for the Aurora cluster and promote it to primary if the primary AZ fails.
C.Create a second VPC in a different AWS Region and set up a second Aurora cluster. Use Route 53 with failover routing to direct traffic to the secondary region if the primary fails.
D.Ensure the Auto Scaling group launches instances in both Availability Zones. Configure the Application Load Balancer to be cross-zone load balancing enabled.
AnswerD

An Auto Scaling group configured to launch instances across two Availability Zones automatically replaces any failed instances, including when an entire AZ becomes unhealthy. With cross-zone load balancing enabled on the Application Load Balancer, traffic is distributed evenly across all healthy targets in both AZs, and the ALB will immediately route all traffic to the remaining AZ if the other is impaired. Because the ALB is a regional service, it continues to operate independently of any single AZ failure.

Why this answer

Distributing EC2 instances across both Availability Zones and enabling cross-zone load balancing on the ALB ensures that traffic is routed to healthy instances in any AZ, providing high availability for the application tier during an AZ failure without changing the database cluster. Option A is incorrect because Multi-AZ for Aurora is already inherent; increasing instance size does not improve resilience. Option B is incorrect because a cross-region read replica does not help with within-region AZ failures and adds cost and complexity.

Option C is incorrect because creating a second VPC in another region with a separate cluster is expensive and unnecessary for an AZ-level failure.

149
Multi-Selecthard

A company wants to back up its on-premises file servers to AWS for disaster recovery. The data changes frequently, and the company needs to minimize data loss. Which THREE steps should the company take? (Select THREE.)

Select 3 answers
A.Use AWS Backup to create backup plans for the Storage Gateway.
B.Use S3 Transfer Acceleration for uploads.
C.Configure the file gateway to cache data locally for frequently accessed files.
D.Set up S3 Cross-Region Replication from the backup bucket.
E.Deploy an AWS Storage Gateway file gateway on-premises.
AnswersA, C, E

AWS Backup natively integrates with Storage Gateway, allowing you to define backup plans that automatically create point-in-time snapshots of the gateway's data and store them in S3. This gives you a fully managed, policy-driven backup solution with retention and lifecycle rules, so you don't need to run manual backup scripts on the on-premises file servers. The snapshots are durable in S3 and can be restored quickly when needed, making this the correct way to automate backup of the gateway.

Why this answer

AWS Backup can create backup plans for Storage Gateway volumes and file shares, providing automated, policy-based backups to Amazon S3. This ensures consistent, scheduled backups of the on-premises data cached in the file gateway, minimizing data loss by capturing frequent changes according to the backup plan's schedule.

Exam trap

The trap here is that candidates often confuse S3 Transfer Acceleration or Cross-Region Replication as backup solutions, but they are data transfer or replication features, not backup services that provide point-in-time recovery and lifecycle management for on-premises data.

150
MCQeasy

A company is using Amazon S3 to store critical data with versioning enabled. The SysOps administrator needs to implement a solution that automatically transitions objects to S3 Glacier Deep Archive after 90 days and permanently deletes them after 7 years. Which S3 feature should be used?

A.S3 Lifecycle policies
B.S3 Intelligent-Tiering
C.S3 Object Lock
D.S3 Cross-Region Replication
AnswerA

S3 Lifecycle policies are the correct answer because they provide exactly the automation needed to manage the data's full lifecycle. You define rules that transition objects to less expensive storage classes like S3 Standard-IA or S3 Glacier after specified durations, and you can also set expiration actions that permanently delete objects or their versions. Lifecycle policies act on prefixes, tags, and object age, giving you precise control over when data moves to colder tiers and when it is eventually removed, which directly accomplishes both transitioning and expiring critical data.

Why this answer

S3 Lifecycle policies are the correct choice because they allow you to define rules that automatically transition objects to S3 Glacier Deep Archive after a specified number of days (90) and then permanently expire (delete) them after a longer period (7 years). This directly meets the requirements for automated, time-based storage class transitions and deletion without manual intervention.

Exam trap

The trap here is that candidates often confuse S3 Intelligent-Tiering with lifecycle policies, thinking it can enforce fixed time-based transitions, when in reality Intelligent-Tiering only responds to access patterns and cannot guarantee a specific schedule for archiving or deletion.

How to eliminate wrong answers

Option B is wrong because S3 Intelligent-Tiering automatically moves objects between access tiers based on changing access patterns, but it does not support a fixed schedule for transitioning to Glacier Deep Archive or permanent deletion after a specific number of days. Option C is wrong because S3 Object Lock is designed to prevent object deletion or overwrites for a fixed retention period (compliance or governance mode), not to automate transitions or scheduled deletions. Option D is wrong because S3 Cross-Region Replication asynchronously replicates objects to a different AWS region for redundancy or compliance, but it does not provide any lifecycle management or automated deletion capabilities.

← PreviousPage 2 of 3 · 205 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Soa Reliability Bc questions.