Courseiva

AWS Certified SysOps Administrator Associate SOA-C02 (SOA-C02) — Questions 1–75

1169 questions total · 16pages · All types, answers revealed

Page 1 of 16

Page 2
1
MCQhard

A SysOps administrator is troubleshooting an application that runs on EC2 instances behind an ALB. Users report intermittent 503 errors. The administrator checks the ALB access logs and finds entries with 'elb_status_code' 503 and 'target_status_code' '-'. What is the most likely cause?

A.The target instances are unhealthy, causing the ALB to return 503.
B.The SSL certificate on the ALB has expired.
C.The target instances have high CPU utilization.
D.The security group on the ALB is blocking traffic.
AnswerA

When every instance in a target group fails consecutive health checks, the ALB marks them unhealthy and has no target to forward client traffic to. Instead of proxying a 5xx error from an overloaded backend, the load balancer itself responds with 503 Service Unavailable because the service is deemed unreachable. The health check interval and threshold settings determine how quickly an unhealthy target is removed from rotation, but the client sees 503 until at least one target passes.

Why this answer

The ALB access log entry with `elb_status_code` 503 and `target_status_code` '-' indicates that the load balancer itself generated the 503 error because it could not establish a connection to any healthy target. The dash for the target status code means the request never reached a target instance, which occurs when all targets in the target group are marked unhealthy by the health checks. This is the most common cause of intermittent 503 errors with an ALB.

Exam trap

The trap here is that candidates often confuse a 503 error with target-side issues (like high CPU or application errors), but the dash in the target_status_code is the key indicator that the ALB itself is rejecting the request due to no healthy targets, not that the request reached a target and failed.

How to eliminate wrong answers

Option B is wrong because an expired SSL certificate on the ALB would cause TLS handshake failures (e.g., 502 or 525 errors), not a 503 with a dash for the target status code. Option C is wrong because high CPU utilization on target instances would still allow the ALB to forward requests to them (resulting in a target_status_code like 200 or 500), but the dash indicates no connection was attempted. Option D is wrong because the ALB's security group controls inbound traffic to the load balancer; if it were blocking traffic, clients would receive a 504 or connection timeout, not a 503, and the access log would show a different elb_status_code.

2
MCQhard

A SysOps administrator is troubleshooting connectivity issues between an Amazon EC2 instance in a VPC and an on-premises data center connected via AWS Direct Connect. The EC2 instance can reach other instances in the same VPC but cannot reach the on-premises network. The virtual private gateway (VGW) is attached to the VPC and the Direct Connect virtual interface is up. Which configuration step should the administrator verify first?

A.Check the security group rules for the EC2 instance
B.Confirm that the Direct Connect virtual interface is associated with the correct VLAN
C.Add a route in the VPC route table for the on-premises CIDR pointing to the virtual private gateway
D.Verify the network ACL inbound and outbound rules for the VPC subnet
AnswerC

The VPC route table must contain a route for the on-premises CIDR targeting the virtual private gateway; without it, traffic to on-premises has no path despite the Direct Connect virtual interface being up. This is the most common cause of this symptom.

Why this answer

For an EC2 instance to reach an on-premises network via Direct Connect, the VPC route table must contain a route for the on-premises CIDR block pointing to the virtual private gateway (VGW). Since the instance can reach other VPC instances, local routing works; the missing piece is the route to the on-premises destination. Without this route, traffic has no path to the VGW and is dropped.

Exam trap

SOA-C02 often tests whether candidates jump to security groups or NACLs (familiar troubleshooting steps) instead of first verifying routing—the most common cause of hybrid connectivity failures is a missing route, not a security rule.

How to eliminate wrong answers

Option A is wrong because security group rules control instance-level traffic; if they blocked traffic, the instance likely couldn't reach other VPC instances either, and the symptom is specifically on-premises reachability. Option B is wrong because the question states the Direct Connect virtual interface is up, implying VLAN association is correct; verifying it is not the first step when the VIF is already operational. Option D is wrong because network ACLs are stateless subnet-level filters; while they could block traffic, the more fundamental issue is the absence of a route, and NACLs would typically affect all traffic, not just on-premises.

3
MCQmedium

A company runs an Amazon RDS for MySQL DB instance in us-east-1. The SysOps administrator needs to implement a disaster recovery solution that can recover from a regional outage with a Recovery Point Objective (RPO) of less than 1 second and a Recovery Time Objective (RTO) of less than 1 minute. Which solution should the administrator use?

A.Multi-AZ deployment
B.Cross-region read replica
C.Aurora Global Database
D.Automated snapshot copy to another region
AnswerC

Aurora Global Database is specifically designed for cross-region disaster recovery. It uses a dedicated replication channel, typically with sub-second replication lag between the primary Region and up to five secondary Regions, which gives an RPO of usually under a second. A secondary Region can be promoted to primary in about a minute, providing a low RTO. This makes Aurora Global Database the correct choice when the requirement is for rapid, cross-region failover while losing almost no committed transactions.

Why this answer

Aurora Global Database is the correct choice because it provides a fully managed cross-region replication solution with a typical RPO of less than 1 second and an RTO of less than 1 minute during a regional failover. It uses a primary cluster in one region and up to five secondary clusters in other regions, with asynchronous replication that is optimized for low latency, meeting the stringent RPO/RTO requirements.

Exam trap

The trap here is that candidates often confuse Multi-AZ deployments (which are for high availability within a region) with cross-region disaster recovery, and they underestimate the replication lag and failover time of standard cross-region read replicas versus the optimized architecture of Aurora Global Database.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability within a single region, not cross-region disaster recovery, and its failover RTO is typically 1-2 minutes, exceeding the required 1 minute. Option B is wrong because a cross-region read replica for RDS MySQL uses asynchronous replication with a typical RPO of seconds to minutes, not less than 1 second, and promoting a read replica to a primary instance can take several minutes, failing the RTO requirement. Option D is wrong because automated snapshot copy to another region has an RPO of at least 5 minutes (the minimum snapshot interval) and restoring from a snapshot can take tens of minutes, both far exceeding the required RPO and RTO.

4
MCQhard

A SysOps administrator receives an alarm that an EC2 instance's status check has failed. The instance is part of an Auto Scaling group behind an Application Load Balancer. The administrator needs to ensure that the instance is automatically replaced and that the root cause is investigated. What is the MOST efficient combination of actions to achieve this?

A.Configure an Auto Scaling lifecycle hook to terminate the unhealthy instance and send the instance system log to an S3 bucket for analysis.
B.Create a CloudWatch alarm that triggers an SNS notification to the administrator to manually replace the instance.
C.Reboot the instance from the AWS Management Console and then review CloudTrail logs.
D.Manually stop and start the instance to recover it, then check the system logs.
AnswerA

A lifecycle hook on the terminating state of an Auto Scaling group intercepts the EC2 instance termination process, enabling a Lambda function to capture the instance's system log (console output) and upload it to an S3 bucket before the instance is destroyed. Because the Auto Scaling group has already marked the instance unhealthy via its health checks, the group will automatically launch a replacement instance after the lifecycle action completes, ensuring fully automated recovery. The hook's timeout and the ability to call complete-lifecycle-action guarantee that the log is securely stored, and the system log provides the diagnostic evidence needed for root cause analysis.

Why this answer

It combines automatic instance replacement via the Auto Scaling group's health check (which marks the instance unhealthy and terminates it) with a lifecycle hook that captures the instance's system log before termination and sends it to S3 for root cause analysis. This is the most efficient approach as it requires no manual intervention and preserves diagnostic data.

Exam trap

The trap here is that candidates may think manual actions (reboot, stop/start) are sufficient for recovery, but the question explicitly requires automatic replacement and root cause investigation, which only a lifecycle hook with data capture provides.

How to eliminate wrong answers

Option B is wrong because it relies on manual replacement via SNS notification, which is inefficient and violates the requirement for automatic replacement. Option C is wrong because rebooting an instance with a failed status check does not address the underlying issue and does not automatically replace the instance; CloudTrail logs record API calls, not system-level diagnostics. Option D is wrong because manually stopping and starting the instance is not automatic and does not guarantee recovery; it also fails to capture diagnostic data for root cause analysis.

5
MCQmedium

An administrator runs the above CloudWatch command to analyze CPU utilization for an EC2 instance. The instance is currently running with a t3.large instance type. The company wants to optimize costs. Based on the data, which action should the administrator take?

A.Stop the instance to save costs immediately.
B.Downsize the instance to t3.medium to reduce costs.
C.Scale up the instance to t3.xlarge to improve performance.
D.Purchase a Reserved Instance for the current instance type.
AnswerB

CloudWatch shows low CPU utilization, indicating the t3.large (or current size) is over-provisioned for the actual workload. Downsizing to t3.medium, which has the same vCPU count but half the memory (4 GiB vs 8 GiB), will reduce the on-demand hourly cost while still providing adequate compute for the observed load. Before resizing, confirm memory, network throughput, and application performance remain above the 95th percentile to avoid an undersized instance.

Why this answer

The average CPU utilization over the week is low (around 15.5%), indicating the instance is over-provisioned. Downsizing to a t3.medium (half the vCPUs and memory) would likely be sufficient and reduce costs. Reserved Instances would be cost-effective only if the instance runs consistently, but the current utilization is low.

Scaling up would increase costs. Stopping the instance is not an option if needed.

6
MCQeasy

A SysOps administrator receives a notification that an EC2 instance's status check has failed. The instance is part of an Auto Scaling group. What is the immediate impact on the application?

A.The instance is still accessible and serving traffic.
B.The instance is immediately terminated.
C.The instance is automatically stopped and started.
D.The Auto Scaling group will launch a new instance to replace the failed one, potentially causing temporary downtime.
AnswerD

The Auto Scaling group detects the failed status check and marks the instance as unhealthy, then terminates it and launches a new instance to maintain the desired capacity. This replacement process involves a brief period during which the new instance is initializing and passing health checks, potentially causing temporary downtime. The behavior aligns with the Auto Scaling lifecycle for handling failed instances.

Why this answer

When an EC2 instance fails a status check, the Auto Scaling group detects the failure and initiates a replacement by launching a new instance. However, the failed instance is not immediately terminated; it may remain in a stopped or impaired state until the replacement is fully in service, which can cause temporary downtime for the application if the instance was actively handling traffic.

Exam trap

The trap here is that candidates assume the Auto Scaling group immediately terminates the failed instance (Option B), but in reality, the group waits for a health check grace period and the replacement process is not instantaneous, causing temporary downtime.

How to eliminate wrong answers

Option A is wrong because a failed status check indicates the instance is impaired (e.g., unreachable due to OS-level issues or hardware problems), so it is not accessible or serving traffic. Option B is wrong because the Auto Scaling group does not immediately terminate the instance; it first waits for the health check grace period and then performs a gradual replacement, and the instance may be terminated only after the new one is ready. Option C is wrong because EC2 status check failures do not automatically stop and start the instance; that action would require a manual or automated recovery via CloudWatch alarms or EC2 auto-recovery, not the Auto Scaling group's default behavior.

7
MCQhard

A SysOps Administrator attempted to update a CloudFormation stack. The stack update failed and is now in UPDATE_ROLLBACK_IN_PROGRESS state as shown in the exhibit. What should the administrator do to recover the stack to a stable state?

A.Wait for the rollback to complete and then investigate the failure reason.
B.Delete the stack and recreate it.
C.Manually update the Auto Scaling group to correct the issue.
D.Execute a change set to fix the failed resource.
AnswerA

While the stack is in UPDATE_ROLLBACK_IN_PROGRESS, CloudFormation is automatically reverting resources to the last known good state. Any attempt to modify the stack—whether via delete, update, or change set execution—will be rejected because the stack is in a transient state. Once the rollback reaches UPDATE_ROLLBACK_COMPLETE, the Events tab in the AWS Console or describe-stack-events will show the exact failure reason (e.g., a parameter validation error or an EC2 resource failure). This diagnostic information is essential for correcting the template and retrying the update safely.

Why this answer

When a CloudFormation stack update fails and enters UPDATE_ROLLBACK_IN_PROGRESS, AWS CloudFormation is automatically rolling back the stack to its last known stable state. The administrator should wait for this rollback to complete, which will result in the stack returning to UPDATE_ROLLBACK_COMPLETE (or UPDATE_ROLLBACK_FAILED if rollback fails). After that, they can investigate the failure reason using stack events and then retry the update with corrections.

Exam trap

SOA-C02 often tests the misconception that manual intervention is needed during an in-progress rollback, when the correct action is to wait for CloudFormation to complete the rollback automatically.

How to eliminate wrong answers

Option B is wrong because deleting and recreating the stack would cause downtime and loss of resources, and it is unnecessary since CloudFormation can recover automatically. Option C is wrong because manually updating the Auto Scaling group would interfere with CloudFormation's management and could cause drift, making the stack inconsistent. Option D is wrong because executing a change set is not possible while the stack is in UPDATE_ROLLBACK_IN_PROGRESS; the stack must first reach a stable state.

8
MCQmedium

A company has the following S3 bucket policy attached to a bucket named 'example-bucket'. A user is unable to download an object from the bucket using an HTTP URL (not HTTPS). What is the cause?

A.The bucket policy does not allow GetObject for anonymous users.
B.The Deny statement blocks all S3 actions when the request is not using HTTPS.
C.The bucket policy requires server-side encryption for all requests.
D.The Deny statement only applies to PutObject, not GetObject.
AnswerB

The Deny statement uses Action 's3:*' with a condition checked via aws:SecureTransport set to false, so any S3 API request sent over HTTP instead of HTTPS is denied. In IAM policy evaluation, an explicit Deny always overrides any Allow, meaning the earlier Allow for GetObject does not help when the request is not encrypted. This is exactly why non-HTTPS access is blocked for all S3 operations.

Why this answer

The bucket policy contains a Deny statement that applies to all s3:* actions when the request does not use HTTPS (SecureTransport is false). Even though there is an Allow statement for GetObject to everyone, the explicit Deny overrides the Allow. Option A is incorrect because the issue is not about anonymous users; the Deny affects all requests.

Option C is incorrect because the policy does not mention server-side encryption. Option D is incorrect because the Deny statement applies to all S3 actions, not just PutObject.

9
MCQhard

A company has a production DynamoDB table with on-demand capacity. They need to ensure business continuity with a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour in case of a regional outage. What is the MOST cost-effective solution?

A.Use AWS Backup to schedule daily backups and restore in another region.
B.Enable DynamoDB global tables for the table.
C.Enable point-in-time recovery (PITR) on the table.
D.Configure cross-region read replicas for the table.
AnswerB

DynamoDB global tables replicate your table automatically and asynchronously across multiple AWS Regions, typically with sub-second latency, which easily satisfies a 5-minute RPO. Because every replica accepts reads and writes, you can fail over instantly by rerouting application traffic (e.g., via Route 53 health checks) to a healthy region, meeting the 1-hour RTO without needing to restore data. Global tables also handle conflict resolution using last-writer-wins, ensuring data consistency across regions and making them the correct managed solution for regional disaster recovery.

Why this answer

DynamoDB global tables provide multi-region, multi-active replication with changes typically propagated within seconds, meeting an RPO of 5 minutes and an RTO of 1 hour by allowing traffic to be redirected to a replica in another region. This is the most cost-effective solution because it eliminates the need for separate backup storage or compute resources, and on-demand capacity scales automatically without provisioning.

Exam trap

The trap here is that candidates confuse point-in-time recovery (PITR) with cross-region disaster recovery, not realizing PITR is single-region only and cannot meet the RTO for a regional outage.

How to eliminate wrong answers

Option A is wrong because AWS Backup daily backups cannot achieve an RPO of 5 minutes (backups are at most daily) and restoring to another region would exceed the 1-hour RTO due to restore time. Option C is wrong because point-in-time recovery (PITR) only protects against accidental writes/deletes within the same region, not a regional outage, and restoring a table from PITR to another region would take longer than 1 hour. Option D is wrong because cross-region read replicas are read-only and cannot accept write traffic during a regional outage, failing the RTO requirement for write continuity.

10
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer across three Availability Zones. The application stores session data in an RDS MySQL database. To improve reliability, the company wants to ensure that a single Availability Zone failure does not impact the application's availability. Which combination of actions should the SysOps administrator take?

A.Configure the ALB to use only healthy instances and enable detailed CloudWatch metrics.
B.Increase the EC2 instance size to handle more traffic in a single AZ.
C.Increase the Auto Scaling group's desired capacity to a larger number.
D.Deploy RDS in Multi-AZ configuration with automatic failover, and enable cross-zone load balancing on the ALB.
AnswerD

Deploying RDS in a Multi-AZ configuration with automatic failover gives the database a synchronous standby replica in a different Availability Zone, so an AZ outage triggers a failover to the standby with minimal downtime. Enabling cross-zone load balancing on the ALB distributes incoming traffic across EC2 instances in multiple AZs, ensuring that if one AZ fails, the ALB continues to route requests to healthy instances in the remaining AZ. Together these actions provide both application-tier and database-tier high availability, which is the core requirement for running the workload across AZs and surviving an AZ failure.

Why this answer

Deploying RDS in Multi-AZ with automatic failover ensures database resilience against an AZ failure by maintaining a synchronous standby replica in a different AZ, while enabling cross-zone load balancing on the ALB distributes traffic across instances in all AZs, preventing a single AZ failure from taking down the entire application. Together, these actions address both the session data persistence and compute layer availability requirements.

Exam trap

The trap here is that candidates often think increasing instance count or size alone provides high availability, but they overlook the need for both database redundancy (Multi-AZ RDS) and cross-AZ traffic distribution (cross-zone load balancing) to survive an entire AZ failure.

How to eliminate wrong answers

Option A is wrong because configuring the ALB to use only healthy instances and enabling detailed CloudWatch metrics only improves monitoring and traffic routing to healthy targets, but does not provide redundancy for the RDS database or ensure compute capacity survives an AZ failure. Option B is wrong because increasing the EC2 instance size to handle more traffic in a single AZ does not eliminate the single point of failure; if that AZ fails, all instances are lost regardless of size. Option C is wrong because increasing the Auto Scaling group's desired capacity to a larger number does not guarantee instances are distributed across multiple AZs; without cross-zone load balancing and Multi-AZ RDS, a single AZ failure can still take down all instances and the database.

11
Multi-Selecteasy

A company wants to audit all API calls made in their AWS account for security analysis. They need to record both management events and data events. Which THREE steps should be taken to set up comprehensive logging? (Choose THREE.)

Select 3 answers
A.Enable AWS CloudTrail to record data events for S3 and Lambda.
B.Enable AWS CloudTrail to record management events.
C.Enable VPC Flow Logs to capture API call metadata.
D.Send the log files to Amazon CloudWatch Logs for real-time analysis.
E.Configure the trail to deliver log files to an S3 bucket.
AnswersA, B, E

Data events capture object-level API operations for S3, such as GetObject, PutObject, and DeleteObject, as well as Lambda function invocations. Unlike management events, data events are not enabled by default; you must explicitly configure the trail to include them, which is essential for auditing access to sensitive content and detecting suspicious data read/write patterns. Without this explicit enablement, the audit will miss critical resource-level activity that management events do not cover.

Why this answer

Option A is correct because CloudTrail data events are not logged by default; you must explicitly enable them for resources like S3 objects and Lambda invocations to capture those API-level operations. Option B is correct because management events (control-plane operations such as CreateBucket or RunInstances) are the core of CloudTrail auditing and must be enabled on the trail to record API activity across the account. Option E is correct because a CloudTrail trail must deliver its log files to an Amazon S3 bucket, which is the required destination for storing and later analyzing the audit logs.

Option C is not correct because VPC Flow Logs capture IP traffic metadata for network interfaces, not API call details, so they do not satisfy the API auditing requirement. Option D is not correct because sending logs to CloudWatch Logs is an optional enhancement for monitoring and alerting, not a required step for setting up comprehensive CloudTrail logging of management and data events.

Exam trap

SOA-C02 often tests the distinction between CloudTrail (API activity) and VPC Flow Logs (network traffic) — candidates pick Flow Logs thinking 'all activity' includes API calls, but Flow Logs cannot see IAM identities or API actions.

12
Multi-Selectmedium

A company is using an Auto Scaling group with a dynamic scaling policy based on average CPU utilization. The SysOps administrator notices that the scaling is not triggering as expected. Which THREE steps should the administrator take to troubleshoot the issue?

Select 3 answers
A.Check the scaling activity history in the Auto Scaling group for any errors or cooldown periods.
B.Ensure that the EC2 instances are passing the ELB health checks.
C.Review the scaling policy's cooldown period and threshold settings.
D.Verify that the CloudWatch alarm associated with the scaling policy is in ALARM state when CPU is high.
E.Manually increase the desired capacity to see if the scaling policy takes effect.
AnswersA, C, D

Scaling activity history is the authoritative log of every scaling action the Auto Scaling group attempted or skipped. It records events such as policy executions, cooldown period blocks, and failures (e.g., insufficient instance capacity or unhealthy instances). If no scaling action was logged despite high CPU, that directly reveals whether the alarm-to-policy path was broken or a cooldown suppressed the action, making it the correct first place to diagnose why the group did not scale out.

Why this answer

The scaling activity history provides a log of all scaling actions, including errors, cooldown periods, and why a scaling event was or was not triggered. By reviewing this history, the administrator can identify if the scaling policy was blocked by a cooldown period, if the alarm state was not reached, or if there were any configuration errors that prevented the scaling action from executing.

Exam trap

The trap here is that candidates may confuse ELB health checks with the metric-based alarm that drives scaling, or think that manually adjusting capacity is a valid diagnostic step, when in fact it bypasses the automated policy logic and does not reveal why the policy failed to trigger.

13
MCQmedium

A company runs a stateful web application on a single Amazon EC2 instance. The application stores session state in memory and writes critical data to an Amazon EBS volume. The SysOps administrator needs to implement a highly available architecture that can tolerate an Availability Zone (AZ) failure. The administrator plans to use an Auto Scaling group and an Application Load Balancer (ALB). Which combination of steps is required to make the application highly available while preserving session and data durability across AZ failures?

A.Create an AMI of the current instance, configure an Auto Scaling group with a launch template that uses the AMI, and attach the existing EBS volume to new instances.
B.Create a multi-AZ Auto Scaling group and use sticky sessions (session affinity) on the ALB to tie users to specific instances.
C.Use an Auto Scaling group across multiple AZs, migrate session storage to Amazon ElastiCache (multi-AZ), and migrate application data from EBS to Amazon EFS (file system mounted across AZs).
D.Use an Auto Scaling group in a single AZ and use a Multi-AZ RDS instance for data storage.
AnswerC

ElastiCache provides a shared, cross-AZ in-memory session store. EFS provides a shared, cross-AZ file system. The Auto Scaling group launches instances in multiple AZs, and the ALB distributes traffic. This architecture survives an AZ failure.

Why this answer

It addresses both session state and data durability across AZ failures. Migrating session storage to ElastiCache (multi-AZ) ensures session data survives instance failure, and migrating application data from EBS to EFS provides a shared, multi-AZ file system that persists independently of any single EC2 instance. This combination allows the Auto Scaling group to launch new instances in any AZ and immediately access both session and application data.

Exam trap

The trap here is that candidates often assume sticky sessions (session affinity) alone are sufficient for high availability, but they fail to realize that sticky sessions do not replicate session state across instances, so an instance failure still loses the session data.

How to eliminate wrong answers

Option A is wrong because attaching the existing EBS volume to new instances is not possible across AZs (EBS volumes are AZ-scoped) and does not provide a shared, durable data layer; it also fails to address session state persistence. Option B is wrong because sticky sessions alone do not preserve session data if the instance fails; they only route traffic to the same instance, and if that instance goes down, the session is lost. Option D is wrong because using a single AZ for the Auto Scaling group cannot tolerate an AZ failure, and while Multi-AZ RDS handles database durability, it does not address the application's in-memory session state or EBS-stored data.

14
MCQhard

A SysOps administrator is troubleshooting an issue where an EC2 instance running a web server is not reachable from the internet. The instance has a public IP and is in a public subnet. The security group allows HTTP and HTTPS from 0.0.0.0/0. The network ACL allows all inbound and outbound traffic. What should the administrator check NEXT?

A.Check that the instance is associated with an Elastic IP address.
B.Verify that the subnet's route table has a route to an internet gateway.
C.Confirm that the instance's operating system firewall is disabled.
D.Review the VPC Flow Logs for the instance's network interface.
AnswerB

A public subnet's route table must contain a default route (0.0.0.0/0) that targets an Internet Gateway (IGW) for outbound and inbound internet traffic to flow. Even with a public IP and permissive security group and NACL rules, if this route is missing or points to a misconfigured target, the instance cannot send or receive packets to/from the internet. Verifying the route table is the correct next step because it directly addresses the network path after you've already confirmed the stateful and stateless filtering layers.

Why this answer

The instance is in a public subnet with a public IP and security group allowing HTTP/HTTPS, and the network ACL allows all traffic. The most likely remaining issue is that the subnet's route table lacks a route to an internet gateway (IGW), which is required for traffic to and from the internet. Without this route, the instance cannot send responses back to internet clients, making it unreachable despite having a public IP.

Exam trap

The trap here is that candidates often assume a public IP and permissive security groups are sufficient for internet access, overlooking the critical requirement of a route table entry pointing to an internet gateway for the subnet.

How to eliminate wrong answers

Option A is wrong because an Elastic IP is not required for internet connectivity; an instance with a public IP (auto-assigned) can already be reached from the internet if routing is correct. Option C is wrong because the question states the instance is not reachable from the internet, and while an OS firewall could block traffic, the more fundamental network-level routing issue should be checked first, and the OS firewall is not the most likely cause given the security group and NACL are permissive. Option D is wrong because VPC Flow Logs are useful for analyzing traffic that has already reached the network interface, but if the route table lacks an IGW route, traffic never reaches the instance, so flow logs would not show the missing route and are not the next logical check.

15
MCQhard

A SysOps administrator notices that an Amazon RDS for MySQL instance's CPU utilization is consistently above 80% during business hours. The administrator wants to identify the queries causing the high load without impacting performance. Which action should be taken?

A.Enable the MySQL slow query log and store it in CloudWatch Logs.
B.Enable Performance Insights on the RDS instance.
C.Enable Enhanced Monitoring to get OS-level metrics.
D.Increase the retention period for CloudWatch metrics to 15 months.
AnswerB

Performance Insights provides a real-time database load dashboard that breaks down DB load by SQL, wait events, and dimensions such as hosts and users, allowing you to pinpoint the exact statements responsible for CPU spikes. It operates with minimal overhead by sampling internal diagnostic data continuously, making it the most direct tool for identifying the SQL causing high RDS CPU usage without requiring additional logging or manual query analysis.

Why this answer

Performance Insights provides a database-specific performance schema that visualizes database load and identifies the SQL queries responsible for high CPU utilization. It operates with minimal overhead by sampling the database engine's internal performance data, making it ideal for diagnosing query performance issues without impacting the production workload.

Exam trap

The trap here is that candidates often confuse Enhanced Monitoring (OS-level metrics) with Performance Insights (database-level query analysis), or assume the slow query log is the best tool for identifying all high-CPU queries despite its threshold-based limitation.

How to eliminate wrong answers

Option A is wrong because the MySQL slow query log captures only queries that exceed a defined execution time threshold, not all queries causing high CPU utilization, and enabling it can add I/O overhead that may impact performance. Option C is wrong because Enhanced Monitoring provides OS-level metrics (CPU, memory, disk I/O) but does not identify which specific SQL queries are consuming CPU resources. Option D is wrong because increasing CloudWatch metric retention to 15 months only preserves historical data for long-term analysis, it does not help identify current queries causing high CPU load.

16
MCQeasy

A company runs a web application on EC2 instances in an Auto Scaling group. The application is behind an Application Load Balancer. The company wants to ensure that the application can handle a sudden spike in traffic without downtime. What should the SysOps administrator do?

A.Use a scheduled scaling policy to add instances during business hours.
B.Configure a target tracking scaling policy based on average CPU utilization.
C.Reduce the number of Availability Zones to lower latency.
D.Manually increase the desired capacity of the Auto Scaling group when traffic increases.
AnswerB

A target tracking scaling policy based on average CPU utilisation automatically adjusts the desired capacity of the Auto Scaling group to maintain a predefined CPU target, such as 50%. This satisfies the requirement to handle a sudden traffic spike without downtime because the policy proactively adds EC2 instances as CPU load increases, preventing performance degradation before the application becomes overwhelmed.

Why this answer

A target tracking scaling policy based on average CPU utilization allows the Auto Scaling group to automatically adjust capacity in response to real-time demand spikes. This dynamic scaling approach maintains a target metric (e.g., 50% CPU) by adding or removing instances, ensuring the application can handle sudden traffic bursts without downtime.

Exam trap

The trap here is that candidates often confuse scheduled scaling (for predictable patterns) with dynamic scaling (for unpredictable spikes), or they mistakenly think reducing Availability Zones improves performance when it actually harms reliability.

How to eliminate wrong answers

Option A is wrong because scheduled scaling policies are designed for predictable traffic patterns (e.g., business hours), not for sudden, unplanned spikes; they cannot react to real-time changes. Option C is wrong because reducing the number of Availability Zones actually decreases fault tolerance and increases the risk of downtime during a zone failure, contradicting the goal of handling spikes without downtime. Option D is wrong because manually increasing desired capacity requires human intervention, which is too slow to respond to sudden spikes and defeats the purpose of automated elasticity.

17
MCQeasy

A SysOps administrator maintains an AWS CloudFormation stack that deploys an Amazon EC2 instance. The administrator needs to change the instance type from t2.micro to t3.micro. The administrator wants to review the proposed changes before applying them to ensure no unexpected resource replacement occurs. Which CloudFormation feature should the administrator use?

A.Use the AWS CloudFormation console to directly update the stack with the new instance type and monitor the events.
B.Create a change set from the updated template, review the changes, and then execute the change set.
C.Use the AWS CloudFormation drift detection feature to check for differences between the stack and the template.
D.Modify the CloudFormation template locally and use the AWS CLI to validate it with 'aws cloudformation validate-template'.
AnswerB

A change set is generated from an updated template against the current stack, providing a detailed, resource-by-resource summary of whether CloudFormation will add, modify, or replace resources. Because changing an EC2 instance type typically requires replacement, the change set would explicitly flag a "Replace" action, letting the administrator assess the impact and even cancel before executing. Only after reviewing and confirming the change set is it executed, and the execution applies the exact changes that were previewed.

Why this answer

A change set allows the administrator to review the proposed modifications (including whether any resource replacement will occur) before applying them. By creating a change set from the updated template, the administrator can inspect the list of changes, such as the instance type update, and confirm that no unexpected resource replacement (e.g., a new EC2 instance being created) will happen. Only after reviewing the change set can the administrator safely execute it to apply the changes.

Exam trap

The trap here is that candidates confuse change sets with drift detection or template validation, not realizing that change sets are specifically designed to preview the impact of stack updates before execution.

How to eliminate wrong answers

Option A is wrong because directly updating the stack via the console applies changes immediately without a review step, so the administrator cannot preview whether resource replacement will occur. Option C is wrong because drift detection compares the current stack resources against the expected template configuration to identify manual changes, not to preview proposed updates before applying them. Option D is wrong because 'aws cloudformation validate-template' only checks the syntax of the template, not the impact of changes on existing resources or whether replacement will occur.

18
MCQmedium

A company's security team requires that all Amazon S3 buckets are encrypted at rest using server-side encryption with Amazon S3 managed keys (SSE-S3). A SysOps administrator needs to automatically detect any S3 bucket that does not have encryption enabled and automatically apply SSE-S3 encryption. The solution should leverage AWS managed services and minimize custom code. Which combination of AWS services should be used?

A.Use AWS Trusted Advisor to identify unencrypted buckets and then manually enable encryption.
B.Use AWS Config managed rule 's3-bucket-server-side-encryption-enabled' with an automatic remediation action using AWS Systems Manager Automation.
C.Use AWS CloudTrail to detect PutBucket operations and trigger a Lambda function that enables encryption.
D.Create an IAM bucket policy that denies any PutObject request that does not include x-amz-server-side-encryption header.
AnswerB

The AWS Config managed rule 's3-bucket-server-side-encryption-enabled' continuously evaluates whether S3 buckets have default encryption such as SSE-S3, SSE-KMS, or DSSE-KMS configured. When a bucket is marked NON_COMPLIANT during either a configuration change or a periodic evaluation, AWS Config can invoke an AWS Systems Manager Automation document, specifically the AWS-EnableS3BucketEncryption runbook, to automatically update the bucket's encryption setting and immediately reevaluate it. This approach provides a closed-loop, automated remediation workflow that satisfies the security team's requirement without custom code or manual intervention.

Why this answer

AWS Config's managed rule 's3-bucket-server-side-encryption-enabled' continuously evaluates S3 buckets for encryption compliance, and its automatic remediation action can invoke an AWS Systems Manager Automation document to enable SSE-S3 encryption on noncompliant buckets without custom code. This fully meets the requirement to automatically detect and remediate unencrypted buckets using AWS managed services.

Exam trap

The trap here is that candidates often confuse enforcing encryption on object uploads (via bucket policies or CloudTrail/Lambda) with ensuring the bucket's default encryption setting is enabled, which is what AWS Config's managed rule and remediation specifically address.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor only provides a manual check and recommendation; it cannot automatically apply encryption, and the requirement specifies automatic detection and remediation. Option C is wrong because AWS CloudTrail logs PutBucket operations but does not detect existing unencrypted buckets, and using a Lambda function introduces custom code, which the solution should minimize. Option D is wrong because an IAM bucket policy that denies PutObject requests without the encryption header only enforces encryption on new object uploads, not on the bucket's default encryption setting, and does not detect or remediate existing unencrypted buckets.

19
Multi-Selectmedium

A SysOps administrator is investigating a performance issue with an Amazon RDS for PostgreSQL instance. The administrator has enabled Performance Insights. Which TWO metrics from Performance Insights can help identify the root cause of a sudden increase in database load? (Choose TWO.)

Select 2 answers
A.Read IOPS and Write IOPS.
B.Average Active Sessions.
C.DB Load by Wait Events.
D.CPUUtilization percentage.
E.Top SQL queries by DB Load.
AnswersC, E

DB Load by Wait Events splits database load across wait categories, exposing which resource contention spiked. It satisfies the stem's need to pinpoint the root cause of a sudden load increase rather than merely confirming that load rose.

Why this answer

The correct answers are C and E. DB Load by Wait Events (C) breaks down the total load into wait event categories (e.g., CPU, I/O, locks), directly identifying the resource causing the bottleneck. Top SQL queries by DB Load (E) shows which specific SQL statements contribute the most to the load, enabling targeted optimization.

Together, they provide a complete root cause analysis. Options A and D (Read/Write IOPS, CPUUtilization) are infrastructure-level metrics that may indicate resource pressure but do not reveal database-level contention or query-level impact. Option B (Average Active Sessions) is the overall load metric itself, not a breakdown, so it does not pinpoint the cause.

Exam trap

The trap here is that candidates confuse 'Average Active Sessions' (the overall load metric) with 'DB Load by Wait Events' (the breakdown), or they mistakenly think raw I/O metrics like IOPS are sufficient to diagnose database-level contention, when in fact wait event analysis is required to isolate the specific resource bottleneck.

20
Multi-Selectmedium

A company is using Amazon CloudFront to deliver content from an S3 bucket. The SysOps administrator wants to restrict access so that only CloudFront can access the S3 bucket. Which TWO steps should be taken?

Select 2 answers
A.Generate presigned URLs for all objects in the S3 bucket.
B.Configure the S3 bucket policy to grant the OAI s3:GetObject permission.
C.Configure CloudFront signed URLs to limit viewer access.
D.Create an Origin Access Identity (OAI) for the CloudFront distribution.
E.Set the S3 bucket policy to allow access only from the CloudFront distribution ID.
AnswersB, D

Configuring the S3 bucket policy to grant the OAI s3:GetObject permission is the critical step that makes the origin access control effective. The policy explicitly identifies the OAI as the only principal allowed to read objects, which permits CloudFront to fetch content on behalf of viewers while denying all direct S3 access requests. This is the recommended pattern because it combines the OAI identity with the necessary authorization, ensuring that the S3 bucket remains private and only CloudFront can serve content.

Why this answer

To restrict access so that only CloudFront can access the S3 bucket, the correct steps are to create an Origin Access Identity (OAI) for the CloudFront distribution (option D) and then configure the S3 bucket policy to grant the OAI s3:GetObject permission (option B). This ensures that only the CloudFront distribution with that OAI can read objects from the bucket, while all other principals are denied access. Option A is incorrect because presigned URLs grant temporary access to individual users, not to CloudFront.

Option C is incorrect because signed URLs control viewer access, not origin access. Option E is incorrect because bucket policies reference the OAI, not the distribution ID.

21
MCQhard

A company uses a centralized logging solution with Amazon OpenSearch Service. The log volume has grown significantly, increasing costs. The logs are retained for 90 days for compliance, but only the last 30 days are frequently accessed. Which combination of actions would reduce costs without compromising compliance?

A.Move the logs to Amazon S3 Glacier and use a Lambda function to query them.
B.Increase the number of data nodes to improve indexing performance.
C.Migrate indices older than 30 days to UltraWarm nodes.
D.Configure an index lifecycle policy to delete indices older than 30 days.
AnswerC

UltraWarm nodes in Amazon OpenSearch Service provide a warm storage tier that is dramatically cheaper per GiB than hot storage but still fully searchable, because it uses a combination of Amazon S3 and a small caching layer. By creating an Index State Management (ISM) policy that moves indices older than 30 days to UltraWarm, the company retains the ability to query compliance-related logs without paying for high-performance EBS-backed hot nodes. This meets both the cost-reduction and the 90-day compliance/retention requirements, since the data remains accessible and the policy can later delete or transition it as needed.

Why this answer

Migrating indices older than 30 days to UltraWarm nodes reduces storage costs because UltraWarm provides cost-effective storage for infrequently accessed data. This retains data for the full 90-day compliance period while keeping the last 30 days in hot storage for fast access. Option A is incorrect because moving logs to S3 Glacier would make querying impractical and is not designed for OpenSearch.

Option B is incorrect because increasing data nodes would increase costs without addressing storage optimization. Option D is incorrect because deleting indices after 30 days would violate the 90-day retention requirement.

22
MCQeasy

A company has a fleet of EC2 instances in an Auto Scaling group behind an Application Load Balancer. The security team requires that all traffic to the instances be encrypted in transit. Currently, the ALB terminates HTTPS and forwards HTTP to the instances. The security team wants to ensure that the traffic between the ALB and the instances is also encrypted. What should the SysOps administrator do to meet this requirement with minimal changes?

A.Replace the ALB with a Network Load Balancer and use TLS termination on the instances.
B.Place a CloudFront distribution in front of the ALB and use HTTPS for all origins.
C.Set up a VPN connection between the ALB and the instances.
D.Change the ALB listener to use HTTPS and configure the target group to use HTTPS with a self-signed certificate on the instances.
AnswerD

This option correctly addresses the requirement by enabling TLS at both ends of the connection: the listener uses HTTPS to encrypt traffic from clients to the ALB, and the target group is configured with the HTTPS protocol so the ALB re-encrypts traffic to the instances. The instances use a self-signed certificate because the ALB does not need to validate it against a public CA; it simply establishes an encrypted channel, and you can optionally configure the ALB to verify the certificate if desired.

Why this answer

It encrypts traffic between the ALB and EC2 instances with minimal changes: you change the ALB listener to HTTPS and configure the target group to use HTTPS, using a self-signed certificate on the instances for encryption. Option A is incorrect because replacing the ALB with a Network Load Balancer is not minimal and adds complexity. Option B is incorrect because adding a CloudFront distribution introduces additional cost and overhead, not minimal.

Option C is incorrect because setting up a VPN connection is overly complex and not necessary for this requirement.

23
MCQmedium

A company uses AWS CodeDeploy to deploy a web application to an Auto Scaling group. The deployment fails with the error 'The overall deployment failed because too many individual instances failed deployment, too few healthy instances are available for deployment, or some instances in your deployment group are experiencing problems.' The deployment group has a minimum of 2 healthy instances. The Auto Scaling group has 4 instances. What should the SysOps administrator check first?

A.Review the deployment configuration to ensure it allows enough time for deployment.
B.Verify that the AppSpec file includes the correct hooks.
C.Check the Auto Scaling group's health check type and ensure it is set to ELB.
D.Check the IAM role for CodeDeploy to ensure it has sufficient permissions.
AnswerC

When CodeDeploy deploys to an Auto Scaling group, it relies on the group's health check type to determine whether each instance is healthy and to maintain the minimum healthy instance count. If the health check type is left as EC2 (the default), an instance is considered healthy solely because it is in the 'running' state — even if the deployed web application is not responding on the Elastic Load Balancer. Setting the health check to ELB forces the Auto Scaling group to use the load balancer's target health checks, which validate the actual application response. This directly affects the 'insufficient healthy instances' error, so checking and correcting this setting is the correct first step.

Why this answer

The error message indicates that too few healthy instances are available for deployment. In an Auto Scaling group integrated with CodeDeploy, the health check type must be set to ELB to allow CodeDeploy to use Elastic Load Balancing health checks to determine instance health. If the health check type is set to EC2 (default), CodeDeploy may consider instances healthy even when they are not passing ELB health checks, causing the deployment to fail due to insufficient healthy instances.

Exam trap

The trap here is that candidates often assume the error is due to a misconfigured AppSpec file or insufficient permissions, overlooking the critical integration between CodeDeploy and the Auto Scaling group's health check type, which directly controls how CodeDeploy counts healthy instances during deployment.

How to eliminate wrong answers

Option A is wrong because the error is about insufficient healthy instances, not about deployment timeout; adjusting the deployment configuration timeout would not resolve a health check mismatch. Option B is wrong because the AppSpec file hooks control lifecycle events (e.g., BeforeInstall, AfterInstall) but do not affect how CodeDeploy determines instance health or the minimum healthy hosts requirement. Option D is wrong because insufficient IAM permissions would typically result in an access denied or authorization error, not the specific 'too few healthy instances' error described.

24
Multi-Selectmedium

A SysOps administrator is designing a VPC with public and private subnets. The private subnets need to access the internet for software updates. Which THREE components are required to achieve this?

Select 3 answers
A.A VPC Gateway Endpoint
B.An Internet Gateway attached to the VPC
C.A NAT Gateway in a public subnet
D.A Site-to-Site VPN connection
E.A route table in the private subnet with a default route to the NAT Gateway
AnswersB, C, E

An Internet Gateway (IGW) is the linchpin of public internet access in a VPC; it is a horizontally scaled, redundant, and highly available target for the 0.0.0.0/0 default route placed in public subnet route tables. It performs NAT for instances that have public IP addresses, enabling bidirectional traffic with the internet, and it is also required for a NAT Gateway in a public subnet to receive and send traffic. Because it connects the VPC directly to the internet, it is essential for any design where private subnets need egress.

Why this answer

Option B is correct because an Internet Gateway must be attached to the VPC to provide the public subnet (and the NAT Gateway within it) with a path to the internet; without an IGW, no traffic can reach external destinations. Option C is correct because a NAT Gateway deployed in a public subnet performs source NAT, allowing instances in private subnets to initiate outbound internet traffic while remaining unreachable from the internet. Option E is correct because the private subnet's route table must contain a default route (0.0.0.0/0) targeting the NAT Gateway so that outbound traffic from private instances is forwarded to it.

Option A is incorrect because a VPC Gateway Endpoint only provides private access to specific AWS services such as S3 or DynamoDB, not general internet access. Option D is incorrect because a Site-to-Site VPN connects the VPC to an on-premises network, not to the public internet for software updates.

Exam trap

The trap is assuming a VPC Gateway Endpoint or a VPN can provide general internet access; candidates often forget that a NAT Gateway alone is insufficient without the Internet Gateway and the correct route table entry.

25
MCQhard

A SysOps administrator needs to ensure that all S3 buckets in the account are logged to CloudTrail for data events. The administrator enables CloudTrail with data events for S3 and selects 'All buckets' in the current account. However, after a week, they notice that some buckets are not being logged. What is the most likely reason?

A.The IAM user who created the trail does not have s3:PutObject permissions on the buckets.
B.The S3 buckets do not have a bucket policy that allows CloudTrail to write the log files.
C.The S3 buckets are in a different AWS Region from the CloudTrail trail.
D.The S3 buckets have server access logging enabled, which conflicts with CloudTrail logging.
AnswerC

CloudTrail trails operate on a regional scope unless explicitly created as multi-region. When a trail is active in only one AWS Region, it records S3 data events exclusively for buckets located in that same Region; bucket-level and object-level operations in other Regions never appear in the trail's delivered log files. To capture data events for all buckets, the trail must be configured as multi-region or separate trails must be created per Region. Thus, having buckets spread across Regions while the trail is single-Region explains why only some buckets are missing logs.

Why this answer

The most likely reason some buckets are not being logged is that those buckets reside in a different AWS Region than the one where the CloudTrail trail is configured. If the administrator creates a single-region trail and selects ‘All buckets’ for data events, the trail will only capture data events for S3 buckets in that specific region. Buckets in other regions will not be logged, resulting in partial logging coverage.

Option C correctly identifies this regional mismatch.

Exam trap

The trap is that candidates may assume selecting ‘All buckets’ automatically includes buckets across all regions, but trails need to be configured as multi-region or appropriate regional trails must be created to cover all regions.

How to eliminate wrong answers

Option A is wrong because the IAM user who created the trail does not need s3:PutObject permissions on the buckets being logged; CloudTrail itself writes the logs to the destination bucket, and the trail creation only requires permissions to create the trail and configure logging, not to write to each source bucket. Option C is wrong because CloudTrail can log data events for S3 buckets in any region as long as the trail is configured with 'All buckets' or a bucket ARN that includes the region; regional mismatch does not prevent logging. Option D is wrong because server access logging and CloudTrail data event logging are independent features that can coexist on the same bucket without conflict; enabling one does not disable the other.

26
Multi-Selecthard

A company runs a stateless web application on EC2 instances behind an Application Load Balancer. The SysOps Administrator needs to ensure the application can withstand the loss of an entire Availability Zone. Which THREE steps should be taken? (Choose THREE.)

Select 3 answers
A.Enable cross-zone load balancing on the ALB.
B.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
C.Ensure the ALB is configured to route traffic to all enabled AZs.
D.Configure the Auto Scaling group to use a dynamic scaling policy based on CPU utilization.
E.Use an Elastic IP address for each EC2 instance.
AnswersB, C, D

Specifying subnets in at least two Availability Zones within the Auto Scaling group causes Amazon EC2 Auto Scaling to launch instances across those zones, ensuring that if one Availability Zone fails, the remaining zones continue serving traffic. The Auto Scaling group also uses health checks to detect failed instances and automatically launches replacements in the other healthy Availability Zones, maintaining capacity and availability. This is the foundational design for a fault-tolerant, stateless web tier and directly addresses the requirement to survive an Availability Zone outage.

Why this answer

Configuring the Auto Scaling group to launch instances in at least two Availability Zones (AZs) ensures that if one entire AZ fails, the remaining AZ(s) still have running instances to serve traffic. This is a fundamental pattern for high availability and fault tolerance in AWS, as it distributes the application across physically separate data centers within a region.

Exam trap

The trap here is that candidates often confuse cross-zone load balancing (which optimizes traffic distribution within healthy AZs) with multi-AZ deployment (which ensures instance availability across AZs), leading them to incorrectly select Option A as a solution for AZ failure.

27
Multi-Selectmedium

A company is designing a disaster recovery strategy for a production RDS for MySQL database. The database is currently single-AZ. The recovery point objective (RPO) is 1 hour, and the recovery time objective (RTO) is 15 minutes. Which steps should the SysOps administrator take to meet these requirements? (Choose THREE.)

Select 3 answers
A.Disable automated backups to reduce performance impact.
B.Take manual DB snapshots every hour.
C.Enable automated backups with a retention period of at least 1 day.
D.Modify the DB instance to be Multi-AZ.
E.Create a read replica in a different Availability Zone.
AnswersB, C, D

Taking manual DB snapshots every hour creates a discrete recovery point every 60 minutes, so in the worst case you lose at most one hour of data, satisfying the stated RPO. Restoring from a manual snapshot provisions a new DB instance from that exact point; the RTO includes the time to restore the snapshot and update connection strings, which can be several minutes but is often acceptable for a recovery plan. Manual snapshots are stored in S3 and persist until you delete them, giving you a reliable backup mechanism independent of automated backup retention settings.

Why this answer

Manual DB snapshots can be taken on demand, and taking them every hour ensures that the recovery point objective (RPO) of 1 hour is met. In the event of a failure, you can restore the database from the latest manual snapshot, which provides a point-in-time recovery point within the RPO window. However, manual snapshots alone do not meet the 15-minute recovery time objective (RTO), so they must be combined with other measures like Multi-AZ and automated backups.

Exam trap

The trap here is that candidates often think a read replica or manual snapshots alone can meet both RPO and RTO, but they fail to recognize that Multi-AZ is required for the low RTO and automated backups are needed for the granular RPO.

28
MCQeasy

A SysOps administrator needs to create a VPC with both public and private subnets. The public subnet will host a NAT gateway and a bastion host. The private subnet will host application servers that need outbound internet access for updates. Which routing configuration should the administrator implement?

A.Public subnet route table: 0.0.0.0/0 -> Internet Gateway; Private subnet route table: 0.0.0.0/0 -> Internet Gateway via the NAT Gateway.
B.Public subnet route table: 0.0.0.0/0 -> Internet Gateway; Private subnet route table: 0.0.0.0/0 -> Internet Gateway.
C.Public subnet route table: 0.0.0.0/0 -> NAT Gateway; Private subnet route table: 0.0.0.0/0 -> Internet Gateway.
D.Public subnet route table: 0.0.0.0/0 -> Internet Gateway; Private subnet route table: 0.0.0.0/0 -> NAT Gateway.
AnswerD

This is the correct setup for a VPC with public and private subnets. The public subnet route table sends all outbound traffic (0.0.0.0/0) to the internet gateway, allowing resources like a bastion host or NAT gateway to reach the internet directly. The private subnet route table sends all outbound traffic to the NAT gateway, which resides in the public subnet and performs source network address translation (SNAT) to forward traffic to the internet while keeping instances in the private subnet unreachable from the internet. This preserves the security of private instances while still enabling them to download updates or access external services.

Why this answer

The public subnet needs a route to the Internet Gateway so the NAT gateway and bastion host are reachable from the internet. The private subnet needs a route to the NAT gateway (which itself lives in the public subnet) so application servers can initiate outbound internet traffic without being directly reachable inbound.

Exam trap

SOA-C02 often tests whether candidates correctly place the NAT gateway in the public subnet and point the private subnet's default route at the NAT gateway — distractors swap the IGW and NAT gateway targets to catch memorized-but-unverified answers.

How to eliminate wrong answers

Option A is wrong because it describes routing the private subnet to the Internet Gateway 'via the NAT Gateway' — a NAT gateway is not a path to an IGW; the private route target must be the NAT gateway's ENI, not the IGW. Option B is wrong because routing the private subnet directly to the Internet Gateway makes it a public subnet, defeating the purpose and exposing the app servers. Option C is wrong because it reverses the roles: a NAT gateway cannot serve as the public subnet's default route (it has no inbound path from the internet), and the private subnet pointed at the IGW would again be public.

29
MCQeasy

A company stores sensitive data in an RDS database. Which AWS service should be used to encrypt the database at rest?

A.AWS Certificate Manager (ACM)
B.AWS Identity and Access Management (IAM)
C.AWS Key Management Service (KMS)
D.AWS CloudHSM
AnswerC

AWS Key Management Service (KMS) is a managed service for creating and controlling customer master keys (CMKs) that encrypt data at rest across AWS services, including Amazon RDS. When you enable encryption on an RDS instance, RDS uses a KMS CMK to encrypt the underlying EBS storage, automated backups, snapshots, and read replicas, with encryption handled transparently by the service. KMS is the only service among these options that natively integrates with RDS for at-rest encryption, making it the correct choice.

Why this answer

AWS Key Management Service (KMS) is the AWS service that manages the customer master keys (CMKs) used to encrypt RDS databases at rest. When you enable encryption on an RDS instance, you select a KMS key, and RDS uses that key to encrypt the underlying storage, snapshots, and read replicas. KMS integrates natively with RDS, EBS, S3, and most other AWS services for at-rest encryption.

Exam trap

SOA-C02 often tests the confusion between encryption in transit (ACM/TLS) and encryption at rest (KMS), so candidates who see 'certificate' or 'key' and pick ACM or CloudHSM instead of KMS lose the point.

How to eliminate wrong answers

Option A is wrong because AWS Certificate Manager (ACM) issues and manages TLS/SSL certificates for encryption in transit, not at-rest data encryption. Option B is wrong because IAM handles authentication and authorization (who can do what), not cryptographic key management or data encryption. Option D is wrong because AWS CloudHSM is a dedicated hardware security module for customers with strict key custody requirements; while it can be used as a custom key store behind KMS, it is not the service you directly select to encrypt an RDS database.

30
Multi-Selecthard

A company runs a web application on EC2 instances in an Auto Scaling group. The application uses an Amazon RDS Multi-AZ DB instance. The SysOps administrator notices that during a recent failover test, the application became unresponsive for several minutes. The administrator wants to improve the application's resilience during failover. Which three actions should the administrator take? (Choose THREE.)

Select 3 answers
A.Configure the Application Load Balancer health checks to have a low threshold (e.g., 2 consecutive failures) and a short interval (e.g., 5 seconds).
B.Implement retry logic in the application to handle transient database connection failures.
C.Change the RDS DB instance to use asynchronous replication instead of synchronous replication.
D.Increase the EC2 instance size to handle more connections during failover.
E.Configure an Amazon RDS Proxy in front of the RDS database to pool and share database connections.
AnswersA, B, E

Configuring ALB health checks with a low threshold (e.g., 2 consecutive failures) and a short interval (e.g., 5 seconds) enables the load balancer to detect and deregister unhealthy instances within about 10 seconds, instead of the default ~60-90 seconds. This rapid detection prevents the ALB from continuing to route user traffic to instances whose database connections have been terminated during a failover. By quickly redirecting requests to healthy instances, the overall impact on users is minimized, and the application appears more resilient during RDS failover events.

Why this answer

Configuring the Application Load Balancer (ALB) health checks with a low threshold (e.g., 2 consecutive failures) and a short interval (e.g., 5 seconds) allows the ALB to quickly detect unhealthy EC2 instances and stop routing traffic to them. This reduces the time the application spends trying to serve requests through failing instances during an RDS failover, improving overall responsiveness.

Exam trap

The trap here is that candidates often assume increasing instance size (Option D) or changing replication mode (Option C) will improve failover resilience, but neither addresses the core issue of connection handling and rapid health check detection during a database failover.

31
MCQmedium

A company requires that all users in an AWS account must authenticate with multi-factor authentication (MFA) before they can perform any actions on Amazon EC2 instances. The SysOps administrator needs to implement this requirement using IAM policies. Which IAM policy condition key should be used to enforce MFA?

A.aws:SourceIp
B.aws:MultiFactorAuthPresent
C.aws:RequestedRegion
D.iam:PassedToService
AnswerB

The aws:MultiFactorAuthPresent condition key evaluates whether the principal authenticated with MFA during the current session. Attaching it with a Deny or Allow condition in IAM policies blocks non-MFA sessions from performing EC2 actions, enforcing the stated requirement.

Why this answer

The `aws:MultiFactorAuthPresent` condition key checks whether the user authenticated using a valid MFA device before making the API request. By setting this condition to `true` in an IAM policy, you can enforce that all actions on EC2 instances require MFA authentication, meeting the company's requirement.

Exam trap

The trap here is that candidates often confuse `aws:MultiFactorAuthPresent` with `aws:SourceIp` or `iam:PassedToService`, thinking IP-based or role-passing conditions can enforce MFA, but only the MFA-specific condition key directly checks authentication strength.

How to eliminate wrong answers

Option A is wrong because `aws:SourceIp` restricts access based on the source IP address, not MFA status. Option C is wrong because `aws:RequestedRegion` limits actions to specific AWS regions, not MFA enforcement. Option D is wrong because `iam:PassedToService` controls which roles can be passed to AWS services, not MFA authentication.

32
MCQmedium

A company uses Amazon S3 to store sensitive customer data. A SysOps administrator needs to ensure that any S3 bucket that is incorrectly configured to allow public read access is automatically remediated within five minutes. The administrator wants to use native AWS services with minimal custom code. Which solution should be used?

A.Use AWS Config with the 's3-bucket-public-read-prohibited' managed rule and configure automatic remediation to block public access.
B.Create an Amazon EventBridge (CloudWatch Events) rule that triggers an AWS Lambda function to check and fix public read access.
C.Apply an S3 bucket policy to each bucket that denies public read access.
D.Use AWS Trusted Advisor to check for public read access and manually remediate when notified.
AnswerA

AWS Config's 's3-bucket-public-read-prohibited' managed rule evaluates every S3 bucket against the defined parameter (blocking public read access) on a continuous basis. Because it is a managed rule, there is no custom code to write or maintain, and when coupled with automatic remediation (using an AWS Systems Manager Automation document that applies the 'block all public access' setting or removes bucket policies), noncompliant buckets are corrected within minutes. This is the only option that provides both automated detection and automated remediation using a pre-built, low-maintenance AWS service, meeting the five-minute requirement without manual intervention.

Why this answer

AWS Config with the 's3-bucket-public-read-prohibited' managed rule can automatically evaluate S3 bucket configurations against the desired state. When a non-compliant bucket is detected, AWS Config can trigger an automatic remediation action (e.g., applying an S3 bucket policy or blocking public access) using AWS Systems Manager Automation documents, all within the required five-minute window and with minimal custom code.

Exam trap

The trap here is that candidates often choose EventBridge + Lambda (Option B) because it seems more flexible, but they overlook the 'minimal custom code' constraint and the fact that AWS Config's managed rule with automatic remediation is a fully native, code-free solution.

How to eliminate wrong answers

Option B is wrong because while EventBridge and Lambda can achieve the goal, they require custom code (Lambda function) and manual setup, which contradicts the 'minimal custom code' requirement. Option C is wrong because applying a bucket policy to each bucket is a manual, one-time action that does not provide automatic detection and remediation of newly created or misconfigured buckets. Option D is wrong because Trusted Advisor provides only manual checks and notifications; it cannot automatically remediate misconfigurations, and relying on manual remediation violates the 'automatically remediated within five minutes' requirement.

33
MCQmedium

An environment has 12 individual CloudWatch metric alarms covering CPU, memory, disk, and network. When one instance degrades, all 12 alarms fire simultaneously and send 12 separate notifications to the on-call engineer. The team wants a single notification per incident regardless of how many individual alarms trigger. What CloudWatch feature addresses this?

A.Create a composite alarm that enters ALARM state when any of the 12 child alarms is in ALARM state, and configure a single SNS action on the composite alarm only
B.Increase the alarm evaluation period on all 12 alarms to 30 minutes so they fire less frequently
C.Use an SNS topic with a delivery policy that batches notifications sent within a 60-second window
D.Configure all 12 alarms to write to the same CloudWatch Events rule and suppress duplicate events with EventBridge deduplication
AnswerA

The composite alarm's rule expression 'ALARM(alarm1) OR ALARM(alarm2) OR ...' triggers when any child fires. By routing all notifications through the composite alarm's action and removing actions from the child alarms, exactly one notification is sent per incident. Child alarm states remain visible in the console for root cause analysis.

Why this answer

A composite alarm in CloudWatch can aggregate multiple child alarms into a single parent alarm. When any of the 12 child alarms enters the ALARM state, the composite alarm transitions to ALARM and triggers a single SNS notification, thereby reducing alert noise to one notification per incident.

Exam trap

The trap here is that candidates may think SNS batching or EventBridge deduplication can consolidate separate alarm notifications, but those services do not aggregate distinct alarm state changes into a single event; only composite alarms provide that logical grouping.

How to eliminate wrong answers

Option B is wrong because increasing the evaluation period to 30 minutes does not consolidate multiple notifications into one; it merely delays the alarms, and all 12 would still fire individually after the longer period. Option C is wrong because SNS delivery policies control retries and message batching for HTTP/HTTPS endpoints, not deduplication or aggregation of separate alarm notifications; each alarm still sends its own message to the topic. Option D is wrong because CloudWatch Events (now EventBridge) can route alarm state changes to targets, but EventBridge deduplication applies to events based on a deduplication ID and is designed for idempotent event processing, not for collapsing multiple distinct alarm events into a single notification.

34
Multi-Selecteasy

A company needs to monitor the CPU and memory utilization of its EC2 instances. Which TWO services can be used to collect and visualize these metrics?

Select 2 answers
A.Amazon CloudWatch
B.AWS CloudTrail
C.Amazon CloudWatch Agent
D.AWS Config
E.AWS Systems Manager
AnswersA, C

Amazon CloudWatch is the core AWS monitoring service that automatically collects and stores EC2 CPU utilization metrics at 5-minute intervals (or 1-minute with detailed monitoring) and provides dashboards, alarms, and API access to those time-series data points. It is the appropriate service for monitoring performance because it aggregates telemetry from AWS infrastructure and can extend to in-guest metrics when a signal is pushed back to it. Without CloudWatch, there would be no central repository or visualization for the CPU and memory data that the CloudWatch agent submits.

Why this answer

Amazon CloudWatch is the native AWS monitoring service that collects and stores metrics such as CPU utilization and memory utilization from EC2 instances. However, by default, CloudWatch only captures hypervisor-level metrics (like CPU) and not in-guest metrics (like memory utilization). To collect memory utilization, you must install the Amazon CloudWatch Agent on the instance, which sends custom metrics to CloudWatch.

Together, CloudWatch and the CloudWatch Agent provide both collection and visualization of CPU and memory metrics.

Exam trap

The trap here is that candidates often assume CloudWatch alone collects all EC2 metrics, but they miss that memory utilization requires the CloudWatch Agent because it is an in-guest metric not provided by the hypervisor.

35
MCQeasy

An application runs on c5.xlarge EC2 instances 24 hours a day, 7 days a week in us-east-1. The workload is stable and will not change instance type for at least 12 months. The team wants to reduce compute costs by 30 to 40 percent compared to On-Demand pricing. Which purchasing option achieves this with the lowest financial risk?

A.Purchase a 1-year Standard Reserved Instance for c5.xlarge in us-east-1 with All Upfront or Partial Upfront payment
B.Use Spot Instances with an interruption tolerance of 5 minutes for the workload
C.Enable EC2 Auto Scaling with a target tracking policy to scale down to zero instances during off-peak hours
D.Purchase a 3-year Convertible Reserved Instance to maximize the discount percentage
AnswerA

A 1-year Standard RI matches the 12-month stability horizon and delivers 30–40 percent savings versus On-Demand. All Upfront provides the deepest discount; Partial Upfront reduces the upfront cash requirement with a slightly lower overall saving. The 1-year commitment limits risk compared to a 3-year commitment for an uncertain future period.

Why this answer

A 1-year Standard Reserved Instance (RI) with All Upfront or Partial Upfront payment offers a 30-40% discount over On-Demand pricing for a stable, always-on workload. This option provides the lowest financial risk because it commits to a fixed instance type and region for only one year, matching the workload's stable nature without the flexibility premium of Convertible RIs or the interruption risk of Spot Instances.

Exam trap

The trap here is that candidates may choose the 3-year Convertible RI (Option D) for its higher discount percentage, overlooking the fact that the longer commitment and unnecessary flexibility introduce greater financial risk for a stable, unchanging workload.

How to eliminate wrong answers

Option B is wrong because Spot Instances can be interrupted with as little as a 5-minute warning, which introduces significant financial and operational risk for a workload that must run 24/7 without interruption. Option C is wrong because scaling down to zero instances during off-peak hours would violate the requirement that the application runs 24/7, and it does not address the need to reduce costs for the always-on baseline. Option D is wrong because a 3-year Convertible Reserved Instance, while offering a higher discount percentage, introduces greater financial risk due to the longer commitment period and the unnecessary flexibility to change instance types, which the workload does not require.

36
MCQmedium

A SysOps administrator needs to deploy a web application across multiple AWS Regions for disaster recovery. The application uses Amazon RDS for MySQL and requires a secondary database in a different Region. What is the MOST cost-effective and automated solution to keep the databases synchronized?

A.Create a cross-Region read replica of the primary RDS instance in the secondary Region
B.Use AWS Database Migration Service (DMS) with ongoing replication
C.Set up a cron job on an EC2 instance to export the database and import it into the secondary Region
D.Enable Multi-AZ on the primary RDS instance and configure a read replica in the secondary Region
AnswerA

Amazon RDS cross-Region read replicas use asynchronous replication to continuously copy data from the primary MySQL instance to a read-only instance in a second AWS Region. In a disaster, you can promote the replica to become a standalone primary, providing a low RPO and RTO without custom scripts or manual database dumps. This is the managed, native, and cost-effective mechanism for cross-Region disaster recovery of RDS.

Why this answer

A cross-Region read replica of the primary RDS instance is the most cost-effective and automated solution because RDS natively replicates data asynchronously to the replica in another Region with minimal configuration. It requires no additional replication infrastructure or ongoing management, and it can be promoted to a standalone database during disaster recovery. This meets the requirement for automated, cost-effective cross-Region synchronization.

Exam trap

SOA-C02 often tests the confusion between Multi-AZ (same-Region HA) and cross-Region read replicas (DR), so candidates must recognize that Multi-AZ does not provide cross-Region replication.

How to eliminate wrong answers

Option B is wrong because AWS DMS with ongoing replication is more complex and costly, requiring replication instances and ongoing management, and is typically used for heterogeneous migrations or when native replication is unavailable. Option C is wrong because a cron job with export/import is manual, error-prone, not automated, and introduces significant data loss and downtime. Option D is wrong because Multi-AZ is a synchronous standby within the same Region for high availability, not cross-Region disaster recovery; adding a read replica in another Region is part of the correct solution but Multi-AZ itself does not provide cross-Region replication.

37
MCQmedium

A company uses AWS CloudFormation to deploy infrastructure. The operations team wants to be notified when a stack update fails. What is the simplest way to achieve this?

A.Enable CloudTrail and create a metric filter for 'UpdateStack' events, then set an alarm.
B.Write a script that periodically checks the CloudFormation console for stack status and sends an email.
C.Create an Amazon EventBridge rule that matches CloudFormation events and triggers a Lambda function to send an SNS notification.
D.Configure an SNS topic in the CloudFormation stack's notification options.
AnswerD

CloudFormation natively supports sending stack lifecycle events (such as stack creation, update, and delete failures) directly to an Amazon SNS topic through the stack's notification options. When you create a stack, you can specify one or more SNS topic ARNs in the 'NotificationARNs' property, and CloudFormation publishes all stack events to those topics in real-time. For example, a 'CREATE_FAILED' or 'UPDATE_FAILED' event triggers a notification that is delivered to the SNS topic, and subscribers (such as email or Lambda) receive it immediately. This is the simplest, most direct, and most reliable way to get notified of stack failures, as it requires no additional services or custom scripting.

Why this answer

CloudFormation natively supports specifying an SNS topic in the stack's notification options, which automatically sends notifications on stack events such as failures, without requiring any additional services or custom code. This is the simplest and most direct method to notify the operations team when a stack update fails.

Exam trap

The trap here is that candidates often over-engineer the solution by choosing EventBridge or CloudTrail-based approaches, overlooking CloudFormation's built-in SNS notification feature as the simplest and most direct option.

How to eliminate wrong answers

Option A is wrong because CloudTrail logs API calls but does not directly trigger notifications; creating a metric filter and alarm adds unnecessary complexity when a built-in notification mechanism exists. Option B is wrong because writing a script to poll the CloudFormation console is inefficient, introduces latency, and violates the principle of using event-driven notifications over polling. Option C is wrong because while EventBridge with Lambda and SNS can work, it is more complex than the native SNS integration and requires custom code, making it not the simplest solution.

38
MCQhard

A team uses AWS CodeDeploy with a deployment configuration of CodeDeployDefault.OneAtATime to deploy a web application to an Auto Scaling group. Instances are behind an Application Load Balancer. The deployment fails with 'The overall deployment failed because too many individual instances failed deployment.' What is the most likely cause?

A.The health check grace period on the Auto Scaling group is too short.
B.The target group deregistration delay is too long.
C.The CodeDeploy agent is not installed on the instances.
D.The deployment group is configured to skip the ELB health check.
AnswerA

The health check grace period on the Auto Scaling group is too short. When a deployment launches new instances, the ASG considers an instance healthy only after the grace period expires; if the period is shorter than the time CodeDeploy needs to install the application and pass its own validation, the ASG will prematurely flag the instance as failing ELB health checks. Auto Scaling then terminates and replaces the instance mid-deployment, which CodeDeploy sees as a failed deployment ("too many individual instances"), and the cycle repeats for each new replacement. The correct fix is to increase the grace period to exceed the typical deployment duration.

Why this answer

The deployment fails because the health check grace period on the Auto Scaling group is too short. When CodeDeploy deploys one instance at a time (CodeDeployDefault.OneAtATime), the instance is taken out of service, updated, and then returned to the load balancer. If the grace period expires before the instance passes its health checks, the Auto Scaling group marks it as unhealthy and terminates it, causing the deployment to fail with 'too many individual instances failed.'

Exam trap

The trap here is that candidates often confuse the health check grace period with the deregistration delay or assume the issue is with the CodeDeploy agent, but the specific error 'too many individual instances failed' points to Auto Scaling terminating instances due to health check failures, not a deployment script or agent problem.

How to eliminate wrong answers

Option B is wrong because a long target group deregistration delay would cause traffic to continue flowing to instances being replaced, but it would not cause instances to be terminated by the Auto Scaling group; it delays the removal of instances from the target group but does not trigger deployment failure. Option C is wrong because if the CodeDeploy agent were not installed, the deployment would fail immediately with an agent connectivity error, not with 'too many individual instances failed' after partial success. Option D is wrong because skipping the ELB health check would prevent the load balancer from routing traffic to the instances, but it would not cause the Auto Scaling group to terminate instances; the deployment would likely succeed but with no traffic, not fail with this specific error.

39
MCQmedium

A company runs a fleet of EC2 instances in a production environment. The instances are part of an Auto Scaling group that uses a launch template with a m5.large instance type. The company's SysOps administrator notices that the instances are often over-provisioned, with average CPU utilization below 20% for the past month. The administrator wants to reduce costs without affecting application performance. The application is stateless and can handle temporary performance degradation. Which action should the administrator take?

A.Purchase Reserved Instances for the current m5.large instances to lower hourly cost.
B.Modify the launch template to use a t3.medium instance type.
C.Enable detailed monitoring on all instances to collect more data.
D.Increase the minimum size of the Auto Scaling group to reduce scale-out events.
AnswerB

Changing the launch template to t3.medium corrects the over-provisioning at the source: t3.medium provides the same 2 vCPU count but only 4 GiB of memory, halving the allocated RAM while remaining eligible for burstable CPU credits. For a workload with low average utilization and occasional spikes, t3 instances can burst using earned CPU credits, then settle back to a low baseline, making the smaller instance type adequate. This directly reduces the per-hour cost for every instance the Auto Scaling group launches, which is the most effective right-sizing action listed.

Why this answer

Modifying the launch template to use a t3.medium instance type is correct because t3 instances are burstable and cost-effective for workloads with low average CPU utilization (below 20%). The application is stateless and can handle temporary performance degradation, making burstable instances suitable. This reduces costs without affecting performance during normal operation, as t3.medium provides a baseline CPU with the ability to burst.

Exam trap

SOA-C02 often tests cost optimization, and candidates may choose Reserved Instances thinking it reduces cost, but it only lowers the hourly rate for existing over-provisioned instances, not addressing the root cause of over-provisioning.

How to eliminate wrong answers

Option A is wrong because purchasing Reserved Instances for over-provisioned m5.large instances locks in costs for underutilized resources, not reducing waste. Option C is wrong because enabling detailed monitoring only provides more granular metrics; it does not reduce costs. Option D is wrong because increasing the minimum size of the Auto Scaling group would increase the number of instances, raising costs further.

40
MCQmedium

A SysOps administrator notices that the monthly bill for Amazon S3 has increased significantly. The company uses S3 for storing application logs and user uploads. The logs are accessed rarely but must be retained for 3 years. User uploads are accessed frequently for the first 30 days, then rarely after. Which S3 lifecycle policy will optimize storage costs?

A.Transition logs to S3 Glacier Deep Archive after 30 days, and transition user uploads to S3 Standard-IA after 30 days, then to Glacier Deep Archive after 90 days.
B.Transition user uploads to S3 Glacier Deep Archive after 30 days, and transition logs to S3 Glacier after 90 days.
C.Move logs to S3 Glacier Deep Archive after 30 days, and delete user uploads after 1 year.
D.Transition logs to S3 Standard-IA after 30 days, and transition user uploads to S3 One Zone-IA after 30 days.
AnswerA

This is correct because logs are typically append-only and rarely accessed after the first 30 days, making S3 Glacier Deep Archive the lowest-cost storage class while satisfying the 3-year retention requirement. User uploads, however, are frequently accessed in the month after upload, so S3 Standard-IA after 30 days reduces cost without sacrificing retrieval performance; then transitioning to Glacier Deep Archive after 90 days aligns with the access drop-off and keeps lifecycle costs minimal.

Why this answer

The correct policy matches storage class to access patterns: logs are rarely accessed but retained for three years, so transitioning them to Glacier Deep Archive after 30 days minimizes cost for long-term archival. User uploads are frequently accessed for 30 days, then rarely, so moving them to Standard-IA after 30 days and then to Glacier Deep Archive after 90 days aligns cost with declining access frequency while preserving durability. This combination addresses both data types with the cheapest suitable tiers.

Exam trap

SOA-C02 often tests whether candidates match access frequency and retention requirements to the correct storage class — the trap is choosing the cheapest tier for data that is still frequently accessed, or deleting data that must be retained.

How to eliminate wrong answers

Option B is wrong because it moves user uploads to Glacier Deep Archive after 30 days even though they are accessed frequently during that period, and it delays log transition to 90 days, wasting cost on Standard storage for rarely accessed logs. Option C is wrong because deleting user uploads after one year violates the implied retention need and destroys data that may still be required. Option D is wrong because Standard-IA for logs is more expensive than Deep Archive for rarely accessed data, and One Zone-IA for user uploads sacrifices durability (single AZ) without addressing the long-term archival requirement.

41
MCQmedium

A company uses AWS CodePipeline to automate the deployment of a web application. The pipeline consists of a source stage (AWS CodeCommit) and a deploy stage (AWS CodeDeploy) that deploys to an Auto Scaling group. The SysOps administrator needs to add a stage to run automated unit tests before the deployment proceeds. The tests must be executed in an isolated environment, and if they fail, the pipeline must stop and notify the development team. Which action should the administrator take?

A.Add a manual approval action between the source and deploy stages. The development team will manually run the tests on their local machines and then approve the pipeline to proceed.
B.Insert a test stage after the source stage with an AWS CloudFormation action that deploys a test stack and runs tests using a custom resource Lambda function.
C.Add a stage between source and deploy that uses an AWS CodeBuild action to run unit tests defined in a buildspec file. The pipeline will automatically stop if the build action fails.
D.Add a Lambda function as an action in the pipeline that runs the unit tests. The Lambda function writes the test results to an S3 bucket, and a subsequent approval action checks the results.
AnswerC

CodeBuild is the ideal service for running automated tests in a controlled environment. It integrates natively with CodePipeline: if the CodeBuild build fails, the pipeline transitions to a failed state, stopping further execution and optionally sending notifications via Amazon SNS.

Why this answer

AWS CodeBuild is natively integrated with CodePipeline to run automated tests defined in a buildspec file. When the build action fails, CodePipeline automatically stops the pipeline execution and can send notifications via Amazon SNS, meeting the requirement for an isolated test environment and automatic failure notification without manual intervention.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing CloudFormation or Lambda, overlooking that CodeBuild is the native, simplest, and most cost-effective service for running automated tests within a CodePipeline.

How to eliminate wrong answers

Option A is wrong because it relies on manual approval and local test execution, which violates the requirement for automated tests in an isolated environment and does not provide automatic pipeline stop on test failure. Option B is wrong because using a CloudFormation action to deploy a test stack and run tests via a custom Lambda function adds unnecessary complexity, cost, and latency; it also does not natively integrate with CodePipeline's failure handling as cleanly as CodeBuild. Option D is wrong because a Lambda function action in CodePipeline cannot directly run unit tests in an isolated environment; it would require custom orchestration, and the subsequent approval action would not automatically stop the pipeline on failure—it would only pause for manual review.

42
Multi-Selecteasy

A SysOps administrator wants to monitor the CPU utilization of an Amazon RDS instance and receive an alert if it exceeds 90% for 5 consecutive minutes. Which TWO AWS services are required to set up this monitoring? (Choose TWO.)

Select 2 answers
A.Amazon Simple Notification Service (SNS)
B.AWS Config
C.Amazon RDS Enhanced Monitoring
D.Amazon CloudWatch Alarms
E.Amazon CloudWatch
AnswersD, E

Amazon CloudWatch Alarms monitor a specified CloudWatch metric over a defined period and can trigger actions such as sending an SNS notification or performing an EC2 Auto Scaling action when a threshold is breached. To monitor RDS CPU utilization, you can create an alarm on the CPUUtilization metric with a threshold, comparison operator, and evaluation periods to alert when CPU usage exceeds a desired level. This directly fulfills the SysOps administrator's goal of being notified about CPU utilization.

Why this answer

Amazon CloudWatch is the service that collects and stores metrics such as CPU utilization from RDS instances. Amazon CloudWatch Alarms allow you to set a threshold (e.g., CPU > 90%) and evaluate it over a specified period (e.g., 5 consecutive minutes) to trigger an action, such as sending a notification via SNS.

Exam trap

The trap here is that candidates often confuse Enhanced Monitoring (which provides OS-level metrics) with the standard CloudWatch metrics (which already include CPU utilization), leading them to incorrectly select Enhanced Monitoring as a required service.

43
MCQmedium

A SysOps administrator wants to be alerted when an EC2 instance's status check fails. The instance is part of an Auto Scaling group. What is the BEST approach?

A.Use Amazon EventBridge to detect status check failures.
B.Create a CloudWatch alarm on the 'StatusCheckFailed' metric.
C.Enable CloudTrail to monitor EC2 instance status changes.
D.Configure an Auto Scaling lifecycle hook to send a notification.
AnswerB

A CloudWatch alarm on the StatusCheckFailed metric is the direct and accurate way to detect EC2 health issues because this metric aggregates both the system status check (for underlying hardware/network problems) and the instance status check (for OS/application-level issues). When the alarm enters ALARM state, it can trigger an SNS notification. This metric is available natively for every EC2 instance without extra configuration. Unlike event-driven logs or lifecycle hooks, it specifically measures the instance's operational health, making it the appropriate alert mechanism.

Why this answer

The 'StatusCheckFailed' metric is automatically published by EC2 to CloudWatch, and a CloudWatch alarm on this metric can directly trigger an SNS notification or other action when the status check fails. This is the simplest and most reliable method for alerting on instance health, regardless of whether the instance is in an Auto Scaling group.

Exam trap

The trap here is that candidates often confuse CloudTrail (API logging) with CloudWatch (metrics and alarms), or assume EventBridge is the best choice for all event-driven monitoring, when in fact CloudWatch alarms on the native 'StatusCheckFailed' metric are the simplest and most direct solution for status check alerts.

How to eliminate wrong answers

Option A is wrong because Amazon EventBridge can detect status check failures via EC2 instance state change events, but it does not natively capture the 'StatusCheckFailed' metric; it would require custom event patterns and is less direct than using CloudWatch alarms. Option C is wrong because CloudTrail records API calls (e.g., StartInstances, StopInstances), not status check results, so it cannot detect status check failures. Option D is wrong because Auto Scaling lifecycle hooks are designed for custom actions during instance launch or termination, not for monitoring ongoing instance health or status check failures.

44
MCQeasy

A company is using AWS Lambda functions to process events from Amazon S3. The functions are invoked several thousand times per minute. The SysOps administrator notices that the functions are taking longer to execute during peak times. Which action should the administrator take to improve performance?

A.Increase the Lambda function timeout.
B.Use a larger EC2 instance type for the Lambda functions.
C.Increase the memory allocation for the Lambda functions.
D.Enable provisioned concurrency on the Lambda functions.
AnswerC

Increasing the memory allocation is the correct approach because Lambda allocates a proportional amount of virtual CPU (vCPU) based on the configured memory value, up to the maximum allowed. More memory translates to more compute power, which directly speeds up CPU-bound and memory-intensive code, thereby reducing execution duration. For example, a function with 1,769 MB receives a full vCPU, while higher values receive more, so raising memory is effectively scaling up the function's processing capacity.

Why this answer

Increasing the memory allocation for a Lambda function also proportionally increases its CPU allocation, which can improve execution speed for compute-bound tasks, such as processing events from S3 during peak times. Option A is incorrect because increasing the timeout only allows the function to run longer, but does not make it faster. Option B is incorrect because Lambda abstracts the underlying infrastructure; you cannot choose EC2 instance types.

Option D is incorrect because provisioned concurrency reduces cold starts but does not affect the execution time of warm functions.

45
Multi-Selecthard

An organization uses AWS CloudFormation to manage its infrastructure. The SysOps administrator is implementing a change management process that requires all stack updates to be reviewed and approved before execution. The administrator wants to use CloudFormation change sets to preview changes. Which THREE steps are necessary to implement this process? (Choose THREE.)

Select 3 answers
A.Use the 'Detect Drift' feature to compare the stack with the new template.
B.Update the stack directly using the updated template.
C.Review the change set in the CloudFormation console.
D.Execute the change set after approval.
E.Create a change set from the updated template.
AnswersC, D, E

Reviewing the change set in the CloudFormation console provides a human-readable summary of each resource action (add, modify, or delete) and whether a resource will be replaced. This step allows the administrator to spot potential risks like data loss from deletion or replacement before approving the change set. It is the required review step in a change management workflow, and it directly informs the decision to execute or reject the change set.

Why this answer

To implement a review-and-approval process using CloudFormation change sets, the necessary steps are: create a change set from the updated template (E), review the change set in the CloudFormation console (C), and execute the change set after approval (D). Option A is incorrect because the 'Detect Drift' feature is used to check if the actual stack resources have drifted from the template, not to preview changes. Option B is incorrect because updating the stack directly bypasses the review process, which defeats the purpose of change management.

46
MCQeasy

A SysOps administrator is troubleshooting an issue where an EC2 instance cannot be accessed via SSH from the internet. The security group allows inbound SSH (port 22) from 0.0.0.0/0. The network ACL (NACL) for the subnet has an inbound rule allowing SSH from 0.0.0.0/0. What else could be blocking access?

A.The NACL inbound rule is blocking traffic.
B.The internet gateway is not attached to the VPC.
C.The security group rule is misconfigured.
D.The NACL outbound rule is blocking return traffic.
AnswerD

NACLs are stateless, so the outbound rule is evaluated independently of the inbound rule. Even if the inbound NACL rule allows SSH (port 22) from the client, the instance's response traffic goes to a random ephemeral port (typically 1024–65535) on the client. If the outbound NACL rule does not allow these ephemeral ports, the return packets are dropped, causing the SSH connection to hang or time out. This is the classic cause of asymmetric traffic failures when using stateless filtering.

Why this answer

Network ACLs are stateless, meaning they evaluate inbound and outbound traffic separately. Even if the inbound rule allows SSH, the outbound rule must also allow the return traffic (ephemeral ports) for the SSH session to work. If the NACL outbound rule is blocking return traffic, the SSH connection will fail.

Exam trap

The trap is forgetting that NACLs are stateless; candidates often focus only on inbound rules and overlook the need for outbound rules to allow return traffic, especially for ephemeral ports.

How to eliminate wrong answers

Option A is wrong because the NACL inbound rule already allows SSH from 0.0.0.0/0, so it is not blocking inbound traffic. Option B is wrong because if the internet gateway were not attached, the instance would have no public IP and would not be reachable at all, but the question implies the instance is reachable via SSH (the security group allows it), so the IGW is likely attached. Option C is wrong because the security group rule is correctly configured to allow SSH from 0.0.0.0/0.

47
Multi-Selectmedium

An organization needs to encrypt data in transit between an Amazon EC2 instance and an Application Load Balancer (ALB). Which THREE actions should be taken?

Select 3 answers
A.Enable encryption at rest on the EC2 instance's EBS volumes.
B.Ensure the EC2 instance has a valid SSL/TLS certificate installed.
C.Configure the security group to allow only encrypted traffic.
D.Configure the ALB listener to use HTTPS protocol.
E.Install an SSL/TLS certificate on the Application Load Balancer.
AnswersB, D, E

If the Application Load Balancer is configured to forward traffic to the target group using HTTPS, the EC2 instance must have a valid SSL/TLS certificate installed to complete the TLS handshake. The instance presents this certificate to the ALB so that the session between them is encrypted, ensuring end-to-end protection from the client to the backend. Without a trusted, valid certificate matching the target's hostname, the ALB cannot establish the encrypted connection.

Why this answer

Encrypting data in transit between an EC2 instance and an Application Load Balancer requires the EC2 instance to present a valid SSL/TLS certificate. This allows the ALB to establish a secure HTTPS connection with the instance over the backend (target group) port, ensuring that traffic between the ALB and the instance is encrypted using TLS.

Exam trap

The trap here is that candidates often confuse encryption at rest (EBS encryption) with encryption in transit, or mistakenly believe security groups can filter based on encryption status, when in reality they only filter at the network layer.

48
MCQhard

A company uses Amazon CloudFront to serve content from an S3 bucket. The bucket is configured as an origin with Origin Access Control (OAC). Users report that they can access the content via CloudFront but also directly via the S3 bucket URL. How can the company restrict direct access to the S3 bucket?

A.Disable OAC and use Origin Access Identity (OAI) instead.
B.Use pre-signed URLs for all S3 requests.
C.Remove the bucket policy and rely on ACLs.
D.Update the S3 bucket policy to deny access to any principal other than the CloudFront service.
AnswerD

This is correct because the S3 bucket policy can include an explicit deny statement that applies to any principal other than the CloudFront service, effectively closing the direct access path. By using a condition such as `aws:SourceArn` or `aws:SourceAccount`, the policy can allow only CloudFront while denying all other IAM users, roles, and anonymous requests. This approach is the recommended companion to OAC because it ensures that the bucket is not publicly accessible, and any attempt to access the object via the S3 website or REST endpoint is rejected before returning content.

Why this answer

With Origin Access Control (OAC), CloudFront signs requests to S3, and the S3 bucket policy must be updated to allow only the CloudFront distribution (via the cloudfront.amazonaws.com service principal with a condition on the distribution ARN) and deny all other principals. This blocks direct access via the S3 URL while preserving CloudFront access.

Exam trap

SOA-C02 often tests whether candidates know that enabling OAC alone is insufficient — the S3 bucket policy must also be updated to deny direct access, and candidates frequently pick 'switch to OAI' thinking it is a security fix rather than a legacy alternative.

How to eliminate wrong answers

Option A is wrong because OAI is the legacy mechanism; switching to OAI does not by itself block direct access unless the bucket policy is also updated, and OAC is the recommended modern approach. Option B is wrong because pre-signed URLs grant temporary access to specific objects and do not restrict direct bucket access — they are a different access pattern entirely. Option C is wrong because removing the bucket policy and relying on ACLs would not restrict direct access and would weaken security; ACLs are also deprecated for most use cases.

49
MCQhard

A company runs a web application on EC2 instances in a private subnet. The application needs to connect to an RDS database in a different VPC. The VPCs are peered. The SysOps Administrator is troubleshooting connectivity issues. The RDS security group allows inbound traffic from the EC2 security group, but connections still fail. What could be the issue?

A.The RDS instance does not have public DNS resolution enabled.
B.The network ACL for the private subnet is blocking inbound traffic.
C.The route tables in each VPC do not have routes to the peered VPC CIDR.
D.The security group outbound rules on the EC2 instance are blocking traffic.
AnswerC

For VPC peering to function, each VPC's route table must contain a route to the CIDR block of the peered VPC, with the peering connection (e.g., pcx-xxxx) as the target. Without these routes, traffic destined for the other VPC is not directed to the peering connection and instead falls back to the local route, making the connection unreachable. This is the most common cause of failed peering connectivity and directly explains why the EC2 instance cannot reach the RDS database.

Why this answer

For traffic to flow between peered VPCs, each VPC's route table must have a route pointing to the CIDR block of the other VPC. Without these routes, packets from the EC2 instance in VPC A destined for the RDS database in VPC B will be dropped, even if security groups and network ACLs are permissive. The SysOps Administrator must add a route in the private subnet's route table for the RDS VPC's CIDR, and a corresponding route in the RDS VPC's route table for the EC2 VPC's CIDR, both pointing to the VPC peering connection.

Exam trap

The trap here is that candidates often assume security groups or network ACLs are the sole cause of connectivity failures in peered VPCs, overlooking the mandatory route table entries required for traffic to traverse the peering connection.

How to eliminate wrong answers

Option A is wrong because public DNS resolution is irrelevant for private connectivity within a VPC peering; RDS instances in a VPC use private DNS names that resolve to private IP addresses, and the EC2 instance can connect using the RDS endpoint without public DNS. Option B is wrong because network ACLs are stateless and must allow both inbound and outbound traffic; however, the question states connections fail, and the most common cause is missing routes, not NACL rules, and NACLs are evaluated before security groups. Option D is wrong because security group outbound rules on the EC2 instance are stateful; if the EC2 security group allows outbound traffic (which it does by default), responses from the RDS database are automatically allowed regardless of outbound rules, so this would not cause a failure.

50
MCQeasy

A CloudFormation template launches an EC2 instance with the user data script shown. The instance launches successfully but the web server does not serve PHP pages. What is the MOST likely reason?

A.The script does not install PHP.
B.The script does not have execute permissions.
C.The CloudFormation template is missing a DependsOn clause for the instance.
D.The user data script is not base64 encoded correctly.
AnswerA

The script does not install PHP. The user data script likely installs and starts the Apache HTTP server, but it omits the installation of the PHP package (or the Apache PHP module). Without the PHP package, Apache cannot interpret and execute PHP files; it will either serve them as plain text or offer them for download, resulting in the PHP page not rendering correctly. The correct fix is to add a command that installs PHP, such as `yum install php` or `apt install php`, and then restart the Apache service so the module loads.

Why this answer

The user data script installs Apache (httpd) and starts the service but never installs PHP or the PHP module for Apache. Without PHP installed, Apache cannot serve .php files — it will either download them as plain text or return a 500 error. This makes option A the most likely reason.

Exam trap

SOA-C02 often tests whether candidates confuse user data execution mechanics (root, auto base64, cloud-init) with actual application configuration gaps — the script ran fine, it just didn't install PHP.

How to eliminate wrong answers

Option B is wrong because user data scripts on EC2 run as root via cloud-init, so execute permissions are not an issue — the script clearly ran since httpd was installed and started. Option C is wrong because DependsOn controls CloudFormation resource creation order, not in-instance software configuration; the instance launched successfully, so ordering is not the problem. Option D is wrong because CloudFormation automatically base64-encodes user data when passed via the AWS::EC2::Instance UserData property — manual encoding errors would prevent the script from running at all, but the script did run.

51
MCQhard

A company runs a critical e-commerce application on Amazon ECS with Fargate launch type, fronted by an Application Load Balancer. The application uses an Amazon ElastiCache for Redis cluster for session state and an Amazon RDS for MySQL Multi-AZ database for persistent data. Recently, during a deployment of a new service version, the application became unresponsive for 15 minutes. The SysOps administrator discovered that the deployment updated the task definition with a new environment variable that pointed to an incorrect ElastiCache endpoint. The ECS service was configured with a rolling update, minimum healthy percent of 50%, and maximum percent of 200%. After the deployment, all tasks failed health checks due to a connection timeout to the wrong Redis endpoint. What is the MOST effective way to prevent this issue in future deployments?

A.Configure a CloudWatch alarm that triggers an automatic rollback if the error rate exceeds 10%.
B.Update the ECS service to use a canary deployment by updating one task at a time.
C.Implement a blue/green deployment strategy using AWS CodeDeploy and test the new task definition before shifting traffic.
D.Enable ECS deployment circuit breaker and set the rollback configuration to automatically roll back failed deployments.
AnswerC

CodeDeploy's blue/green deployment on ECS creates a new 'green' task set running the proposed task definition while the existing 'blue' task set continues to serve full production traffic. This lets you run smoke tests, endpoint checks, or a Lambda-based validation hook against the green task set before shifting any traffic from the blue to the green target group using the ALB's weighted routing. Only after you explicitly confirm the new task definition works can you shift traffic (optionally incrementally), and if issues are detected, you can re-shift back to the original high-fidelity task set with minimal disruption. This pre-validation is exactly what prevents the misconfigured variable from ever affecting end users.

Why this answer

Implementing a blue/green deployment with AWS CodeDeploy allows testing the new task definition in a separate target group before shifting traffic. If the new tasks fail health checks (e.g., due to incorrect ElastiCache endpoint), traffic remains on the blue environment, preventing application downtime. Option A is incorrect because a CloudWatch alarm only triggers an alert or rollback after the issue occurs; it does not prevent the deployment from impacting users.

Option B is incorrect because updating one task at a time (canary) still exposes tasks to the wrong configuration, and since the minimum healthy percent is 50%, at least half the tasks would fail before detection. Option D is incorrect because the ECS deployment circuit breaker rolls back only after the deployment fails, but it does not prevent the initial impact during the rolling update.

52
MCQhard

A company uses AWS Organizations to manage multiple accounts. The security team needs a centralized view of all API calls made across all accounts. Which solution should the SysOps administrator implement?

A.Use AWS Config aggregator to view configuration changes across accounts.
B.Create a CloudTrail trail in the management account that logs events for all accounts in the organization.
C.Use CloudWatch cross-account dashboards to view metrics from all accounts.
D.Enable CloudTrail in each account and have each account send logs to its own S3 bucket.
AnswerB

Creating a CloudTrail trail in the management account with organization-wide logging enabled automatically applies to every account and all regions in your AWS organization. This organization trail centralizes API activity logs into a single S3 bucket in the management account, ensuring you have a complete, consolidated audit record. It also includes new member accounts as they are added, eliminating the need to configure trails individually in each account.

Why this answer

AWS CloudTrail supports an organization trail that, when created in the management account, automatically logs API calls for all member accounts in the AWS Organization. This provides a centralized, single point of access to all API activity across the organization without needing to configure individual trails per account.

Exam trap

The trap here is that candidates may confuse AWS Config (which tracks configuration changes) with CloudTrail (which tracks API calls), or assume that individual account trails are sufficient for a centralized view, overlooking the simplicity and automatic coverage of an organization trail.

How to eliminate wrong answers

Option A is wrong because AWS Config aggregator provides a centralized view of resource configuration changes and compliance status, not API calls (which are logged by CloudTrail). Option C is wrong because CloudWatch cross-account dashboards aggregate metrics (e.g., CPU utilization, latency), not API call logs. Option D is wrong because sending logs to separate S3 buckets in each account does not provide a centralized view; it requires aggregating logs manually or using additional services like S3 replication or Athena, which is less efficient than an organization trail.

53
MCQmedium

A SysOps administrator is troubleshooting an EC2 instance that is unresponsive. The administrator can SSH into the instance but finds that the CloudWatch agent is not sending custom metrics. The CloudWatch agent configuration file is at '/opt/aws/amazon-cloudwatch-agent/etc/amazon-cloudwatch-agent.json'. What should the administrator check first?

A.Verify that the IAM role attached to the EC2 instance has the CloudWatchAgentServerPolicy.
B.Ensure that the IAM user has permissions to access CloudWatch.
C.Check if the security group allows outbound traffic on port 443.
D.Run 'sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl -a status' to check the agent status.
AnswerA

The CloudWatch agent on EC2 uses the instance's IAM role to obtain temporary credentials via the instance metadata service. Without the CloudWatchAgentServerPolicy, which grants PutMetricData, the agent cannot publish custom metrics. This is the first thing to verify because it directly controls the agent's authorization to call CloudWatch.

Why this answer

The correct first check is to verify the IAM role attached to the EC2 instance has the CloudWatchAgentServerPolicy. The CloudWatch agent uses the instance's IAM role to obtain credentials for publishing metrics to CloudWatch. Without this policy, the agent will fail to send custom metrics even if it is running correctly and the instance has network connectivity.

Exam trap

The trap here is that candidates often jump to checking network connectivity (security group rules) or agent status first, overlooking that the IAM role permission is the most common root cause for a CloudWatch agent that is installed and running but not sending metrics.

How to eliminate wrong answers

Option B is wrong because the IAM user's permissions are irrelevant; the EC2 instance uses an IAM role, not a user, to access CloudWatch. Option C is wrong because while outbound HTTPS (port 443) is required for CloudWatch endpoints, the agent typically uses port 443 for TLS connections, but the most common cause of failure is missing IAM permissions, not network connectivity, especially when SSH works. Option D is wrong because checking the agent status is a valid troubleshooting step, but the question asks what to check first; verifying IAM permissions is the more fundamental prerequisite before investigating agent runtime issues.

54
MCQhard

A company runs a critical application on an EC2 instance that stores data on an EBS volume. The SysOps administrator needs to implement a backup strategy that provides the ability to restore the volume to a specific point in time within the last 24 hours, with a recovery time objective (RTO) of less than 15 minutes. Which solution meets these requirements?

A.Configure a RAID 1 mirror of the EBS volume across two Availability Zones.
B.Enable automated backups on the EC2 instance.
C.Use AWS Backup to create backup plans for the EBS volume.
D.Schedule EBS snapshots every hour and keep them for 24 hours.
AnswerD

Scheduling EBS snapshots hourly provides point-in-time recovery points no more than one hour apart, satisfying an RPO of up to one hour, and retaining 24 snapshots covers a full day of rollback options. Restoring is done by creating a new EBS volume from the desired snapshot and attaching it to the instance, which typically takes only minutes and meets the required RTO. EBS snapshots are incremental and stored in Amazon S3, making them a durable, low-cost backup mechanism.

Why this answer

Scheduling EBS snapshots every hour and retaining them for 24 hours provides point-in-time recovery granularity within the last 24 hours. EBS snapshots are incremental, stored in Amazon S3, and can be used to create a new volume or restore an existing one; restoring from a snapshot typically takes only a few minutes, meeting the RTO of less than 15 minutes.

Exam trap

The trap here is that candidates may confuse AWS Backup (Option C) as a service that automatically provides hourly snapshots with 24-hour retention, but AWS Backup requires explicit configuration of a backup plan with the desired schedule and retention—it is not a default behavior, and the question tests whether you know the specific implementation (scheduled snapshots) rather than the service name.

How to eliminate wrong answers

Option A is wrong because RAID 1 mirroring across Availability Zones is not a native EBS feature—it would require software RAID on the EC2 instance, which adds complexity, does not provide point-in-time snapshots, and cannot guarantee a restore to a specific time within 24 hours. Option B is wrong because EC2 instances do not have a native 'automated backups' feature; the term is ambiguous and likely refers to EBS snapshots or AMI backups, but without a defined schedule and retention policy, it cannot ensure point-in-time recovery within 24 hours. Option C is wrong because AWS Backup can create backup plans for EBS volumes, but it does not inherently provide the required granularity of hourly snapshots with 24-hour retention unless explicitly configured; the question specifies a solution that meets the requirements, and AWS Backup is a service that can be used to orchestrate snapshots, but the correct answer is the specific action of scheduling hourly snapshots with 24-hour retention.

55
MCQmedium

A SysOps administrator needs to monitor the CPU utilization of an Amazon EC2 instance and receive an email notification when the metric exceeds 90% for 5 consecutive minutes. The solution should use the least operational overhead. Which combination of AWS services should be used?

A.Create a CloudWatch alarm on the CPUUtilization metric and configure the alarm to send a notification to an Amazon SNS topic with email subscriptions.
B.Create an Amazon EventBridge rule that triggers an AWS Lambda function to check the CPUUtilization metric and send an email via Amazon SES.
C.Configure the EC2 instance to publish CPU logs to Amazon CloudWatch Logs, then create a metric filter to detect high CPU and trigger an SNS notification.
D.Use AWS CloudTrail to monitor EC2 CPU metrics and send notifications to an Amazon SQS queue.
AnswerA

A CloudWatch alarm evaluates the CPUUtilization metric against a 90% threshold over five consecutive one-minute periods, then publishes to an SNS topic whose email subscription delivers the notification. This is fully managed, requiring no agents or custom code, so operational overhead is lowest.

Why this answer

A CloudWatch alarm directly monitors the CPUUtilization metric for an EC2 instance and can be configured to evaluate whether the metric exceeds 90% for 5 consecutive minutes (e.g., 5 evaluation periods of 1 minute each). The alarm then publishes to an Amazon SNS topic, which sends email notifications to subscribed endpoints, requiring no additional infrastructure or code, thus minimizing operational overhead.

Exam trap

The trap here is that candidates may overcomplicate the solution by introducing Lambda or log-based filters, when the simplest and most direct path—a CloudWatch alarm on the existing CPUUtilization metric with an SNS action—is the correct answer for minimal operational overhead.

How to eliminate wrong answers

Option B is wrong because it introduces unnecessary complexity by using an EventBridge rule and a Lambda function to poll or process metrics, which increases operational overhead and latency compared to a native CloudWatch alarm. Option C is wrong because publishing CPU logs to CloudWatch Logs and creating a metric filter is designed for log-based metrics (e.g., parsing log entries), not for the native CPUUtilization metric, which is already available as a CloudWatch metric without logs. Option D is wrong because AWS CloudTrail records API calls and management events, not EC2 CPU utilization metrics, and cannot monitor or trigger notifications based on performance metrics.

56
MCQmedium

A company has a web application running on EC2 instances behind an Application Load Balancer. The application experiences latency spikes during peak hours. Amazon CloudWatch metrics show that CPU and memory are not fully utilized. The SysOps administrator suspects the bottleneck is the database. The database is an RDS for MySQL instance. Which action should the administrator take to improve performance without over-provisioning?

A.Add a Read Replica to offload read queries.
B.Convert the RDS instance to a Multi-AZ deployment.
C.Enable RDS Performance Insights to analyze database load and identify slow queries.
D.Increase the RDS instance size to the next tier.
AnswerC

Enable RDS Performance Insights, which gives you a real-time and historical view of database load dimensioned by waits, SQL statements, and hosts. This is the correct first step because it lets you pinpoint whether the bottleneck is CPU, storage, lock waits, or an inefficient query pattern, rather than guessing. Only after identifying the specific load source should you consider changes like adding indexes, rewriting queries, or scaling the instance.

Why this answer

The administrator needs to diagnose the database bottleneck before making changes. RDS Performance Insights provides detailed visibility into database load, wait events, and top SQL statements, allowing identification of slow queries or resource contention. This aligns with the goal of improving performance without over-provisioning, as it enables targeted optimization rather than guessing at scaling.

The other options either add unnecessary resources or don't address the root cause.

Exam trap

SOA-C02 often tests the difference between diagnosing and resolving performance issues; candidates may jump to scaling or replication without first analyzing the database workload, leading to unnecessary costs or ineffective solutions.

How to eliminate wrong answers

Option A is wrong because adding a Read Replica only helps if the bottleneck is read-heavy and the application can be configured to use it; it doesn't diagnose the issue and may not address the actual cause. Option B is wrong because Multi-AZ is for high availability and failover, not performance improvement; it doesn't reduce load on the primary. Option D is wrong because increasing instance size is over-provisioning and may not fix the bottleneck if it's due to inefficient queries or locking.

57
MCQhard

An organization has a requirement to automatically scale its web application based on a custom metric that measures the number of active user sessions stored in Amazon ElastiCache. The metric is published to CloudWatch every minute. The Auto Scaling group currently uses a simple scaling policy based on CPU utilization. What is the most effective way to implement scaling based on this custom metric?

A.Create a target tracking scaling policy that uses the custom metric as a target.
B.Create a step scaling policy that adjusts capacity based on the magnitude of the metric breach.
C.Create a scheduled scaling policy that increases capacity during peak hours.
D.Create a simple scaling policy that adds instances when the custom metric exceeds a threshold and removes when below.
AnswerA

With a target tracking scaling policy, you first publish a custom metric to CloudWatch (for example, ActiveUserSessions) and then define the policy with a target value such as 1000 sessions per instance. Amazon EC2 Auto Scaling continuously computes the required capacity to keep the metric near that target, proactively adding or removing instances without static thresholds or manually tuned cooldowns. This makes it ideal for a dynamic, session-based workload where the relationship between load and capacity is stable and predictable.

Why this answer

Target tracking scaling policies are purpose-built for metric-driven scaling: you specify a target value for a custom CloudWatch metric (e.g., active sessions per instance) and Auto Scaling automatically creates and manages the required CloudWatch alarms and scaling adjustments. This is the most effective approach because it continuously adjusts capacity to keep the metric at target, requires no manual alarm or step definition, and works with any metric published to CloudWatch, including custom ElastiCache session metrics.

Exam trap

SOA-C02 often tests the misconception that step scaling is 'more granular' and therefore better for custom metrics — the trap is overlooking that target tracking eliminates manual alarm/step management and is the recommended default for metric-driven scaling.

How to eliminate wrong answers

Option B is wrong because step scaling requires you to manually define CloudWatch alarms and step adjustments, adding operational overhead and not automatically tracking a target — it only reacts to breaches you pre-specify. Option C is wrong because scheduled scaling is time-based, not metric-based, so it cannot respond to real-time session counts. Option D is wrong because simple scaling is the legacy policy type that uses a single adjustment and cooldown, lacks the responsiveness of target tracking, and still requires manual alarm configuration.

58
MCQhard

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The SysOps administrator notices that the application's response time is increasing during peak hours. The administrator wants to set up a CloudWatch dashboard that displays the average latency of requests across all instances and the number of healthy hosts. Which metrics should be used?

A.Use the ALB's 'TargetResponseTime' metric and the ALB's 'UnhealthyHostCount' metric.
B.Use the ALB's 'TargetResponseTime' metric and the ALB's 'HealthyHostCount' metric.
C.Use the ALB's 'RequestCount' metric and the EC2 Auto Scaling group's 'GroupInServiceInstances' metric.
D.Use the ALB's 'Latency' metric and the EC2 instance's 'CPUUtilization' metric.
AnswerB

TargetResponseTime is the ALB metric that tracks the time elapsed from when the load balancer receives a request to when it receives a response from a target. HealthyHostCount reports the number of healthy targets registered with the ALB. Together, a sustained increase in TargetResponseTime combined with a decreasing HealthyHostCount indicates that the application is degrading because fewer instances are available to handle traffic, making this the correct pair for a performance-based scaling alarm.

Why this answer

The ALB's 'TargetResponseTime' metric measures the average time (in seconds) that requests are routed to targets, which directly reflects application latency. The ALB's 'HealthyHostCount' metric shows the number of healthy registered targets, which is the exact metric needed to monitor host health. Together, these two metrics provide the required visibility into average latency and healthy host count across all instances.

Exam trap

The trap here is that candidates confuse 'UnhealthyHostCount' with 'HealthyHostCount' or mistakenly use instance-level metrics (like CPUUtilization) instead of ALB-level metrics, failing to recognize that the ALB's own metrics are the authoritative source for request latency and target health.

How to eliminate wrong answers

Option A is wrong because 'UnhealthyHostCount' tracks unhealthy hosts, not healthy hosts; the question specifically asks for the number of healthy hosts. Option C is wrong because 'RequestCount' measures total requests, not latency, and 'GroupInServiceInstances' is an Auto Scaling group metric, not an ALB metric; the ALB's 'HealthyHostCount' is the correct source for healthy host count. Option D is wrong because 'Latency' is not a valid ALB metric (the correct metric is 'TargetResponseTime'), and 'CPUUtilization' measures instance CPU usage, not host health or latency.

59
MCQmedium

A company requires all S3 uploads to use server-side encryption with a specific customer managed KMS key. What is the most direct enforcement mechanism?

A.Add a bucket policy that denies PutObject unless the required SSE-KMS headers and key ID are present.
B.Enable S3 versioning only.
C.Enable S3 Transfer Acceleration.
D.Create an IAM user for every uploader with console access.
AnswerA

A bucket policy with an explicit Deny on s3:PutObject can inspect the s3:x-amz-server-side-encryption and s3:x-amz-server-side-encryption-aws-kms-key-id condition keys. If the request lacks the required SSE-KMS header or uses a different KMS key ID, S3 returns 403 AccessDenied before accepting the upload. This enforces encryption at write time, making it the only option that guarantees every upload is encrypted with your designated KMS key.

Why this answer

A bucket policy with a condition that denies `s3:PutObject` unless the required `x-amz-server-side-encryption` header is set to `aws:kms` and the `x-amz-server-side-encryption-aws-kms-key-id` header matches the specific customer managed KMS key ARN is the most direct enforcement mechanism. This policy-based approach ensures that any upload attempt lacking the required SSE-KMS headers and key ID is rejected at the S3 API level, regardless of the IAM permissions of the uploader.

Exam trap

The trap here is that candidates often confuse IAM permissions with bucket policy conditions, assuming that IAM policies alone can enforce encryption headers, when in fact only a bucket policy with the appropriate condition keys can directly deny uploads that lack the required encryption headers.

How to eliminate wrong answers

Option B is wrong because enabling S3 versioning only preserves object versions but does not enforce any encryption requirements on uploads. Option C is wrong because S3 Transfer Acceleration speeds up uploads over long distances but has no effect on encryption enforcement. Option D is wrong because creating an IAM user for every uploader with console access does not enforce server-side encryption; it only provides authentication and does not mandate the use of a specific KMS key or encryption headers.

60
MCQmedium

A company runs a critical production database on Amazon RDS for MySQL with Multi-AZ deployment. The SysOps administrator needs to be automatically notified when a failover event occurs, and also capture the exact time and reason for the failover for compliance purposes. Which AWS service or feature should be used to capture the failover event details with the least operational overhead?

A.Create an Amazon CloudWatch Events rule that matches the 'RDS DB Instance Event' for 'failover' and sends the event to an Amazon SNS topic for notification and logging.
B.Enable detailed monitoring on the RDS instance and stream the logs to Amazon CloudWatch Logs where a metric filter can detect failover patterns.
C.Configure AWS CloudTrail to log all RDS API calls and analyze the logs for the 'Failover' event type.
D.Use AWS Config to create a config rule that evaluates whether the 'DBInstanceStatus' changes to 'failover' and then trigger a remediation action.
AnswerA

Amazon CloudWatch Events (now part of Amazon EventBridge) natively integrates with RDS event notifications, emitting a structured event whenever a DB instance experiences a failover. By creating a rule that matches the 'RDS DB Instance Event' source and the specific detail type for failover, you can route that event to an SNS topic in near-real time, enabling automated alerting, logging, and downstream remediation. This is the intended, low-overhead approach because RDS already publishes these lifecycle events, and no polling or custom detection logic is required.

Why this answer

Amazon CloudWatch Events (now part of Amazon EventBridge) can match RDS DB Instance events, including 'failover', and route them to an SNS topic for notification and to CloudWatch Logs for logging. This approach requires no custom scripting or polling, providing the least operational overhead while capturing the exact time and reason for the failover directly from the RDS event stream.

Exam trap

The trap here is that candidates confuse CloudTrail (which logs API calls) with RDS events (which log internal service events), leading them to choose CloudTrail even though automatic failovers are not API-driven and thus not recorded by CloudTrail.

How to eliminate wrong answers

Option B is wrong because detailed monitoring on RDS provides enhanced metrics (e.g., CPU, memory) but does not generate failover events or detect failover patterns; metric filters on CloudWatch Logs would require RDS to log failover details to CloudWatch Logs, which RDS does not do by default. Option C is wrong because AWS CloudTrail logs API calls (e.g., FailoverDBInstance), not internal failover events triggered by AWS; a Multi-AZ failover is an automatic process, not an API call, so CloudTrail will not capture it. Option D is wrong because AWS Config evaluates resource configuration changes (e.g., DBInstanceStatus) but does not natively detect a 'failover' status change; the DBInstanceStatus transitions through multiple states (e.g., 'creating', 'available', 'resetting-master-credentials') and 'failover' is not a valid status—Config rules would require custom logic and still not capture the exact reason for the failover.

61
MCQeasy

A SysOps administrator needs to monitor the CPU utilization of an EC2 instance and receive an alert when it exceeds 80% for 10 consecutive minutes. Which AWS service should be used to configure this monitoring and alerting?

A.Amazon EventBridge
B.Amazon CloudWatch Alarms
C.AWS Trusted Advisor
D.AWS Config
AnswerB

Amazon CloudWatch Alarms monitor a CloudWatch metric, such as EC2 CPUUtilization, against a defined threshold over a specified number of consecutive evaluation periods. When the alarm transitions to ALARM, OK, or INSUFFICIENT_DATA, it can trigger actions like Auto Scaling policies or SNS notifications. This allows you to detect sustained high CPU usage and respond automatically, making it the correct choice for monitoring CPU utilization.

Why this answer

Amazon CloudWatch Alarms is the correct service because it allows you to monitor a specific metric, such as EC2 CPUUtilization, and trigger an action (e.g., an SNS notification) when the metric crosses a defined threshold (80%) for a specified number of consecutive evaluation periods (10 minutes, which with the default 1-minute period equals 10 datapoints). This directly fulfills the requirement for threshold-based alerting on a single metric over a sustained duration.

Exam trap

The trap here is that candidates confuse Amazon EventBridge (which can trigger actions based on events but cannot natively evaluate sustained metric thresholds) with CloudWatch Alarms, or mistakenly think AWS Config or Trusted Advisor can monitor real-time performance metrics, when they are designed for configuration compliance and best-practice recommendations respectively.

How to eliminate wrong answers

Option A is wrong because Amazon EventBridge is a serverless event bus used to route events from sources (e.g., AWS services, custom apps) to targets (e.g., Lambda, Step Functions), but it does not natively evaluate metric thresholds over time or generate alarms based on sustained CPU utilization. Option C is wrong because AWS Trusted Advisor provides best-practice checks and recommendations (e.g., underutilized instances, security gaps) but does not perform real-time metric monitoring or alerting on CPU utilization thresholds. Option D is wrong because AWS Config is a service for recording and evaluating resource configuration changes against rules (e.g., ensuring EBS volumes are encrypted), not for monitoring performance metrics like CPU utilization or generating threshold-based alerts.

62
MCQeasy

A company runs a batch processing job on Amazon EC2 instances that runs for 3 hours each night. The job can be interrupted and can resume from the last checkpoint without data loss. The SysOps administrator wants to minimize compute costs for this workload. Which Amazon EC2 purchasing option should be used?

A.On-Demand Instances
B.Reserved Instances
C.Spot Instances
D.Dedicated Hosts
AnswerC

Spot Instances give the company the largest possible EC2 discount, up to 90% off On-Demand, and are ideal for fault-tolerant, interruptible workloads like the described batch job. Because the job uses checkpointing, it can safely stop after a two-minute Spot interruption notice and later resume from the last saved state, so cost savings are maximized without sacrificing data integrity or job completion.

Why this answer

Spot Instances are correct because the batch job is fault-tolerant (can be interrupted and resume from checkpoints) and runs for a fixed 3-hour window nightly. Spot Instances offer significant cost savings (up to 90% off On-Demand) by using spare EC2 capacity, which can be reclaimed by AWS with a 2-minute interruption notice. Since the workload can handle interruptions gracefully, Spot Instances minimize compute costs while meeting the job's requirements.

Exam trap

The trap here is that candidates may choose On-Demand or Reserved Instances because they assume a nightly 3-hour job requires guaranteed availability, overlooking that Spot Instances are ideal for fault-tolerant, interruptible workloads and offer the lowest cost.

How to eliminate wrong answers

Option A is wrong because On-Demand Instances have no discount and would be the most expensive option for a predictable nightly workload. Option B is wrong because Reserved Instances require a 1- or 3-year commitment and are designed for steady-state workloads, not for a short 3-hour nightly job that can be interrupted; they would lock in costs without leveraging the fault tolerance of the job. Option D is wrong because Dedicated Hosts provide physical server isolation for licensing or compliance needs, which is unnecessary here and incurs additional costs without any benefit for a batch processing job.

63
MCQmedium

An application running on EC2 instances sends custom metrics to CloudWatch using the PutMetricData API. The SysOps admin notices that some metrics are missing from the CloudWatch console. What is the most likely cause?

A.The metric data does not include a unit
B.The metric data does not include a dimension
C.The metric data is being sent with a timestamp older than 14 days
D.The namespace in the PutMetricData call does not match the namespace in the CloudWatch console
AnswerD

This is the correct answer. The namespace is a container for metrics; if the application sends custom metrics using a different namespace than the one being viewed in the CloudWatch console (e.g., 'Custom/App' vs 'AWS/EC2'), the metrics will be stored under the namespace specified in the PutMetricData call. The console only displays metrics from the selected namespace, so a mismatched namespace makes the metrics appear 'missing.' The fix is to use a consistent namespace across the publishing and viewing steps.

Why this answer

Custom metrics in CloudWatch are uniquely identified by the combination of namespace, metric name, and dimensions. If the namespace used in the PutMetricData API call does not match the namespace being viewed in the CloudWatch console, the metrics will not appear under that namespace. CloudWatch does not automatically merge or alias namespaces, so mismatched namespaces cause the data to be stored under a different namespace, making it invisible in the console view.

Exam trap

The trap here is that candidates often assume missing metrics are due to timestamp or dimension issues, but the most common real-world cause is a namespace mismatch between the PutMetricData call and the console filter, which CloudWatch does not automatically reconcile.

How to eliminate wrong answers

Option A is wrong because the unit field in PutMetricData is optional; CloudWatch accepts metric data without a unit and displays it without a unit label. Option B is wrong because dimensions are optional for custom metrics; while dimensions help organize metrics, a metric without dimensions is still valid and will appear under the specified namespace. Option C is wrong because CloudWatch accepts metric data with timestamps up to 15 days in the past (not 14), and the question states some metrics are missing, not that all data older than 14 days is missing.

64
MCQhard

Refer to the exhibit. A security group is attached to an Application Load Balancer (ALB) that serves HTTPS traffic on port 443. Users can access the application via HTTPS. However, the ALB's health checks to targets on port 80 are failing. What is the reason?

A.The ALB's security group does not allow HTTPS traffic from the internet.
B.The security group for the target instances does not allow HTTP traffic from the ALB's security group.
C.The ALB's security group does not allow HTTP traffic from the target's IP range.
D.The health check is configured to use HTTPS, but the target only supports HTTP.
AnswerB

This is correct because the ALB sends health check requests from its own network interfaces, using the ALB's security group as the source in the allowed inbound rule on each target. The target instance's security group must explicitly allow inbound TCP on the health check port (HTTP/80) from the ALB's security group ID (or from the VPC CIDR if the security group reference is not used). Without that rule, the OS receives the SYN packet but the security group silently drops it, so the health check times out and the target is marked unhealthy. This is the standard root cause for healthy-app-turned-unhealthy after an ALB change or when targets are in a different security group.

Why this answer

ALB health checks originate from the ALB's nodes and are sent to the target's health-check port (here port 80/HTTP). For the check to succeed, the target instance's security group must allow inbound HTTP from the ALB's security group. Since users can reach the app over HTTPS on 443, the ALB listener and its security group are fine — the failure is on the target-side SG not permitting the health-check traffic.

Exam trap

SOA-C02 often tests the misconception that the ALB's own security group controls health-check success — in reality, the target's security group must allow the health-check port from the ALB.

How to eliminate wrong answers

Option A is wrong because users are already accessing the app via HTTPS, proving the ALB's security group allows inbound 443 from the internet. Option C is wrong because the ALB's security group governs inbound traffic to the ALB, not outbound health-check traffic to targets — and referencing the target's IP range is the wrong direction. Option D is wrong because the scenario states health checks fail on port 80, implying the check is HTTP; the target supports HTTP, so protocol mismatch is not the issue.

65
MCQeasy

A company wants to allow an external auditor to assume an IAM role in their AWS account to review resources. What is the minimum information the auditor needs from the company to do this?

A.The Amazon Resource Name (ARN) of the IAM role to assume.
B.The IAM user name and password of the company's admin user.
C.The IAM policy document that grants the auditor access.
D.The AWS account ID and the region where resources are hosted.
AnswerA

The AssumeRole API in AWS Security Token Service (STS) requires a RoleArn parameter, and the role ARN is the globally unique identifier that the external auditor's account must pass to request temporary credentials. The ARN encodes both the AWS account ID that owns the role and the role name, which allows STS to locate the role and apply its trust policy. Without this ARN, the auditor cannot invoke the role-assumption flow at all.

Why this answer

To assume an IAM role, the auditor must know the role's ARN, which uniquely identifies the role (including the account ID and role name). The auditor then calls sts:AssumeRole with that ARN. The role's trust policy must also allow the auditor's AWS account or principal to assume it, but the ARN is the minimum information the auditor needs to initiate the call.

Exam trap

SOA-C02 often tests the misconception that the auditor needs the IAM policy document or account ID to assume a role, but the key requirement is the role's ARN, which uniquely identifies the role and is used in the AssumeRole API call.

How to eliminate wrong answers

Option B is wrong because sharing an admin user's credentials violates least privilege and is unnecessary; the auditor should use their own credentials to assume a role, not the company's admin user. Option C is wrong because the policy document alone does not identify the role; the auditor needs the role's ARN to assume it, and the policy is attached to the role, not used directly by the auditor. Option D is wrong because the account ID and region are insufficient; the auditor needs the specific role ARN, and the region is not required for global IAM role assumption (STS endpoints are global, though regional endpoints exist).

66
MCQeasy

A company uses Amazon S3 to store critical data. The SysOps administrator needs to protect against accidental deletion of objects. Which combination of actions should the administrator take? (Choose the best answer.)

A.Apply a bucket policy that denies s3:DeleteObject for all principals.
B.Set a lifecycle policy to expire objects after 30 days.
C.Enable S3 Versioning and MFA Delete on the bucket.
D.Configure cross-Region replication to a different bucket.
AnswerC

Enabling S3 Versioning preserves every version of an object, so a deletion request creates a delete marker instead of erasing the underlying data, allowing you to restore the object at any time. Adding MFA Delete requires a valid MFA code to permanently delete object versions or to suspend versioning, which prevents accidental or malicious purges even by the root user. Together, these features provide comprehensive, built-in protection against both overwrites and deletes, making them the correct choice for safeguarding critical data.

Why this answer

Enabling S3 Versioning preserves all versions of an object, allowing recovery from accidental deletion or overwrite. MFA Delete adds an additional authentication layer, requiring multi-factor authentication to permanently delete object versions or suspend versioning, thus preventing unauthorized or accidental deletions.

Exam trap

The trap here is that candidates often choose a bucket policy denying s3:DeleteObject (Option A) thinking it prevents accidental deletion, but they overlook that it does not protect against overwrites or that versioning with MFA Delete is the only comprehensive solution that also covers permanent deletion and versioning state changes.

How to eliminate wrong answers

Option A is wrong because a bucket policy that denies s3:DeleteObject for all principals would also block legitimate administrative deletions, and it does not protect against accidental overwrites (PUT operations) or deletion of versioned objects without versioning enabled. Option B is wrong because a lifecycle policy to expire objects after 30 days would actually delete objects automatically, increasing the risk of data loss rather than protecting against accidental deletion. Option D is wrong because cross-Region replication replicates objects to another bucket but does not prevent deletion in the source bucket; deletions are replicated as well unless a delete marker replication rule is configured, and it does not protect against accidental deletion in the source.

67
MCQeasy

A company wants to receive alerts when their monthly AWS spending exceeds $1,000. Which AWS service should be used?

A.AWS Trusted Advisor
B.AWS CloudWatch
C.AWS Budgets
D.AWS Cost Explorer
AnswerC

AWS Budgets is the purpose-built service for tracking spending and usage against predefined cost thresholds, allowing you to create budgets for monthly, quarterly, or annual periods. You can set multiple alerts—both for actual spend and forecasted spend—with custom thresholds and notification preferences (email, SNS) or even Lambda triggers. This directly satisfies the requirement to receive alerts when the monthly AWS spend exceeds a specific amount, making AWS Budgets the correct answer.

Why this answer

AWS Budgets allows setting cost thresholds and alerts. CloudWatch alarms are for metrics, not budgets. Cost Explorer is for visualization.

Trusted Advisor provides recommendations.

68
MCQeasy

A company wants to monitor the number of messages in an Amazon SQS queue and scale the number of EC2 instance consumers based on queue depth. Which combination of AWS services should be used?

A.Amazon CloudWatch and Amazon EC2 Auto Scaling
B.Amazon Elastic Load Balancing and Amazon EC2 Auto Scaling
C.Amazon CloudWatch and AWS Lambda
D.AWS CloudTrail and Amazon EventBridge
AnswerA

Amazon CloudWatch is the correct monitoring service because it publishes the SQS metric ApproximateNumberOfMessagesVisible from the queue, and you can create a CloudWatch alarm on that metric. When the alarm threshold is breached, it triggers an EC2 Auto Scaling policy that adds or removes instances based on the queue depth, directly matching the requirement to scale compute resources with message volume. CloudWatch also retains metric history for troubleshooting and can trigger multiple actions, but the core pairing here is metric-driven auto scaling.

Why this answer

Amazon CloudWatch monitors the SQS queue depth (ApproximateNumberOfMessagesVisible metric) and triggers an Amazon EC2 Auto Scaling scaling policy based on a CloudWatch alarm. This allows the number of EC2 consumer instances to dynamically scale in or out in response to the queue depth, ensuring efficient processing without over-provisioning.

Exam trap

The trap here is that candidates often confuse Elastic Load Balancing with queue-based scaling, assuming ELB can scale EC2 instances based on SQS depth, but ELB only handles HTTP/HTTPS traffic distribution and cannot read SQS metrics.

How to eliminate wrong answers

Option B is wrong because Elastic Load Balancing distributes incoming traffic to EC2 instances but does not monitor SQS queue depth or trigger scaling actions; it is not designed for queue-based scaling. Option C is wrong because while AWS Lambda can process SQS messages, it is a serverless compute service and does not manage EC2 instance scaling; using Lambda alone would not scale EC2 instances. Option D is wrong because AWS CloudTrail records API activity for auditing, and Amazon EventBridge routes events between services, but neither directly monitors SQS queue depth nor triggers EC2 Auto Scaling adjustments.

69
MCQmedium

A company runs a file-sharing application on AWS. Users upload files to an S3 bucket, which triggers a Lambda function to process the files and store metadata in a DynamoDB table. Recently, users have reported that some uploaded files are never processed. The SysOps Administrator checks the CloudWatch logs and finds no errors from the Lambda function. The S3 bucket is configured to send events to the Lambda function. The DynamoDB table has sufficient write capacity. The administrator suspects that the event notifications are being lost. Which action should the SysOps Administrator take to ensure that every file upload triggers a Lambda function and that the function processes the file successfully?

A.Configure an SQS queue as the event destination for the S3 bucket, and have the Lambda function process messages from the queue.
B.Use DynamoDB Streams to capture file metadata changes instead of Lambda invocation.
C.Increase the Lambda function's reserved concurrency to handle more invocations.
D.Increase the write capacity of the DynamoDB table to avoid throttling.
AnswerA

S3 event notifications can be delivered to SQS, providing a durable buffer. Lambda polls the queue, so messages are not lost if the function is throttled or busy. SQS retains messages until processed, and Lambda event source mapping handles retries and batching. This decouples the upload rate from Lambda's invocation capacity.

Why this answer

Using an SQS queue as the event destination for S3 bucket events provides a durable, reliable mechanism to capture every event. S3 sends event notifications to the SQS queue, and if the Lambda function fails or is throttled, the message remains in the queue for later processing. This decouples the event source from the function and ensures no events are lost.

Option B is incorrect because DynamoDB Streams capture changes to DynamoDB items, not S3 events. Option C is wrong because increasing reserved concurrency only helps with Lambda scaling but does not address potential event loss due to failures or delivery issues. Option D is wrong because DynamoDB write capacity is already sufficient; the issue is with S3 event delivery, not database writes.

70
MCQmedium

A company is deploying a new web application using AWS Elastic Beanstalk. The application requires a custom Amazon Machine Image (AMI) with specific software pre-installed. The SysOps administrator creates a custom AMI and configures Elastic Beanstalk to use it. However, during deployment, the instances fail to pass the health check. The health check endpoint is a simple 'index.html' file. What is the MOST likely cause?

A.The Elastic Beanstalk environment was created before the custom AMI was registered.
B.The custom AMI does not have a web server installed and configured to serve the application.
C.The custom AMI is not registered with the same account that owns the Elastic Beanstalk environment.
D.The custom AMI does not have the latest patches, causing the instance to fail the EC2 status checks.
AnswerB

The health check performed by Elastic Beanstalk is an HTTP request to the environment's health check path (typically / on port 80). If the custom AMI lacks a web server or the web server isn't configured to serve the application, the ELB health check receives a connection refused or non-2xx response, causing the instance to be marked unhealthy. Simply having a running EC2 instance is insufficient; the AMI must include the same web server and configuration as the standard Elastic Beanstalk platform AMI to serve traffic.

Why this answer

Elastic Beanstalk relies on the platform's web server (Apache, Nginx, IIS) to serve the application and respond to the health check endpoint. When you supply a custom AMI, you must ensure it includes the same web server and configuration that the chosen Beanstalk platform expects. If the AMI lacks a running web server on the expected port (e.g., port 80 for the default health check path '/'), the ELB health check will fail and instances will be marked unhealthy.

Exam trap

SOA-C02 often tests the misconception that health check failures are caused by patching, account ownership, or resource ordering, when in fact they almost always stem from the application or web server not responding on the expected port and path.

How to eliminate wrong answers

Option A is wrong because the timing of environment creation relative to AMI registration does not affect whether the instance can serve HTTP traffic; Beanstalk validates the AMI at launch time regardless of when the environment was created. Option C is wrong because AMIs are region-scoped but can be used across accounts within the same region if permissions allow; cross-account AMI usage is supported and would produce a launch error, not a health check failure. Option D is wrong because missing patches do not cause EC2 status checks or ELB health checks to fail — status checks verify hypervisor and network reachability, and the health check endpoint tests application response, neither of which depends on patch level.

71
MCQeasy

A SysOps administrator needs to deploy a new version of a web application to Amazon EC2 instances using AWS Elastic Beanstalk. The administrator wants to deploy the new version with zero downtime and validate the new version before routing production traffic to it. Which deployment policy should be used?

A.All at once
B.Rolling
C.Immutable
D.Traffic splitting
AnswerC

The immutable deployment policy launches a completely new set of instances with the new application version. Once healthy, the environment's CNAME is switched to the new instances, providing zero downtime and the ability to validate the new version before traffic is routed.

Why this answer

Immutable deployment is correct because it launches a completely new set of EC2 instances in a separate Auto Scaling group, deploys the new application version to them, and passes health checks before swapping the environment's CNAME record to point to the new instances. This ensures zero downtime and allows validation of the new version before any production traffic is routed to it, as the old instances remain untouched until the swap is complete.

Exam trap

The trap here is that candidates confuse 'Traffic splitting' with 'canary testing' and assume it allows pre-validation, but in Elastic Beanstalk, traffic splitting immediately routes a percentage of live traffic to the new version, whereas immutable deployment keeps all traffic on the old version until the new version is fully validated and swapped.

How to eliminate wrong answers

Option A is wrong because All at once deploys the new version to all instances simultaneously, causing downtime during the deployment and no ability to validate before traffic is routed. Option B is wrong because Rolling deploys the new version in batches across existing instances, which can cause a brief period of reduced capacity and does not allow full validation of the new version before all traffic is switched; it also does not guarantee zero downtime if health checks fail mid-batch. Option D is wrong because Traffic splitting (canary deployment) routes a percentage of traffic to the new version immediately, which does not allow validation before any production traffic is sent; it is designed for gradual traffic shifting, not pre-validation with zero initial traffic.

72
Multi-Selecteasy

Which TWO AWS services can be used to improve the security of a VPC? (Choose TWO.)

Select 2 answers
A.Security Groups
B.Internet Gateway
C.Route Tables
D.Network ACLs
E.VPC Peering
AnswersA, D

Security Groups are stateful virtual firewalls that operate at the ENI/instance level within a VPC. They evaluate all inbound and outbound traffic against a set of allow rules only—there is no explicit deny rule—and return traffic is automatically permitted regardless of the outbound rule configuration. For example, if you allow inbound HTTP from 0.0.0.0/0, the corresponding outbound response traffic is implicitly allowed, which simplifies security but requires careful rule design to avoid overly permissive configurations.

Why this answer

Security Groups (A) act as a virtual firewall for instances, controlling inbound and outbound traffic at the instance level based on allow rules only. Network ACLs (D) provide a stateless firewall layer at the subnet level, supporting both allow and deny rules, and are evaluated in numeric order. Together, they offer defense-in-depth for VPC traffic filtering.

Exam trap

The trap here is that candidates confuse routing components (Internet Gateway, Route Tables, VPC Peering) with security components, assuming any VPC construct that controls traffic flow also provides security filtering.

73
MCQmedium

An organization requires that all Amazon S3 buckets be encrypted at rest by default. A SysOps administrator needs to enforce this using AWS Config. Which AWS Config managed rule should be used?

A.s3-bucket-encryption-enabled
B.s3-bucket-ssl-requests-only
C.s3-bucket-public-read-prohibited
D.s3-bucket-logging-enabled
AnswerA

The AWS Config managed rule s3-bucket-encryption-enabled evaluates whether an S3 bucket has default encryption enabled, which is satisfied by configuring either SSE-S3 or SSE-KMS. This ensures new objects written to the bucket are automatically encrypted at rest, directly meeting the organization's encryption requirement. Without this rule, a bucket could store plaintext objects, making it the correct choice.

Why this answer

The AWS Config managed rule `s3-bucket-encryption-enabled` checks whether S3 buckets have default encryption enabled (SSE-S3, SSE-KMS, or SSE-C). This directly enforces the requirement that all buckets are encrypted at rest by default, as it evaluates each bucket's encryption configuration and flags non-compliant resources.

Exam trap

The trap here is that candidates often confuse encryption in transit (SSL/TLS) with encryption at rest, leading them to select `s3-bucket-ssl-requests-only` instead of the correct rule for default encryption.

How to eliminate wrong answers

Option B is wrong because `s3-bucket-ssl-requests-only` enforces that bucket policies deny HTTP requests, not encryption at rest. Option C is wrong because `s3-bucket-public-read-prohibited` checks for public read access, not encryption. Option D is wrong because `s3-bucket-logging-enabled` verifies that server access logging is enabled, which is unrelated to encryption at rest.

74
MCQmedium

A company has a production Amazon RDS for MySQL DB instance in a single Availability Zone. The SysOps administrator needs to improve database availability to ensure automatic failover in the event of a database failure or an Availability Zone outage. Which configuration should the administrator enable?

A.Enable Multi-AZ deployment
B.Create a read replica in another Availability Zone
C.Enable automated backups
D.Change the DB instance to a larger instance class
AnswerA

Enabling Multi-AZ on an RDS for MySQL instance provisions a standby replica in a separate Availability Zone and uses synchronous replication to keep it current. If the primary fails or its AZ becomes unavailable, Amazon RDS automatically flips the DNS endpoint to the standby, delivering failover in typically 60–120 seconds with no manual intervention and minimal data loss. This is the only option that provides automatic failover and meets the high availability requirement for a production database.

Why this answer

Enabling a Multi-AZ deployment for Amazon RDS for MySQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of a database failure or an AZ outage, Amazon RDS automatically fails over to the standby replica, typically within 60–120 seconds, without requiring manual intervention. This configuration meets the requirement for automatic failover and improved availability.

Exam trap

The trap here is that candidates often confuse a read replica with a Multi-AZ standby, mistakenly believing that a read replica can provide automatic failover, but read replicas require manual promotion and do not maintain synchronous replication.

How to eliminate wrong answers

Option B is wrong because a read replica is an asynchronous copy used for offloading read traffic, not for automatic failover; while it can be promoted to a standalone instance, this requires manual action and does not provide automatic failover. Option C is wrong because automated backups only enable point-in-time recovery and do not provide any failover capability or high availability. Option D is wrong because changing the DB instance to a larger instance class improves performance and scalability but does not provide redundancy or automatic failover across Availability Zones.

75
MCQeasy

A SysOps administrator wants to automate the creation of an AWS Lambda function and its associated IAM role using infrastructure as code. Which AWS service should be used?

A.AWS CloudFormation
B.AWS Elastic Beanstalk
C.AWS CodeDeploy
D.AWS Systems Manager
AnswerA

AWS CloudFormation is the native infrastructure-as-code service that lets you define the Lambda function, IAM role, and every related resource in a declarative JSON or YAML template. CloudFormation automatically handles resource dependencies, creation order, and rollback on failure, making it ideal for automating repeatable, consistent environments. It is the correct tool because you need to provision the resources themselves, not just deploy code to existing infrastructure.

Why this answer

AWS CloudFormation is the correct service because it allows you to define both the Lambda function and its IAM role as infrastructure as code using a template (JSON or YAML). CloudFormation handles the creation, updating, and deletion of these resources in an orderly, repeatable manner, ensuring the IAM role is created before the Lambda function due to dependency management.

Exam trap

The trap here is that candidates often confuse AWS CodeDeploy (which can deploy Lambda code) with the ability to create the Lambda function and its IAM role, but CodeDeploy does not provision the underlying infrastructure resources—it only handles the deployment of the code to an existing function.

How to eliminate wrong answers

Option B (AWS Elastic Beanstalk) is wrong because it is a PaaS service designed for deploying and scaling web applications, not for creating individual Lambda functions and IAM roles via infrastructure as code. Option C (AWS CodeDeploy) is wrong because it automates code deployments to EC2, Lambda, or on-premises instances, but it does not provision the underlying IAM roles or Lambda function resources; it only deploys the code. Option D (AWS Systems Manager) is wrong because it provides operational management and automation for AWS resources (e.g., patching, runbooks), but it is not designed for declarative infrastructure provisioning of Lambda functions and IAM roles.

Page 1 of 16

Page 2