Courseiva

AWS Certified SysOps Administrator Associate SOA-C02 (SOA-C02) — Questions 76–150

1169 questions total · 16pages · All types, answers revealed

Page 1

Page 2 of 16

Page 3
76
MCQeasy

A company is designing a disaster recovery plan for its on-premises database. They need to replicate the database to AWS with low latency. Which AWS service should they use?

A.Amazon S3 with Cross-Region Replication.
B.AWS Storage Gateway with volume gateway.
C.AWS Database Migration Service (DMS) with ongoing replication.
D.AWS Direct Connect to establish a dedicated network connection.
AnswerC

AWS Database Migration Service (DMS) with ongoing replication is the correct choice because DMS can perform continuous change data capture (CDC) from the source database and apply those changes to a target Amazon RDS instance in another Region. This enables a near-real-time replica of the database with low RPO, supporting disaster recovery by keeping the target transactionally consistent. DMS supports both homogeneous and heterogeneous migrations, and ongoing replication is a key feature for DR scenarios when native replication options are not available.

Why this answer

AWS DMS with ongoing replication (change data capture, CDC) is the correct choice because it can continuously replicate changes from an on-premises database to a target database in AWS with low latency, supporting heterogeneous and homogeneous migrations. This meets the requirement for a disaster recovery plan that keeps the AWS copy nearly synchronized with the on-premises source.

Exam trap

The trap here is confusing network connectivity services (like Direct Connect) or storage replication (like S3 CRR or Storage Gateway) with database-level replication, which requires transaction-consistent change capture and application.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Cross-Region Replication is an object-level replication mechanism for S3 buckets, not designed for database replication; it cannot capture transactional changes or maintain database consistency. Option B is wrong because AWS Storage Gateway with volume gateway provides block-level storage volumes that can be backed up to S3, but it does not offer ongoing database replication with low latency; it is intended for hybrid storage, not for replicating live database transactions. Option D is wrong because AWS Direct Connect establishes a dedicated network connection for improved bandwidth and latency, but it is a connectivity service, not a replication service; it does not replicate the database itself.

77
MCQeasy

A company uses Amazon Route 53 for DNS resolution. The company wants to ensure that if a web server becomes unhealthy, traffic is automatically routed to a healthy server in another Availability Zone. Which routing policy should be used?

A.Latency routing policy
B.Weighted routing policy
C.Geolocation routing policy
D.Failover routing policy
AnswerD

Failover routing policy is specifically designed for active-passive configurations: you create a primary record and a secondary record, and Route 53 uses health checks on the primary to determine its status. When the primary fails its health check, Route 53 automatically returns the secondary record's IP address, enabling DNS-level failover without manual intervention. This matches the company's need to route to a secondary endpoint whenever the primary is unhealthy, making it the correct answer.

Why this answer

The Failover routing policy (D) is correct because it is specifically designed to route traffic to a primary resource (e.g., a web server) and automatically redirect to a secondary resource in a different Availability Zone when health checks fail. Route 53 uses health checks to monitor the primary endpoint; if the primary becomes unhealthy, it returns the secondary record in DNS responses, ensuring automatic failover.

Exam trap

The trap here is that candidates often confuse Failover routing policy with Weighted routing policy, mistakenly thinking that weights can be adjusted dynamically to simulate failover, but Route 53 does not automatically adjust weights based on health—only Failover routing policy provides automatic, health-check-driven failover between a primary and secondary resource.

How to eliminate wrong answers

Option A is wrong because Latency routing policy routes traffic based on the lowest network latency to the client, not on health status or failover. Option B is wrong because Weighted routing policy distributes traffic across multiple resources based on assigned weights, but it does not automatically reroute all traffic to a healthy endpoint when one fails—it continues to send a portion of traffic to unhealthy endpoints unless health checks are manually configured to remove them. Option C is wrong because Geolocation routing policy routes traffic based on the geographic location of the DNS resolver, not on the health of the resources, and it does not provide automatic failover between Availability Zones.

78
MCQhard

A company has a production environment with multiple EC2 instances that send logs to CloudWatch Logs. The operations team wants to search across all log groups for a specific error pattern. What is the most efficient way to achieve this?

A.Use CloudWatch Logs Insights to query across all log groups.
B.Set up a subscription filter to stream logs to an Amazon ES domain.
C.Use CloudWatch Logs filter patterns on each log group.
D.Download all logs to an S3 bucket and use Amazon Athena to query.
AnswerA

CloudWatch Logs Insights is purpose-built for ad-hoc querying across multiple log groups. A single query can reference several log groups (e.g., by specifying logGroupNames with `*` wildcards or enumerating them), enabling you to search, filter, and aggregate fields like `@timestamp` and `@message` without moving data. This is the most direct and efficient way for a SysOps administrator to correlate logs from multiple EC2 instances in a production environment, because it requires no additional infrastructure or data pipelines and returns results within seconds. The query language supports commands like `fields`, `stats`, and `filter` to isolate specific errors, making it ideal for fast troubleshooting.

Why this answer

CloudWatch Logs Insights allows you to run SQL-like queries across multiple log groups in a single query, making it the most efficient way to search for a specific error pattern across all log groups without needing to set up additional infrastructure or manually query each group individually.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing a more complex architecture (like streaming to Elasticsearch or using Athena) when CloudWatch Logs Insights provides a native, serverless, and efficient way to query across multiple log groups directly.

How to eliminate wrong answers

Option B is wrong because setting up a subscription filter to stream logs to an Amazon ES domain adds unnecessary complexity, latency, and cost; it requires provisioning and managing an Elasticsearch cluster, which is overkill for a simple cross-log-group search. Option C is wrong because CloudWatch Logs filter patterns operate on a single log group at a time, so you would need to configure and run separate queries for each log group, which is inefficient and not scalable for searching across all log groups. Option D is wrong because downloading all logs to an S3 bucket and using Amazon Athena introduces significant overhead, including export delays, storage costs, and the need to define a schema; it is not the most efficient approach for real-time or ad-hoc searching across log groups.

79
Multi-Selecteasy

Which TWO are valid methods to secure traffic between a client and an Application Load Balancer?

Select 2 answers
A.Configure a listener on port 443 with an SSL certificate from AWS Certificate Manager.
B.Use a security group that only allows HTTPS traffic from the client's IP.
C.Set up an IPsec VPN connection between the client and the ALB.
D.Configure a network ACL to allow only port 443.
E.Enable the ALB's built-in SSL/TLS encryption without a certificate.
AnswersA, B

An Application Load Balancer (ALB) can terminate TLS by configuring an HTTPS listener on port 443. You must associate a valid SSL/TLS certificate, such as one issued by AWS Certificate Manager (ACM), which the ALB uses to decrypt incoming traffic and establish encrypted sessions with clients. This ensures data in transit is protected against eavesdropping and tampering. ACM integrates natively with ALB, handling certificate renewal automatically.

Why this answer

Configuring a listener on port 443 with an SSL certificate from ACM enables TLS encryption between the client and the ALB. Option B is also correct: using a security group that only allows HTTPS traffic enforces that all traffic must be encrypted, securing the communication by blocking unencrypted HTTP traffic. Options C, D, and E are incorrect: IPsec VPN is not terminated on an ALB (C), network ACLs do not provide encryption (D), and SSL/TLS encryption requires a valid certificate (E).

Exam trap

The trap is that candidates often think only option A (SSL termination) secures traffic, but using a security group to allow only HTTPS (option B) also ensures encryption by blocking unencrypted traffic.

80
MCQeasy

A SysOps administrator manages an Application Load Balancer (ALB) that distributes traffic to an Auto Scaling group of EC2 instances. The administrator needs to receive a notification whenever the number of unhealthy targets in the ALB target group exceeds a threshold of 2 for at least 5 consecutive minutes. Which solution meets this requirement with the least operational overhead?

A.Create a CloudWatch alarm on the 'UnHealthyHostCount' metric for the ALB target group, with a threshold of 2 and an evaluation period of 5 minutes. Configure the alarm to send an Amazon SNS notification.
B.Enable AWS CloudTrail logging for the ALB and create a CloudWatch metric filter for 'UnHealthyHostCount' events. Then create an alarm on that metric to notify via SNS.
C.Use an AWS Config rule to evaluate the health of the ALB target group and trigger an SNS notification when non-compliant.
D.Create an Amazon EventBridge rule that triggers every minute to call the AWS CLI command describe-target-health and send a notification via Lambda if unhealthy count exceeds 2.
AnswerA

CloudWatch automatically receives the UnHealthyHostCount metric from the ALB's target group, so a CloudWatch alarm can directly monitor unhealthy host counts without any custom code. Setting the threshold to 2 and an evaluation period of 5 minutes triggers the alarm when the count exceeds 2 for that duration, and the alarm's SNS action sends a notification to subscribed endpoints. This is the simplest and most reliable approach because it uses native AWS monitoring.

Why this answer

CloudWatch publishes UnHealthyHostCount for ALB target groups. To meet the requirement of >2 unhealthy targets for 5 consecutive minutes, the alarm must use a Period of 1 minute and an EvaluationPeriods value of 5. Setting the Period to 5 minutes would aggregate a 5-minute block and not verify each minute's count exceeded 2.

Option A as written is imprecise and could lead to a configuration that does not meet the requirement. The correct implementation still uses an SNS notification on the CloudWatch alarm, so Option A remains the best choice if corrected to specify the appropriate Period and EvaluationPeriods.

Exam trap

The trap is avoiding overcomplicated solutions like custom polling (Option D) or misapplied services (CloudTrail/Config). However, candidates must also correctly configure the CloudWatch alarm's Period and EvaluationPeriods to detect 5 consecutive minutes.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail logs API calls, not real-time metric data like 'UnHealthyHostCount'; creating a metric filter for 'UnHealthyHostCount' events is invalid as CloudTrail does not emit such events. Option C is wrong because AWS Config rules evaluate resource compliance against desired configurations (e.g., security groups, tags), not real-time health metrics like unhealthy host counts; Config cannot trigger based on dynamic metric thresholds. Option D is wrong because it introduces unnecessary operational overhead by requiring a custom Lambda function and EventBridge rule to poll the describe-target-health CLI command every minute, whereas CloudWatch provides a built-in, simpler solution.

81
MCQhard

A company uses AWS CloudFormation to deploy infrastructure. A SysOps admin wants to receive a notification when a stack update fails. Which approach is the most efficient?

A.Write a script that polls the CloudFormation API and sends notifications
B.Use AWS Config to monitor stack resources
C.Create an EventBridge rule that matches CloudFormation stack events
D.Enable CloudTrail and create a metric filter for stack update failures
AnswerC

Amazon EventBridge natively receives CloudFormation events such as STACK_UPDATE_ROLLBACK_IN_PROGRESS, STACK_CREATE_COMPLETE, and others, because CloudFormation publishes stack events to the default event bus. You create a rule with an event pattern for source 'aws.cloudformation' and detail-type 'CloudFormation Stack Status Change', then target SNS, Lambda, or another service to send notifications. This is the recommended serverless, event-driven approach because it reacts immediately and reliably to stack transitions.

Why this answer

Amazon EventBridge can directly capture CloudFormation stack events (e.g., CREATE_FAILED, UPDATE_FAILED) in real time and trigger a notification via SNS or Lambda. This approach is serverless, requires no polling, and is the most efficient method for reacting to stack update failures as they occur.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing CloudTrail or polling, missing the fact that EventBridge provides native, real-time event capture for CloudFormation stack status changes without additional overhead.

How to eliminate wrong answers

Option A is wrong because polling the CloudFormation API introduces latency, consumes unnecessary compute resources, and is less efficient than an event-driven approach. Option B is wrong because AWS Config is designed to evaluate resource compliance against rules, not to monitor CloudFormation stack lifecycle events or send failure notifications. Option D is wrong because CloudTrail logs API calls, but creating a metric filter for stack update failures requires additional steps (e.g., setting up a CloudWatch alarm) and introduces delay compared to native EventBridge event matching.

82
MCQmedium

Refer to the exhibit. A SysOps administrator runs the command to find 'CreateKeyPair' events in January 2023 but gets an empty list. The administrator knows that key pairs were created during that time. What is the most likely reason?

A.The events occurred in a different AWS region.
B.The start and end times are outside the 90-day retention period.
C.CloudTrail is not enabled in the account.
D.The IAM user does not have 'cloudtrail:LookupEvents' permission.
AnswerA

The `lookup-events` API call is regional in scope. When you issue `aws cloudtrail lookup-events --region us-east-1`, CloudTrail searches only the event history for that specific region. If the trail was configured as a single-region trail in a different region (or the events themselves were generated in another region), the lookup in us-east-1 will return zero results, even though the events exist in CloudTrail's global event history. To find them, you must explicitly specify the region where the trail is logging or where the events occurred.

Why this answer

The `aws cloudtrail lookup-events` command returns events only from the region specified in the AWS CLI configuration (or the `--region` parameter). If the administrator did not specify a region, the command defaults to the region set in the CLI profile. Since `CreateKeyPair` events are regional (each key pair is created in a specific region), the empty result indicates the events occurred in a different AWS region than the one queried.

Exam trap

The trap here is that candidates assume CloudTrail events are globally visible by default, but in reality, `lookup-events` is region-scoped unless the `--region` parameter is explicitly set to the correct region.

How to eliminate wrong answers

Option B is wrong because the 90-day retention period applies to CloudTrail event history, and January 2023 is well within 90 days from the current date (assuming the exam is set in 2023 or later), so the start and end times are not outside the retention period. Option C is wrong because CloudTrail is enabled by default in all AWS accounts, and the `lookup-events` command works with the default event history even without a specific trail. Option D is wrong because if the IAM user lacked `cloudtrail:LookupEvents` permission, the command would return an access denied error, not an empty list.

83
MCQmedium

A SysOps administrator needs to implement a backup strategy for an Amazon RDS for PostgreSQL database. The database is 500 GB and experiences heavy write traffic. Which solution provides the most cost-effective backup with the least impact on database performance?

A.Enable automated backups with a retention period of 7 days.
B.Create a Multi-AZ deployment and use the standby for backups.
C.Use AWS Database Migration Service to continuously replicate data to an S3 bucket.
D.Take manual DB snapshots daily during off-peak hours.
AnswerA

Automated backups are the native RDS backup mechanism: daily snapshots of the database volume are captured, and transaction logs are continuously uploaded to S3 so you can restore to any point within the 7-day retention window. These backup operations are low overhead because they leverage the EBS snapshot facility, so the performance impact is minimal and no long-duration downtime is required. With a 7-day retention, you automatically meet a standard RPO (point-in-time) without manual effort or risk of expiring snapshots.

Why this answer

Automated backups are enabled by default with minimal performance impact and include transaction logs for point-in-time recovery. Option B is wrong; while using a standby for backups in Multi-AZ can reduce impact, automated backups already have minimal impact and are more cost-effective. Option C is wrong because AWS DMS replication to S3 is not a backup solution and adds complexity and cost.

Option D is wrong because manual snapshots cause a brief I/O suspension and are less automated than automated backups.

84
MCQhard

A company runs a critical application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application requires very low latency and high availability. The SysOps administrator notices that the application experiences increased latency during traffic spikes even though the Auto Scaling group is scaling out. Which solution would MOST effectively reduce latency?

A.Change the launch configuration to use a larger instance type with more CPU and memory.
B.Pre-warm the load balancer by contacting AWS Support.
C.Increase the Auto Scaling cooldown period.
D.Deploy the instances in multiple Availability Zones.
AnswerA

Changing the launch configuration to a larger instance type (e.g., from t3.medium to t3.large) vertically scales each EC2 instance, providing more vCPUs and memory to process requests concurrently. This directly reduces per-request CPU queueing and memory pressure during traffic spikes, lowering latency. However, the launch configuration only applies to newly launched instances, so an instance refresh or replacement may be needed to realize the benefit on existing instances.

Why this answer

Using a larger instance type with more CPU and memory directly addresses the root cause of increased latency during traffic spikes: the existing instances are resource-constrained under load. By provisioning instances with higher compute capacity, each instance can handle more requests per second, reducing queueing delays and per-request processing time. This is a more immediate and effective solution than scaling out alone, which adds instances but does not improve the performance of each individual instance.

Exam trap

The trap here is that candidates may assume scaling out (adding more instances) always reduces latency, but the question explicitly states that scaling is already occurring yet latency persists, indicating a per-instance performance bottleneck that only a larger instance type can resolve.

How to eliminate wrong answers

Option B is wrong because pre-warming the load balancer is a manual process typically used for anticipated large traffic events (e.g., product launches) and does not address latency caused by under-provisioned instances; the ALB itself is not the bottleneck here. Option C is wrong because increasing the Auto Scaling cooldown period would actually slow down the scaling response, making latency worse during traffic spikes by delaying the addition of new instances. Option D is wrong because deploying instances in multiple Availability Zones improves fault tolerance and availability, but does not reduce per-instance latency under load; it may even increase cross-AZ data transfer costs without solving the resource contention issue.

85
MCQeasy

A company has an Amazon CloudFront distribution with an S3 bucket as origin. The bucket contains sensitive data. Which configuration ensures that users access the content only through CloudFront and not directly via the S3 URL?

A.Enable S3 server-side encryption
B.Configure an Origin Access Identity (OAI) in CloudFront and update the bucket policy
C.Use CloudFront signed URLs or signed cookies
D.Enable S3 Block Public Access on the bucket
AnswerB

An Origin Access Identity is a special CloudFront principal that S3 recognizes; after creating it, you attach it to your distribution and rewrite the bucket policy to grant s3:GetObject only to that OAI's canonical user ID. This makes the bucket private to everyone except CloudFront, so users cannot bypass CloudFront by hitting the S3 endpoint directly. Combined with the distribution's behavior, it is the standard way to force all traffic through CloudFront for a private S3 origin.

Why this answer

An Origin Access Identity (OAI) is a special CloudFront user that the distribution uses to fetch objects from the S3 bucket. By updating the bucket policy to grant read access only to that OAI and removing public access, direct requests to the S3 URL are denied while CloudFront can still serve the content. This is the standard AWS pattern for locking an S3 origin behind CloudFront.

Exam trap

The trap is assuming that signed URLs or Block Public Access alone secure the origin — candidates forget that without an OAI/OAC and a restrictive bucket policy, the S3 URL remains directly reachable.

How to eliminate wrong answers

Option A is wrong because SSE encrypts objects at rest but does not prevent users from accessing the object via the S3 URL if the bucket policy allows it. Option C is wrong because signed URLs/cookies control who can access content through CloudFront, but they do not stop someone from bypassing CloudFront and hitting the S3 endpoint directly. Option D is wrong because Block Public Access prevents public access but does not by itself grant CloudFront the necessary permissions — without an OAI and bucket policy, CloudFront would also be denied.

86
MCQmedium

A company runs a web application on EC2 instances with Elastic Load Balancing. They notice that costs are higher than expected. What is the MOST cost-effective way to optimize costs while maintaining high availability?

A.Use Spot Instances for all traffic.
B.Use On-Demand instances for all traffic.
C.Use a combination of Reserved Instances for baseline traffic and On-Demand for spikes.
D.Reduce the number of instances to one.
AnswerC

This is the correct approach because Reserved Instances provide a significant hourly discount in exchange for a one- or three-year commitment, making them ideal for the steady, predictable baseline web traffic. On-Demand Instances, though more expensive, are billed per second with no commitment, so they are perfect for absorbing temporary spikes without over-provisioning or risking capacity loss. By covering the minimum long-term load with Reserved Instances and adding On-Demand Instances only when load increases, you minimize spend while maintaining availability and elasticity — the core goal of cost-aware architecture.

Why this answer

Reserved Instances provide a significant discount (up to ~72% versus On-Demand) in exchange for a one- or three-year commitment, making them ideal for the predictable baseline load. On-Demand instances then absorb traffic spikes without a long-term commitment, so the workload stays highly available while the blended cost is minimized. This hybrid model is the standard AWS cost-optimization pattern for steady-state web tiers behind an ELB.

Exam trap

The trap here is confusing Spot Instances with a general cost-saving answer — candidates forget that Spot's interruption model makes it unsuitable when the question explicitly requires high availability.

How to eliminate wrong answers

Option A is wrong because Spot Instances can be reclaimed by AWS with only a two-minute interruption notice, so using them for all traffic would break the high-availability requirement. Option B is wrong because paying full On-Demand rates for the predictable baseline wastes the discount that Reserved Instances or Savings Plans would capture. Option D is wrong because collapsing to a single instance eliminates redundancy and creates a single point of failure, violating the high-availability requirement outright.

87
Multi-Selectmedium

A company is designing a disaster recovery plan for its critical applications. The plan must minimize data loss and recovery time. Which TWO measures should the SysOps administrator implement?

Select 2 answers
A.Perform regular backups to Amazon S3.
B.Set a recovery time objective (RTO) of 24 hours.
C.Use manual procedures to restore from backups.
D.Run all workloads in a single AWS region.
E.Replicate data to another AWS region.
AnswersA, E

Perform regular backups to Amazon S3 is a valid DR component because S3 provides 11 nines of durability, versioning for point-in-time recovery, and lifecycle policies to archive to S3 Glacier. Backups create immutable, restorable copies of critical data, protecting against accidental deletion, corruption, or ransomware. However, backups alone are not a complete DR plan; they establish an RPO tied to backup frequency and must be restored before workloads can resume, which elongates RTO. Therefore, this action is correct as a foundational data-protection measure, not as a full DR strategy.

Why this answer

Regular backups to Amazon S3 are a foundational data protection measure because S3 provides 99.999999999% durability and supports lifecycle policies for cost-effective long-term retention. This directly addresses the requirement to minimize data loss by ensuring point-in-time recovery copies exist independently of the primary infrastructure.

Exam trap

The trap here is that candidates often confuse RTO/RPO definitions with actual implementation measures, or they assume that a single region with backups is sufficient for disaster recovery, ignoring the need for geographic separation to survive a region-wide outage.

88
MCQmedium

A company has deployed a web application behind an Application Load Balancer (ALB) across multiple Availability Zones. Users in some regions report slow page load times. Which action should the SysOps Administrator take to improve performance for all users?

A.Use AWS Global Accelerator to route traffic over the AWS global network.
B.Increase the ALB capacity by adding more target instances.
C.Enable Amazon CloudFront to cache dynamic content.
D.Move the application to a single Availability Zone to reduce network hops.
AnswerA

AWS Global Accelerator assigns two static anycast IP addresses at AWS edge locations and directs traffic onto the AWS global backbone, bypassing congested public internet segments. This reduces the number of intermediate hops and round-trip latency for users worldwide, while also providing automatic failover across healthy ALB endpoints. It is specifically designed for TCP/UDP workloads where each request requires low and stable latency, unlike caching services that only speed up repeated content.

Why this answer

AWS Global Accelerator improves performance by directing traffic over the AWS global network and using edge locations close to users, reducing latency for all users. Option A is correct because it optimizes the path from users to the application. Option B is incorrect because increasing ALB capacity only helps with handling more requests, not latency for geographically distant users.

Option C is incorrect because CloudFront is primarily for caching static content, and the application serves dynamic content that may not be cacheable. Option D is incorrect because moving to a single Availability Zone reduces fault tolerance and does not address latency for users far from that zone.

89
MCQeasy

A company runs a web application on Amazon EC2 instances that run 24/7. The application has a predictable and steady load. The SysOps administrator wants to minimize compute costs while ensuring the required capacity is always available. Which purchasing option should be used?

A.On-Demand Instances
B.Reserved Instances
C.Spot Instances
D.Dedicated Hosts
AnswerB

Reserved Instances are a billing discount that requires a 1- or 3-year commitment to a specific instance family, region, and operating system. In exchange for that commitment, you receive up to a 72% discount off the On-Demand hourly rate, depending on the payment option (All Upfront, Partial Upfront, or No Upfront). For an always-on web application with predictable compute needs, this significantly lowers the monthly cost, making Reserved Instances the most cost-effective choice among the available options.

Why this answer

Reserved Instances (RIs) are the correct choice because the workload runs 24/7 with a predictable and steady load. By committing to a 1- or 3-year term, the company can achieve a significant discount (up to 72%) compared to On-Demand pricing, while ensuring capacity is always available. This aligns with the goal of minimizing compute costs without sacrificing availability.

Exam trap

The trap here is that candidates often choose On-Demand Instances for simplicity, overlooking that Reserved Instances provide substantial cost savings for predictable, always-on workloads without any risk of interruption.

How to eliminate wrong answers

Option A is wrong because On-Demand Instances are billed per hour with no upfront commitment, making them more expensive for a steady 24/7 workload that could benefit from a reservation discount. Option C is wrong because Spot Instances can be interrupted with a 2-minute warning when AWS needs capacity back, making them unsuitable for a production web application that requires always-on availability. Option D is wrong because Dedicated Hosts provide physical server isolation for licensing or compliance needs, not cost savings for a standard web application, and they incur additional per-host charges.

90
MCQmedium

A company stores application logs in Amazon S3. The logs are rarely accessed after the first 30 days, but must be retained for 7 years for compliance. The SysOps administrator wants to minimize storage costs while ensuring logs are available for retrieval within 12 hours if needed. Which S3 lifecycle configuration is the most cost-effective?

A.S3 Standard for 30 days, then transition to S3 Glacier Deep Archive
B.S3 Standard for 30 days, then transition to S3 Glacier (Flexible Retrieval)
C.Use S3 Intelligent-Tiering from the start
D.S3 Standard for 30 days, then transition to S3 One Zone-IA
AnswerA

This is the most cost-effective lifecycle policy for logs that are frequently accessed for the first 30 days but rarely read afterward. A lifecycle rule can automatically transition objects to S3 Glacier Deep Archive after 30 days, where storage costs just $0.00099 per GB-month. Even though standard retrieval can take up to 12 hours, that is acceptable for compliance log analysis that is not time-sensitive.

Why this answer

It transitions logs from S3 Standard (for frequent access during the first 30 days) to S3 Glacier Deep Archive, which offers the lowest storage cost for long-term retention. Since logs must be retained for 7 years and only need retrieval within 12 hours, Glacier Deep Archive's 12-hour standard retrieval time meets the requirement while minimizing costs.

Exam trap

The trap here is that candidates may choose S3 Glacier (Flexible Retrieval) because it is a well-known archival tier, but they overlook the specific 12-hour retrieval requirement and the lower cost of Glacier Deep Archive for long-term retention.

How to eliminate wrong answers

Option B is wrong because S3 Glacier (Flexible Retrieval) has higher storage costs than Glacier Deep Archive for 7-year retention, and its standard retrieval time (1-5 minutes) is faster than needed, making it less cost-effective. Option C is wrong because S3 Intelligent-Tiering incurs a monthly monitoring and automation fee per object, and for data that is rarely accessed after 30 days, the cost savings from tiering do not offset the monitoring fee over 7 years. Option D is wrong because S3 One Zone-IA is not designed for long-term archival; it offers lower durability (99.5% vs 99.999999999%) and is not cost-effective for 7-year retention compared to Glacier Deep Archive.

91
MCQhard

A company has a VPC with a public subnet and a private subnet. An Amazon EC2 instance in the private subnet needs to download security patches from the internet, but the instance must not be directly accessible from the internet. The SysOps administrator configured a NAT gateway in the public subnet and added a route in the private subnet's route table pointing 0.0.0.0/0 to the NAT gateway. The instance's security group allows all outbound traffic. However, the instance still cannot reach the internet. What is the most likely missing configuration?

A.Attach an Elastic IP to the NAT gateway
B.Enable DNS resolution in the VPC
C.Add a route in the public subnet's route table that directs 0.0.0.0/0 traffic to an internet gateway
D.Modify the network ACL of the private subnet to allow inbound ephemeral ports from the NAT gateway's private IP
AnswerC

A NAT gateway must be launched in a public subnet, and that subnet's route table needs a destination of 0.0.0.0/0 pointing to an internet gateway, not to another gateway or target. Without this route, the NAT gateway's network interface cannot send translated packets to the IGW or receive return packets, so all traffic from private instances times out. Adding this route is the correct fix because it establishes the final hop between the NAT gateway and the internet.

Why this answer

The NAT gateway is in the public subnet, but for it to route traffic to the internet, the public subnet must have a route table entry that directs 0.0.0.0/0 traffic to an internet gateway (IGW). Without this route, the NAT gateway cannot forward outbound traffic to the IGW, so the private instance's traffic is dropped. Option C correctly identifies this missing route.

Exam trap

The trap here is that candidates assume configuring the private subnet's route table to point to the NAT gateway is sufficient, forgetting that the NAT gateway itself needs a route to the internet via an internet gateway in its own subnet.

How to eliminate wrong answers

Option A is wrong because a NAT gateway automatically gets an Elastic IP assigned at creation; if it were missing, the NAT gateway would fail to provision, not silently fail to route traffic. Option B is wrong because DNS resolution controls the ability to resolve domain names to IP addresses, not the underlying network path for outbound traffic; the instance can still fail to reach the internet even with DNS working. Option D is wrong because the network ACL of the private subnet must allow outbound ephemeral ports for return traffic, not inbound; the default NACL already allows all inbound/outbound traffic, and the issue is the missing route in the public subnet, not NACL rules.

92
MCQhard

A SysOps administrator updates a CloudFormation stack to change the EC2 instance type from t2.micro to t3.medium. The update fails with the error shown. What is the MOST likely cause?

A.The account does not have service limits to launch a t3.medium instance.
B.The AMI used does not support the t3.medium instance type.
C.The CloudFormation template has a parameter constraint that rejects t3.medium.
D.The t3.medium instance type is not available in the specified Availability Zone.
AnswerD

The error is exactly what AWS returns when you attempt to launch an EC2 instance with an instance type that is not available in the selected Availability Zone. Instance types are rolled out to AZs non-uniformly, so a t3.medium may exist in us-east-1a but not us-east-1b, for example. This is a resource-level launch failure rather than a template, account, or AMI problem.

Why this answer

The error indicates that the t3.medium instance type is not available in the specified Availability Zone (AZ). AWS instance types are offered on a per-AZ basis, and not all instance types are available in every AZ. When a CloudFormation stack update fails with this error, it typically means the template explicitly or implicitly specifies an AZ that does not support the target instance type.

Exam trap

The trap here is that candidates often confuse 'unavailable in AZ' with 'service limit exceeded' or 'AMI incompatibility', but the specific error message about instance type availability points directly to an AZ constraint.

How to eliminate wrong answers

Option A is wrong because service limits would cause a different error (e.g., 'LimitExceeded' or 'InsufficientInstanceCapacity'), not an 'unavailable instance type' error. Option B is wrong because AMI compatibility with instance types is generally about driver support (e.g., ENA or NVMe), and the error message does not reference AMI issues; an incompatible AMI would produce a launch failure, not an 'unavailable' error. Option C is wrong because parameter constraints in CloudFormation templates are evaluated during stack creation or update validation, and a constraint violation would produce a validation error (e.g., 'Value failed to satisfy constraint'), not an AZ availability error.

93
MCQmedium

A company is running a web application on EC2 instances behind an Application Load Balancer. They want to ensure that if an entire Availability Zone fails, the application remains available. Which configuration should they implement?

A.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
B.Use an RDS Multi-AZ deployment for the application.
C.Use a larger EC2 instance type.
D.Enable detailed monitoring on the EC2 instances.
AnswerA

Configuring the Auto Scaling group to span multiple Availability Zones is the correct approach because it distributes the EC2 instances across independent failure domains. If one AZ becomes unavailable, the instances in the other AZs continue to serve traffic, and the ASG automatically replaces the unhealthy instances to maintain desired capacity. This architecture provides high availability for the application tier.

Why this answer

To ensure application availability during an entire Availability Zone (AZ) failure, the Auto Scaling group must be configured to launch EC2 instances across multiple AZs. This distributes the workload so that if one AZ becomes unavailable, the remaining AZs continue serving traffic. The Application Load Balancer (ALB) automatically routes requests only to healthy instances in the surviving AZs, maintaining application uptime.

Exam trap

The trap here is that candidates often confuse high availability at the database layer (RDS Multi-AZ) with compute layer fault tolerance, or they mistakenly believe that scaling vertically (larger instances) or increasing monitoring granularity can compensate for a full AZ outage.

How to eliminate wrong answers

Option B is wrong because RDS Multi-AZ provides high availability for the database layer, not for the compute layer (EC2 instances) handling the web application; it does not address EC2 instance distribution across AZs. Option C is wrong because using a larger EC2 instance type increases compute capacity within a single AZ but does not protect against an AZ failure; the instance would still be unavailable if its AZ fails. Option D is wrong because enabling detailed monitoring on EC2 instances provides more granular CloudWatch metrics (1-minute intervals) but does not affect instance placement or fault tolerance across AZs.

94
MCQeasy

A company wants to visualize the geographic distribution of failed login attempts to their web application. The application runs on EC2 instances behind an ALB. They have access logs enabled for the ALB. Which service should be used to create the visualization?

A.Amazon CloudWatch Dashboard with a custom widget.
B.Amazon Kinesis Data Analytics with a Lambda function.
C.Amazon S3 Select with Athena.
D.Amazon QuickSight with S3 as a data source.
AnswerD

Amazon QuickSight is a fully managed business intelligence service that supports geospatial chart types such as point maps and heat maps, and it can connect directly to data stored in Amazon S3—either through native S3 ingestion or by querying with Athena. With QuickSight, you can prepare and visualize ALB access logs by plotting fields like client IPs or derived location attributes (country, city, or coordinates) to show geographic distribution through an interactive dashboard.

Why this answer

Amazon QuickSight is a fully managed business intelligence service that can directly query ALB access logs stored in Amazon S3, enabling the creation of geospatial visualizations (e.g., heat maps) of failed login attempts. ALB access logs are delivered to S3 in a structured format, and QuickSight can parse and visualize this data without additional processing. This makes QuickSight with S3 as a data source the correct choice for building the required geographic visualization.

Exam trap

The trap here is that candidates confuse data querying (Athena) with data visualization (QuickSight), or assume CloudWatch can handle geospatial log analysis, when in fact QuickSight is the only option that natively provides interactive geospatial dashboards from S3 data.

How to eliminate wrong answers

Option A is wrong because CloudWatch Dashboards with custom widgets are designed for real-time metrics and logs, not for ad-hoc geospatial analysis of historical ALB access logs stored in S3; they lack native geospatial visualization capabilities. Option B is wrong because Kinesis Data Analytics is a real-time stream processing service, not suited for batch visualization of historical log data, and adding a Lambda function introduces unnecessary complexity and cost for a simple query-and-visualize task. Option C is wrong because S3 Select is a server-side filtering tool that returns only a subset of data from an object, not a visualization service; Athena can query the logs but does not create visualizations—it would require an additional BI tool to render the geographic map.

95
Multi-Selectmedium

A SysOps administrator is designing a highly available architecture for a web application using an Application Load Balancer (ALB) with EC2 instances in an Auto Scaling group. Which TWO configurations are required to ensure high availability? (Choose TWO.)

Select 2 answers
A.Launch all EC2 instances in a single Availability Zone to reduce latency
B.Use t2.micro instances to reduce cost
C.Configure the ALB with health checks for the target group
D.Disable health checks to reduce load on the ALB
E.Configure the Auto Scaling group to launch instances in at least two Availability Zones
AnswersC, E

The ALB continuously sends health-check requests (e.g., HTTP GET to a specified path) to each instance in its target group; an instance that fails a set number of consecutive checks is marked unhealthy and automatically deregistered, so the ALB stops forwarding new traffic to it and reroutes incoming requests to healthy instances. Configuring health checks with appropriate interval, timeout, and threshold values detects underlying application or instance failures early, enabling rapid failover and improving overall fault tolerance. This is a critical control point; without meaningful health checks, even a perfectly scaled fleet will serve errors to a portion of requests.

Why this answer

Health checks allow the ALB to monitor the status of each EC2 instance in the target group. If an instance fails health checks, the ALB automatically stops routing traffic to it, preventing user requests from reaching a failed instance. This is essential for maintaining application availability and is a core feature of the ALB's high-availability design.

Exam trap

The trap here is that candidates often confuse cost-saving measures (like using smaller instance types) or performance optimizations (like single-AZ deployment for lower latency) with high-availability requirements, but the exam specifically tests the understanding that high availability requires redundancy across Availability Zones and active health monitoring.

96
MCQmedium

A company's security policy requires that all Amazon EC2 instances must have a specific tag 'Environment' with a value of either 'Production' or 'Development'. The SysOps administrator needs to detect any instance that is missing this tag or has an invalid value, and automatically email the operations team. Which AWS service should be used to achieve this with the least operational overhead?

A.AWS Config with the 'required-tags' managed rule and Amazon SNS
B.Amazon CloudWatch Events with an EC2 instance state change rule and AWS Lambda
C.AWS Trusted Advisor with a custom check
D.Amazon Inspector with a network assessment
AnswerA

AWS Config continuously evaluates EC2 instances against the required-tags managed rule, which verifies that the mandated tag keys (and optionally values) are attached to each instance. On launch or any configuration change, noncompliant resources are recorded, and an Amazon SNS notification is delivered instantly to the security team. This approach is fully managed, requires no custom code, and maintains an audit trail of tag compliance over time.

Why this answer

AWS Config's 'required-tags' managed rule continuously evaluates EC2 instances against the specified tag key and allowed values, triggering an SNS notification when non-compliant resources are detected. This provides automated detection and alerting with minimal operational overhead, as it requires no custom code or infrastructure management.

Exam trap

The trap here is that candidates may confuse AWS Config's continuous compliance evaluation with event-driven services like CloudWatch Events, assuming that a state change rule can also check tags, but Config is purpose-built for resource configuration auditing without custom code.

How to eliminate wrong answers

Option B is wrong because CloudWatch Events with an EC2 instance state change rule only triggers on state transitions (e.g., running, stopped), not on tag compliance; it would require a custom Lambda function to check tags, adding operational overhead. Option C is wrong because AWS Trusted Advisor does not support custom checks; it only provides predefined best-practice checks. Option D is wrong because Amazon Inspector performs network assessments for vulnerabilities and unintended network access, not tag compliance.

97
MCQeasy

A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which action should be taken?

A.Launch a second instance in a different Availability Zone.
B.Assign an Elastic IP address to the instance.
C.Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
D.Place the instance in an Auto Scaling group with a min size of 1.
AnswerC

Creating a CloudWatch alarm on the StatusCheckFailed metric and configuring the recovery action is the correct method because EC2 instance recovery automatically restarts the instance on new hardware when the underlying host fails. This recovery action preserves the instance ID, private IP address, Elastic IP address, and all EBS volumes, so the instance's identity and configuration are maintained. The alarm must monitor the StatusCheckFailed_System metric (or the aggregate StatusCheckFailed metric) and invoke the 'recover' action to trigger the automatic recovery.

Why this answer

A CloudWatch alarm on the StatusCheckFailed metric can be configured with the 'recover' action to automatically restart the EC2 instance on a new underlying host if a hardware failure is detected. This recovery action preserves the instance ID, private IP, Elastic IP, and instance metadata, ensuring minimal disruption. The StatusCheckFailed metric specifically monitors the instance's system status checks, which detect AWS hardware issues.

Exam trap

The trap here is that candidates often confuse Auto Scaling recovery (which replaces the instance) with CloudWatch alarm recovery (which recovers the same instance), leading them to choose Option D despite the requirement to preserve the original instance.

How to eliminate wrong answers

Option A is wrong because launching a second instance in a different Availability Zone does not automatically recover the original instance from hardware failure; it creates a separate instance that requires manual or automated traffic redirection. Option B is wrong because assigning an Elastic IP address only provides a static public IP, but does not trigger any recovery mechanism when the underlying hardware fails. Option D is wrong because placing the instance in an Auto Scaling group with a min size of 1 will replace a failed instance with a new one, but it does not preserve the original instance ID, private IP, or Elastic IP, and the replacement is not a recovery of the same instance.

98
MCQmedium

A company runs a web application on Amazon EC2 instances. The application logs are sent to Amazon CloudWatch Logs. The SysOps administrator needs to monitor the logs for an increasing number of HTTP 500 errors. The administrator wants to create a metric filter that will count the number of lines containing 'HTTP 500' in the log group. Which syntax should the administrator use for the metric filter pattern?

A.[error, HTTP, 500]
B."HTTP 500"
C."HTTP" && "500"
D.[HTTP, 500, ...]
AnswerB

Wrapping 'HTTP 500' in double quotes makes the pattern a single literal phrase. In CloudWatch Logs metric filters, a quoted string with spaces matches the exact substring; so any log line containing the characters 'HTTP 500' in that order will trigger the filter. That is the only correct form among the options.

Why this answer

CloudWatch Logs metric filter patterns use literal string matching by enclosing the exact text in double quotes. The pattern "HTTP 500" will match any log line that contains the exact substring 'HTTP 500', which is the simplest and most reliable way to count occurrences of HTTP 500 errors.

Exam trap

The trap here is that candidates confuse the space-delimited token pattern syntax (square brackets) with literal string matching, leading them to choose options like A or D that only match specific token positions rather than any occurrence of 'HTTP 500' in the log line.

How to eliminate wrong answers

Option A is wrong because the syntax [error, HTTP, 500] is a space-delimited token pattern that would match a log line starting with three space-separated tokens like 'error HTTP 500', not a substring anywhere in the line. Option C is wrong because "HTTP" && "500" is not valid CloudWatch Logs metric filter syntax; the && operator is not supported in metric filter patterns. Option D is wrong because [HTTP, 500, ...] is a token-based pattern that expects 'HTTP' and '500' as the first two tokens in the log line, and it would not match lines where 'HTTP 500' appears later in the line or with additional text before it.

99
MCQmedium

A SysOps administrator is setting up monitoring for an application that runs on Amazon ECS with Fargate launch type. The application's performance degrades when memory utilization exceeds 80%. The administrator wants to receive a notification when memory usage approaches this threshold. What should the administrator do?

A.Install the CloudWatch agent on each Fargate task
B.Use Amazon CloudWatch Container Insights to view ReservedMemory metric
C.Enable Service Auto Scaling with a target tracking policy based on MemoryUtilization
D.Create a CloudWatch alarm on the ECS service's MemoryUtilization metric
AnswerD

The Amazon ECS service publishes the MemoryUtilization metric, which represents the average memory used by all tasks in the service relative to the memory reserved for each task. You can create a CloudWatch alarm on this metric with a threshold that, when breached, sends a message to an Amazon SNS topic subscribed to by the SysOps administrator. This alarm-plus-notification pattern directly addresses the requirement to be notified of memory issues.

Why this answer

Amazon ECS services automatically publish a `MemoryUtilization` metric to CloudWatch for Fargate tasks. By creating a CloudWatch alarm on this metric with a threshold of 80%, the administrator can trigger an SNS notification when memory usage approaches the threshold, enabling proactive remediation before performance degrades.

Exam trap

The trap here is that candidates often assume they need to install an agent or use Container Insights to get memory metrics, but ECS Fargate automatically publishes `MemoryUtilization` and `CPUUtilization` metrics to CloudWatch without any extra setup.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent cannot be installed on Fargate tasks; Fargate is a serverless compute engine that does not allow direct access to the underlying host or installation of agents. Option B is wrong because Container Insights provides aggregated metrics and logs for cluster-level visibility, but it does not expose a `ReservedMemory` metric; the relevant metric for memory usage is `MemoryUtilization`, which is already available without Container Insights. Option C is wrong because Service Auto Scaling with a target tracking policy based on `MemoryUtilization` would automatically adjust the number of tasks to maintain a target utilization, but the question asks for a notification when memory approaches 80%, not for automatic scaling.

100
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer (ALB). The application experiences intermittent 502 errors. The SysOps administrator checks the ALB access logs and sees that the error occurs when the target group has 'unhealthy' targets. What is the MOST likely cause of the 502 errors?

A.The SSL certificate on the ALB is expired.
B.The ALB does not have enough capacity to handle the traffic.
C.The client request exceeds the idle timeout.
D.The target instances are not passing health checks.
AnswerD

If the target instances are failing health checks, the ALB will eventually stop routing traffic to them, but while they are in a transitional state or if health checks are misconfigured, the ALB may still attempt to send requests to a non-responsive instance. When the ALB cannot establish a connection or receives no valid response from the target, it returns 502 Bad Gateway. This is one of the most common causes of 502 errors in ALB deployments, especially when targets intermittently fail health checks.

Why this answer

A 502 Bad Gateway error from an ALB occurs when the load balancer cannot successfully forward a request to a healthy target. If the target group has unhealthy targets, the ALB may have no healthy targets to route to, resulting in 502 errors. The most likely cause is that the target instances are failing health checks, so the ALB marks them as unhealthy and cannot serve requests.

Exam trap

The trap here is that candidates may attribute 502 errors to network or capacity issues, but the key clue is 'unhealthy targets' in the logs, pointing directly to health check failures.

How to eliminate wrong answers

Option A is wrong because an expired SSL certificate on the ALB would cause SSL/TLS handshake failures, typically resulting in 503 or certificate errors, not 502. Option B is wrong because ALB capacity is managed by AWS and scales automatically; insufficient capacity is not a typical cause of 502 errors. Option C is wrong because if the client request exceeds the idle timeout, the ALB may close the connection, but this usually results in a 504 Gateway Timeout, not a 502.

101
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer. The database is an RDS MySQL instance with Multi-AZ enabled. The application experiences intermittent 5xx errors that correlate with database failover events. What is the MOST likely cause and solution?

A.Configure the application to use the RDS reader endpoint.
B.Use a read replica to offload read traffic and reduce failover impact.
C.Increase the database connection pool size to handle retries.
D.Enable DNS caching with a low TTL in the application and use the RDS instance endpoint with a retry mechanism.
AnswerD

The RDS instance endpoint is a CNAME record that AWS automatically points to the current primary instance's IPv4 address. By configuring the application's DNS resolver to cache this record with a low TTL (e.g., 5–10 seconds), the app will quickly re-resolve to the new primary after a failover. Pairing this with a retry loop that catches transient connection errors ensures that any packets sent during the failover window are re-attempted once DNS has propagated. This combination directly minimizes downtime and is the officially recommended client-side failover approach.

Why this answer

During an RDS Multi-AZ failover, the DNS record for the primary instance endpoint is updated to point to the standby instance. If the application caches the old DNS resolution with a high TTL, it continues to send connections to the unreachable primary, causing 5xx errors. Enabling DNS caching with a low TTL (e.g., 5 seconds) and implementing a retry mechanism ensures the application quickly resolves the new endpoint and reconnects, minimizing downtime.

Exam trap

The trap here is that candidates often confuse the reader endpoint with the writer endpoint, or assume that read replicas or connection pooling can mitigate failover errors, when the real issue is DNS caching and the need for a retry mechanism.

How to eliminate wrong answers

Option A is wrong because the RDS reader endpoint is used for read-only traffic from read replicas, not for handling failover of the primary writer instance; during a failover, the writer endpoint is the one that updates its DNS. Option B is wrong because read replicas do not participate in Multi-AZ failover; they are for scaling read traffic and do not provide automatic failover for the primary database. Option C is wrong because increasing the connection pool size does not address the root cause of stale DNS resolution; it only adds more connections that will still fail until the DNS cache is refreshed.

102
MCQmedium

A company's security policy requires that IAM users rotate their access keys every 90 days. The SysOps administrator must automatically identify users whose access keys are older than 90 days and notify the security team. Which combination of AWS services should be used to meet this requirement with the least operational overhead?

A.AWS Config with the 'access-keys-rotated' managed rule and Amazon SNS
B.AWS CloudTrail and Amazon CloudWatch Logs with metric filters and alarms
C.IAM Access Analyzer and AWS Lambda
D.Amazon GuardDuty and Amazon EventBridge
AnswerA

AWS Config's managed rule 'access-keys-rotated' continuously evaluates each IAM user's active access keys against the configured maximum age (default 90 days). When the rule detects a key that exceeds the threshold, it marks the user as noncompliant and emits a compliance change notification through the Config delivery channel, which publishes to an Amazon SNS topic. Subscribing security personnel to that topic via email or SMS gives them a near-real-time alert, making this a purpose-built, scalable solution for enforcing rotation policy.

Why this answer

AWS Config's 'access-keys-rotated' managed rule checks whether IAM user access keys have been rotated within the specified number of days (default 90). When a non-compliant resource is detected, AWS Config can trigger an Amazon SNS notification directly, without any custom code or additional infrastructure. This combination provides a fully managed, serverless solution with the least operational overhead.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing custom Lambda or CloudTrail-based approaches, missing that AWS Config provides a fully managed, built-in rule specifically designed for this exact compliance check with zero custom code.

How to eliminate wrong answers

Option B is wrong because CloudTrail logs API calls but does not evaluate the age of access keys; metric filters and alarms would require custom log parsing and lack the built-in compliance check. Option C is wrong because IAM Access Analyzer focuses on analyzing resource-based policies for external access, not on key rotation age; using Lambda would add custom code and maintenance overhead. Option D is wrong because GuardDuty is a threat detection service for malicious activity, not for tracking key rotation compliance; EventBridge alone cannot perform the age evaluation.

103
MCQmedium

An application running on EC2 instances sends large amounts of data to an S3 bucket. The SysOps administrator wants to reduce data transfer costs while ensuring the traffic stays within AWS. What is the most cost-effective solution?

A.Set up an AWS Direct Connect connection.
B.Use S3 Transfer Acceleration.
C.Create a VPC Endpoint for S3 (Gateway type) and use it from the EC2 instances.
D.Route traffic through a NAT Gateway in a public subnet.
AnswerC

A Gateway VPC Endpoint for S3 allows EC2 instances in a private subnet to reach S3 without an internet gateway, NAT device, or public IP address. It is free of charge, has no data processing fees, and works by adding a service prefix list to the route table, which keeps S3 traffic on AWS's private network. This makes it the most direct and cost-effective solution for large data transfers from EC2.

Why this answer

A VPC Endpoint for S3 (Gateway type) allows EC2 instances to access S3 over the AWS private network without traversing the public internet, eliminating data transfer costs for traffic within the same region. Since the traffic stays within AWS, this is the most cost-effective solution as it avoids NAT Gateway, Direct Connect, or S3 Transfer Acceleration charges.

Exam trap

The trap here is that candidates often confuse Gateway VPC Endpoints with Interface Endpoints or assume that S3 Transfer Acceleration is cheaper for large data volumes, when in fact Gateway Endpoints are free and provide the most cost-effective private connectivity within a region.

How to eliminate wrong answers

Option A is wrong because AWS Direct Connect is a dedicated network connection from on-premises to AWS, which incurs monthly port fees and data transfer costs, and is not designed for traffic between EC2 and S3 within the same region. Option B is wrong because S3 Transfer Acceleration uses AWS edge locations and charges per GB transferred, increasing costs for large data transfers, and it still routes traffic over the public internet. Option D is wrong because routing traffic through a NAT Gateway in a public subnet incurs per-GB data processing charges and does not provide private connectivity to S3, as NAT Gateways are used for outbound internet access, not for optimized S3 access.

104
MCQmedium

A company has a VPC with multiple subnets. An EC2 instance in a public subnet needs to communicate with an RDS database in a private subnet. The RDS security group allows inbound traffic from the EC2 instance's security group. However, the EC2 instance cannot connect. What is the most likely cause?

A.The VPC does not have DNS resolution enabled, so the RDS endpoint cannot be resolved.
B.The network ACL for the private subnet blocks inbound traffic from the public subnet.
C.The security group of the RDS database does not allow outbound traffic.
D.The EC2 instance does not have a public IP address.
AnswerA

Amazon RDS exposes its database via a fully qualified domain name, such as `dbname.xxxxx.rds.amazonaws.com`, which the EC2 instance must resolve to an IP address. In a VPC, DNS resolution is governed by the `enableDnsSupport` attribute; if this is set to false, the VPC's Route 53 Resolver does not answer DNS queries. Without DNS resolution, the RDS endpoint cannot be translated into a private IP, so the connection attempt fails at the name resolution stage. This is the direct cause of the connectivity failure described.

Why this answer

The RDS database is in a private subnet and its endpoint is a DNS name. If DNS resolution is disabled on the VPC, the EC2 instance cannot resolve the RDS endpoint's hostname to an IP address, preventing the TCP connection from being established even though security group rules are correctly configured.

Exam trap

The trap here is that candidates focus on security group or NACL misconfigurations, overlooking that DNS resolution is a prerequisite for connecting to any service using a DNS endpoint, especially when the database is in a private subnet without a direct route to a public resolver.

How to eliminate wrong answers

Option B is wrong because network ACLs are stateless and, by default, allow all inbound and outbound traffic unless explicitly modified; the question does not indicate any custom NACL rules blocking traffic. Option C is wrong because security groups are stateful — if inbound traffic from the EC2 instance is allowed, the RDS database automatically allows outbound return traffic, so no explicit outbound rule is needed. Option D is wrong because the EC2 instance is in a public subnet and can have a public IP or use a NAT gateway to initiate outbound connections; the issue is DNS resolution, not the instance's public IP address.

105
Multi-Selecteasy

Which TWO actions should a SysOps administrator take to ensure high availability of a web application running on EC2 instances? (Choose two.)

Select 2 answers
A.Enable termination protection on all EC2 instances.
B.Launch all EC2 instances in a single Availability Zone.
C.Use a larger instance type for all EC2 instances.
D.Configure an Auto Scaling group with a health check to replace unhealthy instances.
E.Deploy EC2 instances across multiple Availability Zones.
AnswersD, E

An Auto Scaling group with health checks continuously monitors instance state and automatically terminates and replaces unhealthy instances, maintaining the desired capacity. This removes failed nodes from service, satisfying the availability requirement without manual intervention.

Why this answer

Option D is correct because an Auto Scaling group with health checks (EC2 status checks and optionally ELB health checks) automatically detects and replaces unhealthy instances, maintaining the desired capacity and thus high availability. Option E is correct because deploying EC2 instances across multiple Availability Zones ensures the application survives an AZ-level failure, since each AZ has independent power, cooling, and networking. Option A is incorrect because termination protection only prevents accidental instance termination; it does not improve availability or recover failed instances.

Option B is incorrect because placing all instances in a single AZ creates a single point of failure, reducing availability. Option C is incorrect because a larger instance type only increases capacity, not redundancy or fault tolerance.

Exam trap

The trap here is that candidates often confuse termination protection (a safety feature) with high availability, or think that larger instance types inherently provide fault tolerance, when in fact only redundancy across multiple Availability Zones and automated health-based replacement ensure high availability.

106
MCQeasy

A company wants to securely store secrets such as database credentials and API keys used by applications running on Amazon EC2. Which AWS service should be used to manage and rotate these secrets automatically?

A.AWS Identity and Access Management (IAM)
B.AWS Secrets Manager
C.AWS Key Management Service (KMS)
D.AWS Systems Manager Parameter Store
AnswerB

AWS Secrets Manager is a purpose-built service for securely storing and managing database credentials, API keys, and other secrets throughout their lifecycle. It natively supports automatic rotation, either through built-in integration with AWS services like RDS, Redshift, and DocumentDB, or via custom AWS Lambda rotations. Unlike generic parameter storage, Secrets Manager enforces fine-grained IAM access policies and provides audit trails via AWS CloudTrail, making it the recommended choice for production secrets that require rotation and regulated access.

Why this answer

AWS Secrets Manager is designed to securely store, manage, and automatically rotate secrets such as database credentials and API keys. It integrates with AWS services like RDS, Redshift, and DocumentDB to provide built-in rotation. This makes it the correct choice for managing and rotating secrets automatically.

Exam trap

SOA-C02 often tests the confusion between Secrets Manager and Parameter Store, leading candidates to choose Parameter Store for automatic rotation when it does not natively support it.

How to eliminate wrong answers

Option A is wrong because AWS Identity and Access Management (IAM) is for managing access to AWS resources, not for storing and rotating secrets. Option C is wrong because AWS Key Management Service (KMS) is for creating and managing encryption keys, not for storing application secrets; it can be used to encrypt secrets but does not provide secret management or rotation. Option D is wrong because AWS Systems Manager Parameter Store can store secrets, but it does not provide automatic rotation natively; you would need to implement custom rotation logic.

107
MCQeasy

Refer to the exhibit. An IAM policy allows a user to run instances only of type t2.micro. What happens when the user tries to run a t2.small instance?

A.The request is allowed because the policy allows ec2:RunInstances.
B.The request is denied because there is an explicit deny on ec2:RunInstances.
C.The request is allowed because the condition only applies to the resource ARN, not the instance type.
D.The request is denied because t2.small does not match the condition.
AnswerD

The correct outcome is an implicit deny: the request for a t2.small instance fails the policy statement's condition that instanceType equals t2.micro. In IAM, for a request to be allowed, an allow statement must match the action, resource, and any specified conditions; here the action matches, but the condition does not. Since no other statement provides a different allow, the default deny takes effect.

Why this answer

The IAM policy includes a condition that restricts ec2:RunInstances to the t2.micro instance type. When the user attempts to launch a t2.small instance, the condition evaluates to false, so the statement does not apply and the request is implicitly denied. IAM policies are deny-by-default, so the absence of an allow results in denial.

Exam trap

SOA-C02 often tests the misconception that a policy allowing an action (ec2:RunInstances) automatically permits all variations of that action, ignoring the effect of condition keys that narrow the scope.

How to eliminate wrong answers

Option A is wrong because the policy does not grant unconditional ec2:RunInstances — the condition narrows the allowed instance type. Option B is wrong because there is no explicit deny in the policy; the denial is implicit due to the condition not matching. Option C is wrong because the condition applies to the instance type via a condition key (e.g., ec2:InstanceType), not just the resource ARN, so it does restrict the instance type.

108
MCQhard

A company uses Amazon MQ (RabbitMQ) for messaging between microservices. The SysOps administrator needs to ensure the message broker is highly available with automatic failover and no data loss. Which deployment mode should be used?

A.Single-instance broker
B.Active/standby broker
C.Cluster deployment
D.Multi-AZ broker with read replicas
AnswerB

An active/standby broker deploys two broker instances in different Availability Zones, with synchronous replication of messages and metadata from the active broker to the standby. On failure of the active broker, automatic failover promotes the standby to active with minimal downtime, ensuring messages are not lost because they are already replicated synchronously. This is the correct Amazon MQ deployment mode for meeting high-availability and data-loss-prevention requirements.

Why this answer

Amazon MQ for RabbitMQ supports an active/standby deployment mode that provides automatic failover and no data loss. In this mode, one broker instance is active and a second is a synchronous standby; if the active fails, the standby takes over without losing messages because all data is replicated synchronously across both instances. This meets the high availability and data durability requirements specified in the question.

Exam trap

The trap here is that candidates confuse Amazon MQ's cluster deployment (which is for scaling) with active/standby (which is for high availability and data durability), or they incorrectly apply RDS Multi-AZ concepts to Amazon MQ.

How to eliminate wrong answers

Option A is wrong because a single-instance broker has no redundancy or automatic failover; if it fails, all messages are lost until manual recovery. Option C is wrong because RabbitMQ cluster deployment in Amazon MQ is designed for horizontal scaling and throughput, not for automatic failover with zero data loss; it uses asynchronous replication and can lose messages during a node failure. Option D is wrong because Amazon MQ does not support Multi-AZ brokers with read replicas; that concept applies to Amazon RDS, not to message brokers.

109
MCQmedium

A company is using Amazon CloudFront to distribute content globally. The origin is an S3 bucket. The SysOps administrator notices that cache hit ratio is low. Which configuration change would MOST improve the cache hit ratio?

A.Use query string parameters to differentiate content.
B.Configure custom error responses for 404 errors.
C.Set longer Cache-Control max-age headers on the S3 objects.
D.Enable Origin Shield for the distribution.
AnswerC

Setting a longer Cache-Control: max-age header on the S3 objects tells CloudFront how many seconds the object remains fresh in the edge cache. With a longer cache duration, an object is more likely to still be present and valid when subsequent requests arrive, so more requests are served directly from the edge without a round trip to the origin. This directly increases the CloudFront cache hit ratio, making it the optimal choice.

Why this answer

Cache hit ratio measures how often CloudFront serves content from edge caches instead of forwarding requests to the origin. The Cache-Control max-age header (and its s-maxage/Expires equivalents) directly controls how long CloudFront keeps an object in cache before revalidating with S3. Extending max-age on the S3 objects means each cached copy satisfies more subsequent requests, which raises the hit ratio without changing request patterns.

Exam trap

SOA-C02 often tests the misconception that adding more cache-key dimensions (query strings, headers, cookies) improves caching, when in fact every added dimension splits the cache and reduces the hit ratio.

How to eliminate wrong answers

Option A is wrong because adding query string parameters to the cache key fragments the cache — each unique query string creates a separate cached variant, which lowers the hit ratio rather than improving it. Option B is wrong because custom error responses only change what CloudFront returns to the viewer on a 4xx/5xx; they do not affect whether an object is served from cache. Option D is wrong because Origin Shield adds an additional caching layer between edge locations and the origin, reducing origin load and latency, but it does not increase the proportion of viewer requests served from cache — the TTL still governs that.

110
MCQhard

A company is using Amazon CloudFront to serve static content from an S3 bucket. They want to restrict access so that only CloudFront can access the S3 bucket. How should this be configured?

A.Configure Origin Access Control (OAC) with the S3 bucket policy.
B.Use CloudFront signed URLs or cookies.
C.Attach an IAM role to CloudFront that grants S3 read access.
D.Create a bucket policy that allows access only from the CloudFront distribution's IP addresses.
AnswerA

Origin Access Control (OAC) is the modern, recommended way to restrict an S3 bucket to serve content only through CloudFront. OAC uses a service principal of cloudfront.amazonaws.com with a condition that requires the distribution's ID to match, and the bucket policy grants only GetObject to that principal. This blocks direct S3 access from outside CloudFront while also supporting encrypted S3 objects and SSE-KMS, giving verifiable origin security.

Why this answer

Origin Access Control (OAC) is the recommended method to restrict access to an S3 bucket so that only CloudFront can retrieve objects. When OAC is enabled, CloudFront signs requests to S3 using a specific principal, and the S3 bucket policy is configured to allow access only to that principal. This prevents direct access to the bucket via S3 URLs or other AWS services, ensuring that content is served exclusively through CloudFront.

Exam trap

The trap here is that candidates often confuse viewer-side access control (signed URLs) with origin-side access control (OAC/OAI), or mistakenly think that CloudFront can use IAM roles or static IP addresses to authenticate to S3.

How to eliminate wrong answers

Option B is wrong because signed URLs or cookies control access to CloudFront content at the viewer level, not between CloudFront and the S3 origin; they do not restrict the S3 bucket from being accessed directly. Option C is wrong because CloudFront does not support attaching an IAM role directly to the distribution; IAM roles are used for AWS services like EC2 or Lambda, not for CloudFront-to-S3 authentication. Option D is wrong because CloudFront does not have a fixed set of IP addresses that can be used in a bucket policy; its IP addresses are dynamic and shared across distributions, making this approach unreliable and insecure.

111
MCQeasy

A SysOps administrator needs to audit all changes to IAM policies in an AWS account. Which AWS service should be used to record these changes?

A.Amazon CloudWatch Logs
B.AWS Config
C.AWS CloudTrail
D.Amazon S3
AnswerC

AWS CloudTrail is the authoritative service for auditing API activity in an AWS account, capturing every IAM API call such as PutRolePolicy, AttachUserPolicy, and CreatePolicy. Each CloudTrail event includes the IAM user or role, assumed role, session context, source IP, request parameters, response elements, and a timestamp. By default, CloudTrail provides a 90-day viewable event history, and creating a trail delivers immutable log files to an S3 bucket for long-term storage and optional CloudWatch Logs delivery. This makes CloudTrail the correct choice for auditing all IAM policy changes.

Why this answer

AWS CloudTrail is the correct service because it records API activity in an AWS account, including all IAM policy changes such as creating, updating, or deleting policies. CloudTrail captures these events as JSON logs, which can be stored in an S3 bucket for auditing and analysis. This makes it the appropriate tool for auditing changes to IAM policies.

Exam trap

The trap here is that candidates often confuse AWS Config with CloudTrail, thinking Config records API changes, but Config only tracks resource configuration states and compliance, not the API calls that caused those changes.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Logs is used for monitoring, storing, and accessing log files from AWS resources like EC2 instances or Lambda functions, but it does not natively record API calls or IAM policy changes; it requires CloudTrail to deliver logs to it. Option B is wrong because AWS Config is a service for evaluating resource configurations against desired policies and tracking configuration changes over time, but it does not record API-level events; it focuses on resource state rather than who made the change. Option D is wrong because Amazon S3 is an object storage service and cannot record or audit changes itself; it can only store logs delivered by other services like CloudTrail.

112
MCQmedium

A company is using AWS CloudFormation to manage infrastructure. They have a stack that creates an EC2 instance and an Elastic IP. The instance is in a VPC with an internet gateway. The stack creation succeeds, but the instance does not have internet connectivity. What is the most likely cause?

A.The subnet's route table does not have a route to the internet gateway.
B.The instance does not have a public IP address.
C.The instance is in a private subnet.
D.The security group does not allow outbound traffic.
AnswerA

Even when an instance has an Elastic IP and the VPC contains an internet gateway, the subnet's route table must include a default route (0.0.0.0/0) with the internet gateway as the target. Without that route, the instance cannot send traffic out to the internet because the IGW is the only mechanism that forwards VPC traffic to the outside world. In CloudFormation, if you create a custom route table but forget to add the IGW route or forget to associate the route table with the subnet, the subnet will use the VPC's main route table, which may not have the required route. This is the most direct and common cause of unreachable internet connectivity despite having a public IP.

Why this answer

For an EC2 instance in a VPC to reach the internet, three things are required: a public IP (or Elastic IP), an internet gateway attached to the VPC, and a route in the subnet's route table pointing 0.0.0.0/0 to that IGW. The question states the instance has an Elastic IP and the VPC has an IGW, so the missing piece is the route table entry. Without a 0.0.0.0/0 route to igw-xxxx, traffic from the instance has no path off the subnet regardless of the EIP.

Exam trap

SOA-C02 often tests the misconception that attaching an Internet Gateway to a VPC automatically gives subnets internet access — candidates forget that a 0.0.0.0/0 route in the subnet's route table is a separate, mandatory step.

How to eliminate wrong answers

Option B is wrong because the scenario explicitly states an Elastic IP is attached, which provides the public IP needed for internet-bound traffic. Option C is wrong because being in a 'private subnet' is defined by the absence of a route to an IGW — the question already establishes an IGW exists, so the real issue is the missing route, not the subnet classification. Option D is wrong because security groups are stateful and allow all outbound traffic by default; even if egress were restricted, the symptom would be connection timeouts rather than the complete lack of a routing path implied here.

113
MCQmedium

A company runs a critical application on Amazon EC2 instances with data stored on Amazon EBS volumes. The SysOps administrator needs to implement a backup strategy that supports point-in-time recovery with a Recovery Point Objective (RPO) of 1 hour and a Recovery Time Objective (RTO) of 4 hours. Which solution meets these requirements with the least operational overhead?

A.Use AWS Backup to schedule hourly EBS snapshots and restore to a new volume when needed.
B.Use Amazon Data Lifecycle Manager (DLM) to take hourly snapshots and create an AWS CloudFormation template to launch a new instance from the snapshot.
C.Use custom scripts to copy snapshots to an Amazon S3 bucket and restore from there.
D.Use Amazon S3 Lifecycle policies to transition data to Amazon S3 Glacier.
AnswerA

AWS Backup offers a fully managed, policy-based backup service that can create hourly EBS snapshots automatically. It provides centralized backup governance, retention management, and lifecycle policies, with the ability to restore a snapshot to a new EBS volume quickly. This minimizes RTO because restore is a native AWS operation and does not require custom scripting or additional orchestration.

Why this answer

AWS Backup provides a fully managed, policy-based backup service that can schedule EBS snapshots hourly, meeting the 1-hour RPO. Restoring from an AWS Backup snapshot to a new EBS volume and attaching it to an EC2 instance can be completed within the 4-hour RTO, with minimal operational overhead as it eliminates the need for custom scripts or lifecycle management.

Exam trap

The trap here is that candidates may choose DLM (Option B) because it can schedule snapshots, but they overlook the operational overhead of manually creating a CloudFormation template for recovery, whereas AWS Backup provides a fully managed restore workflow that meets the least operational overhead requirement.

How to eliminate wrong answers

Option B is wrong because Amazon Data Lifecycle Manager (DLM) can schedule hourly snapshots, but requiring a CloudFormation template to launch a new instance from the snapshot adds unnecessary operational overhead and complexity, whereas AWS Backup can directly restore the volume and instance. Option C is wrong because using custom scripts to copy snapshots to S3 introduces additional complexity, potential for errors, and does not leverage native AWS backup services, increasing operational overhead. Option D is wrong because Amazon S3 Lifecycle policies are designed for object lifecycle management in S3, not for EBS snapshots or point-in-time recovery of EC2 instances, and S3 Glacier is for archival, not rapid recovery with a 4-hour RTO.

114
MCQmedium

A SysOps administrator notices that an Amazon EC2 instance's CPU utilization is consistently above 90% during business hours. The instance is part of an Auto Scaling group with a simple scaling policy based on average CPU utilization. However, the Auto Scaling group is not launching new instances. What is the most likely cause?

A.The scaling policy is in a cooldown period after a previous scaling activity.
B.The Auto Scaling group has a minimum size equal to the current number of instances.
C.The Auto Scaling group has a scheduled scaling action that is overriding the dynamic policy.
D.The instance is not healthy and is being terminated by the Auto Scaling group.
AnswerA

The cooldown period is a timer that initiates after a simple scaling policy performs an activity, during which all subsequent scaling requests from that policy are ignored until the timer expires—by default 300 seconds. Because the group recently acted, the alarm-driven scale-out request is suppressed even though CPU utilization remains elevated, preventing a rapid series of changes while the newly launched instance passes standard health checks and the alarm evaluation period resets. Once the cooldown timer expires, the policy can immediately respond to any new breach of the CPU utilization threshold.

Why this answer

The simple scaling policy in Auto Scaling has a cooldown period (default 300 seconds) that prevents the group from launching or terminating instances immediately after a previous scaling activity. If the policy triggered a scale-out event recently, the cooldown period is still active, so even though CPU utilization remains above 90%, no new instances are launched until the cooldown expires. This is the most likely cause because the cooldown is designed to stabilize metrics and avoid thrashing.

Exam trap

The trap here is that candidates often assume high CPU utilization always triggers a scale-out immediately, forgetting that simple scaling policies enforce a cooldown period that can delay subsequent scaling actions, even when the metric remains elevated.

How to eliminate wrong answers

Option B is wrong because if the minimum size equals the current number of instances, the Auto Scaling group would still launch new instances to meet the desired capacity set by the scaling policy; the minimum size only prevents scaling below that number, not above it. Option C is wrong because a scheduled scaling action overrides dynamic policies only at the scheduled time, but it does not block the dynamic policy from acting during business hours unless the scheduled action explicitly sets the desired capacity to a value that prevents scaling. Option D is wrong because an unhealthy instance is terminated and replaced by the Auto Scaling group, which would launch a new instance, not block scaling; the group would still respond to high CPU utilization with a new launch.

115
MCQmedium

A SysOps administrator is tasked with encrypting data at rest for an Amazon S3 bucket that stores sensitive customer information. The company requires that the encryption keys be managed by AWS and rotated automatically. Which encryption solution meets these requirements?

A.Use client-side encryption with AWS KMS.
B.Use server-side encryption with customer-provided keys (SSE-C).
C.Use server-side encryption with Amazon S3-managed keys (SSE-S3).
D.Use server-side encryption with AWS KMS (SSE-KMS).
AnswerC

Server-side encryption with Amazon S3-managed keys (SSE-S3) is the simplest way to encrypt data at rest in S3: S3 automatically encrypts each object with a unique key that is itself wrapped by a root key, all managed by AWS. These keys are automatically rotated on a regular basis, so you have no key material to manage or rotate, and there is no additional cost. This directly meets the requirement of encrypting data at rest with AWS managing the keys, making it the correct answer.

Why this answer

SSE-S3 uses Amazon S3-managed keys that are automatically rotated by AWS, meeting the requirement for AWS-managed and automatic rotation. Option A is wrong because client-side encryption is not managed by AWS and does not use server-side encryption. Option B is wrong because SSE-C requires the customer to provide their own encryption keys, which are not automatically rotated.

Option D is wrong because SSE-KMS uses AWS KMS keys that are customer-managed unless automatic rotation is specifically enabled, and the question requires automatic rotation without additional configuration.

116
MCQmedium

A SysOps administrator notices that an EC2 instance's CPU utilization is consistently above 90% during business hours. The instance is part of an Auto Scaling group with a scaling policy based on average CPU utilization. Despite high utilization, no scaling events are triggered. What is the most likely cause?

A.The scaling policy has a cooldown period that is too long, preventing new scaling activities.
B.The instance type is not supported by the Auto Scaling group's launch configuration.
C.The CloudWatch alarm is in the ALARM state but the Auto Scaling group has a suspended process for Add instances.
D.The Auto Scaling group's health check type is set to ELB, causing the instance to be marked unhealthy.
AnswerA

The scaling policy's cooldown period intentionally suppresses scaling actions for a set duration after the previous activity. If that cooldown is excessively long, it overrides the high CPU metric by preventing new scaling operations until the cooldown expires, so the Auto Scaling group cannot add instances even though the alarm remains in ALARM. This directly explains why no new instances launch despite sustained high utilization.

Why this answer

The most likely cause is that the scaling policy has a cooldown period that is too long. After a scaling activity completes, the Auto Scaling group enters a cooldown period that prevents additional scaling activities from being triggered until the cooldown expires. If the cooldown period is set too long (e.g., 600 seconds or more), the group will not launch new instances even if the CloudWatch alarm remains in ALARM state with high CPU utilization, because the scaling policy is blocked from executing.

Exam trap

The trap here is that candidates often assume a scaling policy will always trigger when the CloudWatch alarm is in ALARM state, overlooking the cooldown period as a deliberate throttling mechanism that can prevent scaling activities from being initiated.

How to eliminate wrong answers

Option B is wrong because if the instance type were not supported by the launch configuration, the instance would fail to launch or would be in an impaired state, but the existing instance would still be running and scaling events would still be triggered (though they might fail). Option C is wrong because if the Add instances process were suspended, the Auto Scaling group would not launch new instances at all, but the question states that no scaling events are triggered, which implies the scaling policy itself is not executing; a suspended process would still allow the CloudWatch alarm to trigger a scaling event (which would then be blocked), but the event would still appear in the scaling activity history. Option D is wrong because the health check type being set to ELB would cause the instance to be marked unhealthy only if the ELB health checks fail, but high CPU utilization alone does not cause an ELB health check failure; the instance would still be considered healthy and scaling events would still be triggered based on the CloudWatch alarm.

117
Drag & Dropmedium

Drag and drop the steps to troubleshoot an unhealthy target in an Application Load Balancer target group into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Troubleshooting starts with security group rules, then health check configuration, then instance and application status, then logs, and finally replacement if needed.

118
Multi-Selectmedium

A company wants to automatically remediate an Amazon EC2 instance that becomes unresponsive by rebooting it. The solution should use AWS managed services to minimize custom code. Which combination should a SysOps administrator use? (Choose TWO.)

Select 2 answers
A.Amazon EC2 Auto Scaling and lifecycle hooks
B.Amazon CloudWatch alarm on EC2 status check failures
C.Amazon CloudWatch alarm and AWS Lambda function
D.Amazon EventBridge rule to trigger an SNS notification
E.AWS Systems Manager Automation document to reboot the instance
AnswersB, E

A CloudWatch alarm that watches the StatusCheckFailed (or StatusCheckFailed_System/Instance) metric detects when the EC2 service signals that the instance is unreachable or the OS is unresponsive. You can set the alarm to invoke the 'Reboot' EC2 action directly, requiring no custom code or additional infrastructure. Because the action is native, it is the most straightforward managed remedy for status-check-only failures.

Why this answer

Amazon CloudWatch can monitor EC2 status check failures (both system and instance checks) and trigger an alarm. When the alarm enters the ALARM state, it can directly invoke an AWS Systems Manager Automation document to reboot the instance, which is a managed, code-free remediation approach. This combination minimizes custom code by using built-in AWS services.

Exam trap

The trap here is that candidates often choose Option C (CloudWatch alarm + Lambda) because it is a common pattern, but the question explicitly requires minimizing custom code, making the managed Systems Manager Automation document the correct choice over a custom Lambda function.

119
MCQmedium

A company is using AWS CloudFormation to deploy a stack that includes an Amazon RDS DB instance. The database password is stored in AWS Secrets Manager. The CloudFormation template references the secret using a dynamic reference. However, the stack creation fails with an error that the secret cannot be retrieved. What is the most likely cause?

A.The secret is in a different AWS Region.
B.The stack name does not match the secret name.
C.The template uses the wrong dynamic reference syntax.
D.The CloudFormation service role lacks permissions to read the secret.
AnswerD

When a CloudFormation stack operation uses a service role, all AWS API calls, including reading secrets, are made using that role's credentials. The role must be granted the 'secretsmanager:GetSecretValue' action and, if the secret uses a customer-managed KMS key, the 'kms:Decrypt' permission as well. Without these permissions, CloudFormation is denied access to the secret and returns an error that the secret cannot be retrieved, even though the secret exists and the reference syntax is valid.

Why this answer

To use dynamic references, the CloudFormation service role must have permission to read the secret. The stack name and parameters are not related to secret retrieval. The secret must be in the same region.

The template syntax might be incorrect, but the most common issue is missing permissions.

120
Multi-Selectmedium

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application stores data on an Amazon EBS volume. The SysOps administrator needs to implement a backup strategy that ensures data can be recovered in the event of an AZ failure. Which TWO actions should be taken? (Choose TWO.)

Select 2 answers
A.Increase the EBS volume size to maximize I/O performance.
B.Configure automated snapshots using Amazon Data Lifecycle Manager.
C.Create a lifecycle policy to automatically take snapshots of the EBS volume and copy them to another region.
D.Enable EBS encryption using AWS KMS.
E.Enable EBS Multi-Attach to allow the volume to be attached to instances in another AZ.
AnswersB, C

Amazon Data Lifecycle Manager (DLM) is the native service for automating EBS snapshot creation, retention, and deletion according to a policy-defined schedule. By configuring a DLM policy, you ensure consistent point-in-time backups of the EBS volume without manual intervention, and you can set retention rules to age out old snapshots to control costs. This directly satisfies the backup requirement for a critical application and is the correct answer to the scenario.

Why this answer

Amazon Data Lifecycle Manager (DLM) automates the creation, retention, and deletion of EBS snapshots, providing a scheduled backup mechanism that protects against data loss. Option C is correct because copying snapshots to another region ensures data is recoverable even if an entire AWS Availability Zone (AZ) fails, as the snapshots are stored independently in a different geographic region.

Exam trap

The trap here is that candidates often confuse EBS Multi-Attach (which provides high availability within an AZ) with a cross-AZ backup strategy, or they mistakenly think increasing volume size or enabling encryption alone constitutes a backup plan.

121
Multi-Selecteasy

Which TWO of the following are features of Amazon Route 53? (Select TWO.)

Select 2 answers
A.SSL/TLS termination
B.Health checking of resources
C.SSL certificate management
D.Domain name registration
E.Content caching at edge locations
AnswersB, D

Route 53 supports health checking of endpoints by actively sending periodic TCP, HTTP, or HTTPS requests to configured IPs or domains to verify availability. These health checks integrate with failover routing policies, allowing Route 53 to automatically remove unhealthy resources from DNS responses and direct traffic to healthy ones. This is a core feature that extends beyond basic name resolution.

Why this answer

Amazon Route 53 is a DNS web service that provides domain name resolution, domain registration, and health checking of resources. Health checking monitors the availability and performance of endpoints (e.g., web servers) via HTTP/HTTPS/TCP requests, and can automatically failover DNS responses to healthy resources, ensuring high availability.

Exam trap

The trap here is that candidates confuse Route 53's DNS-level health checking with application-layer features like SSL termination or caching, leading them to select options that belong to other AWS services like CloudFront or ALB.

122
MCQhard

A SysOps administrator reviews the CloudWatch metric data for an EC2 instance. The instance runs a web application that experiences high traffic between 12:00 and 14:00 UTC daily. The administrator wants to optimize costs while maintaining performance. What should the administrator do?

A.Convert the instance to a Reserved Instance to reduce hourly cost.
B.Replace the instance with a larger instance type and enable detailed monitoring.
C.Create an Auto Scaling group with a scheduled scaling policy to add instances during peak hours.
D.Increase the instance size to handle peak load at all times.
AnswerC

An Auto Scaling group with a scheduled scaling policy adds EC2 instances before the peak begins and removes them after it ends, aligning running capacity with the predicted demand. This is the correct solution because it handles the peak utilization while avoiding the cost of keeping that extra capacity running all day. You can define the schedule using a cron expression, and the group can also be configured with dynamic policies to absorb unexpected spikes beyond the scheduled capacity.

Why this answer

The instance has low CPU utilization most of the day but spikes to 90% during peak hours. Using a scheduled Auto Scaling to add instances during peak hours ensures performance without over-provisioning. Option A is wrong because the instance is not constantly at high utilization.

Option B is wrong because upgrading instance size increases cost during off-peak hours. Option D is wrong because a larger instance would be underutilized.

123
MCQeasy

An administrator deploys a CloudFormation template that includes the snippet shown in the exhibit. Later, the administrator deletes the stack. What happens to the S3 bucket?

A.The bucket is deleted only if it contains no objects
B.The bucket is emptied and then deleted
C.The bucket is deleted along with the stack
D.The bucket is retained but no longer managed by CloudFormation
AnswerD

When a CloudFormation stack is deleted, resources with a DeletionPolicy of Retain are left in place in the AWS account. CloudFormation disassociates the resource from the stack, meaning it no longer tracks or manages the bucket through stack operations. The bucket remains fully functional and accessible, but subsequent stack updates or deletions will not affect it.

Why this answer

The DeletionPolicy attribute set to 'Retain' on the S3 bucket resource causes CloudFormation to preserve the bucket when the stack is deleted. The bucket will still exist in the account but will no longer be under CloudFormation management. Therefore, option D is correct.

Option A is incorrect because the bucket is retained regardless of its contents. Option B is incorrect because the bucket is not emptied or deleted. Option C is incorrect because the bucket is not deleted along with the stack.

124
MCQhard

A company has a web application behind an Application Load Balancer (ALB) with sticky sessions enabled. The ALB's target group contains EC2 instances in an Auto Scaling group. After a deployment, users report that they are being logged out frequently. What is the most likely cause?

A.The deregistration delay is set too low.
B.The ALB's stickiness cookie is not configured or is being overwritten.
C.Health checks are too frequent and marking instances unhealthy.
D.Cross-zone load balancing is disabled.
AnswerB

Sticky sessions on an Application Load Balancer rely on the AWSALB cookie being issued and honored by the client. If the stickiness policy is not enabled on the target group, or the application overwrites or strips the Set-Cookie header, the ALB receives no cookie and treats each request as a new session, routing it to any available target. This directly causes the observed behavior of users losing their session.

Why this answer

Sticky sessions (session affinity) rely on a cookie generated by the ALB (AWSALB) to route subsequent requests from a user to the same target. If the cookie is not configured or is overwritten (e.g., by the application after deployment), the ALB treats each request as new and may route to different instances, causing the user to be logged out. Option A is incorrect because deregistration delay controls how long the ALB waits before deregistering an instance, which does not affect existing sessions unless instances are being removed.

Option C is incorrect because while frequent health checks may cause instances to be marked unhealthy, this would result in connection errors or routing to other instances, but it does not explain logouts due to sticky session loss. Option D is incorrect because cross-zone load balancing distributes traffic across zones but does not affect session stickiness.

125
Multi-Selecthard

Which THREE AWS services can be used to improve security and performance for a web application that uses an Application Load Balancer? (Select three.)

Select 3 answers
A.AWS Shield Advanced
B.AWS WAF
C.Amazon Route 53
D.Amazon CloudFront
E.AWS Direct Connect
AnswersA, B, D

AWS Shield Advanced provides always-on network and transport layer DDoS protection with automatic inline mitigations. It offers enhanced detection for sophisticated attacks, access to the AWS DDoS Response Team (DRT), and cost protection against scaling charges. It integrates with CloudFront, Application Load Balancer, and Elastic Load Balancing to monitor traffic patterns and mitigate large-scale volumetric attacks.

Why this answer

AWS Shield Advanced provides enhanced protection against Distributed Denial of Service (DDoS) attacks, including application-layer attacks targeting the Application Load Balancer (ALB). It integrates directly with ALB to offer always-on detection and automatic mitigation, improving security without requiring changes to the application architecture.

Exam trap

The trap here is that candidates often confuse Amazon Route 53's DNS routing features (like latency-based routing or geolocation) with performance improvement for the application itself, but Route 53 does not cache content or accelerate traffic; it only resolves DNS queries, which is a separate concern from web application performance.

126
MCQhard

Instances in a private subnet need outbound internet access for software updates. The route table sends 0.0.0.0/0 to a NAT gateway, but updates fail. Which condition should you check first?

A.Confirm the NAT gateway is in a public subnet whose route table has 0.0.0.0/0 to an internet gateway.
B.Attach an internet gateway directly to the private subnet instances.
C.Replace all security groups with network ACLs.
D.Enable VPC peering to another account.
AnswerA

A NAT gateway only provides outbound internet access when placed in a public subnet whose route table points 0.0.0.0/0 to an internet gateway. If it sits in a private subnet, traffic never reaches the internet, so updates fail.

Why this answer

A NAT gateway must reside in a public subnet with a route table entry directing 0.0.0.0/0 to an internet gateway (IGW). Without this, the NAT gateway cannot translate private IPs to the IGW's public IP, so outbound traffic from private instances fails. This is the most common root cause for failed internet access through a NAT gateway.

Exam trap

The trap here is that candidates assume any subnet with a NAT gateway automatically has internet access, overlooking the requirement that the NAT gateway itself must be in a public subnet with a default route to an internet gateway.

How to eliminate wrong answers

Option B is wrong because attaching an internet gateway directly to a private subnet is not supported; an IGW can only be attached to a VPC and associated with public subnets, and private subnet instances lack public IPs to use it directly. Option C is wrong because replacing security groups with network ACLs does not solve the routing issue; NACLs are stateless and can filter traffic, but they do not provide internet connectivity. Option D is wrong because VPC peering does not provide internet access; it only enables private connectivity between VPCs, and does not route traffic to the internet.

127
MCQmedium

A company uses AWS CodeDeploy to deploy applications to an Auto Scaling group. The deployment fails with the error: 'The overall deployment failed because too many individual instances failed deployment, too few healthy instances are available for deployment, or some instances in your deployment group are experiencing problems.' The SysOps administrator checks the deployment logs and finds that the application installation script exits with a non-zero exit code. What is the MOST likely cause?

A.The Auto Scaling group does not have enough instances to meet the minimum capacity.
B.The security group for the instances blocks outbound traffic to CodeDeploy endpoints.
C.The AppSpec file contains a lifecycle hook that fails.
D.The CodeDeploy agent is outdated on the instances.
AnswerC

A lifecycle hook in the AppSpec file that exits non-zero aborts the deployment on each instance, producing the reported failure threshold error. CodeDeploy surfaces the script's non-zero exit code directly, confirming the hook itself is the failing installation step.

Why this answer

A non-zero exit code from an AppSpec lifecycle hook (e.g., ApplicationStop, BeforeInstall, AfterInstall, ApplicationStart, ValidateService) during the deployment process causes the overall deployment to fail. The error message indicates that individual instances failed deployment, and the installation script exiting with a non-zero exit code is a direct sign of a lifecycle hook failure. Option A is incorrect because insufficient instances in the Auto Scaling group would trigger a different error related to minimum capacity, not a script exit code.

Option B is incorrect because the security group blocking outbound traffic would prevent the CodeDeploy agent from communicating with the service, resulting in a connection error, not a script exit code issue. Option D is incorrect because an outdated CodeDeploy agent would typically produce agent-specific errors or version mismatch warnings, not a non-zero exit code from the installation script.

128
MCQmedium

A company runs a stateless web application on Amazon EC2 instances in an Auto Scaling group with a minimum of 2 and maximum of 10 instances. The instances are behind an Application Load Balancer (ALB). The SysOps administrator needs to ensure that the application can survive the failure of an entire AWS Availability Zone (AZ) in the region. Which configuration is necessary?

A.Configure the Auto Scaling group with subnets in at least two Availability Zones and ensure the ALB has subnets in the same AZs.
B.Increase the Auto Scaling group minimum to 10 instances to absorb the failure.
C.Use larger instance types to handle the load of a failed AZ.
D.Use multiple Application Load Balancers in different AZs.
AnswerA

Correctly designed for failure domain isolation: an Auto Scaling group spanning subnets in at least two Availability Zones (AZs) lets EC2 instances be provisioned across independent infrastructure, and an ALB with subnets in those same AZs can route traffic to healthy instances in any AZ. If one AZ becomes unavailable, the ALB continues distributing requests to instances in the remaining AZs, while Auto Scaling replaces failed instances in the other AZs. This provides high availability because the application is stateless and can serve all traffic from a single AZ when needed.

Why this answer

Deploying the Auto Scaling group across multiple Availability Zones (AZs) and ensuring the ALB has subnets in the same AZs allows the application to continue serving traffic even if one entire AZ fails. The ALB can route requests to healthy instances in the remaining AZs, and the Auto Scaling group will replace failed instances in other AZs as needed, maintaining the minimum instance count. This architecture is a fundamental pattern for high availability in AWS.

Exam trap

The trap here is that candidates often think increasing instance count or size alone provides high availability, but without multi-AZ distribution, a single AZ failure can still cause total application downtime.

How to eliminate wrong answers

Option B is wrong because simply increasing the minimum to 10 instances does not provide AZ resilience; all instances could still be in a single AZ, and a failure of that AZ would take down all 10 instances. Option C is wrong because using larger instance types only increases compute capacity per instance, but does not distribute instances across AZs; a single AZ failure would still eliminate all instances if they are all in that AZ. Option D is wrong because using multiple ALBs in different AZs is unnecessary and adds complexity; a single ALB can already distribute traffic across multiple AZs, and multiple ALBs would require additional DNS routing logic (e.g., Route 53) and do not inherently improve AZ failure survival.

129
Multi-Selectmedium

Which TWO actions should a SysOps administrator take to improve the availability and reduce latency for a web application hosted on EC2 instances behind an Application Load Balancer?

Select 2 answers
A.Use larger EC2 instance types to handle more traffic.
B.Configure the ALB health check to have a shorter interval.
C.Use an Amazon CloudFront distribution in front of the ALB to cache content at edge locations.
D.Implement Auto Scaling to add instances based on CPU utilization.
E.Deploy EC2 instances in multiple Availability Zones.
AnswersD, E

Implementing Auto Scaling with a CPU utilization-based policy enables the fleet to add instances during demand spikes and remove them when load drops, which distributes the request load across more targets and reduces response time. This horizontal elasticity directly addresses latency and performance constraints by expanding capacity in an automated and predictable way. Combined with multi-AZ deployment, it also provides a self-healing, resilient architecture, but its primary benefit here is maintaining performance under changing load.

Why this answer

Auto Scaling based on CPU utilization dynamically adjusts the number of EC2 instances to match demand, improving availability by ensuring sufficient capacity during traffic spikes and reducing latency by distributing load across more instances. Option E is correct because deploying EC2 instances in multiple Availability Zones (AZs) provides fault tolerance: if one AZ fails, the ALB continues routing traffic to healthy instances in other AZs, which also reduces latency by serving users from the closest AZ.

Exam trap

The trap here is that candidates often confuse vertical scaling (larger instances) with horizontal scaling (more instances across AZs), or they mistakenly think that reducing health check intervals always improves availability, when in fact it can cause flapping and reduce stability.

130
MCQhard

A company is running a stateful web application on a single EC2 instance in a public subnet. The instance stores user sessions locally. The company wants to improve availability without rewriting the application. Which design should they use?

A.Create a second EC2 instance in a different AZ and use Route 53 with health checks.
B.Use an Auto Scaling group across multiple AZs but keep sessions on instance.
C.Deploy an Application Load Balancer across multiple AZs, move session storage to ElastiCache, and use an Auto Scaling group.
D.Use an Application Load Balancer with sticky sessions and an Auto Scaling group in a single AZ.
AnswerC

This is correct because an ALB in multiple AZs distributes traffic across healthy instances, while ElastiCache stores session data independently of any individual EC2 instance, making the app stateless. If an instance fails or is terminated by the ASG, other instances can immediately serve users because their sessions are still in ElastiCache. The ASG handles capacity and automatically replaces unhealthy instances based on ALB health checks, giving both elasticity and high availability.

Why this answer

It addresses the core issue of stateful sessions without rewriting the application. By moving session storage to ElastiCache (a centralized, external data store), the application becomes stateless from the instance's perspective, allowing an Auto Scaling group across multiple Availability Zones (AZs) and an Application Load Balancer (ALB) to distribute traffic seamlessly. This design improves availability by enabling horizontal scaling and fault tolerance, as any instance can handle any request since sessions are stored externally.

Exam trap

The trap here is that candidates often assume sticky sessions (session affinity) alone solve the stateful application problem, but they fail to recognize that sticky sessions do not protect against instance failure or AZ outages, and they still require local session storage, which is lost when an instance is replaced.

How to eliminate wrong answers

Option A is wrong because simply adding a second EC2 instance in a different AZ with Route 53 health checks does not solve the session state problem; user sessions stored locally on the original instance would be lost if traffic fails over to the new instance, breaking the stateful application. Option B is wrong because keeping sessions on the instance while using an Auto Scaling group across multiple AZs means that if an instance is terminated or replaced, all local session data is lost, and new instances cannot serve existing sessions, leading to user disruption. Option D is wrong because using an ALB with sticky sessions and an Auto Scaling group in a single AZ still creates a single point of failure at the AZ level; if that AZ goes down, the entire application becomes unavailable, and sticky sessions alone do not persist session data across instance replacements.

131
Multi-Selecteasy

A company uses CloudWatch Logs to monitor application logs. The SysOps administrator wants to search for specific error patterns across multiple log groups. Which THREE AWS services can be used to achieve this?

Select 3 answers
A.CloudWatch Logs Insights
B.Amazon OpenSearch Service
C.Amazon Kinesis Data Analytics
D.Amazon Athena
E.AWS Glue
AnswersA, B, D

CloudWatch Logs Insights is the native, serverless query engine for log data already in CloudWatch Logs. It uses a purpose-built query language with commands like fields, stats, filter, parse, and sort to run interactive, ad-hoc queries across one or multiple log groups in the same AWS account and Region. Because it operates directly on the log data without requiring any export or additional infrastructure, it is the most direct and cost-effective way to query application logs already collected by CloudWatch Logs.

Why this answer

CloudWatch Logs Insights is correct because it is a native AWS service designed specifically for querying and analyzing log data stored in CloudWatch Logs. It allows you to run SQL-like queries (using a query language) across multiple log groups to search for specific error patterns, making it a direct and efficient solution for this use case without requiring data export or additional infrastructure.

Exam trap

The trap here is that candidates may overlook Amazon OpenSearch Service and Amazon Athena as valid options because they require additional configuration (streaming or exporting logs), but the question asks which services 'can be used' to achieve the goal, not which are the most direct or native, so all three (A, B, D) are technically feasible.

132
MCQmedium

A company runs a batch processing job on Amazon EMR every night. The job runs for 6 hours and requires a cluster of 20 m5.xlarge instances. The company wants to reduce costs while ensuring the job completes on time. Which solution is MOST cost-effective?

A.Use On-Demand instances for all nodes.
B.Purchase Reserved Instances for the entire cluster.
C.Use Spot Instances for core and task nodes and an On-Demand instance for the primary node.
D.Use Spot Instances for the primary node and On-Demand for core and task nodes.
AnswerC

This is the AWS-recommended cost-optimization pattern for transient EMR clusters: the primary node runs On-Demand to guarantee stable HDFS NameNode and ResourceManager availability, while core and task nodes run Spot Instances to exploit steep discounts (often 50–90% off On-Demand). Spot interruptions on core and task nodes are recoverable by EMR's instance-group resizing and task-node retries, but losing the primary node would fail the entire cluster. This mix preserves durability and responsiveness for the critical coordinator while dramatically lowering compute cost for the bulk of the cluster.

Why this answer

Amazon EMR clusters consist of a primary node (master), core nodes (which run HDFS and task processes), and task nodes (which only run tasks and can be lost without data loss). Spot Instances are ideal for task nodes because they can be interrupted without affecting HDFS data, and for core nodes if the cluster is resilient to interruptions (e.g., using EMRFS consistent view or if the job can tolerate some core node loss). The primary node must be On-Demand to avoid cluster termination if the Spot Instance is reclaimed.

This mix minimizes cost while ensuring the job completes on time.

Exam trap

SOA-C02 often tests the misconception that Spot Instances can be used for all node types, but the primary node must be On-Demand to avoid cluster termination, and core nodes may risk data loss if not using EMRFS.

How to eliminate wrong answers

Option A is wrong because On-Demand instances for all nodes are significantly more expensive than Spot, and the job can tolerate interruptions on task and core nodes. Option B is wrong because Reserved Instances require a 1- or 3-year commitment, which is not cost-effective for a nightly 6-hour job and does not provide the same savings as Spot for short-term, fault-tolerant workloads. Option D is wrong because using Spot for the primary node risks cluster termination if the Spot Instance is reclaimed, which would cause the job to fail and not complete on time.

133
MCQmedium

A team of developers is deploying a new microservice that uses Amazon DynamoDB as its data store. The SysOps administrator must ensure that the application can handle a sudden spike in read traffic without throttling. Which DynamoDB feature can be used to automatically handle increases in read capacity?

A.DynamoDB Global Tables
B.DynamoDB Time to Live (TTL)
C.DynamoDB Auto Scaling
D.DynamoDB Accelerator (DAX)
AnswerC

DynamoDB Auto Scaling is the correct solution because it continuously adjusts a table's provisioned read and write capacity in response to actual traffic using the Application Auto Scaling target-tracking policy. By monitoring CloudWatch metrics like consumed capacity, it scales out before throttling occurs and scales in when traffic declines, all within user-defined minimum and maximum limits. This gives the microservice a fully automated way to handle variable load without manual capacity planning.

Why this answer

DynamoDB Auto Scaling is the correct feature because it automatically adjusts the provisioned read and write capacity based on actual traffic patterns, preventing throttling during sudden spikes. The SysOps administrator can define a target utilization percentage (e.g., 70%), and DynamoDB Auto Scaling uses the Application Auto Scaling service to increase capacity before requests are throttled, ensuring consistent performance without manual intervention.

Exam trap

The trap here is that candidates often confuse DynamoDB Accelerator (DAX) with a scaling solution, thinking its caching automatically handles capacity increases, but DAX only reduces read load and does not modify provisioned capacity or prevent throttling from a sudden spike in uncached reads.

How to eliminate wrong answers

Option A is wrong because DynamoDB Global Tables provide multi-region replication for disaster recovery and low-latency reads, but they do not automatically handle sudden read traffic spikes within a single region; they require provisioned capacity management per replica. Option B is wrong because DynamoDB Time to Live (TTL) automatically deletes expired items to reduce storage costs, but it has no effect on read capacity or throttling prevention. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read latency for repeated queries, but it does not automatically increase provisioned read capacity; it only reduces the load on the underlying table by caching frequently accessed data.

134
Multi-Selecteasy

A SysOps administrator needs to monitor the disk space usage on an EC2 instance running Windows Server. Which actions are required to collect this metric? (Select TWO.)

Select 2 answers
A.Install the CloudWatch Logs agent to monitor disk usage logs.
B.Enable EC2 status checks to monitor disk health.
C.Use Windows Performance Monitor to track disk space and send to CloudWatch.
D.Create an IAM role with permissions to publish custom metrics and attach it to the instance.
E.Install the CloudWatch agent on the instance and configure it to collect disk metrics.
AnswersD, E

Custom metrics require authorisation to call the CloudWatch PutMetricData API. Attaching an IAM role with cloudwatch:PutMetricData to the instance supplies temporary credentials via instance metadata, satisfying the permissions constraint so the agent can publish disk space data without hard-coded keys.

Why this answer

Option E is correct because the CloudWatch agent is the AWS-supported tool that runs on a Windows Server EC2 instance and can be configured (via the wizard or a JSON config file) to collect disk metrics such as LogicalDisk % Free Space and publish them to CloudWatch. Option D is correct because the agent needs AWS credentials to call cloudwatch:PutMetricData, and attaching an IAM role with permissions to publish custom metrics to the instance provides those credentials without hardcoding keys. Option A is wrong because the CloudWatch Logs agent ships log files, not disk-space metrics, so it cannot collect this metric.

Option B is wrong because EC2 status checks report instance and system reachability (hypervisor and network level), not guest OS disk space. Option C is wrong because Windows Performance Monitor alone does not send data to CloudWatch; you still need the CloudWatch agent to publish those counters as metrics.

Exam trap

The trap here is that candidates often confuse the CloudWatch Logs agent with the CloudWatch agent, or assume that EC2 status checks or Performance Monitor can directly send disk metrics to CloudWatch without additional configuration.

135
Multi-Selectmedium

A SysOps administrator is setting up monitoring for an RDS MySQL database. The administrator needs to be notified when the database connection count exceeds 100. Which steps should be taken to achieve this? (Choose TWO.)

Select 2 answers
A.Configure the CloudWatch alarm to send a notification to an SNS topic.
B.Create a CloudWatch alarm on the 'DatabaseConnections' metric.
C.Enable Enhanced Monitoring for RDS.
D.Enable CloudTrail to log RDS API calls.
E.Create an Amazon EventBridge rule that triggers on RDS events.
AnswersA, B

A CloudWatch alarm alone does nothing until you attach an action; publishing to an SNS topic is the standard action that enables out-of-band notification. When the alarm enters ALARM state, it sends a message to the SNS topic, which then delivers via email, SMS, or Lambda. This is the notification layer of the solution, not the metric collection itself.

Why this answer

Amazon CloudWatch alarms can send notifications to an Amazon SNS topic when the alarm state changes. This allows the SysOps administrator to receive alerts (e.g., via email, SMS, or HTTP) when the database connection count exceeds the threshold. Option B is correct because the 'DatabaseConnections' metric is a standard CloudWatch metric for RDS MySQL that tracks the number of current connections to the database instance.

Creating an alarm on this metric with a threshold of 100 will trigger when the connection count exceeds that value.

Exam trap

The trap here is that candidates often confuse Enhanced Monitoring (which provides OS-level metrics) with CloudWatch metrics (which provide database-level metrics like connection counts), leading them to incorrectly select Enhanced Monitoring as a solution for connection-based alarms.

136
MCQhard

A SysOps administrator needs to monitor Amazon EC2 instances for disk space usage. Disk space metrics are not available by default in Amazon CloudWatch. The administrator wants to collect disk space metrics from all EC2 instances across multiple AWS accounts and aggregate them in a single CloudWatch dashboard. Which combination of steps should the administrator take?

A.Install the CloudWatch agent on each instance using SSM Run Command, configure the agent to collect disk metrics, and use CloudWatch cross-account observability to aggregate metrics from multiple accounts.
B.Enable detailed monitoring on the EC2 instances, create a custom metric in CloudWatch for disk space, and use CloudWatch Logs to forward logs to a central account.
C.Use AWS Config to track disk space and send metrics to CloudWatch.
D.Use AWS Trusted Advisor to monitor disk space and send alerts via SNS.
AnswerA

The CloudWatch agent is the only way to collect OS-level disk space metrics from EC2 because standard EC2 metrics (CPU, network) are hypervisor-level and do not include disk utilization. SSM Run Command can automate deployment of the agent at scale using the AmazonCloudWatch-ManageAgent document or a package installation document. Once the agent publishes custom metrics like disk_used_percent to CloudWatch, cross-account observability (via CloudWatch Observability Access Manager) lets a central monitoring account aggregate and view metrics from all member accounts, enabling fleet-wide dashboards and alarms.

Why this answer

The CloudWatch agent is required to collect custom metrics like disk space from EC2 instances, as these are not available by default. SSM Run Command enables scalable, agentless installation across instances. CloudWatch cross-account observability allows you to aggregate metrics from multiple AWS accounts into a single monitoring account, meeting the requirement for a unified dashboard.

Exam trap

The trap here is that candidates assume detailed monitoring or AWS Config can capture OS-level metrics like disk space, when in fact only an in-guest agent (CloudWatch agent) can collect such data, and cross-account aggregation requires a specific feature (cross-account observability) rather than simple log forwarding.

How to eliminate wrong answers

Option B is wrong because enabling detailed monitoring only provides hypervisor-level metrics (CPU, network, etc.) at 1-minute frequency, not disk space metrics; disk space requires an in-guest agent. Option C is wrong because AWS Config tracks resource configuration changes (e.g., instance type, security groups) and can trigger rules, but it does not collect or emit disk space metrics to CloudWatch. Option D is wrong because AWS Trusted Advisor provides best-practice checks and recommendations, but it does not collect real-time disk space metrics or send them to CloudWatch for dashboard aggregation.

137
MCQeasy

A SysOps administrator needs to monitor the CPU utilization of an Amazon EC2 instance and send an alert when it exceeds 90% for 5 consecutive minutes. Which combination of AWS services should the administrator use to meet this requirement?

A.Amazon CloudWatch metric (CPUUtilization), a CloudWatch alarm, and an Amazon SNS topic.
B.Amazon CloudWatch Logs, a metric filter to extract CPU utilization from logs, and an alarm on that metric.
C.A CloudWatch dashboard and an AWS Lambda function that checks the dashboard periodically.
D.Amazon EventBridge (CloudWatch Events) and a Lambda function that calls the EC2 DescribeInstances API.
AnswerA

EC2 publishes a standard hypervisor-level CPUUtilization metric to CloudWatch every 5 minutes (or 1 minute with detailed monitoring). A CloudWatch alarm can evaluate that metric against a threshold using a period, statistic (e.g., Average), and evaluation periods, then transition to ALARM state and publish a message to an SNS topic, which can fan out to email, SMS, or HTTP endpoints. This is the native, least-effort, and most reliable pattern for triggering on CPU utilization; it requires no custom code, log filtering, or polling.

Why this answer

The correct approach is to use a CloudWatch metric for CPUUtilization, which is automatically published by EC2 instances. A CloudWatch alarm can be configured to evaluate this metric over a period of 5 consecutive minutes with a threshold of 90%, and when the alarm state is triggered, it publishes to an SNS topic to send notifications. This is the native, efficient, and recommended method for monitoring and alerting on EC2 CPU utilization.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs metric filters (used for custom log-based metrics) with the built-in EC2 metrics, or think that EventBridge can directly access CPU utilization data, when in fact CPUUtilization is a CloudWatch metric and must be monitored via CloudWatch alarms.

How to eliminate wrong answers

Option B is wrong because CloudWatch Logs and metric filters are used to extract custom metrics from log data (e.g., application logs), not to monitor the built-in CPUUtilization metric which is already available as a CloudWatch metric without needing log extraction. Option C is wrong because a CloudWatch dashboard is a visualization tool and does not trigger alerts; a Lambda function polling a dashboard periodically is inefficient, introduces latency, and is not a supported pattern for real-time alerting. Option D is wrong because EventBridge and a Lambda function calling DescribeInstances API only retrieves instance metadata and state, not CPU utilization metrics; CPU utilization is a CloudWatch metric, not available via the EC2 DescribeInstances API.

138
MCQeasy

A company runs a web application on a fleet of Amazon EC2 instances that operate 24/7 with a steady and predictable load. The SysOps administrator wants to minimize compute costs while ensuring the required capacity is always available. Which EC2 purchasing option should the administrator use?

A.Reserved Instances
B.Spot Instances
C.On-Demand Instances
D.Dedicated Hosts
AnswerA

Reserved Instances are the most cost-effective choice for a fleet of EC2 instances running a production web application that operates continuously. By committing to 1- or 3-year terms, you can save up to 72% compared to On-Demand pricing, with Flexible size benefits and the ability to change Availability Zones or instance sizes within the same family. Payment options (All Upfront, Partial Upfront, No Upfront) let you further optimize cost based on cash flow, and unused compute capacity is applied automatically. For an always-on workload, the upfront commitment is justified by the substantial savings.

Why this answer

Reserved Instances (RIs) are the correct choice because the workload runs 24/7 with steady, predictable load. RIs provide a significant discount (up to 72%) over On-Demand pricing in exchange for a one- or three-year commitment, ensuring cost minimization while guaranteeing capacity availability in the specified Availability Zone.

Exam trap

The trap here is that candidates often choose Spot Instances for any cost-saving scenario, forgetting that Spot Instances can be terminated at any time, making them unsuitable for steady, predictable workloads that require constant availability.

How to eliminate wrong answers

Option B is wrong because Spot Instances are designed for fault-tolerant, flexible workloads and can be interrupted with a 2-minute warning when EC2 needs capacity back, making them unsuitable for a steady, always-on web application. Option C is wrong because On-Demand Instances offer no discount and are the most expensive option for continuous 24/7 usage, failing to minimize compute costs. Option D is wrong because Dedicated Hosts provide physical servers dedicated for your use, which is unnecessary for cost optimization and typically incurs higher costs; they are used for licensing or compliance requirements, not for general cost savings.

139
MCQeasy

A company wants to receive an alert when its AWS spending exceeds $5,000 in a month. The SysOps administrator needs to set up a proactive alert that monitors actual costs. Which AWS service should be used?

A.AWS Cost Explorer
B.AWS Budgets
C.AWS Trusted Advisor
D.AWS CloudTrail
AnswerB

AWS Budgets is the native service for setting a custom cost budget, such as $5,000 per month, and receiving alerts when actual or forecasted charges reach or exceed that amount. You can define multiple alert thresholds (for example, 80% and 100%) and publish notifications to Amazon SNS, which can email you or trigger automated workflows. This makes AWS Budgets the correct solution for a spending alert.

Why this answer

AWS Budgets is the correct service because it allows you to set a cost budget that proactively sends alerts when actual costs exceed a specified threshold (e.g., $5,000 per month). Unlike Cost Explorer, which is an analytical tool, Budgets provides real-time notifications via Amazon SNS when costs reach or are forecasted to exceed the budgeted amount, enabling proactive monitoring of actual spending.

Exam trap

The trap here is that candidates confuse AWS Cost Explorer's historical analysis capabilities with proactive alerting, assuming it can send notifications, when in fact it only provides dashboards and reports without automated threshold-based alerts.

How to eliminate wrong answers

Option A is wrong because AWS Cost Explorer is a visualization and analysis tool for historical cost data; it does not support proactive alerts based on actual cost thresholds. Option C is wrong because AWS Trusted Advisor inspects your AWS environment for cost optimization, security, and performance best practices, but it does not monitor or alert on actual spending against a specific budget. Option D is wrong because AWS CloudTrail records API activity for auditing and governance, not cost monitoring or alerting.

140
MCQhard

A company uses S3 to store critical data. They need to ensure that data can be recovered in the event of accidental deletion or overwriting by users. Which combination of actions should they take?

A.Enable S3 Cross-Region Replication and S3 Transfer Acceleration.
B.Enable S3 Versioning and S3 Transfer Acceleration.
C.Enable S3 Versioning and MFA Delete.
D.Enable S3 Server Access Logging and S3 Object Lock.
AnswerC

S3 Versioning maintains every version of an object, including the original before any overwrite or deletion, so a deleted or replaced object can be restored by reverting to an older version ID. MFA Delete requires the AWS account root's MFA code to permanently delete a version or to change the bucket's versioning state, preventing an attacker who compromises IAM credentials from executing an unrecoverable purge. Together these features provide both a recovery mechanism and a hardened authorization boundary, directly addressing the requirement to ensure critical data cannot be lost through accidental or malicious deletion.

Why this answer

Enabling S3 Versioning preserves all versions of an object, allowing recovery from accidental deletion or overwriting. MFA Delete adds an extra layer of protection by requiring multi-factor authentication to permanently delete object versions or suspend versioning, preventing unauthorized or accidental permanent data loss.

Exam trap

The trap here is that candidates often think S3 Cross-Region Replication (CRR) or S3 Object Lock alone can prevent accidental deletion, but CRR replicates delete markers and does not protect the source, while Object Lock without Versioning cannot recover overwritten data; the correct combination requires both Versioning and MFA Delete to enable recovery and prevent permanent deletion.

How to eliminate wrong answers

Option A is wrong because S3 Cross-Region Replication (CRR) replicates objects to another region for disaster recovery or compliance, but it does not protect against accidental deletion or overwriting within the source bucket; deleted objects are also replicated as delete markers. S3 Transfer Acceleration speeds up uploads over long distances but provides no data recovery capabilities. Option B is wrong because while S3 Versioning enables recovery, S3 Transfer Acceleration is irrelevant for data recovery; it only improves upload performance.

Option D is wrong because S3 Server Access Logging records requests for auditing but does not enable recovery of deleted or overwritten objects; S3 Object Lock prevents objects from being deleted or overwritten for a fixed retention period, but without Versioning, it cannot recover objects that were overwritten before the lock was applied, and it does not address accidental deletion by users with sufficient permissions.

141
Multi-Selectmedium

A SysOps administrator is automating the deployment of a web application using AWS CloudFormation. The application requires an Application Load Balancer (ALB) and an Auto Scaling group. The administrator wants to ensure that the Auto Scaling group registers instances with the ALB automatically. Which of the following are required? (Choose TWO.)

Select 2 answers
A.A launch template or launch configuration that defines the AMI and instance type.
B.A health check grace period set in the Auto Scaling group.
C.A security group that allows traffic from the ALB.
D.A target group ARN specified in the Auto Scaling group's configuration.
E.An Application Load Balancer created in the same stack.
AnswersA, D

A launch template or configuration is mandatory because the Auto Scaling group must know which Amazon Machine Image (AMI), instance type, key pair, and other instance settings to use when scaling out. Without a template, the ASG has no blueprint for creating EC2 instances, so it cannot launch or register anything with a load balancer. This is the fundamental instance-provisioning requirement, not merely an optional convenience.

Why this answer

A launch template or launch configuration is required because it defines the AMI, instance type, and other configuration details that the Auto Scaling group uses to launch EC2 instances. Without it, the Auto Scaling group cannot provision instances to register with the ALB.

Exam trap

The trap here is that candidates often think the ALB must be in the same CloudFormation stack or that a health check grace period is mandatory, when in fact only the launch template/configuration and target group ARN are required for automatic registration.

142
Multi-Selectmedium

A company needs to restrict access to an S3 bucket so that only users from a specific VPC can read objects. Which THREE configurations are required?

Select 3 answers
A.Create a bucket policy that denies access unless the request comes from a specific VPC endpoint.
B.Update the route table in the VPC to route S3 traffic through the VPC endpoint.
C.Create IAM users and assign them permissions to access the bucket.
D.Create a VPC endpoint for S3 in the specified VPC.
E.Attach a security group to the S3 bucket.
AnswersA, B, D

This bucket policy explicitly denies all S3 access unless the vpc:SourceVpce condition matches the specified VPC endpoint (e.g., vpce-12345678). You must include both an Allow statement for the principal (such as the account root) and a Deny statement with StringNotEquals to prevent all other network paths. Requests from the VPC endpoint will carry the vpcSourceVpce value automatically, so only traffic routed through that endpoint is permitted.

Why this answer

Option D is correct because an S3 gateway VPC endpoint must first exist in the specified VPC to give that VPC private connectivity to S3 and to provide the vpce ID that the bucket policy will reference. Option A is correct because the bucket policy must explicitly deny (or restrict) access unless requests arrive via that VPC endpoint, typically using the aws:sourceVpce condition key with the endpoint ID. Option B is correct because the VPC route table must have a route for the S3 prefix list (pl-xxxxxxxx) pointing to the gateway endpoint so that traffic from the VPC actually traverses the endpoint and satisfies the policy condition.

Option C is not required because the scenario restricts access by network origin (VPC endpoint), not by individual IAM identities, and IAM users alone would not enforce the VPC restriction. Option E is not valid because security groups cannot be attached to S3 buckets; S3 access control uses bucket policies, IAM policies, ACLs, and endpoint policies instead.

Exam trap

SOA-C02 often tests the misconception that a VPC endpoint alone restricts access — candidates forget the bucket policy condition and route table entry, or wrongly assume security groups can be attached to S3.

143
MCQhard

A company has two VPCs in different AWS regions (us-east-1 and eu-west-1) that are peered. Applications in both VPCs need to communicate using private IP addresses. The ping tests are successful, but the latency is significantly higher than expected. Which change is most likely to improve the latency between the VPCs?

A.Enable DNS resolution for the VPC peering connection.
B.Use a Transit Gateway instead of VPC Peering for cross-region connectivity.
C.Increase the MTU on the instances' network interfaces to 9001.
D.Configure ECMP (Equal-Cost Multi-Path) routing on the VPC peering connection.
AnswerA

When a VPC peering connection has DNS resolution enabled in both VPCs, instances can resolve private DNS hostnames of the peer VPC through the Amazon-provided DNS resolver, which returns private IP addresses instead of public ones. This keeps all cross-VPC traffic on the AWS backbone, avoiding the extra latency of routing over the public internet. Without this setting, private DNS names may fail to resolve or map to public endpoints, forcing traffic outside the VPC and adding avoidable round-trip delay.

Why this answer

Enabling DNS resolution for the VPC peering connection allows instances to resolve public DNS hostnames to the private IP addresses of the peered VPC. Without this, DNS queries may return public IP addresses, forcing traffic to traverse the internet or NAT gateways, which adds significant latency. By resolving to private IPs, traffic stays within the AWS backbone, reducing latency.

Exam trap

The trap here is that candidates often assume latency is caused by network path or bandwidth issues (leading them to choose Transit Gateway or MTU changes), but the real culprit is DNS resolution misconfiguration forcing traffic over the public internet instead of the private AWS backbone.

How to eliminate wrong answers

Option B is wrong because using a Transit Gateway instead of VPC Peering for cross-region connectivity does not inherently reduce latency; both use the AWS global backbone, and latency is primarily affected by physical distance and routing, not the service type. Option C is wrong because increasing the MTU to 9001 (jumbo frames) improves throughput for large packets but does not reduce latency; in fact, jumbo frames can increase serialization delay for small packets and are not supported over VPC peering connections (MTU is limited to 1500). Option D is wrong because ECMP routing is not configurable on VPC peering connections; VPC peering does not support multiple paths or load balancing, and ECMP is a feature of Transit Gateway or Direct Connect, not VPC peering.

144
MCQmedium

A SysOps administrator is designing a disaster recovery plan for a web application that runs on EC2 instances with data stored in an RDS MySQL database. The application requires a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour. Which solution meets these requirements most cost-effectively?

A.Use a single RDS instance with daily snapshots and EC2 instance store.
B.Use EC2 Auto Scaling across Regions with an RDS standby instance.
C.Use RDS Multi-AZ with read replicas in another Region.
D.Use RDS Multi-AZ with automated backups and EC2 AMI backups.
AnswerD

RDS Multi-AZ automatically fails over to a standby instance in a different Availability Zone within minutes, satisfying the 1-hour RTO for the database tier. Automated backups on RDS enable point-in-time recovery to any second within the backup retention period, achieving an RPO of approximately 5 minutes. EC2 AMI backups provide a consistent, restorable image of the application servers, allowing rapid relaunch of the compute tier. Together, these components form a practical, cost-effective disaster recovery plan that meets both the RPO and RTO targets without relying on ephemeral storage or manual cross-region promotion.

Why this answer

RDS Multi-AZ provides automatic synchronous replication to a standby instance in a different Availability Zone, ensuring minimal data loss and a fast failover that meets the 5-minute RPO and 1-hour RTO. Automated backups enable point-in-time recovery within the retention period, and EC2 AMI backups allow quick restoration of the application tier, making this the most cost-effective solution that satisfies both objectives without requiring cross-Region resources.

Exam trap

The trap here is that candidates confuse Multi-AZ with cross-Region disaster recovery, assuming that a read replica in another Region is required for a 5-minute RPO, but Multi-AZ within a single Region with automated backups is sufficient and more cost-effective for the given RPO/RTO targets.

How to eliminate wrong answers

Option A is wrong because daily snapshots cannot achieve a 5-minute RPO, as data loss could be up to 24 hours, and EC2 instance store is ephemeral and does not persist data across stops or terminations, failing the RTO/RPO requirements. Option B is wrong because EC2 Auto Scaling across Regions introduces significant complexity and cost for cross-Region data replication, and RDS standby instances in another Region (Multi-Region) are not natively supported without additional replication mechanisms like cross-Region read replicas, which increase latency and cost beyond what is needed. Option C is wrong because RDS Multi-AZ with read replicas in another Region provides read scaling and disaster recovery but incurs higher costs for cross-Region data transfer and replica maintenance, and the read replica is asynchronous, potentially exceeding the 5-minute RPO during a failover scenario.

145
MCQmedium

A company has two Amazon VPCs (VPC-A and VPC-B) in the same AWS Region with non-overlapping CIDR blocks. The SysOps administrator needs to establish private IP connectivity between the two VPCs with high throughput and minimal cost. Which solution should the administrator implement?

A.VPC Peering
B.AWS Transit Gateway
C.AWS VPN CloudHub
D.AWS Direct Connect
AnswerA

VPC peering allows private connectivity between two VPCs using AWS's private network. It is simple to set up, has no bandwidth limitations, and incurs no hourly cost. It is the most cost-effective solution for connecting two VPCs in the same region.

Why this answer

VPC Peering is the correct solution because it allows direct private IP connectivity between two VPCs in the same AWS Region using the AWS global network backbone, with no bandwidth bottlenecks, no single point of failure, and no additional cost beyond data transfer charges. Since the VPCs have non-overlapping CIDR blocks, they can be peered without route conflicts, and traffic flows entirely within AWS without traversing the public internet or requiring a transit hub.

Exam trap

The trap here is that candidates often over-engineer the solution by choosing AWS Transit Gateway for its centralized routing features, forgetting that for a simple two-VPC peering scenario with non-overlapping CIDRs, VPC Peering is the most cost-effective and high-performance option without the overhead of a transit hub.

How to eliminate wrong answers

Option B (AWS Transit Gateway) is wrong because it introduces unnecessary complexity and cost (hourly per-attachment charges and data processing fees) for a simple two-VPC scenario where VPC Peering provides the same high throughput at lower cost. Option C (AWS VPN CloudHub) is wrong because it requires VPN connections over the public internet, which adds latency, reduces throughput, and incurs hourly VPN connection charges, making it less performant and more expensive than VPC Peering. Option D (AWS Direct Connect) is wrong because it is designed for hybrid connectivity between on-premises networks and AWS, not for VPC-to-VPC communication, and involves significant setup costs, long lead times, and monthly port fees that are unnecessary for this use case.

146
MCQmedium

A SysOps administrator needs to monitor memory utilization on an Amazon EC2 instance. Memory metrics are not available by default in Amazon CloudWatch for EC2 instances. Which action should the administrator take to collect memory utilization metrics?

A.Install the CloudWatch agent on the EC2 instance
B.Enable detailed monitoring on the EC2 instance
C.Use an AWS Lambda function to query the EC2 instance for memory metrics
D.Use Amazon Inspector to collect memory metrics
AnswerA

The CloudWatch agent runs on the instance and collects guest-level metrics, including memory utilisation, which the hypervisor cannot see. Installing it satisfies the requirement because EC2's default CloudWatch metrics cover CPU, disk and network only, never memory.

Why this answer

The CloudWatch agent is the correct solution because it can collect custom metrics, including memory utilization, from EC2 instances. Unlike the default hypervisor-level metrics (CPU, network, disk), memory metrics require an in-guest agent to read the operating system's memory counters and publish them to CloudWatch.

Exam trap

The trap here is that candidates confuse 'detailed monitoring' (which increases metric frequency) with the ability to collect new metric types, assuming it will magically include memory metrics when it only affects existing hypervisor-level metrics.

How to eliminate wrong answers

Option B is wrong because enabling detailed monitoring only increases the frequency of default EC2 metrics (e.g., CPU, disk I/O) from 5 minutes to 1 minute; it does not add memory metrics. Option C is wrong because AWS Lambda cannot directly query an EC2 instance's OS-level memory metrics without an agent or API endpoint installed inside the instance. Option D is wrong because Amazon Inspector is a vulnerability assessment service that scans for software vulnerabilities and network exposures, not a tool for collecting OS-level performance metrics like memory utilization.

147
MCQmedium

A company has a VPC with an IPv4 CIDR block of 10.0.0.0/16. They need to add an IPv6 CIDR block to the VPC and ensure that EC2 instances can communicate over IPv6. Which step is necessary?

A.Attach an internet gateway that supports IPv6.
B.Create a new VPC with an IPv6 CIDR block and migrate resources.
C.Associate an Amazon-provided IPv6 CIDR block with the VPC.
D.Enable DNS64 in the VPC.
AnswerC

Associating an Amazon-provided IPv6 CIDR block with the VPC is the correct first step to enable IPv6. The VPC's IPv4 CIDR remains unchanged, and AWS automatically assigns a /56 IPv6 CIDR from Amazon's global unicast address pool. After this association, you must also assign /64 IPv6 CIDRs to subnets, attach an internet gateway, update route tables, and add IPv6 rules to security groups and NACLs to complete native IPv6 support.

Why this answer

To enable IPv6 communication in an existing VPC, you must first associate an Amazon-provided IPv6 CIDR block with the VPC. This is a prerequisite for configuring subnets, route tables, and internet gateways to support IPv6 traffic. Without an IPv6 CIDR block assigned to the VPC, no EC2 instance can obtain an IPv6 address or route IPv6 traffic.

Exam trap

The trap here is that candidates often think attaching an IPv6-capable internet gateway is the first step, but the VPC must first have an IPv6 CIDR block assigned before any IPv6 routing or addressing can occur.

How to eliminate wrong answers

Option A is wrong because attaching an internet gateway that supports IPv6 is necessary only after an IPv6 CIDR block has been associated with the VPC; the gateway itself does not add IPv6 addressing to the VPC. Option B is wrong because you do not need to create a new VPC; you can add an IPv6 CIDR block to an existing VPC without migrating resources. Option D is wrong because DNS64 is used to translate IPv6 DNS queries to IPv4 addresses for NAT64 scenarios, not to add an IPv6 CIDR block or enable native IPv6 communication.

148
MCQhard

A SysOps administrator is managing a fleet of EC2 instances in an Auto Scaling group. The instances are behind an Application Load Balancer. The administrator notices that the 'SurgeQueueLength' metric for the ALB is frequently high. What does this indicate, and what is the BEST remediation action?

A.The targets are unhealthy; decrease the desired capacity to reduce load.
B.The targets are not able to handle the request rate; increase the desired capacity or add scaling policies.
C.The load balancer is accepting too many connections; increase the idle timeout.
D.The load balancer is overloaded; replace it with a Network Load Balancer.
AnswerB

A high SurgeQueueLength metric for an Application Load Balancer signifies that requests are waiting in the load balancer's queue because the registered targets are not processing them fast enough. To resolve this, you should increase the desired capacity of your Auto Scaling group or add scaling policies, such as target tracking based on the ALBRequestCountPerTarget metric, so more targets are available to handle the request rate.

Why this answer

The SurgeQueueLength metric measures the number of requests that are queued by the Application Load Balancer (ALB) because no healthy target is available to process them. A frequently high value indicates that the targets (EC2 instances) are overwhelmed and cannot keep up with the incoming request rate. The best remediation is to increase the desired capacity of the Auto Scaling group or add scaling policies (e.g., based on SurgeQueueLength or RequestCountPerTarget) to automatically add more instances to handle the load.

Exam trap

The trap here is that candidates confuse SurgeQueueLength with connection-level metrics (like idle timeout) or assume the load balancer itself is the bottleneck, when in fact the metric directly indicates insufficient target capacity.

How to eliminate wrong answers

Option A is wrong because decreasing desired capacity would reduce the number of targets, worsening the queue length, and unhealthy targets are indicated by the 'UnHealthyHostCount' metric, not SurgeQueueLength. Option C is wrong because increasing the idle timeout only affects how long the ALB keeps idle connections open, not the rate of incoming requests or the queue depth; SurgeQueueLength is about request backlog, not connection persistence. Option D is wrong because replacing the ALB with a Network Load Balancer (NLB) does not address the root cause—the targets are under-provisioned; an NLB operates at Layer 4 and does not provide the same request queuing behavior, but the underlying capacity issue remains.

149
MCQeasy

A SysOps administrator wants to receive alerts when the estimated charges for an AWS account exceed a certain threshold. Which AWS service should be used?

A.AWS Trusted Advisor
B.AWS Cost Explorer
C.Amazon CloudWatch (billing metric)
D.AWS Budgets
AnswerD

AWS Budgets is the native service for creating cost, usage, and reservation budgets and for sending alerts when actual or forecasted spending exceeds defined thresholds. You can set multiple budget limits, filter by dimensions like service or linked account, and receive notifications via email or SNS. This meets the requirement to receive alerts when estimated charges reach a specified amount, making it the correct answer.

Why this answer

AWS Budgets enables you to set custom cost and usage budgets and receive alerts when your actual or forecasted costs exceed your budget thresholds. This is the correct service for receiving alerts on estimated charges. Option A (AWS Trusted Advisor) provides cost optimization recommendations but does not send billing alerts.

Option B (AWS Cost Explorer) allows you to analyze your costs but does not send proactive alerts. Option C (Amazon CloudWatch billing metric) does have billing metrics, but to set alerts on those metrics you must use AWS Budgets; CloudWatch itself is not the service for budget alerts.

150
MCQeasy

A company uses AWS Key Management Service (KMS) to encrypt data in S3. The security team wants to ensure that only a specific IAM role can decrypt objects in a particular S3 bucket. Which of the following is the MOST effective way to achieve this?

A.Create a separate KMS key for each bucket and assign the IAM role as a key user
B.Configure an S3 bucket policy that allows only the IAM role to perform s3:GetObject
C.Use a KMS key policy with a condition that requires the encryption context to match the bucket ARN, and grant the IAM role decrypt permissions
D.Use an S3 bucket policy that denies s3:GetObject unless the request includes the x-amz-server-side-encryption-aws-kms-key-id header
AnswerC

This is the correct approach because it uses a KMS key policy to enforce that decryption is allowed only when the encryption context matches the specific bucket ARN. When S3 encrypts an object with SSE-KMS, it automatically populates the encryption context with a key-value pair such as `aws:s3:arn` containing the bucket ARN; the key policy can then include a condition like `kms:EncryptionContext:aws:s3:arn` to limit `kms:Decrypt` to requests for that exact bucket. Additionally, granting `kms:Decrypt` to the IAM role in the key policy (or a combination of key and IAM policies) ensures that the role is the only principal allowed to decrypt, while the encryption-context condition prevents the same key from being used to decrypt objects in other buckets. This provides a robust, fine-grained access control that aligns with AWS best practices for SSE-KMS and satisfies the requirement to restrict decryption to the IAM role for the specific bucket.

Why this answer

It uses a KMS key policy with a condition that restricts decryption to requests where the encryption context matches the S3 bucket ARN. This ensures that only the specified IAM role, when making decrypt calls with the correct encryption context, can decrypt objects in that bucket. It directly ties the KMS key's decrypt permission to the bucket, providing a granular and secure access control mechanism.

Exam trap

The trap here is that candidates often confuse S3 bucket policies (which control access to the object) with KMS key policies (which control who can decrypt the underlying data), leading them to pick Option B or D, which address object access but not the decryption permission itself.

How to eliminate wrong answers

Option A is wrong because creating a separate KMS key per bucket and assigning the IAM role as a key user does not restrict decryption to only that bucket; the role could decrypt any object encrypted with that key, even from other locations, and does not enforce the bucket-specific context. Option B is wrong because an S3 bucket policy that allows only the IAM role to perform s3:GetObject controls read access to the object but does not prevent the role from decrypting the object using KMS if it has separate KMS decrypt permissions; the decryption is governed by KMS policies, not S3 bucket policies. Option D is wrong because an S3 bucket policy that denies s3:GetObject unless the request includes the x-amz-server-side-encryption-aws-kms-key-id header only enforces that the object be encrypted with a specific KMS key during upload, but does not control who can decrypt it; the IAM role could still decrypt if it has KMS decrypt permissions on that key.

Page 1

Page 2 of 16

Page 3