Courseiva

AWS Certified SysOps Administrator Associate SOA-C02 (SOA-C02) — Questions 1126–1169

1169 questions total · 16pages · All types, answers revealed

Page 15

Page 16 of 16

1126
MCQmedium

A company uses AWS CloudFormation to deploy its infrastructure. The SysOps administrator needs to be notified when a stack creation fails. Which solution meets this requirement with the LEAST effort?

A.Create a CloudWatch alarm that triggers when the CloudFormation stack status is 'CREATE_FAILED'.
B.Use AWS CloudTrail to monitor CreateStack API calls and trigger an SNS notification.
C.Configure an SNS topic in the CloudFormation stack's 'NotificationARNs' parameter.
D.Write a custom script that polls the CloudFormation API every minute and sends an SNS notification on failure.
AnswerC

The NotificationARNs parameter is a native CloudFormation feature that accepts one or more Amazon SNS topic ARNs, causing CloudFormation to publish all stack lifecycle events—including CREATE_FAILED, ROLLBACK_COMPLETE, and resource-level failures—to the topic. This is the built-in, real-time mechanism designed exactly for this scenario, requiring no custom code, polling, or metric configuration. Simply specify the SNS topic ARN when creating or updating the stack, and ensure the topic policy trusts CloudFormation's publishing role.

Why this answer

CloudFormation natively supports specifying an SNS topic in the 'NotificationARNs' parameter, which automatically sends notifications on stack events such as creation failure. This requires no additional infrastructure, scripting, or monitoring setup, making it the least-effort solution.

Exam trap

The trap here is that candidates often overthink and choose CloudWatch alarms or CloudTrail, not realizing that CloudFormation's built-in SNS notification parameter provides a zero-configuration, event-driven solution for stack failure alerts.

How to eliminate wrong answers

Option A is wrong because CloudWatch cannot directly alarm on CloudFormation stack status; CloudWatch alarms are designed for metrics (e.g., EC2 CPU utilization) and not for CloudFormation stack state changes. Option B is wrong because CloudTrail logs API calls but does not trigger SNS notifications directly; you would need additional services like EventBridge to route the event to SNS, adding complexity. Option D is wrong because writing a custom script to poll the CloudFormation API every minute introduces unnecessary overhead, latency, and maintenance effort, contradicting the 'least effort' requirement.

1127
MCQeasy

A SysOps Administrator needs to allow an EC2 instance in a private subnet to access the internet for software updates. Which AWS service should be used?

A.VPN Connection
B.NAT Gateway
C.Internet Gateway
D.VPC Peering
AnswerB

A NAT Gateway is a managed service that enables instances in a private subnet to initiate outbound IPv4 traffic to the internet while blocking unsolicited inbound connections. It performs source network address translation (SNAT), replacing the instance's private IP with the NAT Gateway's Elastic IP before forwarding traffic to an Internet Gateway. This is the correct architecture because the private subnet's route table points 0.0.0.0/0 to the NAT Gateway, and the NAT Gateway resides in a public subnet to reach the internet.

Why this answer

A NAT Gateway enables EC2 instances in a private subnet to initiate outbound traffic to the internet (e.g., for software updates) while preventing unsolicited inbound connections from the internet. It translates the private IP of the instance to the NAT Gateway's Elastic IP using Source Network Address Translation (SNAT). This is the correct service because the instance is in a private subnet and requires internet access without being directly reachable.

Exam trap

The trap here is that candidates often confuse an Internet Gateway with a NAT Gateway, not realizing that an Internet Gateway requires the instance to have a public IP and be in a public subnet, whereas a NAT Gateway is specifically designed for private subnets to access the internet outbound only.

How to eliminate wrong answers

Option A is wrong because a VPN Connection establishes a secure tunnel between an on-premises network and AWS, not for providing internet access to instances in a private subnet. Option C is wrong because an Internet Gateway is attached to a VPC and allows inbound/outbound internet access, but it requires the instance to have a public IP and be in a public subnet; it cannot be used directly from a private subnet. Option D is wrong because VPC Peering connects two VPCs privately using AWS's internal network, and it does not provide internet access; it also does not support transitive routing or a default route to the internet.

1128
MCQeasy

A SysOps administrator is troubleshooting a Lambda function that does not write logs to CloudWatch Logs. The IAM role attached to the function includes the policy shown. What is the most likely reason the logs are not being created?

A.The log group name in the Resource ARN does not match the actual log group created by the Lambda function.
B.The IAM role is not assigned to the Lambda function's execution role.
C.The Lambda function is in a VPC without a VPC endpoint for CloudWatch Logs.
D.The policy does not include the logs:PutLogEvents permission.
AnswerA

Lambda automatically creates a log group named /aws/lambda/<function-name> in the same Region as the function. If the Resource ARN in the IAM policy points to a different log group name (e.g., a typo or a custom group like /aws/lambda/my-function-v2), CloudWatch Logs rejects the write attempt even though the role and actions are correct. The resulting CloudWatch Logs error typically indicates that the specified log group does not exist or the resource ARN does not match, which matches this root cause.

Why this answer

The IAM policy shown in the question includes a Resource ARN that specifies a specific log group name (e.g., `/aws/lambda/MyFunction`). If the Lambda function is configured to write to a different log group (e.g., `/aws/lambda/MyOtherFunction` or a custom log group), the `logs:CreateLogGroup` and `logs:CreateLogStream` permissions will fail because the ARN does not match. This mismatch prevents the function from creating the log group or stream, so no logs are written to CloudWatch Logs.

Exam trap

The trap here is that candidates often overlook the Resource ARN mismatch and instead focus on missing permissions or VPC connectivity, but the core issue is that the IAM policy's log group ARN does not match the actual log group name Lambda tries to use.

How to eliminate wrong answers

Option B is wrong because the IAM role is already attached to the Lambda function's execution role (the policy is part of that role), so the assignment is not the issue. Option C is wrong because a Lambda function in a VPC can still write logs to CloudWatch Logs via the public internet or a NAT gateway; a VPC endpoint for CloudWatch Logs is not required unless the VPC has no internet access and no NAT gateway. Option D is wrong because the policy shown includes `logs:PutLogEvents` (the question states the policy includes it), so the missing permission is not the cause.

1129
MCQmedium

A company's application running on EC2 instances is experiencing intermittent errors. The SysOps team needs to collect and analyze application logs from all instances centrally. The logs must be stored durably and searchable with minimal latency. Which solution meets these requirements?

A.Enable AWS CloudTrail and store logs in an S3 bucket.
B.Use Amazon Kinesis Data Firehose to send logs directly from each instance to Amazon Redshift.
C.Install the CloudWatch Logs agent on each EC2 instance and stream logs to Amazon CloudWatch Logs.
D.Store logs locally on each instance and periodically copy them to Amazon S3.
AnswerC

Installing the CloudWatch Logs agent on each EC2 instance enables near-real-time streaming of application and system log files to Amazon CloudWatch Logs, providing a centralized, scalable, and searchable log store. With CloudWatch Logs you can create metric filters to trigger CloudWatch alarms, use Logs Insights to run queries across log groups, and perform live tailing to watch logs as they arrive. The agent also handles log rotation, multi-line log records, and timestamp parsing automatically, which simplifies log management and removes the need for manual log harvesting or extra infrastructure.

Why this answer

The CloudWatch Logs agent (or unified CloudWatch agent) installed on each EC2 instance can stream application logs in near real-time to Amazon CloudWatch Logs, which provides durable storage, automatic encryption at rest, and a searchable interface via the console, CLI, or API with minimal latency. This centralized logging solution meets the requirements for collecting logs from all instances, storing them durably, and enabling immediate querying without additional infrastructure.

Exam trap

The trap here is that candidates may confuse CloudTrail (API logging) with application logging, or assume that S3 periodic uploads are sufficient for 'minimal latency' searchability, when in fact CloudWatch Logs is the native AWS service designed for real-time log ingestion and querying from EC2 instances.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity for governance and auditing, not application-level logs generated by processes running on EC2 instances; it cannot capture stdout, stderr, or custom application log files. Option B is wrong because Amazon Kinesis Data Firehose is a streaming ingestion service that delivers data to destinations like S3 or Redshift, but sending logs directly from each instance to Firehose without an agent or SDK is not a standard pattern, and Amazon Redshift is a data warehouse optimized for analytical queries, not a low-latency log search engine; this adds unnecessary complexity and cost. Option D is wrong because storing logs locally on each instance risks data loss on instance termination or failure, and periodically copying logs to S3 introduces latency that prevents real-time searchability, failing the 'minimal latency' requirement.

1130
MCQeasy

A SysOps administrator needs to track changes to security groups in the AWS account. Which AWS service should be used to record configuration changes and provide a history of security group modifications?

A.AWS Trusted Advisor
B.Amazon CloudWatch
C.AWS Config
D.AWS CloudTrail
AnswerC

AWS Config is the correct service because it continuously records configuration items for supported resources, including security groups, and maintains a configuration history. When a security group rule is added or removed, AWS Config generates a configuration item and allows you to review the previous and new state using the configuration timeline. It also enables compliance rules to detect and evaluate changes, making it the definitive service for tracking and auditing security group changes.

Why this answer

AWS Config is the correct service because it provides a detailed inventory of AWS resources, records configuration changes, and maintains a historical timeline of those changes. For security groups, AWS Config can track modifications such as rule additions, deletions, or updates, and it can trigger evaluations against desired configurations. This makes it the ideal service for auditing and compliance use cases involving security group changes.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (which logs API calls) with AWS Config (which records resource configuration state and history), leading them to choose CloudTrail for change tracking when Config is the service designed for configuration history and compliance auditing.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor is an advisory service that inspects your AWS environment and makes recommendations based on AWS best practices, but it does not record or maintain a history of configuration changes to resources like security groups. Option B is wrong because Amazon CloudWatch is a monitoring service for metrics, logs, and alarms; it can detect and alert on changes via CloudWatch Events, but it does not natively store a historical record of configuration changes or provide a timeline of modifications. Option D is wrong because AWS CloudTrail records API calls and events, including those that modify security groups, but it focuses on who made the call and when, not on the state or configuration history of the resource itself; CloudTrail does not provide a point-in-time configuration snapshot or a change timeline for the resource's configuration.

1131
MCQhard

A company runs a web application on EC2 instances behind an Application Load Balancer (ALB). The application stores sensitive data in an S3 bucket. The security team has mandated that all data in transit to the S3 bucket must be encrypted using TLS. The SysOps administrator configured the application to use HTTPS endpoints for S3. However, a security audit reveals that some requests to S3 are still being sent over HTTP. The administrator checks the VPC Flow Logs and sees that the EC2 instances are communicating with the S3 bucket via a VPC endpoint. The company also uses an S3 bucket policy that allows access only from the VPC endpoint. What is the most likely reason that some requests are sent over HTTP?

A.The VPC endpoint is not encrypting traffic between the instances and the endpoint.
B.The VPC endpoint is configured for HTTP instead of HTTPS.
C.The S3 bucket policy does not require HTTPS for requests.
D.The application is not configured to use HTTPS for all S3 requests.
AnswerD

This is the correct root cause. The application must use HTTPS endpoints when calling S3 (e.g., https://bucket.s3.amazonaws.com) to encrypt data in transit between the EC2 instances and S3. Even when using a VPC endpoint, the application's SDK or code explicitly determines whether requests are signed and sent over TLS. If the code uses HTTP URLs or does not enforce TLS, traffic is sent in plaintext, exposing data on the network.

Why this answer

The VPC endpoint for S3 does not automatically encrypt traffic; encryption is handled at the application level. The application must be configured to use HTTPS for all S3 requests. Even though the administrator configured the application to use HTTPS endpoints, some requests may still be sent over HTTP if the application has fallback logic or uses an SDK that defaults to HTTP in certain cases.

Option A is incorrect because the VPC endpoint itself does not encrypt traffic; encryption is an application-layer function. Option B is incorrect because VPC endpoints are not configured for HTTP or HTTPS; they use AWS API calls which can be made over either protocol depending on the client. Option C is incorrect because the S3 bucket policy does not require HTTPS; it only restricts access to the VPC endpoint, not the protocol.

1132
Multi-Selecthard

A SysOps administrator is troubleshooting a Lambda function that is not processing messages from an SQS queue. The function is subscribed to the queue via an event source mapping. The function has a reserved concurrency of 0. Which TWO actions will resolve the issue?

Select 2 answers
A.Add SQS permissions to the Lambda execution role.
B.Configure a dead-letter queue for the Lambda function.
C.Set the reserved concurrency to a value greater than 0.
D.Enable the event source mapping if it is disabled.
E.Increase the batch size in the event source mapping.
AnswersC, D

Reserved concurrency of 0 is an intentional hard throttle that blocks all function invocations, causing Lambda to return a throttling error for every request. By raising reserved concurrency to any positive value, you grant the function a dedicated pool of execution capacity, allowing the event source mapping to successfully invoke it. This is the most direct fix when the function cannot execute at all, because it removes the service-level blocking condition that prevents both synchronous and event source mapping invocations.

Why this answer

Reserved concurrency of 0 means the Lambda function has no available execution capacity, so it cannot process any invocations, including those from SQS. Setting reserved concurrency to a value greater than 0 (e.g., 1 or more) allocates the necessary execution slots for the function to run. This directly resolves the issue because the event source mapping will successfully invoke the function only when concurrency is available.

Exam trap

The trap here is that candidates often overlook reserved concurrency of 0 as a valid configuration that completely blocks invocations, and instead focus on permissions or queue settings, not realizing that a concurrency limit of 0 is a deliberate disablement mechanism.

1133
Multi-Selectmedium

A SysOps administrator needs to design a VPC with public and private subnets for a web application. Which TWO components are required to allow instances in the private subnet to access the internet?

Select 2 answers
A.NAT gateway in a public subnet
B.Route table entry in the private subnet routing 0.0.0.0/0 to the NAT gateway
C.VPC endpoint for S3
D.Internet gateway attached to the VPC
E.Virtual private gateway
AnswersA, B

A NAT gateway is a managed Network Address Translation service that enables instances in a private subnet to initiate outbound IPv4 traffic to the internet and receive replies, while preventing unsolicited inbound connections from the internet. It must be placed in a public subnet with a route to an Internet Gateway and an associated Elastic IP, so it can translate private-source IPs to the public IP. This is the core component that gives private instances internet access in a VPC design with public and private subnets.

Why this answer

A NAT gateway in a public subnet is required because it allows instances in a private subnet to initiate outbound traffic to the internet (e.g., for software updates) while preventing unsolicited inbound connections. The NAT gateway must be placed in a public subnet with an Internet Gateway (IGW) route to translate private IPs to the gateway's Elastic IP. Without the NAT gateway, private instances have no path to the internet.

Exam trap

The trap here is that candidates often think an Internet Gateway alone is sufficient for private subnet internet access, but they overlook the need for a NAT device to translate private IPs, as the IGW only works with public IPs.

1134
MCQmedium

A company uses Amazon CloudWatch Logs to store application logs. The security team needs to be alerted when any log group contains a specific error pattern. The solution must minimize latency and operational overhead. What should a SysOps administrator do?

A.Stream the logs to Amazon Kinesis Data Firehose, which then triggers a Lambda function to check for errors.
B.Create a CloudWatch metric filter on the log group and set an alarm that triggers an SNS notification.
C.Create a Lambda function subscribed to the CloudWatch Logs log group, which checks for the error pattern and publishes to an SNS topic.
D.Use CloudWatch Logs Insights to run a query every minute and send results via SNS.
AnswerC

When you configure a subscription filter on a CloudWatch Logs group, CloudWatch Logs asynchronously invokes a Lambda function as log events are ingested, delivering a gzip-compressed batch of data that you can decode and search for error signatures. The Lambda function can be written in Python, Node.js, or another supported runtime, applying arbitrary regexes, aggregating events, and then directly publishing to an SNS topic for immediate fan-out to email, SMS, or Chatbot. This pattern provides real-time, event-driven alerting with minimal latency and no extra infrastructure, which is why it is the recommended approach for this requirement.

Why this answer

Subscribing a Lambda function directly to a CloudWatch Logs log group allows real-time, low-latency processing of log events as they arrive. The Lambda function can parse each log event for the specific error pattern and publish to an SNS topic to alert the security team, minimizing operational overhead by avoiding additional streaming or polling services.

Exam trap

The trap here is that candidates may choose Option B (metric filter and alarm) because it seems simpler, but they overlook that metric filters only count occurrences over time and cannot trigger immediate, per-event alerts, which is required for minimizing latency in security alerting.

How to eliminate wrong answers

Option A is wrong because streaming logs to Kinesis Data Firehose adds unnecessary latency and operational complexity; Firehose is designed for batch delivery to destinations like S3 or Redshift, not for real-time alerting with minimal latency. Option B is wrong because a CloudWatch metric filter counts occurrences of a pattern but cannot trigger an alarm on a per-log-event basis; alarms are evaluated periodically (e.g., every minute) and require a threshold, introducing latency and potential missed alerts for sporadic errors. Option D is wrong because CloudWatch Logs Insights queries are on-demand or scheduled at intervals (minimum 1 minute), not real-time, and require manual or scheduled execution, increasing latency and operational overhead compared to event-driven processing.

1135
MCQeasy

A SysOps administrator needs to allow an EC2 instance in a private subnet to download patches from the internet. Which AWS service should be used to achieve this securely?

A.Internet Gateway (IGW)
B.NAT Gateway
C.AWS VPN
D.VPC Peering
AnswerB

A NAT Gateway is a fully managed AWS service that enables instances in a private subnet to initiate outbound connections to the internet (for example, to download patches or access external APIs) while preventing unsolicited inbound connections from the internet. It is deployed in a public subnet with an Elastic IP address, and the private subnet's route table points a default route (0.0.0.0/0) to the NAT Gateway. The NAT Gateway translates the private source IP of outbound traffic to its Elastic IP and uses connection tracking to drop inbound packets that are not part of an established outbound flow. This makes it the correct choice for providing outbound-only internet access to a private EC2 instance.

Why this answer

A NAT Gateway is a managed AWS service deployed in a public subnet that allows instances in private subnets to initiate outbound connections to the internet (such as downloading patches) while preventing inbound connections from the internet. It provides the secure, one-way egress path the scenario requires.

Exam trap

The trap is confusing inbound vs. outbound connectivity — candidates pick IGW because they think 'internet access' means IGW, but IGWs are for public subnets and inbound reachability, while NAT is for private-subnet egress.

How to eliminate wrong answers

Option A is wrong because an Internet Gateway provides bidirectional internet connectivity and requires the subnet to have a route to it plus a public IP — placing an instance in a private subnet with an IGW route would defeat the purpose of the private subnet and expose it to inbound traffic. Option C is wrong because AWS VPN connects on-premises networks to a VPC; it does not provide internet egress for private-subnet instances. Option D is wrong because VPC Peering connects two VPCs privately and does not provide any path to the public internet.

1136
MCQeasy

A company's security team requires that all IAM users must use multi-factor authentication (MFA) to access the AWS Management Console. The SysOps administrator needs to create an IAM policy that denies all console actions if the user has not authenticated with MFA. Which IAM condition key should the administrator use?

A.aws:MultiFactorAuthPresent
B.aws:SourceIp
C.iam:PassedToService
D.aws:RequestedRegion
AnswerA

aws:MultiFactorAuthPresent is a boolean condition key that returns true when the caller authenticated with MFA. To enforce MFA for all IAM user console access, you attach an identity policy with a condition like "Bool": {"aws:MultiFactorAuthPresent": "true"}. This works with temporary credentials obtained via the console or sts:GetSessionToken, but it does not automatically apply to long-lived access keys unless the session is explicitly created with MFA.

Why this answer

The `aws:MultiFactorAuthPresent` condition key evaluates to `true` when the user has authenticated using MFA. By using this key in a `Deny` statement, the policy can block all console actions unless MFA is present, enforcing the security team's requirement.

Exam trap

The trap here is that candidates confuse `aws:MultiFactorAuthPresent` with `aws:MultiFactorAuthAge` (which checks how long ago MFA was used) or assume `SourceIp` can enforce MFA, but only the `MultiFactorAuthPresent` key directly evaluates MFA status for console access.

How to eliminate wrong answers

Option B is wrong because `aws:SourceIp` is used to restrict access based on the user's IP address, not MFA status. Option C is wrong because `iam:PassedToService` is used to control which roles can be passed to AWS services, not for MFA enforcement. Option D is wrong because `aws:RequestedRegion` restricts actions to specific AWS regions, not MFA authentication.

1137
MCQhard

A SysOps administrator is troubleshooting a slow-running application on an EC2 instance. CloudWatch metrics show high CPU utilization but low disk I/O. The instance type is t3.medium. Which action would most likely improve performance?

A.Change the instance type to c5.large, which provides dedicated CPU performance.
B.Increase the instance memory by changing to r5.large.
C.Enable EBS-optimized on the instance and use provisioned IOPS SSD volumes.
D.Increase the size of the EBS volume to improve disk throughput.
AnswerA

c5.large instances are compute-optimized with dedicated vCPUs, so they do not rely on CPU credits or burst capacity. When the current burstable instance exhausts its CPU credit balance, it is throttled to the baseline utilization level, causing the application to run slowly. By moving to c5.large, you get consistent, full CPU performance for sustained periods, directly addressing the CPU-bound bottleneck.

Why this answer

The t3.medium is a burstable instance that relies on CPU credits. High CPU utilization with low disk I/O indicates the application is CPU-bound and the instance has likely exhausted its CPU credits, causing performance throttling. Changing to a c5.large provides dedicated, consistent CPU performance without credit-based limitations, directly addressing the bottleneck.

Exam trap

The trap here is that candidates may focus on disk or memory improvements because the application is 'slow,' but the CloudWatch metrics clearly point to a CPU bottleneck, and the t3 family's credit-based performance model is a common exam pitfall.

How to eliminate wrong answers

Option B is wrong because increasing memory (r5.large) does not resolve CPU starvation; the metrics show high CPU utilization, not memory pressure. Option C is wrong because EBS optimization and provisioned IOPS improve disk throughput, but disk I/O is already low, indicating the bottleneck is not storage-related. Option D is wrong because increasing EBS volume size does not inherently improve disk throughput; throughput depends on volume type and IOPS, not size alone, and disk I/O is not the issue.

1138
MCQhard

A company uses AWS CloudFormation to deploy a multi-tier application. The template includes an AWS::RDS::DBInstance resource. The administrator wants to ensure that the database is not deleted when the stack is deleted. Which CloudFormation resource property should be set?

A.Set the 'DeletionPolicy' attribute to 'Retain' on the DBInstance resource.
B.Set the 'DeletionPolicy' attribute to 'Snapshot' on the DBInstance.
C.Set the 'RetainOnDeletion' property to 'true' on the DBInstance.
D.Set the 'DeletionProtection' property to 'true' on the DBInstance.
AnswerA

Setting the DeletionPolicy attribute to Retain on the AWS::RDS::DBInstance resource is the correct way to prevent CloudFormation from deleting the database when the stack is removed. With this attribute, CloudFormation simply abandons the resource, leaving the DB instance intact and running, though it no longer manages it. This is the intended mechanism for preserving RDS instances during stack teardown.

Why this answer

The 'DeletionPolicy' attribute in AWS CloudFormation controls what happens to a resource when its stack is deleted. Setting 'DeletionPolicy' to 'Retain' on the AWS::RDS::DBInstance resource ensures the database is preserved and not deleted when the stack is deleted. This is the standard mechanism for preventing accidental deletion of critical resources during stack teardown.

Exam trap

The trap here is that candidates confuse the RDS-specific 'DeletionProtection' property with CloudFormation's 'DeletionPolicy' attribute, assuming that enabling deletion protection on the database will prevent CloudFormation from deleting it, but CloudFormation's stack deletion bypasses that protection unless the DeletionPolicy is set to 'Retain'.

How to eliminate wrong answers

Option B is wrong because 'Snapshot' on the DeletionPolicy creates a final snapshot before deleting the DB instance, but it does not prevent deletion; the database is still removed. Option C is wrong because 'RetainOnDeletion' is not a valid CloudFormation resource property for AWS::RDS::DBInstance; the correct attribute is 'DeletionPolicy'. Option D is wrong because 'DeletionProtection' is a property of the RDS DB instance itself that prevents deletion via the console or API, but it does not override CloudFormation's stack deletion behavior; CloudFormation can still delete the resource if the DeletionPolicy is not set to 'Retain'.

1139
MCQeasy

A SysOps administrator is troubleshooting an issue where an EC2 instance's CPU utilization is consistently above 90%, but no CloudWatch alarm is triggered. The alarm is configured to monitor the 'CPUUtilization' metric with a threshold of 80% for 2 consecutive periods of 5 minutes. What is the most likely cause?

A.The alarm is in the 'OK' state and not 'INSUFFICIENT_DATA'.
B.The CPUUtilization metric is not enabled by default for EC2 instances.
C.The CPU utilization spikes above 80% for less than 10 minutes at a time.
D.The alarm period is set to 5 minutes, but the metric is reported every 1 minute.
AnswerC

This is correct because CloudWatch alarms evaluate a metric against the threshold over a specified number of consecutive periods. If the alarm is configured with a period of 5 minutes and evaluation periods of 2, the CPU utilization must exceed 80% for the entire 10-minute span covered by two consecutive data points. When the CPU spikes above 80% for less than 10 minutes, it never passes two full evaluation periods, so the alarm state remains OK. Thus, short-lived spikes, even if severe, will not trigger the alarm.

Why this answer

The CloudWatch alarm requires 2 consecutive periods of 5 minutes (i.e., 10 minutes total) where the CPU utilization exceeds 80%. If the CPU utilization spikes above 80% for less than 10 minutes at a time, the alarm will not trigger because it never meets the consecutive evaluation period requirement. The alarm evaluates each 5-minute period independently, and only when both consecutive periods breach the threshold does the alarm state change to ALARM.

Exam trap

The trap here is that candidates often assume any breach of the threshold triggers the alarm immediately, but they overlook the 'consecutive periods' requirement, which means the alarm only fires after the condition persists for the full evaluation window (e.g., 10 minutes for 2 periods of 5 minutes).

How to eliminate wrong answers

Option A is wrong because the alarm being in the 'OK' state is the result of the condition not being met, not the cause of the alarm not triggering; the question asks for the cause of no alarm being triggered, and the alarm state is a symptom, not a root cause. Option B is wrong because the CPUUtilization metric is enabled by default for all EC2 instances and is available in CloudWatch without any additional configuration; it is a standard metric that is automatically sent every 5 minutes (or 1 minute with detailed monitoring). Option D is wrong because the metric being reported every 1 minute (with detailed monitoring) does not prevent the alarm from triggering; the alarm period of 5 minutes means CloudWatch aggregates the 1-minute data points into 5-minute averages, and the alarm evaluates those averages against the threshold, so the reporting interval does not cause the alarm to fail.

1140
MCQhard

An application running on EC2 instances occasionally throws 'Connection refused' errors when connecting to an RDS database. The SysOps administrator needs to determine if the issue is due to database connection limits or network security groups. Which metrics and logs should the administrator examine?

A.Check CloudWatch RDS CPUUtilization and CloudTrail logs for RDS API calls.
B.Review RDS error logs in CloudWatch Logs and check the EC2 instance's system log.
C.Look at the EC2 instance's CloudWatch NetworkIn and NetworkOut metrics and RDS FreeableMemory metric.
D.Examine the RDS CloudWatch metric DatabaseConnections and analyze VPC Flow Logs for the EC2 instance's network interface.
AnswerD

DatabaseConnections shows whether the RDS instance has hit its connection limit, while VPC Flow Logs reveal whether security groups or network ACLs are rejecting traffic. Together they distinguish a connection-limit problem from a network-blocking problem.

Why this answer

'Connection refused' errors typically stem from either the database exhausting its maximum connections or network-level security groups blocking traffic. The RDS CloudWatch metric `DatabaseConnections` directly shows the current number of active connections against the instance's `max_connections` limit, while VPC Flow Logs capture whether packets are being accepted or rejected by security groups or network ACLs, pinpointing network blockages.

Exam trap

The trap here is that candidates confuse aggregate network metrics (like NetworkIn/NetworkOut) or CPU metrics with the specific indicators needed to differentiate between connection limits and security group denials, leading them to choose options that measure volume rather than connection state or packet acceptance.

How to eliminate wrong answers

Option A is wrong because `CPUUtilization` does not indicate connection limits or security group blocks, and CloudTrail logs record API calls (e.g., creating DB instances) not real-time connection or network failures. Option B is wrong because RDS error logs in CloudWatch Logs may show authentication or query errors but not connection limit exhaustion or network-level rejections, and the EC2 instance's system log (console output) does not capture network flow data. Option C is wrong because `NetworkIn`/`NetworkOut` show aggregate traffic volume, not whether connections are accepted or rejected, and `FreeableMemory` indicates memory pressure but not connection count or security group rules.

1141
MCQeasy

A SysOps administrator needs to monitor the health of an Amazon RDS for MySQL DB instance. The administrator wants to receive an alert when the database connection count exceeds a threshold of 500 for more than 5 minutes. Which AWS service should be used to create this alert?

A.Amazon CloudWatch
B.Amazon Simple Notification Service (SNS)
C.AWS CloudTrail
D.AWS Config
AnswerA

Amazon CloudWatch is the native AWS monitoring service that retrieves and stores metrics from RDS, including the DatabaseConnections metric. You can create a CloudWatch alarm that evaluates this metric every minute, and if it exceeds 500 concurrently for 5 consecutive minutes, the alarm transitions to ALARM state and can trigger SNS notifications or Auto Scaling actions. This makes CloudWatch the correct service for detecting and responding to database connection pressure.

Why this answer

Amazon CloudWatch is the correct service because it can monitor RDS metrics such as DatabaseConnections and trigger an alarm when the value exceeds a threshold of 500 for a specified duration (e.g., 5 consecutive evaluation periods). CloudWatch alarms evaluate metric data against a defined threshold and can then publish to an SNS topic to send notifications.

Exam trap

The trap here is that candidates confuse the service that evaluates metrics (CloudWatch) with the service that delivers notifications (SNS), leading them to select SNS because they think of alerts as notifications, but CloudWatch is the service that creates and evaluates the alarm based on the metric threshold.

How to eliminate wrong answers

Option B (Amazon SNS) is wrong because SNS is a notification service, not a monitoring or alert evaluation service; it cannot itself evaluate metric thresholds or create alarms. Option C (AWS CloudTrail) is wrong because CloudTrail records API calls for auditing and governance, not real-time performance metrics like database connection counts. Option D (AWS Config) is wrong because Config evaluates resource configurations and compliance rules, not operational metrics such as connection counts.

1142
Multi-Selecthard

A company is deploying a microservices application on AWS using Amazon ECS with Fargate launch type. The SysOps administrator needs to automate the deployment process so that when a new Docker image is pushed to Amazon ECR, the ECS service is updated with the new image. Which THREE AWS services should be used together to achieve this? (Choose THREE.)

Select 3 answers
A.AWS CloudFormation
B.AWS CodePipeline
C.Amazon Elastic Container Registry (ECR)
D.Amazon Elastic Container Service (ECS)
E.AWS Systems Manager
AnswersB, C, D

AWS CodePipeline is the fully managed CI/CD orchestrator for exactly this scenario. You can configure an ECR source action that listens for new image pushes, then run build/test stages, and finally use the ECS deployment action (with imagedefinitions.json) to update an ECS service. It automatically creates the task definition revision and applies rolling updates to containers, making it the correct component for continuous delivery.

Why this answer

The correct combination to automate deployment when a new Docker image is pushed to Amazon ECR is AWS CodePipeline, Amazon ECR, and Amazon ECS. CodePipeline orchestrates the CI/CD pipeline triggered by an ECR push event. ECR stores the Docker images.

ECS with Fargate runs the containers and updates the service with the new image. Option A (CloudFormation) is for infrastructure provisioning, not continuous deployment. Option E (Systems Manager) is for management and operations, not CI/CD.

1143
MCQeasy

A company runs a web application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application stores session data in an in-memory cache on the EC2 instances. During an instance refresh, users lose their session data. Which action should be taken to improve reliability without major application changes?

A.Use ElastiCache for Memcached with auto-discovery.
B.Move session state to Amazon ElastiCache for Redis.
C.Increase the minimum size of the Auto Scaling group.
D.Enable sticky sessions (session affinity) on the ALB target group.
AnswerB

Moving session state to Amazon ElastiCache for Redis is the correct solution because Redis supports replication across multiple Availability Zones, persistence through snapshots and AOF log, and atomic operations for managing session TTLs. By externalizing sessions, the EC2 instances become stateless: any instance in the Auto Scaling group can serve any user request, and session data survives instance termination, replacement, or scaling events. ElastiCache for Redis can also be configured with cluster mode to scale horizontally, ensuring session capacity grows with application demand while maintaining high availability.

Why this answer

Amazon ElastiCache for Redis provides a fully managed, external, and highly available in-memory data store that can be used to persist session state outside of the EC2 instances. By moving session data to Redis, the session state survives instance refreshes, terminations, or scaling events without requiring any changes to the application's session management logic beyond pointing to the Redis endpoint. This decouples session state from the compute layer, ensuring reliability and data durability during Auto Scaling lifecycle events.

Exam trap

The trap here is that candidates often confuse sticky sessions (session affinity) with session persistence, mistakenly believing that routing requests to the same instance prevents data loss, when in fact sticky sessions do not protect against instance termination or replacement during scaling events.

How to eliminate wrong answers

Option A is wrong because ElastiCache for Memcached is a pure caching solution that does not offer built-in persistence, replication, or failover capabilities; if the Memcached node fails, all session data is lost, and auto-discovery only helps with client connection management, not data durability. Option C is wrong because increasing the minimum size of the Auto Scaling group does not prevent session data loss during an instance refresh; it only ensures more instances are running, but the in-memory cache on each instance is still ephemeral and lost when instances are replaced. Option D is wrong because enabling sticky sessions (session affinity) on the ALB target group only ensures that a user's requests are routed to the same instance during a session, but it does not preserve session data when that instance is terminated or replaced during an instance refresh; the data is still stored in the instance's local memory and is lost upon instance termination.

1144
MCQeasy

A company has an RDS PostgreSQL database with a Multi-AZ deployment. The primary instance fails. What happens to the application connections?

A.The application must reconnect to the same endpoint; it will be redirected to the standby instance.
B.The application will be automatically redirected to a read replica.
C.The administrator must manually change the CNAME to point to the standby.
D.The application must use a new endpoint in a different AWS Region.
AnswerA

In a Multi-AZ RDS deployment, a standby DB instance is provisioned in a different Availability Zone and is kept in sync with the primary via synchronous replication. When a failure occurs, Amazon RDS automatically promotes the standby and updates the DNS record for the same endpoint to point to the new primary. Existing application connections are dropped, so the application must reconnect, but it must reconnect to the same endpoint; the DNS update happens automatically and transparently to the client.

Why this answer

When the primary RDS PostgreSQL instance fails in a Multi-AZ deployment, AWS automatically fails over to the standby instance in a different Availability Zone. The DNS endpoint (CNAME) remains the same, but its resolution is updated to point to the standby instance's IP address. The application must reconnect to the same endpoint, as existing connections to the failed primary are dropped; after reconnection, traffic is seamlessly routed to the new primary.

Exam trap

The trap here is that candidates often confuse Multi-AZ failover with read replicas, assuming the application is automatically redirected to a read replica for writes, but read replicas are asynchronous and cannot accept write traffic.

How to eliminate wrong answers

Option B is wrong because read replicas are used for read scaling, not automatic failover; Multi-AZ failover uses a standby in a different AZ, not a read replica. Option C is wrong because AWS RDS Multi-AZ automatically updates the DNS CNAME to point to the standby instance; no manual administrator intervention is required. Option D is wrong because the endpoint remains the same across the failover; the application does not need a new endpoint in a different AWS Region, as Multi-AZ operates within a single region.

1145
MCQhard

A SysOps administrator is troubleshooting connectivity issues between Amazon EC2 instances in two different VPCs that are connected via a VPC peering connection. The instances can successfully send ICMP (ping) traffic, but TCP connections on port 443 (HTTPS) fail. The security groups of both instances allow all inbound and outbound traffic. What is the most likely cause of the issue?

A.The Network ACL associated with the subnets is blocking the return traffic for TCP connections on ephemeral ports
B.The VPC peering connection is not properly configured for TCP traffic
C.The route tables in the VPCs do not contain a route for the other VPC's CIDR
D.The security group on the EC2 instance is blocking inbound TCP traffic on port 443
AnswerA

Network ACLs are stateless, so they require explicit rules for traffic in both directions. While ICMP may be permitted by the inbound/outbound rules, TCP return traffic (such as the acknowledgment and response packets) arrives on ephemeral ports (typically 1024–65535). If the outbound NACL rule does not explicitly allow these high ports, the TCP handshake or established connections will fail even though ping works. This is the classic symptom of a NACL blocking return traffic.

Why this answer

The Network ACL (NACL) is stateless, meaning it must explicitly allow both inbound and outbound traffic. While ICMP (ping) works because it doesn't rely on ephemeral ports for return traffic, TCP connections on port 443 require the return traffic to come from the target instance on a high ephemeral port (typically 1024-65535). If the NACL's outbound rules block these ephemeral ports, the TCP handshake fails, even though the security groups allow all traffic.

Exam trap

The trap here is that candidates assume security groups are the only firewall layer, overlooking that Network ACLs are stateless and require explicit rules for ephemeral port return traffic, which is why ICMP works but TCP fails.

How to eliminate wrong answers

Option B is wrong because VPC peering connections are transparent to protocols; they operate at Layer 3 and do not differentiate between ICMP and TCP traffic. Option C is wrong because if the route tables lacked a route for the other VPC's CIDR, ICMP (ping) would also fail, as routing is required for all traffic types. Option D is wrong because the question explicitly states that the security groups allow all inbound and outbound traffic, so they cannot be blocking TCP port 443.

1146
MCQeasy

A company stores critical data in an S3 bucket and wants to be notified immediately when any object is deleted from the bucket. Which combination of services should the SysOps administrator use?

A.Configure an S3 event notification for 's3:ObjectRemoved:*' events to send to an SNS topic.
B.Enable S3 server access logging and send logs to CloudWatch Logs, then create a metric filter and alarm.
C.Use S3 event notifications to invoke a Lambda function that checks the object and sends an email.
D.Use AWS CloudTrail to log DeleteObject calls and create a CloudWatch Events rule to send an SNS notification.
AnswerA

S3 event notifications are the direct, native mechanism for reacting to object lifecycle events. Configuring a notification for the `s3:ObjectRemoved:*` prefix covers both individual `DeleteObject` and multi-object `DeleteObjects` API calls, and SNS can deliver an email immediately upon event publication. With an SNS topic subscribed to by email, the notification is pushed in near real-time, typically within seconds, with no polling or compute layer required. This makes it the simplest and most appropriate option for instant alerting.

Why this answer

S3 event notifications can be configured to trigger on 's3:ObjectRemoved:*' events, which cover both permanent and versioned object deletions. These notifications can be sent directly to an SNS topic, enabling immediate notification without additional compute or logging overhead. This is the simplest and most direct approach for real-time alerts on object deletions.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing CloudTrail or Lambda, not realizing that S3 event notifications can directly trigger SNS for immediate alerts without additional services or delays.

How to eliminate wrong answers

Option B is wrong because S3 server access logs are delivered on a best-effort basis, often with delays of several hours, making them unsuitable for immediate notification. Option C is wrong because invoking a Lambda function to check the object and send an email adds unnecessary complexity and latency; the event notification can directly send to SNS without custom code. Option D is wrong because CloudTrail logs are typically delivered within 5-15 minutes, not in real time, and using CloudTrail for this purpose introduces additional cost and complexity compared to native S3 event notifications.

1147
MCQmedium

A SysOps administrator is designing a disaster recovery plan for a critical application that runs on EC2 instances in a single region. The RTO is 1 hour, and the RPO is 15 minutes. The application data is stored on an Amazon EBS volume. Which approach meets these requirements at the lowest cost?

A.Deploy the application across multiple Availability Zones using an Auto Scaling group and an Application Load Balancer.
B.Use AWS Database Migration Service (DMS) for continuous replication to an EC2 instance in the DR region.
C.Take automated EBS snapshots every 15 minutes and copy them to the DR region. Use a pre-configured Amazon Machine Image (AMI) to launch EC2 instances from the latest snapshot.
D.Use S3 Cross-Region Replication to replicate the EBS volume data to an S3 bucket in the DR region.
AnswerC

Automated EBS snapshots taken every 15 minutes and copied to the DR region deliver a consistent, block-level backup of the instance's data with a recoverable point objective of at most 15 minutes. Because EBS snapshots are stored in Amazon S3 and support cross-region copying, you can recreate the volume in the DR region. Launching EC2 instances from a pre-configured AMI that uses the latest restored snapshot as its root volume minimizes the recovery time objective by avoiding manual OS and application provisioning, making this a cost-effective and reliable DR strategy.

Why this answer

Automated EBS snapshots taken every 15 minutes meet the RPO of 15 minutes, and copying them to the DR region allows launching EC2 instances from the latest snapshot using a pre-configured AMI, which can achieve an RTO of 1 hour. This approach is the lowest cost as it only incurs snapshot storage and cross-region data transfer costs, without requiring continuous replication infrastructure or additional compute resources.

Exam trap

The trap here is that candidates may confuse cross-region replication for EBS volumes with S3 Cross-Region Replication, assuming EBS data can be directly replicated to S3, but EBS volumes are block-level storage and cannot be replicated via S3 CRR without an intermediary like AWS Backup or snapshot copy.

How to eliminate wrong answers

Option A is wrong because deploying across multiple Availability Zones within the same region does not provide disaster recovery for a regional failure; it only provides high availability within the region, failing to meet the DR requirement for a separate region. Option B is wrong because AWS DMS is designed for database replication, not for replicating arbitrary EBS volume data; it would require a database engine and continuous replication infrastructure, increasing cost and complexity unnecessarily. Option D is wrong because S3 Cross-Region Replication cannot directly replicate EBS volume data; EBS volumes are block storage, not object storage, and S3 CRR only works with S3 objects, so this approach is technically infeasible.

1148
MCQmedium

A company uses an Amazon DynamoDB table with on-demand capacity mode. The table handles a workload with a steady baseline of 500 writes per second but spikes to 2,000 writes per second for a few hours each day. The SysOps administrator wants to reduce costs without affecting application performance during spikes. Which action should the administrator take?

A.Switch to provisioned capacity with auto scaling to handle the spikes.
B.Enable DynamoDB Accelerator (DAX) to cache reads.
C.Create a global table to distribute write traffic.
D.Use Amazon ElastiCache to buffer write requests.
AnswerA

For a workload with a predictable baseline and only occasional spikes, on-demand capacity's per-request pricing leads to unnecessary cost during the sustained baseline. Provisioned capacity bills a fixed hourly rate per read/write capacity unit, and DynamoDB auto scaling adjusts the provisioned units based on target utilization (e.g., 70%), so you pay for the baseline plus a small margin instead of paying each time a request occurs. This approach significantly reduces costs when the baseline traffic is steady, as the auto-scaling can scale up during known spike windows and scale down afterward. Because the question highlights cost reduction for recurring spikes, this is the correct action.

Why this answer

On-demand capacity mode is ideal for unpredictable workloads but costs more per write than provisioned capacity. Since this workload has a predictable baseline and spikes, switching to provisioned capacity with auto scaling allows you to pay a lower rate for the steady 500 writes per second while auto scaling automatically adds capacity to handle the 2,000 writes per second spikes, reducing overall cost without impacting performance.

Exam trap

The trap here is that candidates assume on-demand is always the cheapest for spiky workloads, but the question specifies a predictable spike pattern, making provisioned with auto scaling more cost-effective; also, candidates may confuse DAX (read cache) with a write optimization tool.

How to eliminate wrong answers

Option B is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for reads, not writes, and does not reduce write costs or handle write spikes. Option C is wrong because creating a global table replicates writes across regions, increasing write costs and complexity without reducing the cost of the write workload in a single region. Option D is wrong because using Amazon ElastiCache to buffer write requests would introduce latency and complexity, and does not reduce DynamoDB write costs; it would only temporarily hold data before writing to DynamoDB, potentially causing data loss if the cache fails.

1149
MCQmedium

A SysOps administrator notices that an Amazon RDS for MySQL instance has high read activity. The application performs many read queries but few writes. The database is currently a single db.r5.large instance. What change would improve read performance and reduce cost?

A.Enable Multi-AZ for high availability and use the standby for reads.
B.Increase the instance size to db.r5.xlarge.
C.Migrate the database to Amazon DynamoDB.
D.Create a read replica and use a smaller instance type for the replica.
AnswerD

Creating a read replica allows you to offload read traffic to a separate, dedicated instance, freeing the primary to handle writes. You can specify a smaller instance type for the replica, such as a db.t3.medium, because read replicas do not need to match the primary's size; this reduces cost while still providing the read scaling needed. This directly addresses the read-heavy workload without migrating or overprovisioning the primary.

Why this answer

Creating a read replica offloads read traffic from the primary instance, improving performance for read-heavy workloads. Using a smaller instance type for the replica reduces cost compared to scaling up the primary. Option A is wrong because Multi-AZ provides high availability via a standby instance that cannot serve reads; it does not improve read performance.

Option B (increasing instance size) is a vertical scaling approach that increases cost without the same cost-efficiency as a read replica. Option C is incorrect because migrating to DynamoDB would require application changes and may not be compatible with the existing MySQL schema and queries.

1150
MCQhard

A SysOps administrator is troubleshooting an issue where an EC2 instance's CloudWatch agent is not sending memory metrics. The agent is installed and configured to collect memory metrics. The IAM role attached to the instance has the following policy: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "cloudwatch:PutMetricData", "Resource": "*" }, { "Effect": "Allow", "Action": "cloudwatch:ListMetrics", "Resource": "*" } ] } What is the most likely reason the memory metrics are not appearing?

A.The CloudWatch agent requires the SSM Agent to be installed.
B.The CloudWatch agent must be configured to send metrics to CloudWatch Logs first.
C.The IAM role is missing the ssm:GetParameter permission.
D.Detailed monitoring must be enabled on the EC2 instance.
AnswerC

This is the root cause because the CloudWatch agent is configured to retrieve its metrics configuration from the SSM Parameter Store. To read that configuration, the EC2 instance's IAM role must include the ssm:GetParameter permission; without it, the agent cannot fetch the JSON that defines which memory metrics to collect. Even though the role includes cloudwatch:PutMetricData, the agent has no configuration to act on, so no memory metrics are published.

Why this answer

The IAM policy includes cloudwatch:PutMetricData and cloudwatch:ListMetrics, but if the CloudWatch agent retrieves its memory metrics configuration from Systems Manager Parameter Store (a common setup), the instance requires the ssm:GetParameter permission. Without this permission, the agent cannot fetch the configuration, and memory metrics will not be sent. Therefore, option C correctly identifies the missing IAM permission.

Exam trap

Candidates often focus solely on the cloudwatch:PutMetricData permission and overlook other required permissions like ssm:GetParameter. The policy already includes PutMetricData, so option C is a distractor if not read carefully; however, the real issue is the missing ssm:GetParameter permission needed for configuration retrieval from Parameter Store.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent does not require the SSM Agent to be installed; the CloudWatch agent can run independently and communicate directly with CloudWatch via HTTPS. Option B is wrong because the CloudWatch agent sends metrics directly to CloudWatch Metrics, not to CloudWatch Logs first; memory metrics are custom metrics, not log data. Option D is wrong because detailed monitoring on the EC2 instance only enables 1-minute frequency for hypervisor-level metrics (CPU, network, disk), not memory metrics; memory metrics are collected by the CloudWatch agent and require the agent to be running and properly configured.

1151
Multi-Selectmedium

A SysOps administrator is troubleshooting DNS resolution issues for a custom domain used by an Application Load Balancer. Which TWO steps should the administrator take to diagnose the issue? (Choose two.)

Select 2 answers
A.Ensure the VPC's CIDR block does not overlap with the ALB's IP range
B.Verify that the Route 53 alias record points to the ALB's DNS name
C.Run 'dig' or 'nslookup' from a client to verify the domain resolves to the correct IP
D.Verify that the ALB's security group allows inbound traffic on port 443
E.Check the health status of the ALB's target group
AnswersB, C

The Route 53 alias record must reference the ALB's canonical DNS name (e.g., myapp-1234567890.us-east-1.elb.amazonaws.com) rather than a static IP address, because ALB IPs are ephemeral and can change during scale operations. If the alias target is misspelled, points to a deleted resource, or uses a non-alias type with an IP, Route 53 returns no valid answer or a stale address. Verifying this alias configuration directly corrects the misconfiguration that causes resolution failures.

Why this answer

A Route 53 alias record must point to the ALB's DNS name (e.g., my-alb-1234567890.us-east-1.elb.amazonaws.com) to properly route traffic to the load balancer. If the alias record is misconfigured or points to an incorrect resource, DNS resolution will fail or resolve to an unintended IP, causing the custom domain not to work.

Exam trap

The trap here is that candidates confuse DNS resolution issues with network connectivity or load balancer health, leading them to select security group or target group checks instead of focusing on the DNS configuration itself.

1152
MCQmedium

Multiple microservices each write structured JSON logs to separate CloudWatch log groups. The operations team needs to find all ERROR-level log entries across all log groups for the past 24 hours and count errors by service name. Which approach achieves this with the least operational overhead?

A.Run a CloudWatch Logs Insights query selecting all relevant log groups, filter where level = 'ERROR', and use stats count(*) by service
B.Export each log group to S3 and run an Athena query joining all exported files
C.Subscribe all log groups to a Kinesis Data Firehose stream and query the aggregated data in OpenSearch
D.Use the AWS CLI to download and grep log events from each log group separately, then sum the results
AnswerA

Logs Insights accepts a comma-separated list of log group names (or a log group name prefix pattern) in the query scope. The filter and stats commands work across all selected groups in a single query execution. No additional pipeline or aggregation layer is needed.

Why this answer

CloudWatch Logs Insights natively supports querying multiple log groups in a single query. By specifying all relevant log groups in the query scope, filtering for `level = 'ERROR'` using the `filter` command, and using `stats count(*) by service`, the operations team can directly aggregate error counts per service without any data movement, additional infrastructure, or manual scripting. This approach has the least operational overhead because it leverages existing CloudWatch capabilities with no setup or maintenance.

Exam trap

The trap here is that candidates may overcomplicate the solution by assuming cross-log-group analysis requires data aggregation pipelines (like Kinesis or S3/Athena), when CloudWatch Logs Insights natively supports querying multiple log groups with a single query, making it the simplest and most cost-effective option.

How to eliminate wrong answers

Option B is wrong because exporting logs to S3 and querying with Athena introduces significant operational overhead: you must set up S3 buckets, configure export schedules (which can take hours for large volumes), and manage Athena table definitions and partitions, all of which are unnecessary when CloudWatch Logs Insights can query the same data directly. Option C is wrong because subscribing all log groups to a Kinesis Data Firehose stream and querying in OpenSearch requires provisioning and managing a Firehose delivery stream, an OpenSearch cluster, and index management, which adds complexity and cost far beyond the simple Insights query. Option D is wrong because using the AWS CLI to download and grep log events from each log group separately is manual, error-prone, and does not scale; it requires scripting to handle pagination, rate limits, and log group enumeration, and it lacks the built-in aggregation and filtering capabilities of CloudWatch Logs Insights.

1153
MCQhard

A SysOps administrator wants to automate the deployment of an application to an EC2 instance. The instance is running, but the deployment script fails because the instance is not reachable via SSH. The administrator checks the instance state as shown in the exhibit. What should the administrator check NEXT to troubleshoot the SSH connectivity issue?

A.Check the security group rules to ensure SSH (port 22) is allowed from the administrator's IP.
B.Verify the instance is in a public subnet with an internet gateway.
C.Verify the instance ID is correct.
D.Check if the instance is in a stopped state.
AnswerA

Security groups act as a stateful virtual firewall at the instance level. If no inbound rule permits TCP 22 from your public IP address, all SSH packets are silently dropped, causing the connection to time out even when the instance is running, has a public IP, and is in a public subnet. The correct remediation is to add a rule that allows SSH (port 22) from your specific IP or CIDR block.

Why this answer

When an EC2 instance is running but SSH is unreachable, the most common cause is that the security group does not permit inbound TCP/22 from the administrator's source IP. Security groups are stateful virtual firewalls evaluated before traffic reaches the instance, so a missing or overly restrictive SSH rule silently drops the connection attempt. Checking the security group's inbound rules is the correct next troubleshooting step after confirming the instance is running.

Exam trap

SOA-C02 often tests the instinct to jump to subnet/IGW checks — the trap is overlooking that security groups are the first and most common blocker for SSH, and candidates pick the 'network architecture' answer when the simpler firewall rule check is correct.

How to eliminate wrong answers

Option B is wrong because although a public subnet with an internet gateway is required for direct SSH from the internet, the exhibit already shows the instance is running and the administrator is troubleshooting connectivity — verifying subnet/IGW is a secondary check and not the most likely cause when the instance is otherwise reachable. Option C is wrong because verifying the instance ID is a basic sanity check that would have surfaced immediately; an incorrect instance ID would mean the administrator is looking at the wrong resource entirely, not a connectivity problem. Option D is wrong because the question states the instance is running, so checking for a stopped state contradicts the given scenario and wastes a troubleshooting step.

1154
MCQeasy

A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which configuration should be used?

A.Use AWS Lambda to periodically check instance health and reboot if necessary.
B.Enable termination protection on the instance.
C.Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
D.Place the instance in an Auto Scaling group with a minimum size of 1.
AnswerC

A CloudWatch alarm on the StatusCheckFailed metric can be configured with the EC2 recovery action, which automatically stops and starts the instance on a different physical host when the underlying hardware or network is impaired. This recovery process preserves the instance ID, private IP address, Elastic IP, and instance metadata, and the EBS root volume remains attached with its existing data. Because the action moves the instance to healthy hardware while maintaining its identity and configuration, it is the correct mechanism to recover from a physical host failure without requiring manual intervention.

Why this answer

A CloudWatch alarm on the StatusCheckFailed metric can be configured with an EC2 recovery action. When the alarm triggers (e.g., due to an underlying hardware failure), the recovery action automatically stops and starts the instance on healthy hardware, preserving the instance ID, private IP, Elastic IP, and EBS attachments. This is the native AWS mechanism for automatic instance recovery from hardware failures.

Exam trap

The trap here is that candidates often confuse termination protection (which only prevents deletion) or Auto Scaling replacement (which creates a new instance) with the native recovery action that preserves the instance's identity and state.

How to eliminate wrong answers

Option A is wrong because using AWS Lambda to periodically check instance health and reboot is an unnecessary, custom workaround that adds complexity and latency; AWS already provides the built-in CloudWatch alarm recovery action for this purpose. Option B is wrong because termination protection only prevents accidental deletion of an instance via the console or API; it does not detect or recover from hardware failures. Option D is wrong because placing the instance in an Auto Scaling group with a minimum size of 1 will replace a failed instance with a new one (different instance ID, IP, and metadata), which does not preserve the original instance's identity and attached resources like the recovery action does.

1155
MCQeasy

A SysOps administrator wants to be alerted when an EC2 instance is terminated unexpectedly. Which CloudWatch event should be used to trigger a notification?

A.A CloudWatch alarm on the CPUUtilization metric dropping to zero.
B.A CloudTrail trail that logs TerminateInstances API calls.
C.A CloudWatch alarm on the StatusCheckFailed metric.
D.An Amazon EventBridge rule that matches EC2 Instance State-change Notification events.
AnswerD

An EventBridge rule can match the EC2 Instance State-change Notification event, which is emitted synchronously as an instance transitions among states such as pending, running, stopping, stopped, and terminated. By configuring an event pattern for detail.state = "terminated" and targeting an SNS topic or Lambda function, you receive a near-real-time notification that is the native, purpose-built mechanism for alerting on lifecycle changes.

Why this answer

Amazon EventBridge can capture EC2 Instance State-change Notification events, which are emitted whenever an EC2 instance transitions between states (e.g., running, stopped, terminated). By creating a rule that matches the 'terminated' state, the administrator can trigger an SNS notification or Lambda function to alert on unexpected termination, providing a real-time, event-driven response.

Exam trap

The trap here is that candidates confuse CloudWatch alarms on metrics (like CPUUtilization or StatusCheckFailed) with event-driven notifications, overlooking that EventBridge rules directly capture state-change events for immediate, precise alerting.

How to eliminate wrong answers

Option A is wrong because a CloudWatch alarm on CPUUtilization dropping to zero is not a reliable indicator of termination; an instance could be idle or stopped without being terminated, and the alarm would not fire immediately upon termination. Option B is wrong because a CloudTrail trail logging TerminateInstances API calls records the API action but does not directly trigger a notification; it requires additional integration (e.g., CloudWatch Logs metric filter or EventBridge rule) to generate alerts, and it only captures API-initiated terminations, not those from Auto Scaling or AWS Health events. Option C is wrong because a CloudWatch alarm on StatusCheckFailed monitors system or instance status checks (e.g., OS-level issues), not termination events; an instance can fail status checks without being terminated, and termination does not necessarily cause a status check failure.

1156
Matchingmedium

Match each AWS backup and disaster recovery service to its feature.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Centralized backup management

Automatic object replication across regions

High availability with standby replica

Read scaling and cross-region disaster recovery

Continuous replication for DR

Why these pairings

The correct matches are: AWS Backup for centralized automation, AWS Elastic Disaster Recovery for continuous replication, AWS Storage Gateway for hybrid cloud storage, and AWS Snowball for offline transfer. Common confusions involve swapping the centralized backup and replication functions.

1157
MCQmedium

A SysOps administrator needs to reduce costs for an Amazon RDS for MySQL DB instance that is used for development. The instance is only needed during business hours (9 AM to 5 PM) on weekdays. Which solution is the MOST cost-effective while maintaining the ability to start and stop the instance on a schedule?

A.Create a read replica and promote it when needed.
B.Purchase a reserved instance for the DB instance.
C.Use a Lambda function to stop the instance at 5 PM and start it at 9 AM on weekdays.
D.Use AWS Instance Scheduler to stop and start the RDS instance on a schedule.
AnswerD

AWS Instance Scheduler is a purpose-built solution that uses a CloudFormation stack and tags to automatically start and stop RDS instances on a defined schedule. It leverages Lambda functions and an Amazon DynamoDB state table to track and apply the schedule without requiring you to build a custom scheduler. This reduces compute charges during off hours while storage and backup costs continue, precisely matching the requirement for a business-hours-only database.

Why this answer

AWS Instance Scheduler is a fully managed solution designed specifically to start and stop RDS instances on a defined schedule, such as weekdays from 9 AM to 5 PM. This eliminates compute costs during non-business hours while preserving the instance's storage and configuration, making it the most cost-effective and reliable approach for a development environment.

Exam trap

The trap here is that candidates often choose a Lambda function (Option C) thinking it is the most flexible, but they overlook that AWS Instance Scheduler is a pre-built, managed solution that reduces operational overhead and is the recommended pattern for scheduled start/stop of RDS instances in the AWS Well-Architected Framework.

How to eliminate wrong answers

Option A is wrong because creating a read replica and promoting it does not reduce costs; it incurs additional charges for the replica instance and storage, and the original instance remains running. Option B is wrong because purchasing a reserved instance commits to a 1- or 3-year term, which is not cost-effective for a development instance that is only used part-time and does not align with the need to stop and start on a schedule. Option C is wrong because while a Lambda function can stop and start an RDS instance, it requires custom code, IAM roles, and CloudWatch Events or EventBridge rules to trigger the schedule, which adds complexity and maintenance overhead compared to the purpose-built AWS Instance Scheduler.

1158
MCQhard

A SysOps administrator is troubleshooting an issue where an Amazon RDS DB instance's storage space is running out. The administrator has enabled CloudWatch alarms for FreeStorageSpace, but the alarm did not trigger before the storage was exhausted. What is the most likely reason?

A.The FreeStorageSpace metric is not available for the selected DB instance class.
B.The alarm was configured to use a static threshold but the metric is not emitted during storage operations.
C.The alarm's evaluation period was too long and the storage filled up faster than the alarm could trigger.
D.The alarm was monitoring the wrong metric, such as 'Storage' instead of 'FreeStorageSpace'.
AnswerC

An evaluation period (the number of consecutive 60-second periods the metric must breach the threshold before the alarm enters ALARM state) lengthens the required detection time. If the database's storage fills up completely within, say, 15 minutes, an alarm configured with 10 evaluation periods at 5-minute periods would need 50 minutes of data and thus never fire before the instance goes read-only. Shortening the evaluation period and using a threshold well above zero (e.g., 10% free space) avoids this pitfall.

Why this answer

CloudWatch alarms evaluate metrics based on a specified evaluation period (e.g., 5 minutes). If the storage fills up faster than the alarm's evaluation period, the alarm may not have enough data points to trigger before the storage is exhausted. This is a common issue when the rate of storage consumption exceeds the alarm's evaluation frequency.

Exam trap

The trap here is that candidates assume CloudWatch alarms trigger instantly when a metric crosses a threshold, but in reality, alarms require multiple data points over the evaluation period to change state, which can delay detection if storage fills rapidly.

How to eliminate wrong answers

Option A is wrong because FreeStorageSpace is a standard metric available for all RDS DB instance classes, including those with General Purpose (gp2/gp3) or Provisioned IOPS (io1/io2) storage. Option B is wrong because the FreeStorageSpace metric is emitted continuously during storage operations, regardless of whether the alarm uses a static threshold or anomaly detection. Option D is wrong because 'Storage' is not a valid CloudWatch metric name for RDS; the correct metric is FreeStorageSpace, and monitoring the wrong metric would not cause the alarm to fail to trigger—it would simply not reflect storage exhaustion.

1159
MCQhard

A company has a VPC with public and private subnets. The private subnets contain RDS databases that should not be accessible from the internet. Which configuration ensures that the databases are only accessible from the application servers in the public subnets?

A.Attach an internet gateway to the VPC and route the private subnet's traffic to it.
B.Attach a NAT gateway to the private subnet and route traffic through it.
C.Configure a network ACL on the private subnet to allow inbound traffic from the public subnet CIDR.
D.Create a security group for the RDS instances that allows inbound traffic from the security group attached to the application servers.
AnswerD

Create a security group for the RDS instances and add an inbound rule that allows the database port from the security group attached to the application servers. Because security groups can reference other security groups as sources, this rule automatically restricts access to only those instances that carry the application security group, regardless of their IP addresses or whether the application tier scales up or down. This stateful, least-privilege approach is the AWS best practice for tiered access within a VPC and directly resolves the connectivity failure without affecting other subnets or relying on broad CIDR ranges.

Why this answer

Security groups act as a virtual firewall at the instance level, and you can reference another security group as a source. By creating a security group for the RDS instances that allows inbound traffic from the security group attached to the application servers, you ensure that only those application servers (regardless of their IP addresses) can reach the databases. This approach is more dynamic and secure than using CIDR-based rules, as it automatically accommodates changes in the application servers' IP addresses or scaling events.

Exam trap

The trap here is that candidates often confuse security groups with network ACLs, assuming that a network ACL rule allowing inbound traffic from the public subnet CIDR is sufficient, but they overlook that network ACLs are stateless and do not provide the same granular, instance-level control as security groups, nor do they automatically adapt to changes in application server IPs.

How to eliminate wrong answers

Option A is wrong because attaching an internet gateway to the VPC and routing private subnet traffic to it would expose the RDS databases to the internet, violating the requirement that they should not be accessible from the internet. Option B is wrong because a NAT gateway is used to allow outbound internet traffic from private subnets, not to control inbound access; it would not restrict inbound traffic to only the application servers. Option C is wrong because a network ACL is stateless and requires explicit allow rules for both inbound and outbound traffic; while it could allow inbound traffic from the public subnet CIDR, it would not restrict access to only the application servers (any instance in that CIDR range could connect), and it would not automatically adapt to changes in application server IPs.

1160
MCQeasy

A company wants to centrally collect and analyze logs from all AWS accounts in an organization. The logs include CloudTrail, VPC Flow Logs, and AWS Config logs. Which solution is the most scalable and cost-effective?

A.Stream all logs to a central CloudWatch Logs account using cross-account subscriptions.
B.Use CloudWatch Logs Insights to query logs from each account individually.
C.Use Amazon Kinesis Data Firehose to deliver logs to an Amazon Elasticsearch Service cluster.
D.Configure each account to deliver logs to a centralized S3 bucket and use Amazon Athena to query them.
AnswerD

This option centralizes logs by having each account deliver log files to a single S3 bucket—either directly or via S3 Cross-Account Replication—and then uses Amazon Athena to run SQL queries across that centralized data lake. S3 offers cheap, durable, and scalable storage for large volumes of logs, while Athena is serverless and charges only for the data actually scanned per query. By partitioning the S3 paths by account, region, and date and using AWS Glue Data Catalog (or a crawler), analysts can run cross-account queries from one place without managing any query infrastructure, making it the most cost-effective and operationally simple solution.

Why this answer

It uses a centralized S3 bucket to aggregate logs from all accounts, which is highly scalable and cost-effective due to S3's low storage costs and lifecycle policies. Amazon Athena then allows serverless, pay-per-query analysis of the logs without needing to provision or manage any infrastructure, making it ideal for ad-hoc and cross-account log analysis.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing managed services like CloudWatch Logs or Elasticsearch, overlooking the simplicity, scalability, and cost-effectiveness of S3 + Athena for centralized log analysis across multiple accounts.

How to eliminate wrong answers

Option A is wrong because streaming all logs to a central CloudWatch Logs account via cross-account subscriptions incurs high ingestion and storage costs, and CloudWatch Logs is not designed for long-term, cost-effective storage of large volumes of logs from multiple accounts. Option B is wrong because CloudWatch Logs Insights can only query logs within a single account and cannot aggregate or query logs across multiple accounts, failing the centralization requirement. Option C is wrong because using Amazon Kinesis Data Firehose to deliver logs to an Amazon Elasticsearch Service cluster introduces significant operational overhead for managing the Elasticsearch cluster, and the cost scales with the volume of data indexed and stored, making it less cost-effective than S3 + Athena for infrequent or ad-hoc queries.

1161
MCQeasy

An organization wants to ensure that no Amazon S3 bucket in the entire AWS Organization can be made public. The security team requires a preventive control that cannot be overridden by individual account administrators. Which AWS service or feature should be used?

A.Create a Service Control Policy (SCP) in AWS Organizations that denies permissions to modify S3 bucket public access settings.
B.Enable AWS Config rules in each account to detect public S3 buckets and automatically remediate them using AWS Lambda.
C.Use an IAM policy attached to all IAM users in each account that denies s3:PutBucketPolicy.
D.Apply Amazon S3 Block Public Access at the account level in each individual AWS account.
AnswerA

A Service Control Policy (SCP) attached at the organization root or an organizational unit (OU) is inherited by every AWS account underneath, and it operates as an allow-list or denial of AWS API actions at the account level. Because SCPs are evaluated by AWS Organizations before IAM policies, even an account root user with full administrative rights cannot override an explicit deny of s3:PutBucketPolicy, s3:PutBucketAcl, or s3:PutBucketPublicAccessBlock, making it a true preventative guardrail across the entire organization. This is why the correct answer is to use SCPs rather than account-local controls.

Why this answer

A Service Control Policy (SCP) in AWS Organizations is a preventive guard that applies to all accounts within the organization. It can explicitly deny actions like s3:PutBucketPublicAccessBlock, s3:PutBucketPolicy, and s3:PutObjectAcl, preventing any principal (including root users) from making S3 buckets public. Unlike detective or account-level controls, SCPs cannot be overridden by individual account administrators, meeting the requirement for a non-overridable preventive control.

Exam trap

The trap here is that candidates often choose account-level S3 Block Public Access (Option D) because it seems like a direct preventive control, but they overlook that it can be overridden by account administrators, whereas an SCP is a centralized, non-overridable guardrail that applies across the entire AWS Organization.

How to eliminate wrong answers

Option B is wrong because AWS Config rules are detective and reactive, not preventive; they detect public buckets after the fact and can auto-remediate, but they do not block the initial action and can be overridden by account administrators. Option C is wrong because IAM policies attached to users do not apply to the root user or to services running with assumed roles, and they can be modified by account administrators, so they are not a non-overridable preventive control across the entire organization. Option D is wrong because S3 Block Public Access at the account level can be disabled or modified by any user with the necessary permissions (including account administrators), so it does not provide a centrally enforced, non-overridable control.

1162
MCQeasy

A company uses Amazon RDS for MySQL and wants to monitor the number of database connections in real time. Which CloudWatch metric should the SysOps administrator use?

A.DatabaseConnections
B.CPUUtilization
C.ActiveConnections
D.ConnectionCount
AnswerA

DatabaseConnections is the correct standard CloudWatch metric because Amazon RDS publishes it in the AWS/RDS namespace as a gauge that reports the number of client network connections to the DB instance. It includes all active sessions from applications, readers, and management tools, and it is the authoritative counter to compare against the instance's max_connections, which varies by instance class and DB engine. This metric directly answers the question of how many concurrent connections are present.

Why this answer

Amazon RDS for MySQL exposes a CloudWatch metric named `DatabaseConnections` that reports the number of client connections to the DB instance. This metric is derived from the MySQL `Threads_connected` status variable and is updated every minute, providing real-time visibility into connection counts for monitoring and alarming.

Exam trap

The trap here is that candidates confuse the MySQL status variable `Threads_connected` (which maps to `DatabaseConnections`) with `Threads_running` (which counts only actively executing queries) or assume a generic name like `ActiveConnections` or `ConnectionCount` exists as a CloudWatch metric.

How to eliminate wrong answers

Option B is wrong because `CPUUtilization` measures the percentage of CPU used by the DB instance, not the number of database connections. Option C is wrong because `ActiveConnections` is not a standard CloudWatch metric for RDS; it may be confused with a MySQL status variable (`Threads_running`) but is not exposed as a CloudWatch metric. Option D is wrong because `ConnectionCount` is not a valid CloudWatch metric name for RDS; the correct metric is `DatabaseConnections`.

1163
Multi-Selecthard

Which THREE of the following are valid options for connecting a VPC to an on-premises network? (Select THREE.)

Select 3 answers
A.Transit gateway with VPN attachment
B.AWS Direct Connect
C.VPC peering
D.AWS Site-to-Site VPN
E.VPC endpoint
AnswersA, B, D

An AWS Transit Gateway acts as a network transit hub, and adding a Site-to-Site VPN attachment lets an on-premises network connect to the gateway over an IPsec tunnel using the public internet. This is a valid hybrid connectivity option because the transit gateway can route traffic between many VPCs and the VPN attachment, simplifying the network while still supporting site-to-site VPN connectivity.

Why this answer

A Transit Gateway with a VPN attachment allows you to connect your VPC to an on-premises network by acting as a central hub that interconnects VPCs and on-premises networks via IPsec VPN tunnels. This is a valid option because the Transit Gateway can terminate multiple VPN connections, enabling hybrid connectivity with centralized routing and scalability.

Exam trap

The trap here is that candidates confuse VPC peering (which only connects VPCs) with hybrid connectivity options, or mistakenly think VPC endpoints can extend to on-premises networks, when they are strictly for accessing AWS services privately within a VPC.

1164
Multi-Selectmedium

A company has an S3 bucket that stores sensitive data. The security team requires that all data be encrypted at rest and that all access be logged. Which TWO actions should the SysOps administrator take to meet these requirements? (Choose TWO.)

Select 2 answers
A.Enable S3 Transfer Acceleration.
B.Enable S3 Replication to replicate objects to another bucket.
C.Enable default encryption on the S3 bucket.
D.Enable S3 Object Lock.
E.Enable S3 server access logs.
AnswersC, E

Enabling default encryption on the bucket ensures that every new object stored is automatically encrypted at rest using SSE-S3 (AES-256) or an SSE-KMS key, even when the upload request does not specify an encryption header. This protects sensitive data from physical media theft or unauthorized access to underlying storage infrastructure. It is a fundamental confidentiality control that should be paired with access logging and IAM policies to fully secure the data.

Why this answer

Option C is correct because enabling default encryption on the S3 bucket (using SSE-S3 or SSE-KMS) ensures that every object is automatically encrypted at rest when written, satisfying the requirement that all stored data be encrypted. Option E is correct because enabling S3 server access logs delivers detailed records of every request made to the bucket to a target logging bucket, which fulfills the requirement that all access be logged. Option A is incorrect because S3 Transfer Acceleration only speeds up uploads/downloads over long distances using edge locations; it does not provide encryption or logging.

Option B is incorrect because S3 Replication copies objects to another bucket for durability, latency, or compliance purposes but does not itself encrypt data at rest or log access. Option D is incorrect because S3 Object Lock enforces WORM protection to prevent deletion or modification of objects, which is unrelated to encryption at rest or access logging.

Exam trap

SOA-C02 often tests whether candidates confuse adjacent S3 features (Transfer Acceleration, Replication, Object Lock) with the specific controls for encryption at rest and access logging, causing them to pick a feature that addresses a different requirement.

1165
MCQmedium

A company runs a web application on Amazon EC2 instances that are part of an Auto Scaling group. The application's traffic is predictable with regular peaks during business hours and low traffic at night. The SysOps administrator wants to optimize costs while ensuring that performance meets demand. The administrator also needs to minimize manual intervention. Which scaling policy should be used?

A.Scheduled scaling
B.Target tracking scaling
C.Simple scaling
D.Manual scaling
AnswerA

Scheduled scaling allows you to define a recurring or one-time schedule to adjust the desired capacity of an Auto Scaling group at a future time. Because the company's workload follows a predictable pattern (e.g., a morning peak), scheduled scaling can proactively add EC2 instances before the traffic arrives, eliminating the lag inherent in reactive methods. After the initial configuration, it runs automatically without manual intervention, making it the most efficient choice for a known, consistent demand curve.

Why this answer

Scheduled scaling is the correct choice because the traffic pattern is predictable with regular peaks during business hours and low traffic at night. This policy allows the administrator to define specific times to increase or decrease the desired capacity of the Auto Scaling group, matching capacity to demand without manual intervention and optimizing costs by reducing instances during off-peak hours.

Exam trap

The trap here is that candidates often confuse target tracking scaling with scheduled scaling, assuming dynamic metric-based policies are always optimal, but for predictable patterns, scheduled scaling provides more precise cost control and avoids unnecessary scaling events.

How to eliminate wrong answers

Option B is wrong because target tracking scaling adjusts capacity dynamically based on a real-time metric (e.g., CPU utilization) and is designed for unpredictable or variable traffic patterns, not for a predictable schedule. Option C is wrong because simple scaling requires manual definition of alarms and cooldown periods, and it does not handle predictable time-based changes efficiently, often leading to over-provisioning or under-provisioning. Option D is wrong because manual scaling requires direct human action to change the desired capacity, which contradicts the requirement to minimize manual intervention.

1166
MCQeasy

A SysOps administrator needs to ensure that an Amazon S3 bucket can withstand the loss of an entire AWS Availability Zone. What is the SIMPLEST configuration to meet this requirement?

A.Enable cross-region replication to a bucket in another Region.
B.Use S3 Standard storage class.
C.Use S3 One Zone-IA storage class.
D.Enable MFA Delete on the bucket.
AnswerB

The S3 Standard storage class is the correct choice because it automatically stores each object redundantly across a minimum of three Availability Zones in the same AWS Region. S3 Standard is engineered for 99.999999999% (11 nines) object durability and 99.99% availability, so if one Availability Zone becomes unavailable, S3 can continue serving reads and writes from the remaining AZs without any administrative action. This built-in synchronous replication across AZs is precisely what satisfies the requirement for AZ failure resilience.

Why this answer

S3 Standard automatically stores objects across a minimum of three Availability Zones (AZs) within the same AWS Region. This design ensures that the bucket can withstand the loss of an entire AZ without any additional configuration, making it the simplest option to meet the requirement.

Exam trap

The trap here is that candidates often overthink and choose cross-region replication (Option A) for high availability, but the question specifically asks for resilience against an AZ loss, which S3 Standard already provides within a single Region without the complexity and cost of CRR.

How to eliminate wrong answers

Option A is wrong because cross-region replication (CRR) adds complexity and cost, and is not the simplest way to withstand an AZ loss—S3 Standard already provides AZ resilience within a single Region. Option C is wrong because S3 One Zone-IA stores data in only a single AZ, so the loss of that AZ would result in permanent data loss, failing the requirement. Option D is wrong because MFA Delete is a security feature that adds extra authentication for delete operations; it does not provide any data durability or AZ resilience.

1167
MCQmedium

A company runs a production web application on a single Amazon EC2 instance. The application experiences a predictable and steady workload 24/7. The SysOps administrator wants to minimize compute costs for this instance while ensuring it remains available during the expected workload. Which EC2 purchasing option should the administrator use?

A.On-Demand Instances
B.Reserved Instances
C.Spot Instances
D.Dedicated Hosts
AnswerB

Reserved Instances offer a significant hourly discount in exchange for a one- or three-year commitment, and a Standard RI is best suited for steady-state, predictable production workloads like this always-on web app. By paying all or part of the cost upfront, you can reduce the effective hourly price by up to 72% compared to On-Demand. Because the workload is constant, the utilization will easily justify the commitment, making this the optimal cost-optimization strategy.

Why this answer

Reserved Instances (RIs) are the most cost-effective option for a predictable, steady-state workload running 24/7. By committing to a 1- or 3-year term, you receive a significant discount (up to 72%) compared to On-Demand pricing, while still ensuring the instance remains available for the expected workload. This matches the requirement to minimize compute costs without sacrificing availability.

Exam trap

The trap here is that candidates often choose Spot Instances for cost savings, overlooking the critical requirement of 'remaining available during the expected workload' — Spot Instances can be interrupted at any time, making them unsuitable for production workloads that need consistent availability.

How to eliminate wrong answers

Option A (On-Demand Instances) is wrong because, while they provide full availability, they are the most expensive option for a steady 24/7 workload and do not minimize costs. Option C (Spot Instances) is wrong because they can be terminated by AWS with a 2-minute notification when capacity is reclaimed, making them unsuitable for a production web application that must remain available during the expected workload. Option D (Dedicated Hosts) is wrong because they are designed for regulatory or licensing requirements (e.g., per-socket or per-core licensing) and are significantly more expensive than Reserved Instances, offering no cost benefit for a standard single-instance workload.

1168
MCQhard

A SysOps administrator is investigating why a CloudWatch alarm did not trigger an SNS notification. The alarm state changed to ALARM, but the notification was not sent. The SNS topic has a subscription to an email endpoint. What is the most likely cause?

A.The alarm's evaluation period is too short.
B.The email subscription to the SNS topic has not been confirmed.
C.The alarm's actions are not configured to send to the SNS topic.
D.The SNS topic is encrypted with a KMS key that the alarm does not have permissions to use.
AnswerB

For SNS email subscriptions, the subscriber must confirm the subscription by clicking the link in the initial confirmation message that Amazon SNS sends to the email address. Until that confirmation is completed, the subscription remains in the 'Pending Confirmation' state and Amazon SNS will not deliver messages published to the topic, even if the CloudWatch alarm successfully executes its action. This is a common and often overlooked cause of missing email notifications, and it directly explains the symptom described in the question.

Why this answer

The most likely cause is that the email subscription to the SNS topic has not been confirmed. When an SNS topic has an email subscription, AWS sends a confirmation email to the endpoint, and the subscriber must click the confirmation link before notifications can be delivered. Until the subscription is confirmed, the SNS topic will not send any messages to that endpoint, even if the CloudWatch alarm enters the ALARM state and successfully publishes to the topic.

Exam trap

The trap here is that candidates assume configuring the alarm to publish to an SNS topic is sufficient, overlooking the mandatory subscription confirmation step for email endpoints, which is a distinct and separate requirement.

How to eliminate wrong answers

Option A is wrong because the evaluation period affects how many consecutive data points must breach the threshold before the alarm state changes, but the alarm did change to ALARM, so the evaluation period is not the issue. Option C is wrong because if the alarm's actions were not configured to send to the SNS topic, the alarm would not have published to the topic at all, but the question states the alarm state changed to ALARM, implying the action was configured; the problem is on the subscription side. Option D is wrong because if the SNS topic were encrypted with a KMS key that the alarm did not have permissions to use, the alarm would fail to publish to the topic, and the alarm state would not change to ALARM; the alarm successfully published, so KMS permissions are not the issue.

1169
MCQmedium

A company uses AWS Organizations with multiple member accounts. The SysOps administrator needs to deploy a common AWS CloudFormation template that creates an IAM role across all member accounts in the organization. Which AWS service should be used to deploy this template across accounts?

A.AWS CloudFormation StackSets
B.AWS CodePipeline with cross-account deployment actions
C.AWS CloudFormation cross-stack references
D.AWS Service Catalog
AnswerA

AWS CloudFormation StackSets is the correct answer because it is the native, purpose-built service for deploying the same CloudFormation template across multiple AWS accounts and Regions from a single operation. StackSets uses a delegated administrator account to create stack instances in target accounts with optional automatic deployment and drift detection, and it supports organizational unit (OU) targeting directly through AWS Organizations. This makes it the most efficient and maintainable way to enforce consistent infrastructure, such as security baselines, across an entire organization.

Why this answer

AWS CloudFormation StackSets is the correct service because it extends CloudFormation functionality to deploy templates across multiple accounts and regions from a single management account. StackSets uses a self-managed or service-managed permission model, and with AWS Organizations, it can automatically deploy to all member accounts in the organization or specified organizational units (OUs), making it ideal for deploying a common IAM role across all accounts.

Exam trap

The trap here is that candidates confuse AWS Service Catalog's ability to launch templates in individual accounts with automatic multi-account deployment, overlooking that StackSets is the only service designed for bulk, automated deployment across all organization accounts.

How to eliminate wrong answers

Option B (AWS CodePipeline with cross-account deployment actions) is wrong because CodePipeline orchestrates CI/CD pipelines and, while it can deploy to multiple accounts using cross-account actions, it requires manual setup of each target account and does not natively scale to all member accounts in an organization without additional custom logic. Option C (AWS CloudFormation cross-stack references) is wrong because cross-stack references (using Fn::ImportValue) allow sharing outputs between stacks within the same account or region, not deploying a template across multiple accounts. Option D (AWS Service Catalog) is wrong because Service Catalog enables end users to launch pre-approved products (CloudFormation templates) in their own accounts, but it does not automatically deploy a template across all member accounts; it requires users to provision the product individually.

Page 15

Page 16 of 16