Courseiva

AWS Certified SysOps Administrator Associate SOA-C02 (SOA-C02) — Questions 901–975

1169 questions total · 16pages · All types, answers revealed

Page 12

Page 13 of 16

Page 14
901
MCQeasy

A company uses S3 standard storage for all data. They have data that is accessed rarely but must be retained for 7 years. Which storage class would be MOST cost-effective?

A.S3 One Zone-IA
B.S3 Standard
C.S3 Glacier Deep Archive
D.S3 Intelligent-Tiering
AnswerC

S3 Glacier Deep Archive is the lowest-cost S3 storage class, offering a storage price of about $0.00099 per GB-month and a standard retrieval time of 12 hours or more. It is explicitly designed for long-term retention of data that is accessed at most once a year, such as regulatory and compliance archives. With a 180-day minimum storage duration and a retrieval fee structure, it perfectly matches the requirement for durable, long-term, low-cost archival.

Why this answer

S3 Glacier Deep Archive is the lowest-cost storage for long-term archival. S3 Standard is expensive, S3 IA is for infrequent access but not long-term, and S3 One Zone IA is less durable.

902
MCQmedium

Regulatory requirements mandate that all RDS and EBS backups are replicated to a secondary AWS region within 24 hours of creation. The company has workloads in us-east-1 and must replicate backups to eu-west-1. Restoring from the secondary region must be possible without manual copying steps during a disaster. What service and configuration implements this requirement?

A.Create an AWS Backup plan with a cross-Region copy rule that replicates recovery points to a backup vault in eu-west-1 within 24 hours
B.Schedule a Lambda function that calls CreateDBSnapshot and CopyDBSnapshot to replicate RDS snapshots, and CreateSnapshot and CopySnapshot for EBS volumes to eu-west-1
C.Enable RDS automated backups with cross-region replication and configure EBS snapshot copy separately using Data Lifecycle Manager
D.Use S3 Cross-Region Replication to replicate the backup bucket containing RDS and EBS snapshots to eu-west-1
AnswerA

AWS Backup's cross-Region copy rule runs automatically after each successful backup job. The copy is encrypted with the destination vault's KMS key. In a disaster, operators restore directly from the eu-west-1 vault — no manual cross-region data transfer is needed. A single backup plan can cover multiple resource types (RDS and EBS), satisfying the consolidated requirement.

Why this answer

AWS Backup is the correct service because it natively supports cross-Region copy rules that automatically replicate recovery points (including RDS snapshots and EBS snapshots) to a backup vault in a secondary Region within a specified time window. This meets the 24-hour replication requirement and enables direct restores from the secondary Region without manual copying, as the backup vault in eu-west-1 contains the replicated recovery points ready for use.

Exam trap

The trap here is that candidates often assume they need to use separate services (like Lambda or DLM) for each resource type, missing that AWS Backup provides a unified, managed solution that handles both RDS and EBS snapshots with cross-Region replication and direct restore capabilities.

How to eliminate wrong answers

Option B is wrong because while a Lambda function could technically replicate snapshots, it requires custom code, error handling, and scheduling, and does not provide the native, managed cross-Region restore capability without manual steps; it also lacks the built-in compliance tracking of AWS Backup. Option C is wrong because RDS automated backups with cross-Region replication only apply to RDS, not EBS volumes, and Data Lifecycle Manager (DLM) for EBS snapshots does not support cross-Region copy natively; DLM only copies within the same Region, so EBS snapshots would not be replicated to eu-west-1. Option D is wrong because S3 Cross-Region Replication replicates objects in an S3 bucket, but RDS and EBS snapshots are not stored as S3 objects by default; they are stored in AWS-managed snapshot storage, and even if you manually copy snapshots to S3, the replication would not create usable snapshots in the secondary Region for direct restore.

903
MCQmedium

A company is using AWS CloudFormation to deploy a multi-tier web application. After updating the stack template, the update fails with a stack creation rollback in progress error. The SysOps administrator needs to identify the specific resource that caused the failure. What is the MOST efficient way to accomplish this?

A.Use the AWS Management Console to view the stack status and check the stack policy.
B.Use the aws cloudformation describe-change-set command to review the proposed changes.
C.Check the CloudTrail logs for the UpdateStack API call to see the error message.
D.Run the AWS CLI command aws cloudformation describe-stack-events --stack-name <stack-name> and review the resource status reason.
AnswerD

Running aws cloudformation describe-stack-events --stack-name <stack-name> retrieves every event in the stack's lifecycle, including the most recent update attempt. Each event includes the LogicalResourceId, ResourceStatus (e.g., UPDATE_FAILED), and the ResourceStatusReason field, which contains the specific error message from the underlying AWS service that caused the failure. Reviewing the events in reverse chronological order lets you pinpoint exactly which resource failed and why, making this the definitive troubleshooting command for failed CloudFormation operations.

Why this answer

The describe-stack-events command returns a chronological list of every stack event, including each resource's status and the 'ResourceStatusReason' field that contains the exact error message from the failed resource. This is the fastest, most direct way to pinpoint which resource caused the rollback without digging through unrelated logs.

Exam trap

SOA-C02 often tests the confusion between change sets (which preview proposed changes) and stack events (which record actual execution results) — candidates pick describe-change-set thinking it shows failures, but it only shows what was planned.

How to eliminate wrong answers

Option A is wrong because the console stack status only shows a high-level state (e.g., ROLLBACK_IN_PROGRESS) and the stack policy governs update protections — neither reveals the specific failing resource. Option B is wrong because describe-change-set shows what changes were proposed before execution, not what actually failed during the update. Option C is wrong because CloudTrail records the UpdateStack API call itself but does not surface per-resource failure reasons; you would see the API invocation, not the resource-level error.

904
MCQmedium

A DevOps engineer is designing a CI/CD pipeline for a microservices application. The application consists of several Docker containers that run on Amazon ECS with Fargate launch type. The engineer wants to automate the deployment of new container versions. Which AWS service should be used to orchestrate the build, test, and deployment stages?

A.AWS CodeDeploy
B.AWS CodePipeline
C.AWS CloudFormation
D.AWS CodeBuild
AnswerB

AWS CodePipeline is the correct choice because it is a fully managed continuous delivery service that orchestrates the build, test, and deploy phases of a release process. It lets you model the entire workflow as a series of stages — Source, Build, Test, Deploy — with sequential or parallel actions, and automatically triggers each stage when the previous one succeeds. CodePipeline integrates natively with AWS CodeCommit, CodeBuild, CodeDeploy, Elastic Beanstalk, ECS, and third-party tools like GitHub and Jenkins, making it the central coordinator for a CI/CD pipeline rather than just a tool that executes a single step. Its pipeline structure and transition gates provide the end-to-end automation a DevOps engineer needs for a microservices delivery workflow.

Why this answer

AWS CodePipeline is a fully managed continuous delivery service that orchestrates the build, test, and deploy phases of the release process. Option A is wrong because AWS CodeDeploy is for deploying applications to compute services but does not orchestrate the entire pipeline. Option C is wrong because AWS CloudFormation is for infrastructure as code, not CI/CD orchestration.

Option D is wrong because AWS CodeBuild is for building and testing code, not for orchestrating the entire pipeline.

905
MCQhard

A company runs a critical application on a fleet of EC2 instances that process real-time financial transactions. The application requires consistent low latency. The SysOps administrator notices that the application's latency increases periodically due to noisy neighbors. The administrator wants to optimize performance predictability. Which instance type should the administrator choose?

A.Burstable Performance Instances (T3)
B.Dedicated Instances
C.Spot Instances
D.Reserved Instances
AnswerB

Dedicated Instances run on hardware that is physically isolated from other AWS accounts, meaning there are no 'noisy neighbors' that can monopolize CPU, memory, or I/O resources. This single-tenant environment provides strong performance consistency and can also satisfy strict compliance or licensing requirements. For a critical application running on a fleet, this isolation ensures that the behavior of other customers cannot interfere with your instances, making it the correct choice.

Why this answer

Dedicated Instances run on hardware dedicated to a single customer, eliminating the noisy-neighbor problem because no other AWS accounts share the underlying physical host. This provides the performance predictability required for consistent low-latency financial transaction processing.

Exam trap

SOA-C02 often tests the distinction between tenancy (Dedicated Instances/Hosts) and pricing models (Reserved/Spot) — candidates incorrectly pick Reserved Instances thinking the discount implies isolation, when only dedicated tenancy removes noisy neighbors.

How to eliminate wrong answers

Option A is wrong because T3 burstable instances rely on CPU credits and can be throttled when credits are exhausted, which directly harms latency consistency. Option C is wrong because Spot Instances can be reclaimed by AWS with a two-minute warning, making them unsuitable for a critical always-on application. Option D is wrong because Reserved Instances are a billing discount model, not a distinct hardware isolation model — they do not prevent noisy neighbors.

906
MCQhard

A SysOps administrator is troubleshooting a CloudFormation stack that failed to create. The stack includes an Amazon RDS DB instance. The error message indicates that the DB instance name already exists. The stack uses a parameter for the DB instance identifier. What should the administrator do to resolve this issue and create the stack?

A.Delete the failed stack, change the DB instance identifier parameter to a unique name, and recreate the stack.
B.Manually delete the DB instance from the AWS Management Console and then retry the stack creation.
C.Use the AWS CLI command aws cloudformation update-stack with a new parameter value.
D.Execute ContinueUpdateRollback on the stack to retry the creation.
AnswerA

Deleting the failed stack is the correct first step because CloudFormation creation failures can leave resources in a transient or partially rolled-back state, and the stack cannot be updated or reused while it is in ROLLBACK_COMPLETE. Then changing the DB instance identifier parameter to a globally/regionally unique name avoids the RDS naming conflict that caused the failure; RDS DB instance identifiers must be unique per account within a Region. Finally, recreating the stack with the new parameter allows CloudFormation to provision a fresh DB instance without colliding with an existing resource, so this is the only approach that directly addresses the root cause.

Why this answer

When a CloudFormation stack creation fails due to a naming conflict (e.g., DB instance identifier already exists), the failed stack must be deleted because it cannot be updated or continued. Changing the parameter to a unique name and recreating the stack resolves the conflict. Option B is incorrect because, while deleting the conflicting DB instance would resolve the name conflict, the failed stack still exists and must be deleted before recreating the stack; in addition, deleting an existing DB instance is not the recommended approach — you should use a unique identifier.

Option C is incorrect because `update-stack` cannot be applied to a failed stack creation; the stack is in a failed state and must be recreated. Option D is incorrect because `ContinueUpdateRollback` is used for update rollbacks, not for failed creations.

907
MCQmedium

A company has an Amazon DynamoDB table that stores historical data. The table is accessed infrequently but when queried requires consistent single-digit millisecond latency. The SysOps administrator wants to minimize storage costs while maintaining the required performance. Which DynamoDB table class should the administrator use?

A.DynamoDB Standard
B.DynamoDB Standard-IA (Infrequent Access)
C.DynamoDB On-Demand
D.DynamoDB Provisioned
AnswerB

DynamoDB Standard-IA (Infrequent Access) is a table class designed specifically for data that is read infrequently but still requires DynamoDB's single-digit millisecond latency and full durability. It reduces storage costs by roughly 60% compared to Standard, while adding a small per-GB retrieval fee that only applies when data is actually read. This makes it the ideal choice for historical data that is stored long-term and accessed occasionally, since the storage savings outweigh the modest retrieval charges.

Why this answer

DynamoDB Standard-IA (Infrequent Access) is designed for tables that are accessed less than once per month, offering lower storage costs than DynamoDB Standard while maintaining the same single-digit millisecond latency for queries. Since the table stores historical data with infrequent access but requires consistent performance, Standard-IA minimizes storage costs without sacrificing latency.

Exam trap

The trap here is confusing DynamoDB table classes (Standard vs Standard-IA) with billing modes (On-Demand vs Provisioned), leading candidates to choose a billing mode instead of the correct table class for storage cost optimization.

How to eliminate wrong answers

Option A is wrong because DynamoDB Standard is optimized for frequently accessed data and has higher storage costs, making it suboptimal for infrequently accessed historical data. Option C is wrong because DynamoDB On-Demand is a billing mode (not a table class) that charges per request and is typically more expensive for unpredictable workloads, but it does not address storage cost optimization for infrequent access. Option D is wrong because DynamoDB Provisioned is also a billing mode (not a table class) that requires capacity planning and does not inherently reduce storage costs; it focuses on throughput rather than storage efficiency.

908
MCQhard

An application running on Amazon ECS (Fargate) is experiencing intermittent failures. The logs show 'CannotPullContainerError: error pulling image configuration: download failed after attempts=6'. The SysOps team has verified that the image exists in Amazon ECR and the task role has permissions to pull from ECR. What is the most likely cause?

A.The container image is corrupted.
B.The ECR repository policy is not allowing the task role.
C.The ECS task definition has an incorrect memory allocation.
D.The ECS tasks are in a private subnet without a NAT gateway or VPC endpoints for ECR and S3.
AnswerD

Fargate tasks running in a private subnet do not have public IP addresses, so they cannot reach the internet without a NAT gateway. To pull images from ECR, the task also needs access to the Amazon S3 endpoints that store image layers, which requires either a NAT gateway or VPC endpoints for both ECR and S3. Without these routes, the image pull attempts time out or fail with a network connectivity error, exactly matching the scenario. This is the correct cause.

Why this answer

The error 'CannotPullContainerError: error pulling image configuration: download failed after attempts=6' indicates that the ECS task is unable to download the image layers from Amazon ECR. Since the image exists and the task role has permissions, the most likely cause is a network connectivity issue. When ECS tasks run in a private subnet without a NAT gateway or VPC endpoints for ECR and S3, they cannot reach the public ECR API endpoints or the S3 buckets that store image layers, causing the pull to fail after multiple retries.

Exam trap

The trap here is that candidates often assume the error is due to missing IAM permissions or a corrupted image, overlooking the fact that ECS tasks in private subnets require explicit network paths (NAT gateway or VPC endpoints) to reach ECR and S3, even when permissions are correctly configured.

How to eliminate wrong answers

Option A is wrong because a corrupted image would typically produce a different error, such as a manifest or layer integrity failure, not a download failure after multiple attempts. Option B is wrong because the task role already has permissions to pull from ECR, and the repository policy is a separate mechanism that would cause an authorization error (e.g., 'AccessDeniedException') rather than a download failure. Option C is wrong because incorrect memory allocation would cause the task to fail at launch with a resource-related error (e.g., 'CannotStartContainerError: ResourceInitializationError') or OOM kill, not an image pull failure.

909
MCQhard

A company has a legacy application that requires access to an S3 bucket using an IAM user's access keys. The security team wants to rotate the access keys every 90 days automatically. What is the MOST efficient way to achieve this?

A.Use AWS Lambda with a scheduled CloudWatch Events rule to rotate the keys.
B.Create a script that runs on an EC2 instance using cron to rotate the keys.
C.Store the access keys in AWS Secrets Manager and use automatic rotation.
D.Enable IAM access key rotation in the IAM console.
AnswerC

AWS Secrets Manager natively supports automatic rotation for IAM user access keys: you configure a rotation schedule (e.g., every 30 or 90 days), and the service creates a new access key, updates the secret with the new credentials, and deletes the old key after a safe handoff—all with versioning so you can stage the new secret separately from the current one. It does not require you to write any custom scheduling code because Secrets Manager uses an AWS-provided Lambda rotation function template for IAM user keys, and it can alternate between two active keys to avoid downtime for downstream systems. This directly meets the need to 'require access to the service' without manual intervention, and it also gives you fine-grained IAM permissions to control who can access or rotate the secret, plus CloudTrail audit logs of rotation events.

Why this answer

AWS Secrets Manager provides built-in automatic rotation for IAM user access keys, allowing you to set a 90-day rotation schedule without custom code. Option A is incorrect because while AWS Lambda with a scheduled CloudWatch Events rule could rotate keys, it requires custom code and is less efficient than the managed rotation in Secrets Manager. Option B is incorrect because a script on an EC2 instance using cron would require additional infrastructure and maintenance.

Option D is incorrect because there is no built-in IAM access key rotation in the IAM console; you must manually rotate keys each time.

910
Multi-Selecteasy

A company wants to protect its data in Amazon S3 from accidental deletion. Which TWO methods should the SysOps administrator use? (Choose TWO.)

Select 2 answers
A.Set up cross-Region replication.
B.Enable S3 Transfer Acceleration.
C.Configure S3 event notifications.
D.Enable S3 Versioning on the bucket.
E.Enable MFA Delete on the bucket.
AnswersD, E

Enabling S3 Versioning makes the bucket store every object version, including all overwrites and deletes. When a delete request is made, S3 does not physically remove the object; instead, it inserts a null version delete marker, leaving all prior versions intact and recoverable at any time. This allows you to easily undo accidental deletions by removing the delete marker and restoring an older version, making it a fundamental protection mechanism.

Why this answer

Enabling S3 Versioning preserves all versions of an object, including overwrites and deletes. When versioning is enabled, a delete operation does not permanently remove the object; instead, it adds a delete marker, allowing the object to be restored by removing the marker. This directly protects against accidental deletion.

Exam trap

The trap here is that candidates often confuse replication (CRR) or notifications as protective measures, but they do not prevent deletion; only versioning and MFA Delete directly safeguard against accidental or malicious permanent data loss.

911
MCQhard

A company runs a stateful application on EC2 instances behind a Network Load Balancer (NLB). The application uses sticky sessions (session affinity) to maintain client state. During a deployment, the SysOps administrator needs to replace instances without disrupting active sessions. Which approach should be used?

A.Stop the NLB, replace instances, and restart the NLB
B.Deregister the old instances from the target group with connection draining enabled, then register new instances
C.Update the target group health check to remove old instances faster
D.Terminate the old instances immediately and launch new ones
AnswerB

Deregistering the old instances from the target group with connection draining (the deregistration delay attribute on the NLB target group) causes the load balancer to stop sending new connections to those instances while continuing to allow existing connections to complete up to the configured timeout, which defaults to 300 seconds. This preserves active user sessions for the stateful application, and only after the drain period expires are the instances fully removed. Registering new instances after the old ones have fully drained provides a clean handoff without abrupt termination; for even higher availability, you can register the new instances before starting the drain so fleet capacity is never reduced.

Why this answer

Deregistering instances from the target group with connection draining enabled allows existing connections to complete gracefully before the instances are removed. The Network Load Balancer (NLB) continues to route new connections to the remaining healthy instances, and once draining finishes, new instances can be registered without disrupting active sessions. This approach maintains session affinity (sticky sessions) by ensuring that in-flight requests are completed before the instance is taken out of service.

Exam trap

The trap here is that candidates may think connection draining is only for Application Load Balancers (ALBs) or that stopping the NLB is required for maintenance, but NLB supports connection draining via target group deregistration delay, which is the correct method for zero-downtime deployments.

How to eliminate wrong answers

Option A is wrong because stopping the NLB would terminate all active connections and disrupt all sessions, defeating the purpose of maintaining session affinity. Option C is wrong because updating the health check interval only affects how quickly unhealthy instances are detected; it does not provide a graceful shutdown mechanism for active sessions. Option D is wrong because terminating instances immediately would abruptly drop all active connections, causing session loss and application errors.

912
Multi-Selecteasy

A company uses AWS Shield Advanced to protect against DDoS attacks. Which of the following are benefits of AWS Shield Advanced? (Choose TWO.)

Select 2 answers
A.24/7 access to the DDoS Response Team (DRT)
B.Free for all AWS accounts
C.Integration with AWS WAF to create custom rules
D.Automatic scaling of resources during an attack
E.Enhanced DDoS protection for resources like EC2, ELB, CloudFront, and Route 53
AnswersA, E

Shield Advanced provides 24/7 access to the AWS DDoS Response Team (DRT), a specialized crew that can assist during an ongoing attack, perform proactive architectural reviews, and create tailored mitigations. This human-level support is a unique, paid benefit that distinguishes Shield Advanced from the free Shield Standard, making it crucial for high-risk, mission-critical applications.

Why this answer

AWS Shield Advanced provides enhanced DDoS protection for resources like EC2, ELB, CloudFront, and Route 53 (option E), and includes 24/7 access to the DDoS Response Team (DRT) (option A). Option B is incorrect because Shield Advanced is not free; it has a monthly cost. Option C is incorrect because integration with AWS WAF is not a direct benefit of Shield Advanced; WAF is a separate service that can be used alongside Shield but is not a benefit of Shield Advanced itself.

Option D is incorrect because automatic scaling is not a direct benefit of Shield Advanced; it is handled by other AWS services like Auto Scaling.

913
Multi-Selectmedium

A company is using Amazon Route 53 with a private hosted zone for internal DNS resolution within a VPC. The VPC is connected to an on-premises network via a VPN. On-premises resources cannot resolve DNS names in the private hosted zone. Which TWO actions should be taken to resolve this issue? (Choose two.)

Select 2 answers
A.Configure route propagation from the VPN to the VPC's route table.
B.Associate the private hosted zone with the on-premises network.
C.Enable DNS resolution and DNS hostnames for the VPC.
D.Create a public hosted zone with the same name and associate it with the VPC.
E.Create a Route 53 inbound resolver endpoint in the VPC.
AnswersC, E

This is a required VPC setting: the VPC must have both DNS resolution (the Amazon-provided DNS server at the VPC CIDR + 2) and DNS hostnames enabled. When DNS resolution is enabled, Route 53 can answer queries against private hosted zones from instances or the VPC resolver; DNS hostnames is necessary for AWS to assign and use internal DNS names for instances. Without these settings, the VPC's resolver behavior is disabled and private-hosted-zone lookups fail.

Why this answer

To allow on-premises resources to resolve DNS names in a private hosted zone, you need to create a Route 53 inbound resolver endpoint in the VPC (option E). This endpoint allows DNS queries from on-premises to be forwarded to Route 53. Additionally, you must enable DNS resolution and DNS hostnames for the VPC (option C) to ensure that the VPC's DNS settings support internal resolution.

Option A is incorrect because route propagation affects network routing, not DNS resolution. Option B is incorrect because private hosted zones cannot be associated with on-premises networks directly. Option D is incorrect because a public hosted zone is used for public DNS and does not resolve private DNS queries.

914
MCQhard

An organization uses AWS CloudTrail to log API calls across multiple accounts in AWS Organizations. The logs are delivered to a central S3 bucket. The security team wants to receive near-real-time notifications whenever an IAM user creates a new access key. Which solution is the MOST operationally efficient?

A.Create an Amazon EventBridge rule that matches the CreateAccessKey event from CloudTrail and publishes to an Amazon SNS topic.
B.Enable S3 Event Notifications on the CloudTrail bucket to trigger a Lambda function that scans new objects for access key creation.
C.Configure CloudTrail to send logs to Amazon CloudWatch Logs and set up a metric filter that triggers an alarm.
D.Use Amazon CloudWatch Logs Insights to run a query every minute and send results via SNS.
AnswerA

EventBridge is fully integrated with CloudTrail and receives management events as soon as they occur, typically within seconds. A rule with an event pattern such as `source: aws.iam` and `eventName: CreateAccessKey` will match the IAM API call and trigger an SNS topic as a push target, delivering a notification in near-real-time with no polling, log scanning, or waiting for CloudTrail's 5-minute S3 log delivery.

Why this answer

Amazon EventBridge can directly consume CloudTrail events (including `CreateAccessKey`) in near-real-time without polling or custom code. By creating a rule that matches this specific event and targets an SNS topic, the security team gets immediate notifications with minimal operational overhead. This approach is serverless, event-driven, and requires no intermediate services or custom functions.

Exam trap

The trap here is that candidates often default to CloudWatch Logs metric filters or S3 Event Notifications because they are familiar, but they overlook that EventBridge provides the most direct, low-latency, and operationally efficient path for CloudTrail event-driven notifications.

How to eliminate wrong answers

Option B is wrong because S3 Event Notifications are object-creation events, not real-time; they introduce latency (typically minutes) and require a Lambda function to parse each log file, which is less efficient and adds complexity. Option C is wrong because sending CloudTrail logs to CloudWatch Logs and using metric filters adds latency (logs are delivered in batches, metric filters poll every minute) and requires additional configuration for alarms, making it less near-real-time than EventBridge. Option D is wrong because running a CloudWatch Logs Insights query every minute is polling-based, incurs query costs, and is not event-driven; it also introduces up to 60 seconds of delay and is operationally inefficient compared to a push-based EventBridge rule.

915
MCQmedium

A SysOps administrator is automating the deployment of a three-tier web application using AWS CloudFormation. The administrator wants to ensure that the database tier is created before the application tier. How should the administrator define this dependency in the CloudFormation template?

A.Use the Conditions section to check if the database exists before creating the application tier.
B.Use the Outputs section to export the database endpoint and import it in the application tier.
C.Use the DependsOn attribute on the application tier resources to reference the database tier resources.
D.Use the Parameters section to pass the database instance identifier to the application stack.
AnswerC

The DependsOn attribute is the explicit way to tell AWS CloudFormation that one resource must be created before another, overriding the template's otherwise optional logical ordering. When you apply DependsOn to the application-tier resources and reference the database-tier logical IDs, CloudFormation guarantees the database stack resources are created first, and will also roll back the application tier if the database creation fails. This is necessary when the dependencies are not implicit, such as when the application code only knows the database endpoint from a parameter or discovery service rather than a Ref/GetAtt call. Therefore, DependsOn is the correct and direct mechanism for enforcing the creation sequence.

Why this answer

CloudFormation supports the DependsOn attribute to specify resource dependencies. The application tier resources must wait for the database tier resources to be created. Option A is wrong because the Conditions section determines whether resources are created based on conditions, not the creation order.

Option B is wrong because the Outputs section exports values for use in other stacks, but does not control the order of creation within the same stack. Option D is wrong because the Parameters section defines input values passed to the template, not dependencies.

Exam trap

A common trap is to confuse the purpose of Conditions and Outputs with dependency management. Conditions only control if a resource is created, not when. Outputs are for exporting values, not for ordering.

916
MCQmedium

An organization is using AWS CloudFormation to deploy infrastructure. The SysOps administrator needs to ensure that if a stack update fails, the stack automatically rolls back to the last known good state. Which stack update option should be configured?

A.Disable rollback
B.Change sets
C.Rollback on failure
D.Stack policy
AnswerC

The 'Rollback on failure' stack setting is the direct mechanism that causes CloudFormation to revert resources to the last successfully deployed state if an update operation fails. When enabled, CloudFormation automatically initiates a rollback by undoing any resource changes that occurred during the failed update, keeping the stack consistent with its previous known-good template. This setting is the only option listed that actively restores the prior stack state after an unsuccessful update.

Why this answer

CloudFormation's 'Rollback on failure' option, enabled by default, automatically reverts a stack to its last known good state if a stack update fails. This ensures that failed updates do not leave the infrastructure in an inconsistent or partially deployed state, maintaining reliability and business continuity.

Exam trap

The trap here is that candidates may confuse 'Rollback on failure' with 'Disable rollback' or think that change sets or stack policies handle rollback behavior, when in fact only the rollback configuration directly controls automatic recovery from failed updates.

How to eliminate wrong answers

Option A is wrong because 'Disable rollback' prevents automatic rollback on failure, leaving the stack in a failed state, which is the opposite of what the requirement asks. Option B is wrong because change sets allow you to preview how changes will affect a stack before execution, but they do not provide automatic rollback behavior on failure. Option D is wrong because a stack policy defines which stack resources can be updated during a stack update, not the rollback behavior on failure.

917
MCQeasy

A SysOps administrator deploys the above CloudFormation template. The stack creation fails with an error. What is the most likely reason?

A.The EBS volume must specify a SnapshotId.
B.The template uses a deprecated AWSTemplateFormatVersion.
C.The instance and volume are in different Availability Zones.
D.The VolumeAttachment resource is missing the Device property.
AnswerD

The AWS::EC2::VolumeAttachment resource requires the Device property, which specifies the logical device name (such as /dev/sdh) presented to the instance. Without Device, CloudFormation cannot complete the attachment resource and will fail to create it, even though the volume and instance are otherwise correctly configured. This is the actual defect in the template.

Why this answer

The VolumeAttachment resource in AWS CloudFormation requires the Device property to specify the device name (e.g., /dev/sdf) for the attached EBS volume. Without this property, CloudFormation cannot determine the mount point, causing the stack creation to fail. The error occurs because the template omits this required field.

Exam trap

The trap here is that candidates often assume the Device property is optional or that CloudFormation will auto-assign a device name, but AWS requires explicit specification for EBS volume attachments.

How to eliminate wrong answers

Option A is wrong because EBS volumes can be created without a SnapshotId; they can be empty volumes or created from snapshots, but a SnapshotId is not mandatory. Option B is wrong because AWSTemplateFormatVersion '2010-09-09' is the current and only valid version, not deprecated. Option C is wrong because CloudFormation automatically places the EC2 instance and EBS volume in the same Availability Zone when they are in the same template without explicit AZ specification, so this is not the cause of failure.

918
MCQhard

Refer to the exhibit. A SysOps administrator reviews the account password policy. Which of the following is true based on this output?

A.Passwords do not expire
B.Users cannot reuse their last 5 passwords
C.The maximum password age is 120 days
D.Users cannot change expired passwords
AnswerB

This statement is correct. The IAM password policy includes PasswordReusePrevention configured to 5, which prevents a user from reusing any of their previous 5 passwords when changing or resetting a password. This means the specified number of previous passwords is stored and cannot be reused, directly matching the statement.

Why this answer

The output shows MaxPasswordAge: 90 and ExpirePasswords: true, meaning passwords expire after 90 days. PasswordReusePrevention: 5 means users cannot reuse the last 5 passwords. Option B is correct.

Option A is wrong because password expiration is enabled (ExpirePasswords: true). Option C is wrong because MaximumPasswordAge is 90 days. Option D is wrong because HardExpiry is false, meaning users can change expired passwords.

919
MCQhard

A DevOps engineer is designing a CI/CD pipeline for a microservices application hosted on Amazon ECS with Fargate. The team wants to deploy updates to the services without downtime. The current pipeline builds a Docker image, pushes it to Amazon ECR, and updates the ECS service using AWS CodeDeploy with a blue/green deployment. However, during the deployment, the new tasks fail to start due to an incorrect environment variable. The engineer wants to validate the task definition before the actual deployment. What should the engineer do?

A.Use Amazon CloudWatch Synthetics canaries to monitor the health of the new tasks after deployment.
B.Run the Docker container locally using 'docker run' with the same environment variables to verify the configuration.
C.Use Amazon ECS Service Auto Scaling to gradually increase the number of tasks and monitor CPU utilization.
D.Configure CodeDeploy to use a validation hook with an AWS Lambda function that tests the new task definition before shifting traffic.
AnswerD

This is correct because CodeDeploy for Amazon ECS uses an AppSpec file that can define a `BeforeAllowTraffic` lifecycle hook, which invokes an AWS Lambda function after the new task set is registered with the target group but before any production traffic is shifted. The Lambda can perform an HTTP health check against the new task's endpoint, verify the container is listening on the expected port, or check internal state, and if it fails, CodeDeploy aborts the deployment and rolls back to the original task set. This provides a true pre-flight validation gate for the task definition within the actual Fargate environment, ensuring that only valid task definitions ever receive traffic.

Why this answer

AWS CodeDeploy for ECS supports lifecycle event hooks, including a BeforeAllowTraffic hook that runs an AWS Lambda function before traffic is shifted to the replacement task set. The Lambda can inspect the new task definition, verify environment variables, and fail the deployment if validation fails, preventing the bad configuration from ever receiving production traffic. This directly addresses the requirement to validate the task definition before the actual deployment.

Exam trap

The trap is that candidates focus on 'monitoring' or 'scaling' as the safety mechanism, missing that the question explicitly asks for pre-deployment validation, which only a CodeDeploy lifecycle hook can provide.

How to eliminate wrong answers

Option A is wrong because CloudWatch Synthetics canaries monitor endpoints after deployment — they detect problems post-traffic-shift, not before, so downtime could still occur. Option B is wrong because running the container locally does not validate the ECS task definition, IAM roles, secrets, or Fargate-specific configuration, and it is not part of the automated pipeline. Option C is wrong because ECS Service Auto Scaling adjusts task count based on load metrics; it has nothing to do with validating a task definition and would not catch a bad environment variable.

920
MCQeasy

A company uses AWS OpsWorks to manage a stack of web servers. They need to deploy a configuration change that updates the /etc/nginx/nginx.conf file on all instances. Which OpsWorks feature should be used?

A.Custom Chef recipes
B.OpsWorks layers
C.Lifecycle events
D.Custom cookbooks
AnswerA

A custom Chef recipe is the executable configuration unit in OpsWorks Stacks: when assigned to a layer lifecycle event, it runs on the instance and uses Chef resources (file, template, package, service, execute) to actually change configuration, such as updating an application configuration file. This is why the scenario—applying a specific configuration change to web servers—maps directly to custom recipes, not to broader constructs.

Why this answer

Custom Chef recipes are the correct choice because OpsWorks uses Chef to manage configuration, and deploying a specific file change like /etc/nginx/nginx.conf requires a custom recipe that directly modifies the file using Chef resources (e.g., template or file resource). This allows precise, idempotent configuration management across all instances in the stack.

Exam trap

The trap here is confusing lifecycle events (the trigger) with the actual configuration logic (the recipe), leading candidates to pick 'Lifecycle events' when the question asks for the feature that performs the file update.

How to eliminate wrong answers

Option B is wrong because OpsWorks layers define the structure and services of a stack (e.g., load balancer, application server), but they do not directly execute configuration changes to specific files like nginx.conf. Option C is wrong because lifecycle events (e.g., Setup, Configure, Deploy) trigger Chef runs but are not themselves a feature that deploys configuration changes; they are the timing mechanism, not the content. Option D is wrong because custom cookbooks are the collection of recipes, attributes, and templates, but the question asks for the feature used to deploy the change, and the specific executable unit within a cookbook is a recipe.

921
MCQmedium

A company runs a critical application on EC2 instances in an Auto Scaling group. The application processes messages from an Amazon SQS queue. The SysOps administrator notices that during periods of high load, the SQS queue depth increases significantly, and the application takes a long time to recover. The administrator wants to improve the application's ability to handle spikes in traffic without over-provisioning resources. The application is stateless and can scale horizontally. What should the administrator do?

A.Change the SQS queue from standard to FIFO to ensure messages are processed in order.
B.Configure an auto scaling policy for the Auto Scaling group based on the SQS queue depth (ApproximateNumberOfMessagesVisible).
C.Use a larger EC2 instance type with enhanced networking to process messages faster.
D.Increase the EC2 instance size to a larger type with more CPU and memory.
AnswerB

Configuring the Auto Scaling group to scale based on the SQS queue depth (ApproximateNumberOfMessagesVisible) is the correct approach because it directly matches worker capacity to the unprocessed backlog. An Amazon CloudWatch alarm or target tracking policy on this metric can add EC2 instances as messages accumulate and terminate instances as the queue drains. For best results, the policy should use a per-instance backlog metric (queue depth divided by desired capacity) to avoid over-scaling or flapping when the queue is briefly busy.

Why this answer

Scaling on the SQS queue depth (ApproximateNumberOfMessagesVisible) directly ties Auto Scaling capacity to the backlog, so the group adds instances when messages accumulate and removes them when the queue drains. This is the canonical target-tracking or step-scaling pattern for queue-driven, stateless workloads and avoids over-provisioning during normal load.

Exam trap

SOA-C02 often tests whether candidates default to vertical scaling (bigger instances) or queue-type changes when the correct answer is elastic, metric-driven horizontal scaling based on queue depth.

How to eliminate wrong answers

Option A is wrong because switching to a FIFO queue changes ordering semantics but does not improve throughput or scaling — FIFO queues actually have lower throughput limits (300 TPS without batching) and would worsen the spike problem. Option C is wrong because a larger instance with enhanced networking speeds up a single instance but doesn't add capacity dynamically, so the queue would still back up under high load. Option D is wrong because increasing instance size is vertical scaling, which is capped and doesn't respond elastically to traffic spikes.

922
MCQhard

A company runs a critical web application on a fleet of EC2 instances behind an Application Load Balancer (ALB). The instances are in an Auto Scaling group. The operations team uses CloudWatch alarms to monitor the application's health. Recently, they noticed that the application's error rate has increased sporadically, but the CPU utilization and memory usage remain normal. The team suspects that the issue is related to a specific HTTP endpoint returning 5xx errors. They want to set up monitoring that will alert them when the error rate exceeds 5% of total requests over a 5-minute period. The application logs are already sent to CloudWatch Logs. Which combination of steps should the SysOps administrator take to meet this requirement?

A.Create a metric filter in CloudWatch Logs to extract error codes and total requests from the application logs. Create two custom metrics: one for error count and one for total requests. Then create a CloudWatch alarm using a math expression that calculates error rate (error count / total requests) and triggers when >0.05 for 5 minutes.
B.Enable AWS X-Ray on the application to trace requests and identify error patterns. Create a CloudWatch alarm on the X-Ray error rate metric.
C.Install the CloudWatch agent on the EC2 instances to collect application-level metrics. Configure the agent to emit a custom metric for error rate. Then create an alarm on that metric.
D.Enable detailed monitoring on the ALB and create a CloudWatch alarm on the HTTPCode_ELB_5XX metric with a threshold of 5% of the request count. Use the ALB's RequestCount metric to compute the percentage.
AnswerA

This is the correct approach because the application logs are already flowing into CloudWatch Logs, and a metric filter can parse them in real time to extract both the number of error codes (e.g., status codes or application-specific errors) and the total request count. By creating two custom metrics—ErrorCount and TotalRequests—you can then define a CloudWatch alarm using a metrics math expression such as e1/e2, with the alarm triggering when the ratio exceeds 0.05 for a 5-minute period. This leverages the existing log data without requiring additional instrumentation or external services, and it accurately reflects application-level error rates as observed in the logs.

Why this answer

The requirement is to alert when the 5xx error rate exceeds 5% of total requests over a 5-minute window, and the logs are already in CloudWatch Logs. A metric filter extracts the error count and total request count into custom metrics, and a CloudWatch alarm with a math expression (errors/requests) evaluates the ratio against the 0.05 threshold. This is the only option that produces a true percentage-based alarm from the existing log data.

Exam trap

SOA-C02 often tests whether candidates know that ALB's HTTPCode_ELB_5XX only counts load-balancer-generated errors, and that CloudWatch alarms need metric math (not a simple threshold) to evaluate a ratio like error percentage.

How to eliminate wrong answers

Option B is wrong because X-Ray traces requests for latency and dependency analysis but does not natively emit an 'error rate' CloudWatch metric that can be alarmed on as a percentage of total requests. Option C is wrong because the CloudWatch agent collects OS-level and application metrics but does not parse application logs to derive an error-rate metric; it would require custom application instrumentation that isn't described. Option D is wrong because HTTPCode_ELB_5XX counts only errors generated by the load balancer itself (e.g., 502/503/504 from unhealthy targets), not application-generated 5xx responses, and CloudWatch alarms cannot directly compute a percentage of another metric without a math expression.

923
MCQhard

A company has an S3 bucket that stores sensitive customer data. The security team requires that all objects uploaded to the bucket must be encrypted at rest using AWS KMS with a specific customer managed key. Which bucket policy condition should be used to enforce this?

A."Condition": {"StringEquals": {"s3:x-amz-server-side-encryption": "aws:kms"}}
B."Condition": {"StringEquals": {"s3:x-amz-server-side-encryption": "aws:kms", "s3:x-amz-server-side-encryption-aws-kms-key-id": "arn:aws:kms:us-east-1:123456789012:key/abc123"}}
C."Condition": {"StringEquals": {"s3:x-amz-server-side-encryption-aws-kms-key-id": "arn:aws:kms:us-east-1:123456789012:key/abc123"}}
D."Condition": {"Null": {"s3:x-amz-server-side-encryption": "false"}}
AnswerB

Combines both conditions to enforce KMS encryption and the specific customer managed key, meeting the requirement.

Why this answer

It uses both conditions: 's3:x-amz-server-side-encryption' set to 'aws:kms' ensures that objects are encrypted with SSE-KMS, and 's3:x-amz-server-side-encryption-aws-kms-key-id' set to the specific key ARN ensures that only the designated customer managed key is used. Option A enforces KMS encryption but does not restrict which KMS key, allowing any managed key. Option C enforces a specific key ARN but does not require the encryption header to be present, which could allow objects without encryption if the key ID header is omitted (though in practice, the key ID is only valid with SSE-KMS, the condition alone is not sufficient to guarantee encryption).

Option D uses a 'Null' condition incorrectly and would not properly enforce encryption.

924
MCQeasy

A company is designing a highly available web application using an Application Load Balancer (ALB) with EC2 instances in an Auto Scaling group across two Availability Zones. Which configuration ensures that the application remains available if one Availability Zone fails?

A.Use a single large EC2 instance instead of multiple instances
B.Disable health checks on the ALB to avoid false positives
C.Configure the Auto Scaling group to launch instances in at least two Availability Zones
D.Launch all EC2 instances in a single Availability Zone
AnswerC

Configuring the Auto Scaling group to launch instances in at least two Availability Zones is the correct approach for high availability because it isolates the application from a single-AZ failure. If one AZ becomes unavailable, the ASG automatically launches new instances in another AZ and the ALB routes traffic only to healthy instances in the remaining AZs. This distribution also supports even scaling and aligns with the AWS Well-Architected Framework's reliability pillar.

Why this answer

Configuring the Auto Scaling group to launch instances in at least two Availability Zones ensures that if one AZ fails, the remaining AZ continues to serve traffic. The ALB distributes incoming requests across healthy targets in all enabled AZs, and Auto Scaling replaces failed instances in the remaining AZs automatically, maintaining application availability.

Exam trap

The trap here is that candidates may think a single large instance or disabling health checks improves availability, but in reality, these actions introduce single points of failure or prevent automatic failure detection, which is exactly what the SOA-C02 exam tests for multi-AZ resilience.

How to eliminate wrong answers

Option A is wrong because using a single large EC2 instance creates a single point of failure; if that instance or its AZ fails, the application becomes unavailable, and it does not leverage the ALB's multi-AZ load balancing. Option B is wrong because disabling health checks on the ALB prevents it from detecting unhealthy targets, causing traffic to be routed to failed instances, which degrades availability and violates the goal of high availability. Option D is wrong because launching all EC2 instances in a single Availability Zone means that a failure of that AZ will take down all instances, making the application unavailable regardless of the ALB or Auto Scaling configuration.

925
MCQeasy

A SysOps administrator wants to receive alerts when the root user performs an action in the AWS account. Which service should be used?

A.AWS Identity and Access Management (IAM)
B.Amazon CloudWatch Metrics
C.AWS Config
D.AWS CloudTrail and Amazon CloudWatch Logs
AnswerD

AWS CloudTrail captures a full history of API activity and management events, including the root user's sign-in attempt, which is recorded as a 'ConsoleLogin' event with a 'userIdentity.type' of 'Root'. Sending that CloudTrail trail to Amazon CloudWatch Logs allows you to create a CloudWatch Logs metric filter that matches the JSON pattern for root-level sign-in events, and then attach a CloudWatch alarm to that metric with a threshold of one or more events. When the metric filter sees a matching root sign-in, the alarm transitions to ALARM and sends a notification to the SNS topic you configured. Together, CloudTrail supplies the detailed event data and CloudWatch Logs/alarms provide the real-time detection and alerting, making this the correct and recommended AWS solution.

Why this answer

AWS CloudTrail captures all API calls made by the root user as events. By sending these events to Amazon CloudWatch Logs, you can create a metric filter that matches root user activity and trigger an alarm via CloudWatch Alarms. This combination enables real-time notification when the root user performs any action.

Exam trap

The trap here is that candidates often choose AWS Config because it monitors resource changes, but they fail to realize that root user actions are API calls, not configuration changes, and thus require CloudTrail and CloudWatch Logs for detection.

How to eliminate wrong answers

Option A is wrong because AWS IAM manages users, roles, and permissions but does not provide event monitoring or alerting capabilities for root user actions. Option B is wrong because Amazon CloudWatch Metrics alone cannot capture or alert on specific API actions; it requires CloudTrail logs and metric filters to detect root user activity. Option C is wrong because AWS Config evaluates resource configurations and compliance rules, not API call activity; it cannot detect when the root user performs an action.

926
MCQeasy

A SysOps administrator needs to monitor the CPU utilization of an Amazon EC2 instance and receive an alert if it exceeds 80% for 10 consecutive minutes. The instance is in a VPC with no Internet access. What is the MOST efficient way to meet these requirements?

A.Use AWS Systems Manager to run a script on the instance that checks CPU and sends an SNS notification.
B.Create a CloudWatch alarm on the CPUUtilization metric with a period of 5 minutes and an evaluation period of 2.
C.Enable detailed monitoring on the EC2 instance and create a CloudWatch alarm on the CPUUtilization metric.
D.Install the CloudWatch agent on the EC2 instance to collect CPU metrics and create a CloudWatch alarm.
AnswerB

This is the correct approach because EC2 instances automatically emit CPUUtilization metrics every 5 minutes under basic monitoring, so a CloudWatch alarm with a period of 5 minutes will have the data it needs without any extra configuration. An evaluation period of 2 means the alarm only enters ALARM state after two consecutive data points breach the threshold, which reduces false positives from transient spikes. This is the simplest, most cost-effective way to monitor CPU utilization on an EC2 instance.

Why this answer

A CloudWatch alarm with a period of 5 minutes and an evaluation period of 2 means the alarm evaluates two consecutive 5-minute data points, totaling 10 minutes. Since the EC2 instance is in a VPC with no Internet access, CloudWatch metrics are still reported via the CloudWatch service endpoint within the VPC (or via VPC endpoints), so no additional agent or script is needed. This is the most efficient approach as it uses native CloudWatch functionality without requiring any custom scripts or additional software.

Exam trap

The trap here is that candidates often assume they need detailed monitoring or the CloudWatch agent to meet a specific time window, but the default 5-minute period with multiple evaluation periods can achieve the same result more efficiently and at lower cost.

How to eliminate wrong answers

Option A is wrong because using AWS Systems Manager to run a script on the instance that checks CPU and sends an SNS notification introduces unnecessary complexity and overhead; it requires the instance to have outbound internet access or a VPC endpoint for Systems Manager, and it is not the most efficient native solution. Option C is wrong because enabling detailed monitoring (1-minute metrics) is not required for this scenario; a 5-minute period with 2 evaluation periods already meets the 10-minute requirement, and detailed monitoring would incur additional cost without benefit. Option D is wrong because installing the CloudWatch agent is unnecessary; the EC2 instance already publishes the CPUUtilization metric by default (basic monitoring) without any agent, and the agent is only needed for custom or OS-level metrics, not for standard CPU utilization.

927
MCQmedium

A company wants to ensure that only specific IAM roles within the same AWS account can encrypt and decrypt data using an AWS KMS customer managed key. Which type of policy must be configured to achieve this restriction?

A.IAM policy attached to the roles
B.KMS key policy
C.Service control policy (SCP)
D.Resource policy attached to the KMS key
AnswerB

A KMS key policy is the resource-based policy attached directly to the customer master key (CMK) that defines which principals can use the key. For roles within the same AWS account, you can either specify the role ARNs directly in the key policy or allow the account root to delegate permissions via IAM policies. This is the authoritative mechanism to restrict key usage to only the designated IAM roles.

Why this answer

A KMS key policy is the primary mechanism to control access to a customer managed key. By default, a KMS key policy must explicitly grant the necessary permissions (kms:Encrypt, kms:Decrypt) to IAM roles, and it can restrict those permissions to specific roles within the same account using the `aws:PrincipalArn` condition key. This ensures that only the designated IAM roles can encrypt and decrypt data with that key.

Exam trap

The trap here is that candidates often think an IAM policy alone is sufficient to grant KMS key access, but they forget that KMS key policies act as a resource-based policy that must explicitly allow the IAM principal, otherwise the IAM policy is ignored.

How to eliminate wrong answers

Option A is wrong because an IAM policy attached to the roles alone is insufficient; KMS requires a key policy that explicitly allows the IAM roles to use the key, and without such a key policy, the IAM policy has no effect (the key policy acts as a resource-based policy that must grant access). Option C is wrong because a Service Control Policy (SCP) is used in AWS Organizations to set permission boundaries across accounts, not to grant or deny specific IAM roles access to a KMS key within a single account. Option D is wrong because while a KMS key policy is technically a resource policy, the term 'resource policy attached to the KMS key' is redundant and misleading; the correct and specific term is 'KMS key policy', and the question asks for the type of policy, not a generic description.

928
MCQeasy

A company runs a batch processing application on Amazon EC2 that runs for 2 hours every night. The workload can tolerate interruptions. Which EC2 purchasing option provides the lowest cost for this use case?

A.On-Demand Instances
B.Reserved Instances
C.Spot Instances
D.Dedicated Hosts
AnswerC

Spot Instances operate using spare EC2 capacity that AWS makes available at a significantly reduced hourly rate—often up to 90% off On-Demand pricing. This is the best fit here because the nightly 2-hour batch is both short and fault-tolerant: if capacity is reclaimed, work can be re-queued or resumed without violating the batch window. You can further reduce interruption risk by using a Spot Fleet with multiple instance types and by implementing checkpointing so progress is saved between runs. The result is a dramatic cost reduction for a workload that would otherwise be idling and paying full price.

Why this answer

Spot Instances are the correct choice because the workload is fault-tolerant, runs for a fixed 2-hour window nightly, and can tolerate interruptions. Spot Instances offer significant cost savings (up to 90% off On-Demand) by using spare EC2 capacity, which aligns perfectly with a batch job that can be retried if interrupted.

Exam trap

The trap here is that candidates may choose Reserved Instances because they see a predictable nightly schedule, but they overlook that Reserved Instances are cost-effective only for 24/7 workloads, not for short, interruptible batch jobs where Spot Instances provide far greater savings.

How to eliminate wrong answers

Option A is wrong because On-Demand Instances provide no discount and are not cost-optimal for a predictable, interruptible workload. Option B is wrong because Reserved Instances require a 1- or 3-year commitment and are designed for steady-state, always-on workloads, not a short 2-hour nightly batch job. Option D is wrong because Dedicated Hosts are a physical server dedicated to a single customer, incurring high costs for licensing or compliance needs, and are overkill for a batch processing application that can tolerate interruptions.

929
MCQmedium

A company uses an Application Load Balancer (ALB) to distribute traffic to EC2 instances. The SysOps team wants to reduce costs by ensuring that idle capacity is minimized. Which configuration should they implement?

A.Configure an Auto Scaling group with a step scaling policy based on CPU utilization.
B.Configure an Auto Scaling group with a target tracking scaling policy based on ALB request count per target.
C.Increase the number of EC2 instances in the Auto Scaling group to handle peak load.
D.Set the Auto Scaling group desired capacity to the maximum expected load.
AnswerB

A target tracking scaling policy with the ALBRequestCountPerTarget metric is purpose-built for web workloads behind an ALB because it continuously adjusts capacity to maintain a specified target value (e.g., 1000 requests per instance), automatically creating and managing the required CloudWatch alarms and scaling activities. This policy responds directly to the metric that reflects actual user traffic per instance, so it scales out quickly when request volume increases and scales in when traffic decreases, significantly reducing idle instance capacity. AWS recommends this approach over CPU or memory based scaling for HTTP(S) applications because it aligns scaling with the real bottleneck—request throughput per instance.

Why this answer

A target tracking scaling policy based on ALB request count per target automatically adjusts the number of EC2 instances to keep the request count per instance at a specified target value. This directly ties scaling to actual traffic demand, minimizing idle capacity while maintaining performance. It is the most cost-efficient and responsive configuration for an ALB-fronted workload.

Exam trap

SOA-C02 often tests the difference between reactive scaling (step/CPU-based) and demand-based scaling (target tracking on request count) — candidates pick CPU-based step scaling out of habit, missing that request count per target is the most direct and cost-efficient metric for ALB workloads.

How to eliminate wrong answers

Option A is wrong because a step scaling policy based on CPU utilization requires manual definition of CloudWatch alarms and scaling adjustments, and CPU utilization is a lagging indicator that may not correlate well with request load — it can leave idle capacity during low-traffic periods or fail to scale quickly enough during spikes. Option C is wrong because increasing the number of EC2 instances to handle peak load provisions for the worst case at all times, which maximizes idle capacity and increases cost — the opposite of the stated goal. Option D is wrong because setting desired capacity to the maximum expected load statically over-provisions and guarantees idle capacity during normal and low-traffic periods, directly contradicting the cost-reduction objective.

930
MCQeasy

A SysOps administrator needs to grant an application running on an Amazon EC2 instance access to an Amazon S3 bucket without embedding long-term AWS credentials in the application code. The EC2 instance is in a private subnet and must not use an IAM user access key. Which solution should the administrator use?

A.Store the AWS credentials in an encrypted Amazon S3 bucket and have the application download them at startup.
B.Create an IAM user with programmatic access, generate an access key, and store it in AWS Systems Manager Parameter Store as a SecureString parameter.
C.Configure the S3 bucket policy to allow access from the EC2 instance's private IP address using the aws:SourceIp condition.
D.Attach an IAM role to the EC2 instance with a policy that allows the required S3 actions, and let the application use the instance metadata service to obtain temporary credentials.
AnswerD

Attaching an IAM role to an EC2 instance provides temporary credentials through the instance metadata service (IMDS). The application can retrieve these credentials automatically without embedding long-term keys. This is the recommended best practice for granting AWS permissions to EC2 instances. The role's policy can be scoped to only the necessary S3 actions, following least privilege.

Why this answer

Attaching an IAM role to the EC2 instance allows the application to obtain temporary credentials from the instance metadata service. This eliminates the need for long-term access keys and follows AWS best practices for least privilege and credential management. The role can be scoped to only the required S3 actions.

Exam trap

The trap here is assuming that storing credentials in Parameter Store or an encrypted S3 bucket solves the problem, but those still involve long-term credentials and additional bootstrap challenges.

931
MCQeasy

A company uses AWS Elastic Beanstalk to deploy a web application. After updating the environment configuration, the deployment fails and the environment health turns red. The SysOps administrator checks the logs and finds a permission error related to the EC2 instance profile. What should the administrator do to resolve the issue?

A.Rebuild the environment from scratch using a saved configuration template.
B.Update the IAM instance profile associated with the environment to include the required permissions.
C.Modify the security group attached to the environment to allow outbound traffic.
D.Update the application version to the latest build.
AnswerB

Elastic Beanstalk environments use an IAM instance profile to grant permissions to the underlying EC2 instances. To resolve permission errors where the application cannot access required AWS services or resources, you must attach a policy that includes the necessary actions to the instance profile role. After updating the role, perform an environment update or restart so the running instances pick up the new permissions; this directly addresses the root cause because instance profile permissions govern what the application can do on AWS.

Why this answer

Elastic Beanstalk uses an IAM instance profile for the EC2 instances. The instance profile must have the necessary permissions to access resources like S3 buckets or DynamoDB tables. Updating the instance profile with the required permissions resolves the issue.

Option A is wrong because rebuilding the environment from scratch using a saved configuration template does not address the underlying permission issue; it would only recreate the same problem. Option C is wrong because the security group controls network access, not IAM permissions; modifying it would not resolve the permission error. Option D is wrong because updating the application version does not fix permission issues; the application version itself does not grant or modify IAM permissions.

932
MCQhard

Refer to the exhibit. A SysOps administrator created this IAM policy for an application that sends custom metrics to CloudWatch and writes logs to CloudWatch Logs. The application reports that it cannot publish logs. What is the most likely reason?

A.The resource ARN for the logs actions is incorrect; it should include 'log-group:' before the wildcard.
B.The policy requires a condition key to restrict access to specific log groups.
C.The policy does not allow the cloudwatch:PutMetricData action for the specific metric.
D.The application must assume an IAM role to write logs.
AnswerA

The ARN for CloudWatch Logs actions must follow the format arn:aws:logs:region:account-id:log-group:log-group-name:*. Omitting the 'log-group:' prefix produces an invalid ARN that cannot match any log group, so the policy would not grant the necessary permissions. Without this prefix, IAM cannot resolve the resource to a specific log group, causing the logs actions to fail at runtime.

Why this answer

The IAM policy uses `arn:aws:logs:us-east-1:123456789012:*` for the `Resource` element of the `logs:PutLogEvents` and `logs:CreateLogGroup` actions. For CloudWatch Logs, the resource ARN must include the `log-group:` prefix before the log group name or wildcard, such as `arn:aws:logs:us-east-1:123456789012:log-group:*`. Without this prefix, the ARN does not match any valid CloudWatch Logs resource, causing the application to fail when attempting to publish logs.

Exam trap

The trap here is that candidates often assume a wildcard resource ARN like `arn:aws:logs:region:account:*` is sufficient for CloudWatch Logs actions, but they overlook the required `log-group:` prefix in the ARN structure, leading them to incorrectly suspect missing conditions or role assumption issues.

How to eliminate wrong answers

Option B is wrong because the policy does not require a condition key to restrict access to specific log groups; the issue is the malformed resource ARN, not the absence of conditions. Option C is wrong because the policy includes `cloudwatch:PutMetricData` with a wildcard resource (`*`), which is correct for CloudWatch custom metrics, and the application's failure is specifically about publishing logs, not metrics. Option D is wrong because the application can use IAM user credentials or an IAM role attached to an EC2 instance profile; the policy itself does not require assuming a role, and the error is due to the incorrect resource ARN, not the authentication method.

933
MCQeasy

A company requires that all AWS account activity be recorded and the logs be stored in a centralized S3 bucket for analysis. Which two AWS services should be used together to meet this requirement?

A.Amazon GuardDuty and Amazon S3
B.AWS CloudTrail and Amazon S3
C.AWS Config and Amazon S3
D.Amazon Inspector and Amazon S3
E.VPC Flow Logs and Amazon S3
AnswerB

AWS CloudTrail is the correct service for recording all AWS account activity. It captures every API call made in the account, including the identity of the caller, the time of the call, the source IP address, and the requested action, delivering these events as log files. These logs can be delivered to an Amazon S3 bucket for long-term, tamper-evident storage, which directly satisfies the compliance and auditing requirement.

Why this answer

The correct answer is B: AWS CloudTrail and Amazon S3. CloudTrail records all AWS account activity as API call events, and it can be configured with a trail that delivers those event logs to a centralized S3 bucket for storage and later analysis. The other options do not fit: GuardDuty is a threat-detection service that generates findings rather than recording all account activity, AWS Config tracks resource configuration changes and compliance rather than full API activity, Amazon Inspector assesses vulnerabilities on workloads, and VPC Flow Logs capture IP traffic metadata for network interfaces, not AWS account API activity.

934
Multi-Selectmedium

Which TWO actions can a SysOps administrator take to secure an Amazon S3 bucket that contains sensitive data? (Choose TWO.)

Select 2 answers
A.Configure a cross-origin resource sharing (CORS) policy.
B.Enable default encryption using AWS KMS (SSE-KMS) on the bucket.
C.Enable cross-region replication for the bucket.
D.Block all public access to the bucket using the S3 Block Public Access feature.
E.Enable MFA Delete on the bucket to require multi-factor authentication for delete operations.
AnswersB, D

Enabling default encryption with SSE-KMS ensures that every object uploaded without an explicit encryption header is automatically encrypted at rest using a KMS-managed customer key. This provides envelope encryption and allows you to control key access through IAM policies, while CloudTrail records key usage for auditing. It directly addresses data-at-rest confidentiality, making it a necessary and effective security action.

Why this answer

Enabling default encryption with SSE-KMS ensures that all objects uploaded to the S3 bucket are automatically encrypted at rest using AWS KMS-managed keys. This protects sensitive data even if the uploader forgets to specify encryption, meeting security and compliance requirements for data-at-rest protection.

Exam trap

The trap here is that candidates often confuse operational features like replication or MFA Delete with security controls that prevent unauthorized access or ensure encryption, leading them to select options that do not directly secure sensitive data.

935
MCQeasy

A company uses AWS Systems Manager to automate patching of EC2 instances. The instances are in an Auto Scaling group. The company wants to ensure that patching does not affect application availability. Which feature should be used?

A.State Manager
B.Maintenance Windows
C.Patch Manager
D.Run Command
AnswerB

Maintenance Windows are the Systems Manager feature designed to schedule time-bounded administrative actions, such as patching, during approved maintenance periods. They let you register Patch Manager tasks, Run Command commands, or Automation workflows, and control execution with rate limits and concurrency thresholds. This scheduling capability directly answers the need for automated patching with minimal business impact.

Why this answer

Systems Manager Maintenance Windows allow scheduling patching during specific time windows, and can be configured to work with Auto Scaling to maintain availability by ensuring instances are patched without affecting the overall application availability. Option A is wrong because State Manager is used for maintaining consistent configuration of instances, not for scheduling patching. Option C is wrong because Patch Manager is the service that applies patches, but it is typically run within a Maintenance Window to control timing and availability.

Option D is wrong because Run Command is for executing ad-hoc commands and scripts, not for scheduled patching with availability considerations.

936
MCQmedium

A SysOps Administrator is configuring a Network Load Balancer (NLB) for a TCP-based application. The application requires that clients see the original source IP address of the request. Which configuration should the Administrator use?

A.Use the NLB default behavior; no additional configuration needed.
B.Use an Application Load Balancer instead, which preserves the source IP.
C.Enable cross-zone load balancing on the NLB.
D.Enable Proxy Protocol v2 on the target group.
AnswerA

A Network Load Balancer operates at Layer 4 and forwards packets as-is to the registered targets, so the client IP address is preserved in the original packet headers by default. Unlike Layer 7 load balancers, the NLB does not terminate the TCP connection or perform NAT on the source address. Therefore, no additional settings, headers, or protocols are required to make the client source IP visible to the backend.

Why this answer

Network Load Balancers (NLBs) preserve the original source IP address of clients by default when forwarding TCP traffic to targets. This is because NLBs operate at Layer 4 and do not terminate the TCP connection; instead, they pass packets directly to the backend, allowing the target to see the client's IP. No additional configuration is required for this behavior.

Exam trap

The trap here is that candidates often confuse NLB and ALB behavior, assuming that preserving source IP requires a special configuration like Proxy Protocol, when in fact NLBs do this by default for TCP traffic.

How to eliminate wrong answers

Option B is wrong because an Application Load Balancer (ALB) terminates the client connection and re-establishes a new connection to the target, which by default replaces the source IP with the ALB's private IP; ALBs require the X-Forwarded-For header to convey the original client IP, not direct preservation. Option C is wrong because cross-zone load balancing distributes traffic across targets in multiple Availability Zones but does not affect source IP preservation. Option D is wrong because Proxy Protocol v2 is an optional header that can be added to preserve client IP information when using TCP listeners, but it is not required for NLB default behavior; enabling it would add an extra header, not fix a missing IP.

937
MCQmedium

A company runs an application on Amazon EC2 instances behind an Application Load Balancer (ALB). The ALB terminates SSL/TLS and forwards traffic to the instances over HTTP. The SysOps administrator needs to capture the original client IP address in the instance logs. How should the administrator configure this?

A.Enable stickiness on the ALB target group.
B.Enable the X-Forwarded-For header on the ALB.
C.Configure the ALB to use Proxy Protocol v2.
D.Enable access logs on the ALB and store them in Amazon S3.
AnswerB

The ALB automatically adds the X-Forwarded-For header to each HTTP/HTTPS request as it passes through, containing the original client IP address in a comma-separated list. Since the ALB terminates the client's TLS connection and opens a new connection to the target, the EC2 instance must read this header to record the client IP in its logs. By default, the ALB overwrites any existing X-Forwarded-For header to prevent client spoofing, and you should configure your web server or application to log the first IP in the header, which is the true client IP.

Why this answer

When an Application Load Balancer terminates SSL/TLS and forwards traffic to EC2 instances over HTTP, the original client IP address is preserved by the ALB in the X-Forwarded-For header. By enabling this header on the ALB, the SysOps administrator ensures that the web server or application can log the true client IP, which is essential for analytics, security, and troubleshooting.

Exam trap

The trap here is that candidates confuse Proxy Protocol v2 (used for NLB TCP/UDP listeners) with the X-Forwarded-For header (used for ALB HTTP/HTTPS listeners), leading them to select option C even though it is not applicable to ALB's HTTP-based forwarding.

How to eliminate wrong answers

Option A is wrong because enabling stickiness (session affinity) on the ALB target group only ensures that requests from the same client are routed to the same target instance; it does not capture or forward the original client IP address. Option C is wrong because Proxy Protocol v2 is used with Network Load Balancers (NLB) or TCP listeners, not with Application Load Balancers (ALB) which use HTTP/HTTPS listeners and rely on the X-Forwarded-For header for client IP preservation. Option D is wrong because enabling ALB access logs and storing them in Amazon S3 captures request details including client IP, but it does not inject the original client IP into the instance logs; the instance logs still see the ALB's private IP unless the X-Forwarded-For header is used.

938
MCQeasy

A company wants to reduce data transfer costs for traffic between EC2 instances in the same AWS Region. Which action should the SysOps administrator take?

A.Use Elastic IP addresses for all instances
B.Place instances in public subnets and route traffic through a NAT Gateway
C.Ensure instances communicate using private IP addresses within the same VPC
D.Use VPC endpoints to communicate between instances
AnswerC

Ensuring instances communicate using private IP addresses within the same VPC is the correct and most cost-effective approach. Traffic sent over private IPv4 addresses stays entirely inside the VPC, never crossing the internet gateway, so it is not subject to public internet data transfer rates. For instances in the same Availability Zone, this traffic is absolutely free; even across Availability Zones, the per-GB charge is only $0.01 each way, which is far cheaper than public IP or gateway-based alternatives.

Why this answer

Traffic between EC2 instances in the same VPC using private IP addresses does not incur public internet data transfer costs. If the instances are in the same Availability Zone, the traffic is free; if they are in different Availability Zones, standard inter-AZ data transfer charges apply. Using private IPs is still the most cost-effective option compared to using public IPs (Elastic IPs) or routing through a NAT Gateway.

Exam trap

The trap is confusing cost reduction with security or availability. Also, be aware that 'same Region' does not guarantee 'same Availability Zone'; inter-AZ traffic is charged.

How to eliminate wrong answers

Option A is wrong because Elastic IP addresses are public IPv4 addresses; traffic sent to or from an Elastic IP address traverses the internet gateway, incurring standard data transfer charges for both inbound and outbound traffic. Option B is wrong because placing instances in public subnets and routing traffic through a NAT Gateway would force traffic to go through the NAT Gateway, which adds per-GB data processing charges and data transfer costs for traffic leaving the VPC, increasing costs unnecessarily. Option D is wrong because VPC endpoints are designed for private connectivity to AWS services (e.g., S3, DynamoDB) and cannot be used for communication between EC2 instances; they do not replace the need for private IP routing within a VPC.

939
Multi-Selecthard

A company deploys microservices on Amazon ECS using Fargate. The deployment is managed by AWS CodePipeline. The administrator notices that deployments sometimes fail because the new task definition is not registered before the deployment. Which THREE steps should the administrator take to resolve this issue? (Choose THREE.)

Select 3 answers
A.Ensure that the task definition is registered in the CodePipeline build stage before the deploy stage.
B.Manually update the ECS service with the new task definition after the pipeline runs.
C.Use the ECS deploy action in CodePipeline which automatically registers the task definition.
D.Add a step in the pipeline to register the task definition using the AWS CLI or SDK.
E.Store the task definition in Amazon ECR alongside the container image.
AnswersA, C, D

In CodePipeline, the Deploy stage's ECS deployment action expects an ARN of an already-registered task definition, typically passed via artifacts from the Build stage. If the build stage does not explicitly register the task definition (e.g., using aws ecs register-task-definition or the ECS deploy action's automatic registration), the deploy action may reference a stale or nonexistent revision, causing the deployment to fail or roll back. Registration must happen before the deploy stage consumes the artifact, ensuring the action has a valid family:revision to run.

Why this answer

The issue is that the task definition must be registered before the ECS service update. Option A is correct because registering the task definition in the build stage ensures it is available for the deploy stage. Option C is correct because the ECS deploy action in CodePipeline automatically handles task definition registration and service update.

Option D is correct because adding a CLI or SDK step explicitly registers the task definition. Option B is incorrect because the task definition should be registered automatically, not manually. Option E is incorrect because ECR stores container images, not task definitions.

940
MCQmedium

A company is running a critical application on EC2 instances in an Auto Scaling group. The application experiences occasional CPU spikes. The SysOps administrator needs to configure a scaling policy that reacts quickly to increased load but avoids unnecessary scaling actions due to short bursts. Which scaling policy type should be used?

A.Manual scaling
B.Scheduled scaling policy
C.Simple scaling policy
D.Target tracking scaling policy with a CPU utilization target of 70%
AnswerD

Target tracking scaling policy automatically creates and manages CloudWatch alarms and adjusts the desired capacity proportionally to the deviation from the 70% CPU utilization target. It continuously evaluates the metric, allowing it to react to sudden spikes by adding instances in a dynamic manner, while its built-in cooldown and scale-in safeguards prevent oscillation. This policy is ideal for the scenario because it directly addresses the goal of maintaining CPU utilization at a defined level without manual or predetermined intervention.

Why this answer

Target tracking scaling policy with a CPU utilization target of 70% is correct because it dynamically adjusts the Auto Scaling group's desired capacity to maintain a specified metric (CPU utilization) at the target value. It uses a built-in algorithm that reacts quickly to sustained load increases while smoothing out short bursts by applying a cooldown and proportional control logic, preventing unnecessary scaling actions from transient spikes.

Exam trap

The trap here is that candidates often choose simple scaling (Option C) thinking it reacts quickly, but they overlook that simple scaling lacks the smoothing and proportional control needed to avoid unnecessary actions from short bursts, which is the exact requirement in the question.

How to eliminate wrong answers

Option A is wrong because manual scaling requires human intervention to change capacity, which cannot react quickly to CPU spikes and defeats the purpose of automation. Option B is wrong because scheduled scaling is based on predictable time patterns, not real-time load, so it cannot respond to occasional, unpredictable CPU spikes. Option C is wrong because simple scaling policies have a fixed cooldown period and only scale based on a single alarm breach, which can lead to either over-reaction to short bursts or slow response to sustained load, lacking the proportional and smoothing logic of target tracking.

941
MCQmedium

A company deploys a web application on EC2 instances behind an Application Load Balancer. The SysOps administrator needs to allow inbound traffic only from the ALB to the EC2 instances. Currently, the EC2 security group allows inbound HTTP from 0.0.0.0/0. Which security group configuration should the administrator apply?

A.Keep the existing rule that allows inbound HTTP from 0.0.0.0/0, but add a network ACL to block traffic from the internet.
B.Modify the EC2 security group to allow inbound HTTP from the ALB's security group.
C.Modify the EC2 security group to allow inbound HTTP from the ALB's private IP addresses.
D.Modify the EC2 security group to allow inbound HTTP from the ALB's public IP addresses.
AnswerB

Referencing the ALB's security group as the source in the EC2 instance security group creates a highly precise trust boundary: only traffic originating from network interfaces associated with that ALB security group is permitted, so direct internet access to the instance is impossible. Because the rule references the security group ID rather than any IP address, it automatically follows the ALB's elastic network interfaces as the load balancer scales or moves between Availability Zones, requiring no manual updates. This leverages the stateful nature of security groups — return traffic is automatically allowed — and it is the documented, recommended pattern for application load balancer to target communication.

Why this answer

Security groups can reference other security groups as sources, so allowing inbound HTTP on the EC2 instances' security group from the ALB's security group automatically permits traffic from all current and future ALB nodes. This is the AWS-recommended pattern because ALB IP addresses change dynamically and cannot be reliably hardcoded. It also ensures only traffic that passed through the load balancer reaches the instances.

Exam trap

SOA-C02 often tests the misconception that you must whitelist the ALB's IP addresses; the trap is not knowing that security group referencing is the correct, dynamic-safe method.

How to eliminate wrong answers

Option A is wrong because leaving 0.0.0.0/0 open still allows direct internet access to the instances; a network ACL is stateless and subnet-level, so it does not cleanly restrict traffic to only the ALB. Option C is wrong because ALB private IP addresses are not static and change as the load balancer scales, so rules based on them will break. Option D is wrong because instances should never receive traffic from the ALB's public IPs; the ALB forwards traffic from its private nodes, and public IPs are not stable either.

942
MCQmedium

A SysOps administrator uses AWS CloudFormation to manage a stack that includes an Amazon RDS DB instance. The administrator needs to update the stack by changing a parameter that, if applied directly, would replace the database. The administrator wants to prevent accidental replacement during the update. Which CloudFormation feature should the administrator use?

A.Change sets
B.Stack policy
C.Resource-level permissions
D.Stack sets
AnswerB

A stack policy is the correct safeguard because it is a JSON-based policy attached to the CloudFormation stack that explicitly denies specific update actions on specific resources. By adding a rule such as a Deny on 'Update:Replace' for the RDS logical resource ID, the stack update will fail and the RDS instance will be protected from replacement even if an S3 event or trusted user attempts an update. This provides a hard guard that is enforced by CloudFormation before any changes are applied, without preventing other non-replacing updates.

Why this answer

A stack policy is a CloudFormation feature that defines which stack resources can be updated or replaced during a stack update. By setting a stack policy that explicitly denies replacement updates on the RDS DB instance, the administrator can prevent accidental replacement while still allowing other updates. This is the correct approach because it directly controls the update behavior at the resource level without requiring manual intervention.

Exam trap

The trap here is that candidates often confuse change sets (which only preview changes) with stack policies (which actually enforce update restrictions), leading them to select change sets as a safety mechanism when they only provide visibility, not prevention.

How to eliminate wrong answers

Option A is wrong because change sets allow you to preview the changes that will be made to a stack, including whether a resource will be replaced, but they do not prevent the replacement from occurring; they only show what will happen. Option C is wrong because resource-level permissions (via IAM policies) control who can perform actions on resources, not what specific update actions (like replacement) are allowed or denied during a CloudFormation stack update. Option D is wrong because stack sets are used to deploy stacks across multiple accounts and regions, not to control update behavior or prevent replacement of individual resources within a single stack.

943
MCQmedium

Refer to the exhibit. The alarm has been in INSUFFICIENT_DATA state for several hours. What is the most likely cause?

A.The alarm evaluation period is too long.
B.The EC2 instance is stopped or terminated.
C.The instance has no CloudWatch agent installed.
D.The instance is running but the CPU utilization is below the threshold.
AnswerB

If the instance is stopped, no metrics are emitted.

Why this answer

The INSUFFICIENT_DATA state for several hours indicates that CloudWatch has not received any metric data points for the specified period. If the EC2 instance is stopped or terminated, the CloudWatch agent stops sending metrics, and the default CPU utilization metric (which is published by AWS, not the agent) also ceases because the instance is no longer running. This causes the alarm to remain in INSUFFICIENT_DATA indefinitely until the instance is started again or the metric resumes.

Exam trap

The trap here is that candidates often confuse INSUFFICIENT_DATA with ALARM or OK states, mistakenly thinking low CPU utilization or missing CloudWatch agent would cause this state, when in fact INSUFFICIENT_DATA strictly means no metric data has been received at all for the evaluation period.

How to eliminate wrong answers

Option A is wrong because the alarm evaluation period being too long would only delay transitions between states, but it would not cause a permanent INSUFFICIENT_DATA state; data would still be collected and eventually evaluated. Option C is wrong because the CPU utilization metric is a default EC2 metric published by AWS automatically without requiring the CloudWatch agent; the agent is only needed for custom or OS-level metrics. Option D is wrong because if the instance is running and CPU utilization is below the threshold, the alarm would be in ALARM or OK state (depending on the comparison operator), not INSUFFICIENT_DATA; INSUFFICIENT_DATA specifically means no data points are available, not that data exists but is below a threshold.

944
MCQhard

A company uses AWS CloudTrail to log API calls in a multi-account environment. The security team wants to be alerted when an IAM user in the production account modifies a security group to allow inbound SSH from 0.0.0.0/0. Which combination of actions should be taken to meet this requirement?

A.Use AWS Config managed rule 'restricted-ssh' to detect the security group change and trigger an SNS notification.
B.Enable AWS Security Hub and configure a custom insight to detect the security group modification.
C.Create an AWS Lambda function that is triggered by CloudTrail events and publishes to SNS.
D.Stream CloudTrail logs to CloudWatch Logs, create a metric filter for the specific API call, and set a CloudWatch Alarm that sends a notification to an SNS topic.
AnswerD

Streaming CloudTrail logs to CloudWatch Logs is the foundation for real-time monitoring of API activity. You can then create a CloudWatch Logs metric filter that matches the specific API call, such as AuthorizeSecurityGroupIngress or RevokeSecurityGroupIngress, and increments a custom metric. Finally, set a CloudWatch alarm on that metric with a threshold of one, which triggers an SNS notification to the designated topic. This is the standard, event-driven method that provides immediate alerting with minimal overhead and is the recommended pattern for API call monitoring.

Why this answer

CloudTrail logs can be streamed to CloudWatch Logs, where a metric filter can be created to match the specific API call (e.g., AuthorizeSecurityGroupIngress with a CIDR of 0.0.0.0/0 and port 22). A CloudWatch Alarm based on that metric can then trigger an SNS notification, providing a real-time alert for the exact security group modification.

Exam trap

The trap here is that candidates may confuse AWS Config rules (which are reactive and evaluate configuration state) with CloudWatch metric filters (which provide real-time alerting on API calls), leading them to choose Option A despite its inability to trigger immediate notifications on the specific event.

How to eliminate wrong answers

Option A is wrong because the AWS Config managed rule 'restricted-ssh' only checks whether security groups allow unrestricted SSH access at the time of evaluation; it does not provide real-time alerting on the API call itself and cannot trigger an SNS notification directly without additional configuration. Option B is wrong because Security Hub custom insights are used for querying and visualizing findings, not for real-time alerting; they do not directly send notifications to SNS. Option C is wrong because CloudTrail events cannot directly trigger a Lambda function; CloudTrail delivers events to an S3 bucket or CloudWatch Logs, and Lambda can be triggered from those sources, but the option states 'triggered by CloudTrail events' which is technically incorrect without an intermediary.

945
MCQhard

A company hosts a multi-tier web application on AWS. The application consists of an Application Load Balancer (ALB), a fleet of Amazon EC2 instances running in an Auto Scaling group, and an Amazon RDS for MySQL database. The application is accessed by users worldwide. Recently, the company has expanded to new geographic regions, and users in those regions are experiencing high latency. The SysOps administrator is tasked with optimizing performance for global users while keeping costs low. The administrator has already implemented Amazon CloudFront as a CDN for static content. However, dynamic content that requires database queries is still slow. The application's Auto Scaling group is configured with a dynamic scaling policy based on average CPU utilization, but the scaling is not responsive enough during traffic spikes, causing performance degradation. Additionally, the database is a single db.r5.large instance in the us-east-1 region, and all traffic must hit that database, causing high latency for remote users. The administrator needs to propose a comprehensive solution that addresses both compute and database performance issues globally, while considering cost. Which solution is MOST effective?

A.Increase the minimum and maximum size of the Auto Scaling group and use a step scaling policy based on memory utilization.
B.Use Amazon Aurora Global Database to create read replicas in other regions, and configure the Auto Scaling group with a target tracking scaling policy based on request count per target.
C.Implement Amazon ElastiCache for Redis to cache database queries, and use predictive scaling for the Auto Scaling group.
D.Use larger EC2 instances (e.g., c5.2xlarge) for the application tier and provision a Multi-AZ RDS instance for better performance.
AnswerB

Amazon Aurora Global Database replicates data across AWS Regions with a typical replication lag of under one second, allowing application reads to be served from regional read replicas. This dramatically reduces cross-region network latency for global users, making database reads fast regardless of user location. Configuring the Auto Scaling group with a target tracking policy based on Application Load Balancer request count per target directly aligns compute capacity with incoming traffic patterns, allowing the web tier to scale quickly during demand spikes without waiting for memory or CPU alarms to fire.

Why this answer

Amazon Aurora Global Database provides low-latency read replicas in other regions for dynamic content, reducing latency for global users. Additionally, a target tracking scaling policy based on request count per target is more responsive to traffic spikes than CPU-based scaling, as it directly reflects application load. Option A is wrong because using step scaling based on memory utilization does not address global latency and memory may not be the bottleneck.

Option C is wrong because ElastiCache caching reduces database load but does not reduce latency for users far from the primary database; predictive scaling may not handle sudden spikes well. Option D is wrong because using larger instances and Multi-AZ does not reduce global latency and increases cost.

946
MCQeasy

A SysOps administrator needs to send a notification when an EC2 instance's CPU utilization exceeds 90% for 5 consecutive minutes. Which AWS service should be used to create the alarm?

A.AWS Config
B.Amazon CloudWatch Alarms
C.AWS Trusted Advisor
D.Amazon EventBridge
AnswerB

Amazon CloudWatch Alarms are the native mechanism for monitoring a single metric or a metric math expression over a specified time period. When the metric breaches a defined threshold for consecutive periods, the alarm transitions to the ALARM state and can trigger an action, such as publishing to an SNS topic to send a notification. This directly fulfills the requirement to notify when a metric condition is met.

Why this answer

Amazon CloudWatch Alarms are the correct service because they allow you to monitor a specific metric (e.g., EC2 CPUUtilization) and trigger an action when the metric crosses a defined threshold for a specified number of consecutive evaluation periods. In this case, you can create a CloudWatch alarm with a statistic of 'Average', a threshold of 90%, and set the 'Datapoints to Alarm' and 'Evaluation Periods' to 5 (with a period of 1 minute) to achieve the '5 consecutive minutes' requirement. The alarm can then send a notification via Amazon SNS.

Exam trap

The trap here is that candidates often confuse Amazon EventBridge with CloudWatch Alarms, thinking EventBridge can directly evaluate metric thresholds, but EventBridge requires a CloudWatch Alarm to generate the event, and it cannot perform the metric evaluation itself.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating and auditing the configuration of AWS resources against desired policies (e.g., compliance rules), not for monitoring real-time performance metrics like CPU utilization. Option C is wrong because AWS Trusted Advisor provides best-practice recommendations for cost optimization, security, fault tolerance, and performance, but it does not create metric-based alarms or send notifications for threshold breaches. Option D is wrong because Amazon EventBridge is a serverless event bus for routing events between AWS services and custom applications, but it cannot directly evaluate a metric over a time window; it relies on CloudWatch Alarms or other sources to generate the events that trigger its rules.

947
MCQeasy

A company uses AWS Elastic Beanstalk to deploy a web application. The environment is running low on memory, and the administrator needs to change the instance type from t2.micro to t3.small. What is the correct way to perform this change with minimal downtime?

A.Modify the instance type in the Elastic Beanstalk environment's configuration.
B.Terminate the environment and create a new one with the desired instance type.
C.Create a new environment and perform a swap URL.
D.Manually modify the Auto Scaling group's launch configuration.
AnswerA

Modify the instance type by updating the 'Instances' or 'Capacity' configuration in the Elastic Beanstalk environment's management console, EB CLI, or API. Elastic Beanstalk then performs a rolling update, replacing EC2 instances in batches to keep the application available, and automatically updates the Auto Scaling group's launch configuration. This is the fully supported path for resizing compute resources in a running environment.

Why this answer

Modifying the instance type in the Elastic Beanstalk environment's configuration triggers a rolling update or immutable update, which replaces instances with the new type while keeping the environment running. This approach minimizes downtime because Elastic Beanstalk manages the instance replacement process automatically, ensuring that traffic continues to be served during the transition.

Exam trap

The trap here is that candidates often think manual changes to the Auto Scaling group (Option D) are acceptable, but Elastic Beanstalk treats such manual modifications as configuration drift, which can cause the environment to become out of sync and fail subsequent managed updates.

How to eliminate wrong answers

Option B is wrong because terminating the environment and creating a new one causes complete downtime and loses environment configuration, which is unnecessary when a simple configuration change can be applied. Option C is wrong because creating a new environment and performing a swap URL (CNAME swap) is a blue/green deployment strategy that introduces additional complexity and cost, and is not the minimal-downtime approach for a simple instance type change. Option D is wrong because manually modifying the Auto Scaling group's launch configuration does not update running instances; it only affects new instances launched in the future, and it bypasses Elastic Beanstalk's managed updates, potentially causing drift and inconsistent state.

948
MCQmedium

A company uses AWS CodeDeploy to deploy a new version of an application to EC2 instances in an Auto Scaling group behind an Application Load Balancer. The company requires zero downtime during the deployment. Which deployment configuration should be used?

A.CodeDeployDefault.AllAtOnce
B.CodeDeployDefault.OneAtATime
C.CodeDeployDefault.EC2/OnPremises: BlueGreenDeployment
D.Create a blue/green deployment by configuring CodeDeploy to launch new instances and shift traffic after validation.
AnswerD

A blue/green deployment with CodeDeploy provisions a new, separate fleet of EC2 instances (green) and installs the new revision on them while the original fleet (blue) continues to serve traffic. After validation of the green instances—through health checks, tests, or a manual approval—you can shift traffic from the blue fleet to the green fleet using an Application Load Balancer. This approach isolates the new version from production traffic until it is verified, and if a problem arises you can instantly reroute traffic back to the blue fleet, ensuring zero downtime.

Why this answer

A blue/green deployment with CodeDeploy, where new instances are launched and traffic is shifted only after validation, ensures zero downtime by keeping the old environment (blue) fully serving traffic until the new environment (green) is verified healthy. This approach avoids any in-place updates that could temporarily reduce capacity or cause service disruption, meeting the requirement for zero downtime during deployment.

Exam trap

The trap here is that candidates confuse the predefined deployment configurations (like AllAtOnce or OneAtATime) with the blue/green deployment method, not realizing that blue/green is a separate deployment type configured in the deployment group settings, not a predefined configuration name.

How to eliminate wrong answers

Option A is wrong because CodeDeployDefault.AllAtOnce deploys to all instances simultaneously, which can cause downtime if the new version fails or requires a restart, as there is no gradual traffic shifting or rollback capability. Option B is wrong because CodeDeployDefault.OneAtATime deploys to one instance at a time, which minimizes risk but still involves in-place updates that can cause brief interruptions if the application requires a restart or health check failure during deployment. Option C is wrong because CodeDeployDefault.EC2/OnPremises: BlueGreenDeployment is not a valid predefined deployment configuration name; CodeDeploy does not have a built-in configuration with that exact string, and blue/green deployments must be explicitly configured via the deployment group settings, not selected as a predefined configuration.

949
MCQmedium

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The SysOps administrator needs to monitor the application's HTTP 5xx error rate and set an alarm when the error rate exceeds 5% over a 5-minute period. The alarm must trigger an Amazon SNS notification. Which metric should be used for the alarm?

A.HTTPCode_ELB_5XX_Count
B.HTTPCode_Target_5XX_Count
C.RequestCount
D.TargetResponseTime
AnswerB

The HTTPCode_Target_5XX_Count metric reports the number of HTTP 5xx responses returned directly by the registered EC2 instances, capturing errors such as 500 Internal Server Error from the application. Because these are the actual responses sent to clients from your web application, this metric is the correct measurement for an alarm that detects application-level failures. You can then create a CloudWatch alarm on this metric to trigger when the count exceeds a threshold, possibly combined with RequestCount to derive an error rate.

Why this answer

The alarm must monitor the error rate from the application targets (EC2 instances) behind the ALB. HTTPCode_Target_5XX_Count tracks HTTP 5xx responses generated by the targets themselves, which directly reflects application-level errors. To calculate the error rate, you would divide this metric by RequestCount, but the metric itself is the correct source for target-side 5xx errors.

Exam trap

The trap here is that candidates confuse HTTPCode_ELB_5XX_Count with HTTPCode_Target_5XX_Count, assuming all 5xx errors originate from the load balancer, when in fact the ALB separates its own errors from target-generated errors to provide precise fault isolation.

How to eliminate wrong answers

Option A is wrong because HTTPCode_ELB_5XX_Count tracks 5xx errors generated by the ALB itself (e.g., due to load balancer failures or configuration issues), not the application targets, so it would not reflect the application's error rate. Option C is wrong because RequestCount is a count of all requests processed by the ALB, not a measure of error rate; it is used as a denominator in rate calculations but cannot trigger an alarm on error percentage alone. Option D is wrong because TargetResponseTime measures the time taken for targets to respond, not error codes, and is unrelated to HTTP 5xx error rate monitoring.

950
MCQeasy

A company stores log files in Amazon S3. The logs are accessed frequently for the first 30 days, then rarely after that. The company wants to automatically transition objects to a lower-cost storage class after 30 days. Which S3 feature should be configured?

A.S3 Lifecycle rule
B.S3 Versioning
C.S3 Transfer Acceleration
D.S3 Object Lock
AnswerA

S3 Lifecycle rules automatically transition objects to lower-cost storage classes (e.g., from S3 Standard to S3 Standard-IA or S3 Glacier Deep Archive) based on age or other criteria. For logs that are accessed infrequently after 30 days and must be retained for years, a lifecycle rule is the most cost-effective way to automate the transition and expiration of objects while still meeting retention requirements.

Why this answer

An S3 Lifecycle rule is the correct feature because it allows you to define a transition action that automatically moves objects from a higher-cost storage class (e.g., S3 Standard) to a lower-cost storage class (e.g., S3 Standard-IA or S3 Glacier) after a specified number of days. This directly meets the requirement to transition logs after 30 days without manual intervention, optimizing storage costs based on access patterns.

Exam trap

The trap here is that candidates may confuse S3 Versioning or Object Lock as tools for cost optimization, but they are governance features, not lifecycle management features, and do not automate storage class transitions.

How to eliminate wrong answers

Option B is wrong because S3 Versioning is used to preserve, retrieve, and restore every version of an object, not to automate storage class transitions; it does not provide any cost optimization based on age. Option C is wrong because S3 Transfer Acceleration is a feature that speeds up uploads over long distances using AWS edge locations, and it has no role in managing storage class transitions or lifecycle policies. Option D is wrong because S3 Object Lock is designed to prevent objects from being deleted or overwritten for a fixed retention period, and it does not automate transitions to lower-cost storage classes.

951
MCQeasy

A company uses AWS CloudTrail to log API activity. The SysOps administrator needs to ensure that log files are protected from accidental deletion and are available for compliance audits for at least 7 years. Which service should be used to meet these requirements?

A.Enable S3 Object Lock in Compliance mode.
B.Enable S3 Versioning on the CloudTrail S3 bucket.
C.Move CloudTrail logs to Amazon S3 Glacier after 90 days.
D.Store logs in Amazon CloudWatch Logs with an expiration policy of 7 years.
AnswerA

S3 Object Lock in Compliance mode enforces a write-once-read-many (WORM) model, preventing any user — including the AWS account root user — from deleting or overwriting the log objects for the specified retention period. Because CloudTrail log files become immutable once written, this directly satisfies the need to preserve audit activity for 7 years in an unmodifiable state.

Why this answer

S3 Object Lock in Compliance mode provides a write-once-read-many (WORM) model that prevents any user, including the AWS account root user, from deleting or overwriting objects for the specified retention period. This meets both the protection from accidental deletion and the 7-year compliance audit requirement, as the retention mode cannot be shortened or removed once applied.

Exam trap

The trap here is that candidates often confuse versioning (which only preserves copies) with immutability (which prevents deletion entirely), or they assume that moving data to a cold storage tier like Glacier automatically protects it from deletion, when in fact Glacier objects are still deletable without an additional lock mechanism.

How to eliminate wrong answers

Option B is wrong because S3 Versioning alone does not prevent deletion; it only preserves previous versions of objects, but a user with s3:DeleteObject permission can still delete the current version, and versioned delete markers can be removed. Option C is wrong because moving logs to S3 Glacier after 90 days does not inherently protect them from deletion; Glacier objects can still be deleted unless additional controls like Object Lock are applied. Option D is wrong because CloudWatch Logs expiration policies only control log retention and automatic deletion, but they do not provide a WORM lock to prevent accidental or malicious deletion before the expiration date.

952
MCQhard

A SysOps administrator needs to route traffic to multiple AWS regions for disaster recovery using Amazon Route 53. The primary region should receive all traffic unless it becomes unhealthy. Which routing policy should be used?

A.Failover routing policy
B.Geolocation routing policy
C.Latency routing policy
D.Weighted routing policy
AnswerA

Route 53 failover routing policy implements active-passive failover by using two records with the same name and type: a primary record associated with a health check and a secondary record. When the primary health check fails, Route 53 automatically returns the secondary record's value in DNS responses, directing traffic to the second resource. This policy is expressly designed to keep services available if the primary target becomes unhealthy, which matches the requirement to route traffic to multiple AWS endpoints with automatic failover.

Why this answer

Failover routing policy is correct because it allows you to configure an active-passive setup where all traffic is directed to a primary resource (e.g., an Elastic Load Balancer in the primary region) unless Route 53 health checks determine that the primary is unhealthy. When the primary fails, Route 53 automatically routes traffic to the secondary (disaster recovery) resource in another region. This directly meets the requirement of sending all traffic to the primary region unless it becomes unhealthy.

Exam trap

The trap here is that candidates often confuse failover routing with weighted or latency routing, thinking they can achieve disaster recovery by distributing traffic, but only failover routing provides the required active-passive health-based failover behavior.

How to eliminate wrong answers

Option B (Geolocation routing policy) is wrong because it routes traffic based on the geographic location of the user, not based on the health of the resource; it does not provide automatic failover to a disaster recovery region. Option C (Latency routing policy) is wrong because it routes traffic to the region with the lowest latency for the user, which does not guarantee that all traffic goes to a single primary region unless it becomes unhealthy. Option D (Weighted routing policy) is wrong because it distributes traffic across multiple resources based on assigned weights, not on health status; it cannot ensure that all traffic goes to the primary region unless it fails.

953
Multi-Selectmedium

Which THREE AWS features can be used to improve the performance of an Amazon DynamoDB table that is experiencing high read latency? (Choose THREE.)

Select 3 answers
A.Enable DynamoDB Accelerator (DAX).
B.Use DynamoDB global tables.
C.Enable Auto Scaling for read capacity.
D.Use Time to Live (TTL) to delete old items.
E.Increase the provisioned read capacity units.
AnswersA, B, E

DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of your DynamoDB table, serving reads at microsecond latency by avoiding disk access and the complexity of managing a separate caching tier. As a justified correct answer, it directly addresses read latency for even the most frequently accessed items, offloading repeated read traffic from the table's provisioned capacity and reducing the time each request takes from the database engine itself.

Why this answer

DynamoDB Accelerator (DAX) is a fully managed, highly available, in-memory cache that can reduce DynamoDB response times from milliseconds to microseconds. By caching frequently read items, DAX offloads read requests from the underlying table, directly addressing high read latency without requiring additional read capacity units or table modifications.

Exam trap

The trap here is that candidates often confuse Auto Scaling (which prevents throttling) with a performance improvement feature, but Auto Scaling does not reduce latency for individual read requests—it only ensures sufficient capacity to avoid throttling.

954
MCQmedium

A company runs an application on Amazon EC2 instances in private subnets of a VPC. The application needs to upload files to an Amazon S3 bucket in the same AWS Region. The SysOps administrator wants to ensure that traffic to S3 does not traverse the internet and minimizes data transfer costs. Which solution should the administrator implement?

A.Create a NAT gateway in a public subnet and route private subnet traffic to it.
B.Create an S3 Gateway Endpoint and add a route in the private subnet route table pointing to it.
C.Create an S3 Interface Endpoint and assign a security group.
D.Use AWS PrivateLink to connect to S3.
AnswerB

An S3 Gateway Endpoint is a free, highly available gateway object attached to a VPC that uses a prefix list to route S3 traffic from a private subnet without going over the internet. You must add a route in the private subnet's route table with the destination as the S3 prefix list and the target as the gateway endpoint; traffic stays entirely inside the AWS network. This is the recommended pattern for private subnets because it requires no NAT gateway, no IGW, and no data transfer charges.

Why this answer

An S3 Gateway Endpoint is the correct solution because it provides private connectivity from a VPC to S3 without traversing the internet, using AWS's internal network. By adding a route in the private subnet's route table pointing to the gateway endpoint, traffic to S3 stays within the AWS backbone, minimizing data transfer costs (no NAT gateway charges) and avoiding internet egress fees.

Exam trap

The trap here is that candidates often confuse Gateway Endpoints with Interface Endpoints, assuming Interface Endpoints are always better because they use security groups, but for S3, Gateway Endpoints are free and more cost-effective, while Interface Endpoints incur additional charges.

How to eliminate wrong answers

Option A is wrong because a NAT gateway routes traffic through the internet to reach S3, incurring data transfer costs and NAT gateway hourly charges, and it still traverses the internet, violating the requirement to avoid internet traversal. Option C is wrong because an S3 Interface Endpoint (powered by AWS PrivateLink) incurs hourly charges and per-GB data processing fees, making it more expensive than a Gateway Endpoint for S3, and it is typically used for services that don't support Gateway Endpoints (e.g., DynamoDB, API Gateway). Option D is wrong because AWS PrivateLink is the underlying technology for Interface Endpoints, not a separate solution; using PrivateLink directly would still involve Interface Endpoint costs and complexity, and it is not the optimal choice for S3 when a Gateway Endpoint is available.

955
MCQhard

A company's security team notices that an IAM user has been generating multiple access keys and deleting them within a short period. The SysOps administrator needs to detect and alert on this behavior. Which solution is the MOST effective?

A.Enable AWS Trusted Advisor security checks and review the report weekly.
B.Enable IAM Access Analyzer to analyze user activity and send alerts.
C.Enable AWS CloudTrail and create a CloudWatch Events rule that triggers on iam:CreateAccessKey events and sends a notification to an SNS topic.
D.Use AWS Config to track IAM user configuration changes and trigger an alert when an access key is created.
AnswerC

AWS CloudTrail is the correct foundation because it records management events, including the iam:CreateAccessKey API call, as a JSON audit log. A CloudWatch Events rule (or Amazon EventBridge rule) can use an event pattern to match the eventSource, eventName, and other fields, then invoke an SNS topic to send an email or SMS notification. This combination provides near-real-time detection of the exact API action—something a periodic report or static policy scanner cannot deliver.

Why this answer

The behavior described—creating and deleting access keys rapidly—is an API-level event that must be captured in real time. AWS CloudTrail logs all IAM API calls, including CreateAccessKey and DeleteAccessKey. By creating a CloudWatch Events (now Amazon EventBridge) rule that matches the iam:CreateAccessKey event, you can trigger an SNS notification immediately, enabling the security team to detect and respond to the suspicious activity.

This solution is the most effective because it provides near-real-time detection and alerting based on the specific API call.

Exam trap

SOA-C02 often tests the difference between services that monitor configuration changes (AWS Config) and those that capture API activity (CloudTrail), and candidates may incorrectly choose AWS Config because it can track IAM changes, but it lacks real-time event-driven alerting on specific API calls.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor security checks focus on best practices and configuration weaknesses (e.g., MFA on root, public S3 buckets), not on real-time API activity like access key creation; weekly reviews are too slow. Option B is wrong because IAM Access Analyzer analyzes resource policies to identify external access, not user API activity; it does not monitor or alert on CreateAccessKey events. Option D is wrong because AWS Config tracks resource configuration changes and can alert on access key creation, but it is not real-time and does not capture the API call details or the rapid deletion pattern as effectively as CloudTrail/CloudWatch Events.

956
MCQmedium

A company uses AWS Lambda functions that process data from an Amazon SQS queue. The Lambda function is failing intermittently due to timeouts. The SysOps administrator needs to be notified immediately when the function times out. What is the most efficient way to achieve this?

A.Modify the Lambda function to catch the timeout exception and log it to CloudWatch Logs
B.Create a CloudWatch alarm on the Lambda Errors metric that sends an SNS notification
C.Set up a CloudWatch Logs subscription filter to send error logs to an SNS topic
D.Configure an Amazon SNS topic as a Lambda destination for the function
AnswerB

A CloudWatch alarm on the AWS/Lambda Errors metric directly monitors every failed invocation, including timeouts, and transitions to ALARM when the error count exceeds the configured threshold. The alarm then publishes to an SNS topic, which can deliver email, SMS, or trigger a webhook or incident-management tool. This is the native, low-latency path because the metric is emitted automatically by the Lambda service without any code changes or log processing.

Why this answer

A CloudWatch alarm on the Lambda Errors metric directly monitors function invocations that result in errors, including timeouts. When the alarm state changes to ALARM, it can immediately trigger an SNS notification, providing the fastest and most efficient notification mechanism without requiring code changes or additional infrastructure.

Exam trap

The trap here is that candidates often assume they can catch a timeout exception inside the function code (Option A) or that Lambda destinations (Option D) are a catch-all for all errors, when in fact they only apply to specific invocation types and do not cover all timeout scenarios.

How to eliminate wrong answers

Option A is wrong because catching a timeout exception inside the Lambda function is not possible — Lambda enforces a hard timeout at the configured limit, and the runtime cannot catch it; the function simply terminates with an error. Option C is wrong because a CloudWatch Logs subscription filter requires the function to first write a log entry, which may not occur reliably on timeout, and it adds latency and complexity compared to a direct metric alarm. Option D is wrong because Lambda destinations are triggered only on successful invocation or explicit failure states like 'OnFailure' for asynchronous invocations, but they do not capture all timeout errors reliably and require additional configuration; also, SNS as a destination is not supported for synchronous invocations like those from SQS.

957
MCQhard

A company has a VPC with public and private subnets. A NAT Gateway is deployed in the public subnet to allow instances in the private subnet to access the internet. However, private instances cannot reach an external service at 203.0.113.50:443. What should be checked first?

A.The route table for the private subnet has a route 0.0.0.0/0 pointing to the NAT Gateway.
B.The NAT Gateway has an Elastic IP assigned.
C.The security group for the NAT Gateway allows inbound traffic from the private subnet.
D.The internet gateway is attached to the VPC.
AnswerA

Private-subnet instances reach the internet only if their route table has 0.0.0.0/0 targeting the NAT Gateway. Verifying this route first confirms whether traffic can leave the subnet at all before investigating security groups, NACLs or the NAT Gateway itself.

Why this answer

The most common cause of private instances failing to reach the internet via a NAT Gateway is a missing or incorrect route in the private subnet's route table. The private subnet must have a route for 0.0.0.0/0 pointing to the NAT Gateway; otherwise, traffic has no path out. This is the first thing to verify because it is the fundamental routing requirement for NAT Gateway functionality.

Exam trap

SOA-C02 often tests the order of troubleshooting steps: candidates may jump to checking the NAT Gateway's Elastic IP or security groups, but the most common and first thing to verify is the private subnet's route table.

How to eliminate wrong answers

Option B is wrong because a NAT Gateway always requires an Elastic IP at creation, so it is not a likely misconfiguration; checking it first is less efficient. Option C is wrong because NAT Gateways do not use security groups; they are managed by AWS and do not have security group associations. Option D is wrong because if the internet gateway were not attached, the NAT Gateway itself would not function, but the question specifies the NAT Gateway is deployed, implying the IGW is likely present; also, the private subnet route is more directly related to the private instances' failure.

958
MCQhard

A company has deployed a global web application using AWS CloudFront with an Application Load Balancer (ALB) as the origin. The ALB is in a single AWS region. Users in different geographic regions report high latency, and some users are unable to access the application. The SysOps administrator verifies that the CloudFront distribution is configured correctly and that the ALB is healthy. The administrator also confirms that the ALB's security group allows traffic from the CloudFront IP ranges. What is the most likely cause of the issue?

A.The ALB is overwhelmed by the number of concurrent connections from CloudFront
B.CloudFront is not caching content, causing all requests to go to the origin
C.The CloudFront distribution is using TCP instead of HTTP, causing higher latency
D.The SSL/TLS certificate on the ALB is not trusted by CloudFront
AnswerA

CloudFront's global network of edge locations each establishes a pool of persistent (keep-alive) TCP connections to the ALB origin. In a busy distribution, the aggregate of these connections across all edges can exceed the ALB's concurrent connection capacity (MaxConnectionIdleTime, target group limits, or instance/scale limits), causing SYN queue saturation, latency, and timeouts. The fix is to scale the ALB and adjust keep-alive timeouts, not to assume caching or TLS errors.

Why this answer

The ALB in a single region can become overwhelmed by the high volume of concurrent connections from CloudFront's global edge locations, even though the security group allows traffic from CloudFront IP ranges. This can cause high latency and access failures for users in different regions. Option B is incorrect because CloudFront caching typically reduces the load on the origin by serving cached content at edge locations.

Option C is incorrect because CloudFront distributions support HTTP and HTTPS protocols, not TCP/UDP, and the protocol used does not explain the regional latency issue. Option D is incorrect because the SSL/TLS certificate on the ALB must be trusted by CloudFront for HTTPS connections, but this would not cause intermittent access issues across regions if the distribution is configured correctly.

959
MCQeasy

A SysOps administrator needs to track changes made to an Amazon S3 bucket policy and receive notifications when changes occur. Which AWS service should be used?

A.AWS Trusted Advisor
B.Amazon CloudWatch Events
C.AWS Config
D.AWS CloudTrail
AnswerC

AWS Config is the correct choice because it continuously records and evaluates configuration changes to supported AWS resources, including S3 bucket policies. You can set up AWS Config rules to detect specific changes and configure Amazon SNS to send notifications when a resource deviates from a desired configuration or when a change occurs. It maintains a configuration history and timeline, enabling both auditing and alerting for modifications, which directly matches the requirement.

Why this answer

AWS Config is designed to record configuration changes to AWS resources, including S3 bucket policies, and can evaluate them against desired configurations. It can send notifications via Amazon SNS when changes occur, and it provides a history of configuration changes. This directly meets the requirement to track changes and receive notifications.

Exam trap

The trap is confusing CloudTrail (API activity logging) with AWS Config (configuration change tracking); candidates often pick CloudTrail because it 'tracks changes', but it does not provide configuration state or compliance notifications.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor provides best-practice checks and recommendations, not change tracking or notifications for specific resource policies. Option B is wrong because Amazon CloudWatch Events (now Amazon EventBridge) can react to API calls via CloudTrail, but it does not track configuration changes or provide a configuration history; it's an event bus. Option D is wrong because AWS CloudTrail records API activity (who made the change) but does not track resource configuration state or send notifications on configuration changes by itself; it logs events but requires additional services for alerting.

960
MCQmedium

Operators have been making direct changes to AWS resources (security group rules, IAM policy modifications) that were originally created by CloudFormation stacks. The team wants to identify which stacks and specific resources have drifted from their template definitions. What is the correct tool and operation sequence?

A.Run drift detection on each CloudFormation stack; review the results in the Drift status panel to see which resources have MODIFIED or DELETED status
B.Enable AWS Config conformance packs that check CloudFormation stack compliance against desired template states
C.Re-deploy all stacks with the original templates using CloudFormation update-stack to overwrite any manual changes
D.Use AWS Trusted Advisor to identify resources that have been modified outside of their originating CloudFormation stacks
AnswerA

Drift detection calls AWS APIs to read the current configuration of each resource and compares it to the template. Resources with live configurations differing from the template are marked MODIFIED. Deleted resources outside the stack are marked DELETED. The results show the exact property-level differences, enabling targeted remediation.

Why this answer

AWS CloudFormation drift detection is the correct tool because it directly compares the current state of resources in a stack (including security group rules and IAM policies) against the stack's template definitions. Running drift detection on each stack and reviewing the Drift status panel reveals which resources have been modified or deleted outside of CloudFormation, providing the exact identification the team needs.

Exam trap

The trap here is that candidates may confuse drift detection with compliance checks (AWS Config) or remediation actions (update-stack), but the question specifically asks for identification of drifted stacks and resources, not remediation or compliance evaluation.

How to eliminate wrong answers

Option B is wrong because AWS Config conformance packs evaluate resource compliance against rules, not against CloudFormation template states; they cannot detect drift from a specific stack template. Option C is wrong because re-deploying stacks with update-stack overwrites manual changes but does not identify which stacks or resources have drifted; it is a remediation action, not a detection tool. Option D is wrong because AWS Trusted Advisor checks for best practices and cost optimization, not for drift between CloudFormation templates and actual resource configurations.

961
MCQeasy

A company wants to ensure that all data in Amazon S3 is encrypted at rest using server-side encryption with AWS KMS managed keys (SSE-KMS). Which bucket policy statement should be used to deny any PUT request that does not include the 'x-amz-server-side-encryption' header with value 'aws:kms'?

A.Condition: { StringNotEquals: { 's3:x-amz-server-side-encryption': 'aws:kms' } }
B.Condition: { StringNotEquals: { 's3:x-amz-server-side-encryption-aws-kms-key-id': 'alias/aws/s3' } }
C.Condition: { StringNotEquals: { 's3:ServerSideEncryption': 'KMS' } }
D.Condition: { StringEquals: { 's3:x-amz-server-side-encryption': 'aws:kms' } }
AnswerA

This Deny policy uses StringNotEquals on the s3:x-amz-server-side-encryption request header to block any PUT that does not explicitly specify aws:kms. Because it's an explicit deny, it overrides all allows, so objects uploaded with AES256, no encryption header, or any other value are rejected. That directly enforces the requirement that all S3 data be encrypted with SSE-KMS.

Why this answer

It uses the condition key 's3:x-amz-server-side-encryption' and denies the request when the header value is not 'aws:kms', enforcing SSE-KMS. Option B is incorrect because it checks the specific KMS key ID via 's3:x-amz-server-side-encryption-aws-kms-key-id', not the encryption type. Option C is incorrect because 's3:ServerSideEncryption' is not a valid condition key for S3; the correct key is 's3:x-amz-server-side-encryption'.

Option D is incorrect because using StringEquals with this condition key would only match requests that include the header with value 'aws:kms', but a deny policy with this condition would not block requests that omit the header entirely; additionally, the question requires denying requests that do not include the specified encryption header.

962
MCQmedium

A company is running a web application on EC2 instances behind an Application Load Balancer. The application experiences intermittent latency spikes. The SysOps administrator needs to identify the root cause. Which set of CloudWatch metrics should be analyzed first?

A.ALB TargetResponseTime and EC2 CPUUtilization
B.EC2 CPUUtilization and NetworkIn
C.EC2 StatusCheckFailed and ALB UnhealthyHostCount
D.ALB RequestCount and HealthyHostCount
AnswerA

TargetResponseTime isolates ALB-to-target latency, showing whether spikes originate at the application tier, while CPUUtilization reveals host saturation driving slow responses. Together they distinguish backend compute pressure from network or client-side delay, the fastest first step before drilling into ELB 5xx counts or request queues.

Why this answer

Intermittent latency spikes in a web application behind an Application Load Balancer (ALB) are most directly investigated by correlating ALB TargetResponseTime (which measures the time taken for the target to respond to the ALB) with EC2 CPUUtilization (which indicates whether the instance is under compute pressure). A spike in TargetResponseTime alongside high CPUUtilization suggests the EC2 instance is struggling to process requests, pointing to a compute bottleneck as the root cause.

Exam trap

The trap here is that candidates often confuse latency metrics with availability metrics, choosing options like C or D that indicate failures or traffic volume, rather than the performance-specific metrics needed to diagnose intermittent slowness.

How to eliminate wrong answers

Option B is wrong because while EC2 CPUUtilization is relevant, NetworkIn alone does not directly indicate latency; high network input could be normal traffic and does not measure response time or processing delays. Option C is wrong because EC2 StatusCheckFailed and ALB UnhealthyHostCount indicate instance or health check failures, not intermittent latency spikes; these metrics would show binary health states, not gradual performance degradation. Option D is wrong because ALB RequestCount and HealthyHostCount measure traffic volume and target health, not response latency; high request count alone does not explain why responses are slow.

963
MCQeasy

A company has an application running on EC2 instances in a VPC. The application needs to access an S3 bucket in the same AWS region. Which configuration provides the MOST secure and cost-effective access?

A.Make the S3 bucket publicly accessible and use the public endpoint from the EC2 instances.
B.Set up a NAT Gateway in a public subnet and route traffic from the EC2 instances through it to the S3 endpoint.
C.Create a VPC Gateway Endpoint for S3 and update the route tables for the private subnets.
D.Create an Internet Gateway and route traffic from the EC2 instances through it to a public S3 endpoint.
AnswerC

A VPC Gateway Endpoint routes S3 traffic privately over the AWS network, bypassing NAT Gateways and internet gateways entirely. This removes NAT data-processing charges and keeps traffic off the public internet, satisfying both the security and cost-effectiveness constraints for same-region S3 access.

Why this answer

A VPC Gateway Endpoint for S3 allows EC2 instances in private subnets to access S3 directly over the AWS network without traversing the internet, eliminating the need for a NAT Gateway or Internet Gateway. This provides the most secure and cost-effective access by keeping traffic within the AWS backbone and avoiding data transfer costs associated with NAT Gateways or public endpoints.

Exam trap

The trap here is that candidates often confuse VPC Gateway Endpoints with VPC Interface Endpoints (powered by AWS PrivateLink), but for S3, a Gateway Endpoint is the correct and most cost-effective choice because it does not require an Elastic Network Interface or incur hourly charges, unlike an Interface Endpoint.

How to eliminate wrong answers

Option A is wrong because making the S3 bucket publicly accessible exposes it to the entire internet, violating security best practices and potentially leading to unauthorized access or data breaches. Option B is wrong because a NAT Gateway incurs hourly charges and data processing costs, and it routes traffic through the internet unnecessarily, making it less cost-effective and less secure than a VPC Gateway Endpoint. Option D is wrong because an Internet Gateway is designed for public internet access, and routing EC2 traffic through it to a public S3 endpoint exposes the traffic to the internet, increasing latency and security risks while adding unnecessary complexity and cost.

964
MCQmedium

A company uses AWS CloudFormation to deploy a three-tier web application. The template includes an Amazon RDS DB instance. The SysOps administrator needs to ensure that the database password is not exposed in the template or in the stack outputs. The password should be stored securely and rotated automatically every 90 days. Which solution should the administrator use?

A.Store the password as a plaintext parameter in the CloudFormation template and mark it as NoEcho.
B.Use AWS Systems Manager Parameter Store to store the password as a SecureString and reference it using the dynamic reference {{resolve:ssm-secure:password}} in the template.
C.Use AWS Secrets Manager to store the password and reference it using the dynamic reference {{resolve:secretsmanager:secretId:secretString:password}} in the CloudFormation template. Enable automatic rotation.
D.Hardcode the password in a userdata script that is passed to the EC2 instances.
AnswerC

Secrets Manager is a purpose-built secret management service with managed automatic rotation via a configurable Lambda rotation function, so the database password is rotated on a schedule without custom code. The dynamic reference {{resolve:secretsmanager:secretId:secretString:password}} fetches the current secret value at CloudFormation stack creation/update time without ever putting the password in the template. This combines secure storage, seamless retrieval, and automated rotation, making it the correct choice.

Why this answer

AWS Secrets Manager is designed to securely store secrets like database passwords, supports automatic rotation (including a 90-day schedule), and can be referenced in CloudFormation templates using the dynamic reference {{resolve:secretsmanager:secretId:secretString:password}}. This ensures the password is never exposed in the template or stack outputs, and rotation is handled automatically without manual intervention.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store (which can store SecureStrings but lacks native rotation) with AWS Secrets Manager (which is purpose-built for secrets with automatic rotation), leading them to choose Option B instead of C.

How to eliminate wrong answers

Option A is wrong because marking a parameter as NoEcho only hides it from console output and logs, but the plaintext value is still stored in the template and can be retrieved by anyone with access to the template or stack metadata; it does not provide secure storage or automatic rotation. Option B is wrong because AWS Systems Manager Parameter Store (SecureString) stores the password securely but does not natively support automatic rotation; you would need to build a custom rotation solution, and the dynamic reference {{resolve:ssm-secure:password}} does not trigger rotation. Option D is wrong because hardcoding the password in a userdata script exposes it in plaintext within the EC2 instance metadata and logs, violating security best practices and providing no rotation capability.

965
MCQeasy

A SysOps administrator is setting up a backup plan for an RDS MySQL database. The database is 500 GB in size and is used for a critical application. The company requires a Recovery Point Objective (RPO) of 5 minutes and a Recovery Time Objective (RTO) of 1 hour. Which solution meets these requirements?

A.Deploy the RDS instance in a Multi-AZ configuration with automatic failover.
B.Use AWS Backup to copy snapshots to a different AWS Region.
C.Take manual snapshots every 5 minutes and store them in Amazon S3.
D.Configure a cross-region read replica and promote it during failover.
AnswerA

Multi-AZ RDS uses synchronous replication to a standby instance in a different Availability Zone; every transaction is committed on both the primary and standby before being acknowledged, giving an RPO of zero. Should the primary fail, AWS automatically flips the DNS endpoint to the standby, typically completing failover in 60–120 seconds, comfortably meeting a 1-hour RTO. This makes it the only option that satisfies both the 5-minute RPO and the 1-hour RTO without manual intervention.

Why this answer

Multi-AZ RDS with automatic failover meets the RPO of 5 minutes and RTO of 1 hour because it synchronously replicates data to a standby instance in a different Availability Zone. In the event of a failure, Amazon RDS automatically fails over to the standby, typically completing within 60–120 seconds, which satisfies the RTO. The synchronous replication ensures zero data loss (RPO of effectively 0), well within the 5-minute requirement.

Exam trap

The trap here is that candidates confuse Multi-AZ (synchronous, zero data loss, automatic failover) with read replicas (asynchronous, potential data loss, manual promotion), and assume cross-region replicas can meet tight RPO/RTO when they cannot due to replication lag and promotion time.

How to eliminate wrong answers

Option B is wrong because AWS Backup cross-region snapshot copies are asynchronous and typically run on a schedule (e.g., hourly/daily), which cannot achieve a 5-minute RPO. Option C is wrong because manual snapshots cannot be taken every 5 minutes—Amazon RDS enforces a minimum interval of 5 minutes between manual snapshots, and the process itself takes time, making it impossible to meet a 5-minute RPO consistently. Option D is wrong because a cross-region read replica uses asynchronous replication, which can introduce lag exceeding 5 minutes, and promoting it requires manual intervention or automation that often takes longer than 1 hour to complete, failing the RTO.

966
MCQmedium

A SysOps administrator needs to monitor the disk usage on Amazon EC2 instances running Linux. The administrator wants to collect disk utilization metrics every 5 minutes and set up an alarm when disk usage exceeds 80%. Which solution meets these requirements?

A.Use the EC2 detailed monitoring feature to collect disk metrics.
B.Install the Amazon CloudWatch Agent on the instances and configure it to collect disk metrics.
C.Use AWS Systems Manager Patch Manager to check disk space.
D.Configure an Amazon CloudWatch metric filter on the system log.
AnswerB

The Amazon CloudWatch Agent runs inside the guest OS and gathers custom metrics from the operating system, including disk utilization metrics like disk_used_percent, disk_free, and inode usage. You define the metrics to collect in the agent configuration file, and the agent sends them to CloudWatch, where you can create alarms on thresholds. This is the standard method for monitoring filesystem space on EC2 instances.

Why this answer

The Amazon CloudWatch Agent is required to collect custom metrics like disk utilization from EC2 instances. It can be configured to gather disk space metrics every 5 minutes and publish them to CloudWatch, where an alarm can be set to trigger when usage exceeds 80%. EC2 detailed monitoring only collects hypervisor-level metrics (CPU, network, disk I/O), not guest OS-level disk usage.

Exam trap

The trap here is that candidates confuse EC2 detailed monitoring (which collects hypervisor-level metrics) with the ability to collect guest OS metrics like disk usage, leading them to choose Option A incorrectly.

How to eliminate wrong answers

Option A is wrong because EC2 detailed monitoring provides hypervisor-level metrics such as CPU, network, and disk I/O, but it does not collect guest OS-level disk utilization (e.g., filesystem usage percentage). Option C is wrong because AWS Systems Manager Patch Manager is used for patching operating systems and applications, not for monitoring disk space or setting CloudWatch alarms. Option D is wrong because CloudWatch metric filters parse log data from log groups to create metrics, but they cannot extract disk usage metrics from system logs unless the logs contain structured disk usage data, and they do not replace the need for a CloudWatch agent to collect guest OS metrics.

967
MCQmedium

A SysOps administrator manages an Amazon RDS for MySQL instance that experiences high CPU utilization during business hours. The application is read-heavy. Which action will most effectively improve performance and reduce cost?

A.Enable Multi-AZ deployment.
B.Scale up the instance size to a larger instance class.
C.Add a read replica.
D.Enable automated backups.
AnswerC

A read replica is an asynchronous MySQL replica that continuously syncs changes from the primary and can serve read-only traffic, including SELECT queries and reporting workloads. By routing non-critical reads to the replica, the primary's CPU cycles are freed up for write operations, reducing overall CPU utilization on the primary instance. This is a cost-effective scale-out approach because you add a smaller replica instance rather than resizing the primary, and read replicas can be promoted or removed as demand changes.

Why this answer

Adding a read replica offloads read traffic from the primary RDS for MySQL instance, directly addressing the read-heavy workload and high CPU utilization. This improves performance by distributing SELECT queries to the replica, and reduces cost because you can use a smaller primary instance and only pay for the replica's resources when needed, rather than scaling up the entire instance.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which is for high availability) with read replicas (which are for read scaling), and assume that any scaling must involve resizing the instance rather than adding a separate read-only endpoint.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability and automatic failover, but does not offload read traffic or reduce CPU utilization on the primary instance; it only maintains a standby replica that cannot serve reads. Option B is wrong because scaling up to a larger instance class increases cost significantly and may still leave the instance underutilized during off-peak hours, whereas a read replica allows cost-effective scaling of read capacity. Option D is wrong because enabling automated backups adds overhead to the primary instance during backup windows, potentially increasing CPU utilization, and does not improve read performance or reduce cost.

968
Multi-Selecthard

A company runs a production database on Amazon RDS for PostgreSQL. The SysOps administrator wants to improve query performance for a read-heavy application without increasing costs significantly. Which THREE actions should the administrator take? (Choose three.)

Select 3 answers
A.Enable Multi-AZ deployment
B.Increase the allocated storage size
C.Optimize slow queries by reviewing the slow query log
D.Implement an Amazon ElastiCache cluster to cache frequent query results
E.Add one or more Read Replicas
AnswersC, D, E

Reviewing the slow query log is a direct diagnostic step because it captures SQL statements that exceed a configured execution-time threshold, allowing you to identify exactly which queries are problematic. Once identified, you can then add indexes, rewrite the query, or adjust parameters like work_mem (for PostgreSQL) to reduce execution time. This addresses the root cause of the performance issue rather than adding more resources to mask it.

Why this answer

Reviewing the slow query log in RDS for PostgreSQL allows the administrator to identify and optimize poorly performing queries, which directly improves query performance without incurring additional infrastructure costs. This is a standard performance tuning practice that targets the root cause of read-heavy application slowdowns.

Exam trap

The trap here is confusing Multi-AZ deployment with read scaling; candidates often assume Multi-AZ improves read performance, but it only provides a standby replica for failover, not for serving read traffic.

969
MCQeasy

A SysOps administrator wants to be notified when an Auto Scaling group launches a new instance. Which AWS service can be used to capture the Auto Scaling lifecycle events and send a notification?

A.AWS Config
B.Amazon CloudWatch Logs
C.Amazon Simple Notification Service (SNS)
D.AWS CloudTrail
AnswerC

Amazon Simple Notification Service (SNS) is a fully managed pub/sub messaging service that Auto Scaling can directly integrate with for lifecycle events. You create an SNS topic, subscribe endpoints like email or Lambda, and then configure the Auto Scaling group to publish notifications such as 'autoscaling:EC2_INSTANCE_LAUNCH' and 'autoscaling:EC2_INSTANCE_TERMINATE'. This provides immediate, push-based alerts to administrators, making it the correct choice.

Why this answer

Amazon SNS is the correct choice because it can receive lifecycle notifications from Auto Scaling groups via Amazon EventBridge (formerly CloudWatch Events) and then deliver those notifications to subscribers (e.g., email, SMS, HTTP endpoints). Auto Scaling groups emit lifecycle events (e.g., `EC2 Instance-launch Lifecycle Action`) that can be captured by EventBridge rules, which then invoke an SNS topic to send the notification.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (which records API calls) with EventBridge (which captures service events), leading them to choose CloudTrail for event-driven notifications, but CloudTrail does not handle lifecycle events or push notifications directly.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating resource configurations against desired policies and tracking configuration changes, not for capturing real-time lifecycle events or sending notifications. Option B is wrong because Amazon CloudWatch Logs is used to store, monitor, and access log files from AWS resources; it does not natively send notifications for Auto Scaling lifecycle events without additional integration (e.g., metric filters to SNS). Option D is wrong because AWS CloudTrail records API calls for auditing and compliance, but it does not capture Auto Scaling lifecycle events (which are not API calls) and cannot directly send notifications.

970
MCQeasy

A company uses Amazon CloudWatch Logs to store application logs. The SysOps administrator needs to detect when the number of log entries containing the string 'ERROR' exceeds 100 in any 5-minute window. When this threshold is breached, an email should be sent to the operations team. Which combination of AWS services should be used with the least operational overhead?

A.CloudWatch Logs Insights scheduled query with SNS action.
B.Create a metric filter on the log group for 'ERROR', then create a CloudWatch alarm on that metric with an SNS action to send email.
C.Use a Lambda function that reads the log stream and sends an email via Amazon Simple Email Service (SES) when errors exceed 100.
D.Install an agent on the application server that sends logs to Amazon SQS, then poll the queue with a Lambda function to trigger an email.
AnswerB

This is the native, fully managed approach: a metric filter is attached to the log group and processes new log events in real-time as CloudWatch Logs ingests them, extracting each occurrence of the pattern 'ERROR' and incrementing a custom CloudWatch metric. A CloudWatch alarm continuously evaluates that metric against the specified threshold (e.g., more than 100 errors in a 5-minute period); when the alarm transitions to the ALARM state, it invokes an SNS topic that sends the email notification. This pattern uses only built-in CloudWatch features and requires no custom code, external agents, or additional queueing services.

Why this answer

It uses CloudWatch metric filters to extract the count of 'ERROR' log entries as a custom metric, then a CloudWatch alarm on that metric triggers an SNS topic to send email notifications. This approach requires no custom code or additional infrastructure, minimizing operational overhead while meeting the requirement of detecting >100 errors in any 5-minute period.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing a Lambda-based or custom agent approach, not realizing that CloudWatch metric filters and alarms provide a fully managed, serverless way to monitor log patterns with minimal operational overhead.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights scheduled queries do not natively support triggering actions like SNS; they are designed for ad-hoc analysis and can only output results to S3 or other destinations, not directly invoke SNS. Option C is wrong because using a Lambda function to read log streams and send email via SES introduces unnecessary complexity, custom code, and potential latency, increasing operational overhead compared to the native metric filter and alarm approach. Option D is wrong because installing an agent to send logs to SQS and polling with Lambda adds significant operational overhead, requires managing custom infrastructure, and is not a native CloudWatch solution for real-time log monitoring.

971
MCQeasy

A company uses Amazon CloudFront to distribute content globally. The operations team notices that the data transfer costs are higher than expected. The origin server is an S3 bucket in us-east-1. Which change would reduce data transfer costs?

A.Use Lambda@Edge to resize images on the fly.
B.Increase the default TTL for objects.
C.Use multiple S3 buckets in different regions as origins.
D.Enable compression for compressible content.
AnswerD

Enabling CloudFront's automatic compression for compressible content (HTML, CSS, JavaScript, JSON, etc.) reduces the number of bytes sent to viewers when the request includes an Accept-Encoding header. Since CloudFront egress is billed per gigabyte delivered, compressing responses directly lowers the data transfer cost, often by 60-70%. It also improves latency and page load times, and it is a simple configuration change that does not require code changes or additional AWS services.

Why this answer

Enabling compression reduces the amount of data transferred from CloudFront to viewers, directly lowering data transfer costs. Option A (Lambda@Edge to resize images) can reduce image sizes but adds compute costs and may not be as effective as compression. Option B (increasing default TTL) reduces origin requests but does not reduce data transfer from CloudFront to viewers.

Option C (multiple S3 buckets in different regions) increases complexity and may increase costs due to cross-region replication or multiple origins. Thus, D is the correct choice.

972
MCQhard

A SysOps administrator is using AWS OpsWorks to manage a stack of web servers. The administrator wants to automate the installation of custom software on all new instances that are added to the layer. What is the best approach?

A.Assign a custom Chef recipe to the layer's Setup lifecycle event.
B.Use AWS CloudFormation to install software on new instances.
C.Create a custom AMI with the software pre-installed and use that in the layer.
D.Use EC2 user data scripts in the layer configuration.
AnswerA

In AWS OpsWorks, the Setup lifecycle event runs on every new instance immediately after it finishes booting and before the Deploy event. Assigning a custom Chef recipe to Setup gives the OpsWorks agent an idempotent, declarative way to install and configure required software, ensuring consistent state across all instances. This is the native OpsWorks mechanism for bootstrapping software and is fully integrated with the stack's configuration management.

Why this answer

AWS OpsWorks uses Chef to manage configuration. Assigning a custom Chef recipe to the layer's Setup lifecycle event ensures the recipe runs automatically on every new instance when it boots, installing the custom software consistently. This is the native, best-practice approach within OpsWorks for automating software installation on new instances.

Exam trap

The trap here is that candidates may confuse OpsWorks lifecycle events with EC2 user data or CloudFormation, not realizing that OpsWorks has its own built-in Chef-based automation for instance configuration.

How to eliminate wrong answers

Option B is wrong because AWS CloudFormation is an infrastructure-as-code service for provisioning resources, not for running configuration management on instances within an existing OpsWorks stack; it would require additional orchestration and does not integrate with OpsWorks lifecycle events. Option C is wrong because while a custom AMI pre-installs software, it bypasses OpsWorks's configuration management capabilities and makes updates harder to manage; it is not the 'best approach' for automation within OpsWorks. Option D is wrong because EC2 user data scripts are executed only at first boot of an EC2 instance, but OpsWorks manages instances through Chef and lifecycle events; user data is not the intended mechanism for OpsWorks-managed instances and would not integrate with OpsWorks's lifecycle hooks.

973
MCQeasy

A SysOps administrator notices that an Amazon RDS instance's CPU utilization is consistently above 90% during peak hours. The application is read-heavy and can tolerate eventual consistency. Which action would MOST effectively reduce CPU load?

A.Increase the instance size from db.r5.large to db.r5.xlarge.
B.Enable Multi-AZ deployment for the RDS instance.
C.Create one or more read replicas and direct read traffic to them.
D.Change the storage type to Provisioned IOPS.
AnswerC

Creating one or more read replicas and redirecting SELECT queries to them offloads a significant portion of the read workload from the primary DB instance, directly lowering its CPU utilization. The primary continues to handle writes and any critical read-your-own-writes traffic, while the replicas asynchronously apply changes and can serve large numbers of reads. This horizontal scaling approach is the standard and most cost-effective solution for read-heavy or OLTP workloads experiencing high CPU. Read replicas can even be placed across Availability Zones or regions to also improve read latency.

Why this answer

The application is read-heavy and can tolerate eventual consistency, making read replicas the ideal solution. By offloading read traffic to one or more read replicas, the primary RDS instance's CPU load is reduced because it no longer has to process all read queries. This directly addresses the high CPU utilization during peak hours without requiring a larger instance or other changes.

Exam trap

The trap here is that candidates often confuse Multi-AZ standby replicas with read replicas, assuming Multi-AZ can also offload read traffic, but AWS explicitly prohibits using the Multi-AZ standby for reads—it only supports failover.

How to eliminate wrong answers

Option A is wrong because increasing the instance size (e.g., from db.r5.large to db.r5.xlarge) adds more CPU and memory capacity, but it does not offload read traffic; it only scales the existing single instance, which may still be overwhelmed by the same read-heavy workload. Option B is wrong because enabling Multi-AZ deployment provides high availability and automatic failover via a synchronous standby replica, but that standby replica cannot serve read traffic (it is not a read replica), so it does not reduce CPU load on the primary instance. Option D is wrong because changing the storage type to Provisioned IOPS improves I/O performance and reduces latency, but it does not directly reduce CPU utilization; CPU load is driven by query processing, not storage throughput.

974
Multi-Selecteasy

A SysOps administrator needs to monitor the CPU and memory utilization of an EC2 instance running a legacy application that cannot be modified. Which TWO methods can be used to collect this information? (Choose TWO.)

Select 2 answers
A.Enable detailed monitoring on the instance to get memory metrics.
B.Install the CloudWatch agent on the instance and configure it to collect memory metrics.
C.Use a custom script to push memory data to CloudWatch via the PutMetricData API.
D.Use the EC2 hypervisor metrics available from CloudWatch.
E.Use AWS Systems Manager Inventory to collect memory utilization.
AnswersB, C

The CloudWatch agent runs inside the EC2 instance's guest OS and reads memory utilization directly from OS sources, such as /proc/meminfo on Linux or performance counters on Windows. After installation, you configure the agent with a JSON file or the console wizard, and it publishes memory metrics under the reserved CWAgent namespace—for example, mem_used_percent or mem_available. This is the standard, fully supported method for collecting memory metrics because the agent has privileged access to the OS internals, something the hypervisor cannot provide.

Why this answer

The CloudWatch agent can be installed on an EC2 instance to collect custom metrics, including memory utilization, which is not available by default from the hypervisor. The agent sends these metrics to CloudWatch using the PutMetricData API, enabling monitoring of in-guest resources like memory and disk.

Exam trap

The trap here is that candidates often assume detailed monitoring or hypervisor metrics include memory utilization, not realizing that memory is an in-guest metric requiring an agent or custom script to collect.

975
MCQmedium

A company runs a web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The application stores session state in memory on each instance. The SysOps administrator wants to make the application highly available across multiple Availability Zones without losing session data when instances are terminated or replaced. The solution must minimize application changes. Which approach should the administrator take?

A.Use sticky sessions (session affinity) on the ALB and configure the Auto Scaling group with a larger min size.
B.Store session data in a shared Amazon ElastiCache cluster and modify the application to read/write session state to ElastiCache.
C.Deploy the application in multiple AWS Regions and use Amazon Route 53 with latency-based routing.
D.Store session data in an Amazon RDS for MySQL database and configure the application to read/write session state to the database.
AnswerB

ElastiCache provides a centralized, in-memory data store (such as Redis) that can be shared by all EC2 instances in the Auto Scaling group. By moving session state to ElastiCache, the application becomes stateless at the instance level, so any instance can serve any user request without losing session data. ElastiCache supports replication and automatic failover, making session data highly available across Availability Zones. This directly satisfies the HA requirement and is the best practice for a decoupled web tier.

Why this answer

Storing session state in a shared Amazon ElastiCache cluster decouples session data from individual EC2 instances, allowing any instance in the Auto Scaling group to serve any user request without losing session data when instances are terminated or replaced. This approach requires minimal application changes (only modifying the session handler to point to ElastiCache) and supports high availability across multiple Availability Zones by using a replicated ElastiCache cluster (e.g., Redis with replication).

Exam trap

The trap here is that candidates often choose sticky sessions (Option A) because they seem to solve session affinity without code changes, but they fail to realize that sticky sessions do not persist session data across instance terminations, which is the core requirement for high availability without data loss.

How to eliminate wrong answers

Option A is wrong because sticky sessions (session affinity) bind a user's session to a specific EC2 instance; if that instance is terminated or replaced, the session data stored in memory is lost, violating the requirement to not lose session data. Option C is wrong because deploying across multiple AWS Regions with Route 53 latency-based routing does not address session state persistence within a single region; it introduces cross-region latency and complexity without solving the fundamental issue of in-memory session loss on instance termination. Option D is wrong because while storing session data in Amazon RDS for MySQL would persist session state, it introduces significant overhead (e.g., database connection management, schema design, and slower read/write compared to in-memory caching) and requires more extensive application changes than using ElastiCache, which is purpose-built for session storage.

Page 12

Page 13 of 16

Page 14