Courseiva

CCNA Continuous Improvement for Existing Solutions Questions

75 of 228 questions · Page 1/4 · Continuous Improvement for Existing Solutions · Answers revealed

1
MCQmedium

A company uses AWS CodePipeline to deploy a web application to an Elastic Beanstalk environment. The deployment pipeline includes a source stage, a build stage using CodeBuild, and a deploy stage. Recently, deployments have been failing in the deploy stage with the error: 'The environment is in an invalid state for this operation.' The developer confirms the build artifacts are correct. What is the MOST likely cause?

A.The environment's load balancer is not available
B.The environment's Auto Scaling group has insufficient capacity
C.The Elastic Beanstalk environment uses a t2.micro instance type which is not supported by CodePipeline
D.Another deployment or configuration update is already in progress on the environment
AnswerD

Elastic Beanstalk permits only one operation at a time; a concurrent deployment or configuration update leaves the environment in an invalid state, so the deploy stage fails until that operation completes. This matches the stem's error precisely, not artifact or build problems.

Why this answer

Elastic Beanstalk rejects a new deployment when the environment is already processing another deployment or configuration update, returning the 'invalid state for this operation' error. Because the build artifacts are confirmed correct, the most likely cause is a concurrent operation still in progress. Waiting for the in-flight operation to complete or canceling it resolves the failure.

Exam trap

SAP-C02 often tests whether candidates blame infrastructure (load balancer, Auto Scaling, instance type) for a state-machine error that is actually caused by a concurrent deployment, so recognize the InvalidState signature.

How to eliminate wrong answers

Option A is wrong because an unavailable load balancer would surface as health-check failures or 5xx errors, not as an invalid-state error during the deploy stage. Option B is wrong because insufficient Auto Scaling capacity typically causes launch failures or degraded health, not a state-validation rejection from the Elastic Beanstalk API. Option C is wrong because t2.micro is a supported instance type for Elastic Beanstalk environments and CodePipeline does not restrict instance families.

2
MCQmedium

A company is using AWS Lambda functions to process data from an S3 bucket. Recently, the function has been timing out. The function has a 5-minute timeout configured. What is the most likely cause of the timeout?

A.The Lambda function was moved to a different VPC.
B.The Lambda function's reserved concurrency is set too low.
C.The Lambda function's memory is too low.
D.The Lambda function is processing larger files than before.
AnswerD

Larger files increase the time spent reading and processing within the same invocation, so execution duration grows until it exceeds the configured 5-minute limit. The timeout is a symptom of longer processing, not of memory, permissions or concurrency settings.

Why this answer

The most likely cause of the timeout is that the Lambda function is processing larger files than before. Lambda timeout is the maximum execution time, and if the function takes longer than the configured 5 minutes due to increased processing time (e.g., larger files), it will time out. Other options like VPC change, concurrency, or memory are less directly related to timeout duration.

Exam trap

SAP-C02 often tests the distinction between timeout and throttling; candidates may confuse concurrency limits with execution time, or overlook that memory affects performance but is not the direct cause of timeout when file size increases.

How to eliminate wrong answers

Option A is wrong because moving the Lambda function to a different VPC could cause network connectivity issues, but it would typically result in connection errors or increased latency, not necessarily a timeout if the function can still access resources; also, it's not the most likely cause given the scenario. Option B is wrong because reserved concurrency being too low would cause throttling (TooManyRequestsException) rather than timeouts; the function would not be invoked. Option C is wrong because insufficient memory can cause longer execution times and potentially timeouts, but the scenario specifically mentions processing larger files, which directly increases processing time; memory might be a factor, but the most likely cause is the increased file size.

3
MCQeasy

A DevOps engineer notices that a CloudFormation stack update fails with the error: 'UPDATE_ROLLBACK_FAILED'. The stack is in a state where some resources were updated, but others failed to update. The engineer needs to fix the stack and complete the update. What should the engineer do FIRST?

A.Add a new resource to the stack to force a new update
B.Manually correct the resources that are preventing rollback, then use 'ContinueUpdateRollback'
C.Submit another stack update with the original template to overwrite the changes
D.Delete the stack and recreate it with the same template
AnswerB

UPDATE_ROLLBACK_FAILED means CloudFormation cannot roll back because a resource is in an unusable state. The engineer must first manually repair those resources outside CloudFormation, then run ContinueUpdateRollback so the stack can resume rolling back before the update is retried.

Why this answer

When a CloudFormation stack update fails and rollback also fails, the stack enters UPDATE_ROLLBACK_FAILED. The documented recovery path is to manually fix the underlying resource issue that is blocking rollback, then call ContinueUpdateRollback to resume the rollback and return the stack to a stable state. This is the first corrective action before any further updates can be attempted.

Exam trap

The trap is reaching for destructive or workaround actions (delete and recreate, force a new update) instead of the documented recovery API, ContinueUpdateRollback, which is the only supported first step in UPDATE_ROLLBACK_FAILED.

How to eliminate wrong answers

Option A is wrong because you cannot perform a new stack update while the stack is in UPDATE_ROLLBACK_FAILED; CloudFormation rejects updates until the stack is returned to a stable state. Option C is wrong because submitting another update with the original template is not possible in this state and does not address the blocked rollback. Option D is wrong because deleting the stack is a destructive last resort that loses stack history and may fail if resources are still in a bad state; it is not the first step and is generally discouraged.

4
MCQmedium

Refer to the exhibit. A solutions architect runs the AWS CLI command to check the state of an EC2 instance. The output shows the instance is running. However, the application team reports that the instance is unreachable over SSH. What is the MOST likely cause?

A.The CLI command is querying the wrong instance
B.A security group rule blocks inbound SSH traffic
C.The instance is in a 'stopped' state
D.The instance does not have EBS optimization enabled
AnswerB

The instance state and SSH reachability are independent: a running instance with no inbound TCP/22 allowance silently drops connection attempts. A security group rule blocking port 22 explains the unreachable symptom while the instance remains running, unlike host-level or key-pair issues.

Why this answer

The instance state is 'running', so it is not stopped or terminated. The most likely cause for being unreachable over SSH is that a security group rule blocks inbound SSH traffic (port 22). Option B is correct.

Option A is wrong because the query is for the correct instance. Option C is wrong because the instance is running. Option D is wrong because EBS optimization does not affect network connectivity.

5
MCQeasy

A company runs a web application on EC2 instances behind an Application Load Balancer (ALB). The application experiences periodic spikes in traffic. The operations team wants to ensure that the application can handle the spikes without manual intervention. What is the MOST cost-effective solution?

A.Use a scheduled scaling policy to add instances during predicted peak hours.
B.Create a target tracking scaling policy using the ALB RequestCountPerTarget metric.
C.Manually add instances when traffic spikes are expected.
D.Use a simple scaling policy based on CPU utilization.
AnswerB

Target tracking on the ALB RequestCountPerTarget metric scales EC2 capacity directly with incoming request load, satisfying the no-manual-intervention requirement during traffic spikes. Because it reacts to demand rather than a fixed schedule, instances are added only when needed and removed afterwards, delivering the most cost-effective elasticity for periodic, unpredictable spikes.

Why this answer

A target tracking scaling policy using the ALB RequestCountPerTarget metric automatically adjusts the number of EC2 instances to maintain a target value for requests per target, which directly correlates with traffic spikes. This is the most cost-effective because it scales in and out dynamically based on actual load, avoiding over-provisioning. It requires no manual intervention and responds to traffic changes in near real-time.

Exam trap

SAP-C02 often tests the choice between target tracking and simple/scheduled scaling; candidates may pick CPU-based simple scaling out of habit, but the exam expects recognition that RequestCountPerTarget target tracking is more direct and cost-effective for ALB-fronted web apps.

How to eliminate wrong answers

Option A is wrong because a scheduled scaling policy only adds instances during predicted peak hours, which may not align with actual spikes and can lead to over-provisioning (cost) or under-provisioning if spikes occur outside the schedule. Option C is wrong because manual scaling requires intervention and is not automated, contradicting the requirement. Option D is wrong because a simple scaling policy based on CPU utilization is less direct for a web application behind an ALB; CPU may lag behind traffic spikes, and simple scaling has cooldowns that can delay response, whereas target tracking on RequestCountPerTarget is more responsive and directly tied to load.

6
Drag & Dropmedium

Drag and drop the steps to deploy a serverless application using AWS SAM in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct order for deploying a serverless application with AWS SAM is to first write the SAM template (defining resources and configuration), then build the application (compiling code and dependencies), then package the built artifacts (uploading to S3), then deploy the stack (using CloudFormation), and finally test the deployed application to ensure it works as expected. Common mistakes include swapping build and package, writing the template after building, or deploying before packaging, which lead to errors or incomplete deployments.

7
Multi-Selecthard

A company is using AWS CodePipeline to automate deployments of a web application. The pipeline includes a build stage using AWS CodeBuild and a deploy stage using AWS CodeDeploy to an Auto Scaling group. Recently, deployments have been failing during the deploy stage with an error indicating that the target instances are not in a healthy state. The CodeDeploy agent logs show that the agent is running but the application validation scripts are failing. Which THREE actions should the solutions architect take to troubleshoot and resolve the issue?

Select 3 answers
A.Test the validation script manually on a healthy instance to confirm it works as expected.
B.Increase the deployment timeout in the CodeDeploy deployment group to allow more time for validation.
C.Review the CodeDeploy agent logs on a failing instance to identify the specific error in the validation script.
D.Verify that the AppSpec file includes the correct lifecycle event hooks (e.g., ValidateService).
E.Configure an Auto Scaling lifecycle hook to perform health checks before the instance is placed in service.
AnswersA, C, D

Running the validation script manually on a healthy instance isolates whether the script itself is faulty, independent of CodeDeploy orchestration. This directly addresses the stem's constraint that the agent runs but validation scripts fail, confirming whether the script logic or its environment causes the failure.

Why this answer

Option A is correct because running the validation script manually on a healthy instance isolates whether the script itself is broken or whether the failure is environmental (permissions, dependencies, or runtime context), which is the fastest way to reproduce and diagnose the error. Option C is correct because the CodeDeploy agent logs on a failing instance (typically /var/log/aws/codedeploy-agent/codedeploy-agent.log) contain the exact stderr/stdout and exit code from the failing lifecycle hook, pinpointing the specific validation error. Option D is correct because CodeDeploy only executes scripts mapped to lifecycle event hooks defined in the AppSpec file's hooks section (for EC2/on-premises, e.g., BeforeInstall, AfterInstall, ApplicationStart, ValidateService); if ValidateService is missing or misspelled, the validation script never runs as intended and the deployment fails validation.

Option B is not appropriate because increasing the deployment timeout only masks a script that is failing outright rather than slow, and the logs already show the validation scripts are failing, not timing out. Option E is not appropriate because Auto Scaling lifecycle hooks govern instance launch/termination and do not address CodeDeploy application validation script failures during the deploy stage.

Exam trap

SAP-C02 often tests the misconception that increasing timeouts or adding health checks at the Auto Scaling level will resolve application-level validation failures, when the root cause is typically within the script or its configuration.

8
MCQmedium

A company runs a critical Java application on Amazon EC2 instances behind an Application Load Balancer. The application stores session state in local instance memory, and the Auto Scaling group is configured to terminate instances when CPU utilization drops below 20%. During a recent scaling event, users were unexpectedly logged out and lost shopping cart contents. A solutions architect needs to make the application stateless so that instances can be terminated without impacting users, while minimizing changes to the application code. Which solution meets these requirements?

A.Change the Auto Scaling group termination policy to OldestInstance and enable connection draining on the load balancer.
B.Store session state in an Amazon ElastiCache for Redis cluster and update the application to use the ElastiCache endpoint for session data.
C.Enable sticky sessions on the Application Load Balancer and increase the deregistration delay to 300 seconds.
D.Configure the Auto Scaling group to use a lifecycle hook that backs up session data to Amazon S3 before termination, and restore it on new instances.
AnswerB

ElastiCache for Redis provides a highly available, in-memory data store that is external to the EC2 instances. By moving session state to Redis, the application no longer relies on local instance memory, so any instance can be terminated without losing session data. This is the standard approach to make a stateful Java application stateless with minimal code changes, as only the session management layer needs updating.

Why this answer

The application is stateful because it stores session data in local instance memory. To allow instances to be terminated without user impact, session state must be externalized. ElastiCache for Redis is a fully managed, in-memory store that supports high-performance session management and is a common solution for making Java applications stateless.

The other options do not remove the dependency on local instance memory.

Exam trap

The trap here is assuming that sticky sessions or connection draining can preserve session state when an instance is terminated.

9
MCQhard

A company runs a critical web application on EC2 instances behind an Application Load Balancer (ALB). During a recent deployment, users experienced errors. The team wants to automatically roll back the deployment if the error rate exceeds 5% within 10 minutes after deployment. Which solution meets these requirements with minimal operational overhead?

A.Configure the Auto Scaling group to use ELB health checks and replace instances if the error rate increases.
B.Use CodeDeploy with manual approval gates and a script that checks error rates.
C.Use CodeDeploy with a CloudWatch alarm on the ALB error rate that triggers a deployment rollback.
D.Use a custom Lambda function that monitors ALB error rates and triggers a rollback via CodeDeploy API.
AnswerC

CodeDeploy integrates CloudWatch alarms directly into deployment monitoring, automatically rolling back to the previous revision when the alarm breaches its threshold. This meets the 5% error-rate and 10-minute window with minimal operational overhead, requiring no custom automation.

Why this answer

AWS CodeDeploy integrates natively with CloudWatch alarms to automatically roll back a deployment when the alarm enters the ALARM state. By creating a CloudWatch alarm on the ALB's HTTP 5xx error rate (or target group error metric) with a threshold of 5% over a 10-minute period, and associating it with the CodeDeploy deployment group, CodeDeploy will trigger an automatic rollback if the alarm fires. This requires minimal operational overhead because it uses built-in integration.

Exam trap

SAP-C02 often tests whether candidates choose custom Lambda or manual gates for rollback automation, when the native CodeDeploy-CloudWatch alarm integration is the lowest-overhead, fully automated solution.

How to eliminate wrong answers

Option A is wrong because Auto Scaling group health checks replace unhealthy instances but do not roll back a deployment or act on application error rates. Option B is wrong because manual approval gates require human intervention and a custom script, adding operational overhead and delay, which violates the 'minimal operational overhead' and automatic rollback requirements. Option D is wrong because a custom Lambda function introduces custom code to maintain and monitor, increasing operational overhead compared to the native CodeDeploy-CloudWatch integration.

10
MCQhard

A company has a legacy application that runs on a single EC2 instance. The application writes logs to a local file. The company wants to centralize log management without modifying the application code. Which solution is MOST operationally efficient?

A.Use AWS CloudTrail to capture log file changes.
B.Modify the application to write logs to stdout and use the awslogs driver.
C.Install and configure the Amazon CloudWatch agent on the EC2 instance.
D.Set up an Amazon S3 bucket and use an AWS Lambda function to periodically copy log files.
AnswerC

Installing the CloudWatch agent lets it tail the existing local log file and stream entries to CloudWatch Logs, so centralised management is achieved with zero application changes — satisfying the no-code-modification constraint. This is more operationally efficient than custom forwarding scripts, since the agent handles rotation, buffering and delivery natively.

Why this answer

The Amazon CloudWatch agent can be installed on the EC2 instance without modifying application code. It reads the local log file and sends the logs to Amazon CloudWatch Logs for centralized management, making it the most operationally efficient solution.

Exam trap

The trap here is that candidates may think modifying the application to use stdout with the awslogs driver is simpler, but that requires code changes, which the question explicitly prohibits.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail captures API activity and management events, not log file changes on an EC2 instance. Option B is wrong because it requires modifying the application code to write logs to stdout, which violates the requirement to not modify application code. Option D is wrong because setting up an S3 bucket and Lambda function to periodically copy log files introduces unnecessary complexity and latency compared to the real-time streaming provided by the CloudWatch agent.

11
MCQmedium

A company runs a batch processing application on a scheduled EC2 instance that starts every night. The instance processes a large number of files from an S3 bucket and writes results to another S3 bucket. The job takes approximately 6 hours to complete. Recently, the job has been failing after 4 hours with an error indicating that the instance's EBS root volume is full. The instance type is t3.medium with a 20 GB gp2 root volume. The application writes temporary files to the root volume. The company wants to fix this with minimal changes to the application and infrastructure. What should a solutions architect recommend?

A.Create an additional EBS volume and mount it to the instance.
B.Change the instance type to one with instance store volumes.
C.Increase the size of the EBS root volume to 100 GB.
D.Modify the application to compress temporary files.
AnswerC

The application writes temporary files to the root volume, exhausting 20 GB during the six-hour job. Enlarging the gp2 root volume to 100 GB adds capacity without changing the application or instance type, satisfying the minimal-change constraint.

Why this answer

Increasing the root volume size provides more space for temporary files without requiring application changes. Option A is wrong because creating an additional EBS volume and mounting it would require application changes to write to a different path. Option B is wrong because instance store volumes are ephemeral and may not be available on t3 instances, and would also require application changes.

Option D is wrong because compressing temporary files may not be sufficient and requires code changes.

12
MCQmedium

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application stores session state in a local file on each instance. Users report that they are randomly logged out when the ALB routes their requests to a different instance. The company wants to improve the user experience and ensure that sessions persist across multiple instances. The company wants a solution that requires minimal changes to the application code. Which solution meets these requirements?

A.Enable cross-zone load balancing on the ALB and increase the number of instances.
B.Store session state in an Amazon ElastiCache for Redis cluster and update the application to use it.
C.Enable sticky sessions (session affinity) on the ALB target group.
D.Configure the ALB to use a Network Load Balancer (NLB) with sticky sessions.
AnswerC

Enabling sticky sessions on the ALB target group uses a cookie to bind a user's session to a specific instance. This ensures that all requests from that user go to the same instance, so local session files remain accessible. It requires no application code changes and is a quick fix. However, it can lead to uneven load distribution if sessions are long-lived and may not be ideal for high availability, but it meets the minimal-change requirement.

Why this answer

Sticky sessions on the ALB bind a user's session to a specific target instance using a cookie, ensuring that all requests from that user go to the same instance where their session file is stored. This requires no application code changes and directly solves the random logout issue. While not the most scalable long-term solution, it meets the minimal-change requirement and improves user experience.

Exam trap

The trap here is assuming that a Network Load Balancer can provide HTTP cookie-based sticky sessions, which it cannot because it operates at Layer 4.

13
MCQeasy

A company's security team wants to ensure that all S3 buckets are encrypted at rest. They have thousands of existing buckets. Which approach should a Solutions Architect use to identify noncompliant buckets?

A.Use AWS Trusted Advisor to check bucket encryption.
B.Analyze AWS CloudTrail logs for PutBucketEncryption API calls.
C.Enable S3 Inventory to list all objects and their encryption status.
D.Create an AWS Config rule to evaluate S3 bucket encryption settings.
AnswerD

AWS Config continuously evaluates resource configurations against desired settings, so an S3 bucket-encryption rule flags every noncompliant bucket across thousands of existing accounts without manual auditing. This satisfies the requirement to identify, not remediate, noncompliant buckets at scale.

Why this answer

AWS Config provides managed rules such as 's3-bucket-server-side-encryption-enabled' that continuously evaluate the encryption configuration of every S3 bucket in the account and flag noncompliant resources. Because it works at the bucket-configuration level and scales across thousands of buckets, it is the correct tool for identifying which buckets lack default encryption.

Exam trap

SAP-C02 often tests the distinction between detective services — candidates confuse CloudTrail (API activity logging), Trusted Advisor (best-practice checks), S3 Inventory (object listing), and AWS Config (resource configuration compliance evaluation).

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor only surfaces a limited set of security checks (and the S3 bucket permissions check), not a per-bucket encryption compliance evaluation across thousands of buckets. Option B is wrong because CloudTrail logs only record PutBucketEncryption API calls that were made — it cannot tell you the current encryption state of buckets that were never explicitly configured or that were created before logging was enabled. Option C is wrong because S3 Inventory reports object metadata (including per-object encryption status) but does not evaluate bucket-level default encryption settings and is not a compliance-evaluation service.

14
MCQeasy

A company wants to monitor CPU utilization of their EC2 instances and receive an alert when utilization exceeds 80% for 10 minutes. Which AWS service should be used?

A.Amazon Inspector
B.AWS Config
C.Amazon CloudWatch Alarms
D.AWS CloudTrail
AnswerC

CloudWatch Alarms evaluates metric thresholds over defined periods and triggers actions when breached. Setting a CPUUtilization alarm above 80% with a ten-minute evaluation period matches the stem's duration and threshold constraints exactly, which raw metrics or dashboards alone cannot enforce.

Why this answer

CloudWatch Alarms can monitor metrics and trigger actions when a threshold is breached.

15
MCQmedium

A company uses Amazon DynamoDB as its primary database. The operations team is seeing increased read latency during peak hours. The table has a provisioned read capacity of 1000 RCU, but CloudWatch metrics show that consumed read capacity frequently reaches 1000 RCU. The application uses eventually consistent reads. What is the MOST cost-effective way to reduce read latency?

A.Switch to strongly consistent reads to improve consistency.
B.Enable DynamoDB Accelerator (DAX) to cache frequently read items.
C.Create a global secondary index (GSI) on the table to offload reads.
D.Increase the provisioned read capacity to 2000 RCU.
E.Use Amazon ElastiCache for Memcached as a read cache.
AnswerB

DAX sits in front of DynamoDB and serves eventually consistent reads from an in-memory cache, absorbing repeated item reads so they bypass the table entirely. This reduces latency during peaks without raising provisioned read capacity, keeping cost low.

Why this answer

DAX is an in-memory cache purpose-built for DynamoDB that sits in front of the table and serves eventually consistent reads with microsecond latency, offloading read traffic from the table. Since the application already uses eventually consistent reads and the table is at its RCU ceiling during peaks, caching frequently read items in DAX reduces consumed RCU and latency without over-provisioning capacity. It is the most cost-effective fix because it addresses the root cause (repeated reads of hot items) rather than adding raw capacity.

Exam trap

SAP-C02 often tests whether candidates reflexively choose 'increase capacity' for latency problems — the trap is recognizing that caching (DAX) is more cost-effective than over-provisioning when reads are repetitive and eventually consistent.

How to eliminate wrong answers

Option A is wrong because switching to strongly consistent reads doubles RCU consumption and increases latency — the opposite of the goal. Option C is wrong because a GSI does not offload reads from the base table; it adds another read path that itself consumes RCU and only helps if queries can be served by different partition/sort keys. Option D is wrong because doubling RCU to 2000 addresses capacity but does not reduce latency for hot-item reads and costs more than caching.

Option E is wrong because ElastiCache for Memcached requires application-side cache logic, invalidation handling, and does not integrate natively with DynamoDB APIs like DAX does.

16
MCQmedium

A company is running a production web application on AWS Auto Scaling EC2 instances behind an Application Load Balancer. Recent deployments have caused intermittent errors. The team wants to implement a deployment strategy that minimizes downtime and allows for quick rollback. Which strategy should they use?

A.Deploy a new version to a single instance, test, then scale out.
B.Use blue/green deployment with a second Auto Scaling group and switch the ALB target group.
C.Perform rolling updates with a single Auto Scaling group, updating a few instances at a time.
D.Use an immutable deployment by launching a new Auto Scaling group and terminating the old one.
AnswerB

Blue/green deployment with a second Auto Scaling group satisfies the zero-downtime and quick-rollback constraints: the new version runs in a parallel environment, and the ALB listener shifts traffic by swapping target groups. Rollback is a single target-group switch back, avoiding the gradual instance replacement and mixed-version window of rolling updates.

Why this answer

Blue/green deployment with a second Auto Scaling group and an ALB target group switch provides near-zero downtime and instant rollback. The green environment runs the new version fully warmed up; once validated, the ALB listener rule is flipped to the green target group. If errors appear, the listener is flipped back to blue in seconds, restoring the previous version without redeploying.

Exam trap

SAP-C02 often tests the distinction between immutable and blue/green — candidates pick immutable because it also launches a new ASG, but immutable replaces instances in the same group without an instant traffic-switch rollback path.

How to eliminate wrong answers

Option A is wrong because testing on a single instance and then scaling out is not a controlled deployment pattern — the new version is not validated under production load, and rollback requires re-deploying the old code across all instances. Option C is wrong because rolling updates with a single Auto Scaling group mix old and new versions during the update, so rollback requires another rolling pass and downtime risk if the new version is broken. Option D is wrong because immutable deployment replaces instances in place within the same Auto Scaling group — it does not provide the instant traffic-switch rollback that blue/green with separate target groups delivers, and it still requires the new fleet to be healthy before old instances are terminated.

17
MCQeasy

A company uses AWS CloudFormation to manage infrastructure. They want to detect drift from the intended template configuration. Which service should they use?

A.AWS Config
B.AWS Service Catalog
C.AWS CloudTrail
D.CloudFormation Drift Detection
AnswerD

Drift detection compares each stack's actual resource configurations against the expected template values, reporting per-resource differences. This directly satisfies the requirement to detect deviation from the intended template configuration, without rebuilding or redeploying the stack.

Why this answer

CloudFormation Drift Detection compares the actual configuration of stack resources against the expected template configuration and reports resources that have been modified outside of CloudFormation. It is the native feature designed specifically to detect drift for CloudFormation-managed resources. AWS Config can detect configuration changes but is not tied to the template's intended state in the same way.

Exam trap

SAP-C02 often tests the confusion between AWS Config (compliance/configuration history) and CloudFormation Drift Detection (template-vs-actual comparison) — candidates must pick the service purpose-built for template drift.

How to eliminate wrong answers

Option A is wrong because AWS Config records resource configuration changes and evaluates compliance against Config rules, but it does not compare against the CloudFormation template's declared state — it lacks template-aware drift semantics. Option B is wrong because AWS Service Catalog is for governing and provisioning approved products, not for detecting drift. Option C is wrong because AWS CloudTrail logs API activity (who did what), not the current configuration state versus template.

18
MCQmedium

A company runs a batch processing application on AWS. The application reads input files from an S3 bucket, processes them on EC2 instances, and writes results to another S3 bucket. The processing job runs once a day and takes approximately 3 hours. The company wants to reduce costs and operational overhead. The Solutions Architect suggests using AWS Lambda for processing, but the processing time per file can exceed the Lambda maximum execution time of 15 minutes. The architect also considers using AWS Batch. The company wants to minimize the need for infrastructure management. Which solution should the Solutions Architect recommend?

A.Provision a fleet of EC2 instances and use Auto Scaling to manage the processing.
B.Use AWS Lambda with a larger memory allocation to increase CPU and reduce processing time.
C.Use AWS Batch with a managed compute environment that uses Spot Instances and a job queue.
D.Use Amazon ECS with Fargate launch type and run the processing as a task.
AnswerC

AWS Batch with a managed compute environment removes server infrastructure management, satisfying the minimal-overhead constraint, while Spot Instances cut cost for the 3-hour daily batch. It also lifts the 15-minute Lambda execution ceiling that blocked per-file processing.

Why this answer

AWS Batch with a managed compute environment using Spot Instances and a job queue is the best fit: it handles job scheduling, provisioning, and scaling automatically, eliminating infrastructure management, while Spot Instances reduce cost for the daily 3-hour batch job. It supports long-running jobs that exceed Lambda's 15-minute limit and is purpose-built for batch workloads. This satisfies both the cost-reduction and operational-overhead-minimization requirements.

Exam trap

SAP-C02 often tests the Lambda 15-minute limit versus long-running batch workloads, and candidates may pick Lambda with more memory or ECS/Fargate without recognizing that AWS Batch is the managed service designed for queue-based, long-running, cost-optimized batch processing.

How to eliminate wrong answers

Option A is wrong because provisioning and auto-scaling a fleet of EC2 instances requires significant infrastructure management and does not provide batch job scheduling or Spot integration out of the box. Option B is wrong because increasing Lambda memory does not overcome the hard 15-minute execution limit, so files that take longer will still fail. Option D is wrong because ECS with Fargate runs containers but lacks native batch job queuing, dependency management, and Spot Instance cost optimization that AWS Batch provides for batch workloads.

19
MCQeasy

A company uses AWS CodePipeline to deploy a web application. They want to automatically roll back the deployment if the new version fails CloudWatch alarm-based health checks. Which feature should they use?

A.AWS Lambda function invoked by CloudWatch Events.
B.Amazon Route 53 health checks with failover routing.
C.AWS CodeBuild with post-build actions.
D.CodeDeploy automatic rollback configuration with CloudWatch alarm.
AnswerD

CodeDeploy automatic rollback redeploys the last known-good revision when a specified CloudWatch alarm enters ALARM state during deployment. This directly satisfies the requirement to roll back automatically on failed alarm-based health checks, without manual intervention or custom pipeline logic.

Why this answer

AWS CodeDeploy natively supports automatic rollback triggered by CloudWatch alarms. When a deployment causes a CloudWatch alarm to enter an ALARM state, CodeDeploy can automatically revert to the previous working version. This is the most straightforward and integrated solution for the requirement.

Option A (Lambda + CloudWatch Events) is possible but not the primary recommended feature. Option B (Route 53 health checks) is for DNS-level failover, not deployment rollback. Option C (CodeBuild post-build actions) is for build phase, not post-deployment monitoring.

20
MCQhard

A company runs a multi-account AWS environment managed with AWS Organizations. Each account sends VPC Flow Logs, AWS CloudTrail logs, and application logs to a central Amazon S3 bucket in the logging account. The security team needs to query up to 5 years of logs with ad-hoc SQL, correlate events across accounts, and minimize ongoing storage cost for logs older than 90 days. The logs must remain immediately queryable without restoration. Which solution meets these requirements MOST cost-effectively?

A.Deliver all logs to Amazon S3 and query them with Amazon CloudWatch Logs Insights after exporting each account's logs to CloudWatch Logs. Set a 5-year retention policy on the log groups.
B.Deliver all logs to Amazon S3, crawl them with AWS Glue crawlers, and query with Amazon Athena. Transition objects older than 90 days to S3 Glacier Deep Archive.
C.Deliver all logs to Amazon S3, load them nightly into Amazon Redshift Spectrum external tables, and query with Amazon Redshift. Compress older data with columnar storage.
D.Deliver all logs to Amazon S3, register the bucket with AWS Glue Data Catalog, and query with Amazon Athena. Transition objects older than 90 days to S3 Glacier Instant Retrieval using an S3 Lifecycle rule.
AnswerD

Athena queries S3 data in place using the Glue Data Catalog, so no loading or cluster management is needed, and it supports ad-hoc SQL across accounts. S3 Glacier Instant Retrieval keeps archived objects millisecond-retrievable, satisfying the immediate query requirement while cutting storage cost for logs older than 90 days. This combination is the most cost-effective fit for the stated query pattern and retention.

Why this answer

Athena with the AWS Glue Data Catalog queries S3 logs in place using standard SQL, providing ad-hoc, cross-account analysis without managing servers. Moving older objects to S3 Glacier Instant Retrieval reduces storage cost while preserving millisecond retrieval, so historical logs stay immediately queryable. Together these services satisfy the query, retention, and cost requirements more effectively than cluster-based or Deep Archive alternatives.

Exam trap

The trap here is assuming the cheapest archival storage class (Glacier Deep Archive) is always best, overlooking that its multi-hour retrieval latency breaks the immediate-query requirement.

21
MCQmedium

A company has a serverless application that uses AWS Lambda functions. The functions are invoked by Amazon API Gateway and write to an Amazon DynamoDB table. The company wants to improve the existing solution to reduce latency for read-heavy workloads and reduce DynamoDB costs. The application reads the same items repeatedly. Which solution meets these requirements with the LEAST development effort?

A.Enable DynamoDB Accelerator (DAX) for the table and modify the Lambda functions to use the DAX client.
B.Increase the read capacity units (RCUs) of the DynamoDB table to handle the read load.
C.Implement an Amazon ElastiCache for Redis cluster and modify the Lambda functions to cache reads from DynamoDB.
D.Use Amazon DynamoDB global tables to replicate the table to multiple Regions and direct reads to the nearest Region.
AnswerA

DAX is a fully managed, in-memory cache for DynamoDB that provides microsecond latency for read-heavy workloads. By using the DAX client in the Lambda functions, reads are served from the cache, reducing latency and lowering DynamoDB read capacity costs. The development effort is minimal because it only requires changing the DynamoDB client to the DAX client, and DAX is compatible with the existing DynamoDB API. This meets the requirements effectively.

Why this answer

DAX is an in-memory cache specifically designed for DynamoDB, providing microsecond read latency and reducing the need to provision high read capacity. Integrating DAX requires only changing the DynamoDB client in the Lambda functions to the DAX client, which is a minimal code change. This directly addresses the read-heavy workload by caching frequently accessed items, lowering both latency and DynamoDB read costs.

Other options either increase costs, require more development effort, or do not solve the latency and cost issues.

Exam trap

The trap here is assuming that increasing read capacity units reduces latency for repeated reads.

22
MCQhard

A company has a monolithic application running on a single EC2 instance. The application experiences performance issues during peak hours. The company decides to migrate to a microservices architecture using AWS Lambda and Amazon API Gateway. The migration must be done incrementally without downtime. What strategy should the company use?

A.Deploy all microservices in a new VPC and cut over DNS after testing.
B.Create a new version of the monolith that calls Lambda functions as backend.
C.Use AWS CodeDeploy to perform a blue/green deployment of the monolith to Lambda.
D.Use the strangler fig pattern: implement API Gateway to route traffic to new Lambda functions for specific endpoints while keeping the monolith for others.
AnswerD

The strangler fig pattern incrementally routes selected endpoints through API Gateway to Lambda while the monolith still serves the rest, allowing gradual migration with no downtime. Traffic shifts endpoint by endpoint until the monolith is fully replaced.

Why this answer

The strangler fig pattern allows incremental migration by routing specific API requests to new Lambda functions via API Gateway while keeping the monolithic application for the rest. This approach avoids downtime. Option A is incorrect because deploying all microservices in a new VPC and cutting over DNS risks downtime and is not incremental.

Option B is incorrect because creating a new version of the monolith that calls Lambda functions still leaves the monolith in place and does not fully utilize API Gateway for routing. Option C is incorrect because CodeDeploy blue/green deployment is for deploying to EC2 or Lambda, but it does not provide a pattern for incremental migration of a monolith to microservices; the strangler fig pattern is more appropriate.

23
MCQhard

A company operates a microservices platform on Amazon EKS. During incidents, engineers manually inspect CloudWatch metrics, logs, and traces to find the root cause, which takes hours. The company wants to reduce mean time to resolution by automating anomaly detection and correlating metrics, logs, and traces across services. Which approach should a solutions architect recommend?

A.Configure AWS Config rules on the EKS cluster and use AWS CloudTrail to record API calls, then query the data with Amazon Athena during incidents.
B.Deploy the AWS Distro for OpenTelemetry to collect metrics and traces, send them to Amazon CloudWatch with embedded metric format, and enable CloudWatch anomaly detection and ServiceLens for correlated analysis.
C.Enable Amazon GuardDuty for EKS and configure findings to trigger AWS Lambda functions that restart unhealthy pods automatically.
D.Install the CloudWatch agent on each node to collect system metrics, and create static CloudWatch alarms with fixed thresholds for CPU and memory on every service.
AnswerB

AWS Distro for OpenTelemetry collects metrics and traces from EKS workloads and forwards them, while CloudWatch embedded metric format ingests high-cardinality metrics cost-effectively. CloudWatch anomaly detection models normal behavior and flags deviations, and ServiceLens correlates metrics, logs, and traces through a service map. Together they automate detection and cross-signal correlation, directly cutting time to resolution.

Why this answer

Reducing time to resolution requires automated anomaly detection plus correlation across metrics, logs, and traces. AWS Distro for OpenTelemetry feeds EKS telemetry into CloudWatch, where embedded metric format keeps high-cardinality data affordable, anomaly detection learns normal baselines, and ServiceLens ties signals together through a service map. Security, audit, or static-threshold tools do not provide this correlated observability.

Exam trap

The trap here is confusing security and audit services such as GuardDuty, AWS Config, and CloudTrail with observability tooling, when only metrics, logs, and traces correlation addresses runtime troubleshooting.

24
MCQeasy

A company is using Amazon DynamoDB as the primary database for a web application. The application experiences occasional throttling on writes. The company wants to implement a solution that automatically increases write capacity during traffic spikes. Which solution should they use?

A.Switch to DynamoDB On-Demand capacity mode.
B.Implement DynamoDB Accelerator (DAX) for caching.
C.Use DynamoDB Global Tables to distribute writes.
D.Enable DynamoDB Auto Scaling for write capacity.
AnswerD

DynamoDB Auto Scaling adjusts provisioned write capacity in response to CloudWatch-consumed capacity metrics, scaling up during spikes and down afterwards. This automatically raises write throughput when traffic surges, directly addressing the occasional write throttling described in the stem.

Why this answer

DynamoDB Auto Scaling adjusts provisioned write (and read) capacity automatically in response to actual utilization, using CloudWatch alarms and Application Auto Scaling to raise capacity during spikes and lower it when traffic subsides. This directly addresses occasional write throttling while keeping the table in provisioned mode. It is the intended mechanism for elastic capacity on provisioned tables.

Exam trap

SAP-C02 often tests the confusion between read-side tools (DAX) and write-side scaling — candidates must recognize that DAX caches reads and that Global Tables address multi-region, not single-region write capacity.

How to eliminate wrong answers

Option A is wrong because switching to on-demand mode is a capacity-mode change, not an auto-scaling solution, and it can be more expensive for steady workloads; the question asks for a solution that automatically increases write capacity, which Auto Scaling does within provisioned mode. Option B is wrong because DAX caches reads, not writes — it reduces read latency and read capacity consumption but does nothing for write throttling. Option C is wrong because Global Tables replicate writes across regions for multi-region active-active access; they do not increase write capacity in a single region and can actually add replication write costs.

25
Multi-Selectmedium

A company is deploying a new three-tier application on AWS. The application consists of a web tier, an application tier, and a database tier. The company wants to ensure that the application can withstand the failure of a single Availability Zone and that the database tier can fail over automatically. The company also wants to minimize operational overhead. Which two solutions should a solutions architect recommend? (Choose two.)

Select 2 answers
A.Deploy the web and application tiers in a single Availability Zone with a larger instance size to reduce the chance of failure.
B.Use Amazon Route 53 with a latency-based routing policy to distribute traffic across multiple Regions.
C.Deploy the database tier on Amazon RDS with a Multi-AZ configuration.
D.Deploy the database tier on Amazon EC2 instances with a primary and standby replica, and use a third-party clustering solution for automatic failover.
E.Deploy the web and application tiers across multiple Availability Zones using an Auto Scaling group and an Application Load Balancer.
AnswersC, E

Amazon RDS Multi-AZ maintains a synchronous standby replica in a different Availability Zone and automatically fails over to it if the primary fails. This meets the requirement for automatic database failover and AZ resilience. It requires minimal operational effort because AWS manages replication and failover. It is the standard solution for high availability of relational databases on AWS.

Why this answer

Deploying the web and application tiers across multiple Availability Zones with an Auto Scaling group and Application Load Balancer provides AZ resilience and automatic recovery. Using Amazon RDS Multi-AZ ensures automatic database failover with synchronous replication. Together, these solutions meet the high availability and low operational overhead requirements.

Self-managed databases on EC2 increase overhead, single-AZ deployments are not fault-tolerant, and multi-Region routing is unnecessary for AZ-level resilience.

Exam trap

The trap here is assuming that a larger instance in a single Availability Zone or multi-Region routing is needed for AZ resilience, when in fact multi-AZ deployments within a Region are sufficient and simpler.

26
Multi-Selectmedium

An e-commerce company runs its application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses an Amazon Aurora MySQL DB cluster with one writer and two reader instances. During a sales event, the database CPU utilization is high, and read replicas show high replica lag. The company needs to improve the read scalability and reduce replica lag. Which THREE actions should the company take? (Choose THREE.)

Select 3 answers
A.Add more reader instances to the cluster to distribute the read traffic.
B.Enable Multi-AZ for the cluster to improve read availability.
C.Increase the instance size of the writer instance to improve write throughput.
D.Increase the instance size of the reader instances to larger instance types.
E.Enable Aurora Auto Scaling for the reader instances.
AnswersA, D, E

Adding reader instances spreads read connections across more Aurora replicas, lowering per-node CPU and query load. This directly improves read scalability, though it does not by itself address replica lag, which stems from write volume on the writer.

Why this answer

Adding more reader instances (Option A) distributes the read workload across additional nodes, reducing the load on each reader and helping to lower replica lag. Aurora Auto Scaling (Option E) automatically adjusts the number of reader instances based on metrics like CPU utilization or replica lag, providing dynamic scaling during traffic spikes. Increasing the instance size of reader instances (Option D) provides more CPU and memory resources to each reader, enabling them to process more read queries and apply changes from the writer faster, which directly reduces replica lag.

Exam trap

The trap here is that candidates may confuse Multi-AZ with read scaling, but Multi-AZ in Aurora is for high availability only and does not distribute read traffic, while the real solutions involve adding more readers, scaling readers up, or using Auto Scaling to handle variable load.

27
MCQmedium

A solutions architect deployed the above CloudFormation template. However, the Lambda function is not triggered when objects are uploaded to the S3 bucket. What is the most likely cause?

A.The BucketNotification resource depends on MyLambdaFunction, but the notification configuration is incorrect.
B.The Lambda function lacks a resource-based policy that allows S3 to invoke it.
C.The Lambda execution role does not have permission to access S3.
D.The Lambda function code does not read the S3 object content.
AnswerB

S3 invokes Lambda asynchronously, which requires a resource-based policy granting s3.amazonaws.com permission to call InvokeFunction. Without this policy, uploads succeed silently but no invocation occurs, so the function never triggers despite any execution role permissions.

Why this answer

For S3 to invoke a Lambda function, the function must have a resource-based policy (permission) that grants the s3.amazonaws.com service principal permission to call lambda:InvokeFunction. Without this permission, S3's notification configuration is accepted but invocation fails silently. This is the most common cause of 'Lambda not triggered' in CloudFormation deployments.

Exam trap

SAP-C02 often tests the confusion between the Lambda execution role (what the function can access) and the Lambda resource-based policy (who can invoke the function) — candidates pick the execution role fix when the real issue is the missing invoke permission.

How to eliminate wrong answers

Option A is wrong because the BucketNotification resource dependency is not the issue — even with correct dependency ordering, the invocation fails without the Lambda permission. Option C is wrong because the Lambda execution role governs what the function can do (e.g., read S3), not whether S3 can invoke it; the execution role is irrelevant to the trigger. Option D is wrong because the function code reading the object is a runtime concern after invocation — the question states the function is not triggered at all.

28
MCQmedium

A company uses Amazon S3 to store critical data. They need to ensure that data is automatically replicated to another AWS Region for disaster recovery. Which configuration meets this requirement with minimal operational overhead?

A.Enable S3 Versioning on the source bucket.
B.Use S3 Transfer Acceleration to upload objects to both Regions.
C.Configure an S3 Lifecycle policy to transition objects to S3 Glacier.
D.Enable S3 Cross-Region Replication (CRR) on the source bucket.
AnswerD

S3 Cross-Region Replication asynchronously copies objects to a bucket in another Region, satisfying the disaster-recovery constraint. Once configured on the source bucket with versioning and an IAM role, AWS handles ongoing replication, requiring minimal operational overhead.

Why this answer

S3 Cross-Region Replication (CRR) is the native S3 feature that asynchronously replicates objects from a source bucket in one Region to a destination bucket in another Region. Once configured with a replication rule and the required IAM role, new objects are copied automatically with no ongoing operational effort. This satisfies the disaster-recovery requirement with minimal overhead.

Exam trap

SAP-C02 often tests whether candidates confuse Versioning (same-Region durability) with CRR (cross-Region DR), so picking Versioning when the question explicitly says 'another AWS Region' is the classic wrong-answer trap.

How to eliminate wrong answers

Option A is wrong because S3 Versioning only preserves multiple versions of objects in the same bucket/Region; it does not replicate data across Regions. Option B is wrong because S3 Transfer Acceleration speeds up uploads to a single bucket using edge locations; it does not replicate objects to a second Region. Option C is wrong because a Lifecycle policy transitions objects to Glacier storage classes for cost optimization, not cross-Region replication.

29
MCQhard

A company has a multi-account AWS organization with hundreds of accounts. The security team wants to ensure that all accounts have AWS Config enabled with a specific set of rules. They also want to automatically remediate non-compliant resources. Which solution is MOST scalable and operationally efficient?

A.Use AWS CloudFormation StackSets to deploy Config rules to all accounts.
B.Use AWS Config rules in each account with AWS Lambda functions for remediation.
C.Use AWS Config conformance packs deployed via AWS Organizations with automatic remediation using Systems Manager Automation.
D.Use an AWS Config aggregator in the management account to view compliance across accounts.
AnswerC

Conformance packs bundle Config rules and remediation actions into a single deployable entity, and AWS Organizations pushes them to every account automatically. This satisfies the hundreds-of-accounts scalability constraint without per-account scripting, while Systems Manager Automation executes the automatic remediation the security team requires.

Why this answer

AWS Config conformance packs can be deployed at the organization level using AWS Organizations, enabling centralized management of rules across hundreds of accounts. Automatic remediation is achieved by associating Systems Manager Automation documents with non-compliant resources. Option A is incorrect because CloudFormation StackSets require per-account deployment and management, which does not scale as efficiently as conformance packs.

Option B is incorrect because Config rules in each account with Lambda functions lack centralized deployment and management. Option D is incorrect because an AWS Config aggregator only provides a cross-account compliance view, not enforcement or remediation.

30
MCQmedium

A company runs a stateless web application on EC2 instances in an Auto Scaling group. The application is deployed across multiple Availability Zones. The team notices that during a recent traffic spike, some instances were terminated and replaced, causing a temporary drop in performance. How can the team improve the resilience of the application?

A.Purchase Reserved Instances to ensure capacity.
B.Use lifecycle hooks to wait for instance termination.
C.Increase the instance size to handle more traffic.
D.Configure a warm pool for the Auto Scaling group.
AnswerD

A warm pool maintains pre-initialised instances in a stopped or running state, ready to enter service immediately. This directly addresses the performance drop caused by cold-start bootstrapping during scale-out, since replacement instances bypass lengthy application initialisation. It satisfies the resilience constraint by reducing the time the Auto Scaling group operates below desired capacity during traffic spikes.

Why this answer

A warm pool pre-initializes instances, reducing the time needed for new instances to become ready during scale-out events. Option A is wrong because Reserved Instances guarantee capacity but do not reduce the initialization delay. Option B is wrong because lifecycle hooks can delay termination but do not accelerate instance readiness.

Option C is wrong because larger instance size does not prevent the temporary drop in performance caused by instance replacement delays.

31
MCQmedium

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application writes logs to ephemeral instance storage. The operations team needs to retain these logs for 90 days for compliance and wants to search them using a central service. The logs are currently lost when instances are terminated. Which solution meets these requirements with the LEAST operational overhead?

A.Modify the application to write logs directly to an Amazon EFS file system mounted on each instance.
B.Install and configure a third-party log shipper to send logs to an Amazon S3 bucket with a lifecycle policy for 90-day retention.
C.Configure the CloudWatch agent to stream logs to Amazon CloudWatch Logs with a 90-day retention period.
D.Create a cron job on each instance that copies logs to an Amazon S3 bucket every hour.
AnswerC

The CloudWatch agent can collect logs from instance storage and send them to CloudWatch Logs, where you can set a retention policy of 90 days and use CloudWatch Logs Insights to search. This requires minimal operational overhead because it is a managed service and the agent can be installed via user data or AWS Systems Manager.

Why this answer

The CloudWatch agent provides a managed, low-overhead way to collect logs from EC2 instances and send them to CloudWatch Logs. Setting a 90-day retention period satisfies compliance, and CloudWatch Logs Insights allows searching. Other options involve custom scripts, third-party tools, or additional services that increase operational burden.

Exam trap

The trap here is overlooking that CloudWatch Logs can enforce retention and provide search, assuming that S3 is always the go-to for log retention, which adds search complexity.

32
MCQmedium

A company runs a microservices application on Amazon ECS with the Fargate launch type. The services communicate over HTTP and are deployed across multiple Availability Zones. The company wants to improve the application's resilience by implementing automatic retries and circuit breaking for inter-service communication. The services are registered in AWS Cloud Map. Which solution should a solutions architect recommend to meet these requirements with the LEAST operational overhead?

A.Use an Application Load Balancer with health checks and configure the ECS services to use it for service-to-service communication.
B.Deploy an AWS App Mesh service mesh and configure retry policies and circuit breakers for the virtual services.
C.Configure Amazon Route 53 with latency-based routing and health checks for each service, and rely on DNS failover for resilience.
D.Implement retry logic and circuit breakers in each microservice using an AWS SDK and the Cloud Map API for service discovery.
AnswerB

AWS App Mesh provides a service mesh that can automatically handle retries and circuit breaking without requiring changes to application code. It integrates with ECS and Cloud Map, allowing you to define virtual services and routes. By configuring retry policies and outlier detection, you can improve resilience with minimal operational overhead. App Mesh also provides observability, making it easier to monitor inter-service communication.

Why this answer

AWS App Mesh is a managed service mesh that provides retry policies, circuit breaking, and traffic routing without application code changes. It integrates with ECS and Cloud Map, reducing operational overhead. Other options require custom development or do not provide the required resilience features.

Exam trap

The trap here is assuming that a load balancer or DNS failover can provide application-level retries and circuit breaking, which they cannot.

33
MCQeasy

A company has a web application running on Amazon EC2 instances in an Auto Scaling group. The application stores user-uploaded files in an Amazon S3 bucket. The company wants to improve the security of the application by ensuring that the EC2 instances can access the S3 bucket without embedding long-term AWS credentials in the application code or on the instances. Which solution meets these requirements with the LEAST operational overhead?

A.Generate a pre-signed URL for each file upload using AWS credentials stored in the application, and have the application provide the URL to users.
B.Create an IAM user with programmatic access and an access key, store the credentials in AWS Secrets Manager, and have the application retrieve them at startup.
C.Create an IAM role with the necessary S3 permissions and attach it to the EC2 instances via an instance profile. The application uses the instance metadata service to obtain temporary credentials.
D.Store the S3 bucket credentials in an encrypted Amazon EBS volume attached to each EC2 instance, and have the application read them from the volume.
AnswerC

Attaching an IAM role to EC2 instances via an instance profile provides temporary credentials automatically rotated by AWS. The application can use the AWS SDK to retrieve credentials from the instance metadata service without embedding secrets. This is the least operational overhead because it requires no credential management and follows AWS best practices for secure access.

Why this answer

Attaching an IAM role to EC2 instances via an instance profile allows the application to obtain temporary credentials from the instance metadata service automatically. This eliminates the need to embed or manage long-term credentials, reduces operational overhead, and follows AWS security best practices. The other options either use long-term credentials, require secret management, or do not eliminate credential handling.

Exam trap

The trap here is assuming that storing credentials in Secrets Manager or on an encrypted volume is secure enough, when the best practice is to avoid long-term credentials entirely by using IAM roles.

34
MCQeasy

A company uses Amazon S3 to store sensitive data. The security team requires that all objects be encrypted at rest. The company currently uses server-side encryption with S3-managed keys (SSE-S3). The security team wants to ensure that only authorized users can access the decryption keys. What should the company do?

A.Configure an S3 bucket policy to allow only specific IAM roles to put objects.
B.Continue using SSE-S3 and enable S3 Block Public Access.
C.Use client-side encryption with an AWS KMS key.
D.Change the default encryption to server-side encryption with AWS KMS (SSE-KMS) and apply IAM policies to control key usage.
AnswerD

SSE-KMS encrypts objects with AWS KMS keys, and IAM policies on those keys restrict which principals may decrypt. This replaces S3-managed keys, where key access cannot be scoped to authorised users, meeting the stated requirement.

Why this answer

SSE-KMS encrypts objects with keys managed in AWS KMS, and access to those keys is governed by IAM policies and KMS key policies. This satisfies the requirement that only authorized users can access decryption keys, because KMS enforces key-level permissions separately from S3 bucket access. SSE-S3 uses AWS-managed keys that customers cannot control or restrict per-user.

Exam trap

SAP-C02 often tests the misconception that bucket policies or Block Public Access control encryption key access — they control object access, not KMS key usage, which requires IAM/KMS key policies.

How to eliminate wrong answers

Option A is wrong because a bucket policy controlling who can PutObject does not restrict access to the encryption keys — SSE-S3 keys are fully AWS-managed and not subject to customer IAM control. Option B is wrong because SSE-S3 with Block Public Access still leaves keys entirely under AWS control with no way to restrict decryption to specific users. Option C is wrong because client-side encryption with a KMS key is not the standard server-side approach the question asks for, and it shifts encryption responsibility to the application rather than configuring S3 default encryption.

35
MCQmedium

A social media startup uses AWS Lambda functions to process user-uploaded images. The Lambda function resizes images and stores them in Amazon S3. The function uses the S3 SDK to put objects. Recently, the team noticed that the function sometimes fails with 'Timeout' errors for large images. The Lambda function has a timeout of 5 seconds and 256 MB of memory. The team wants to improve the solution to handle larger images reliably and cost-effectively. Which solution should the team implement?

A.Migrate the image processing to a dedicated Amazon EC2 instance with an EBS volume.
B.Increase the Lambda function's timeout to 15 minutes and allocate more memory (e.g., 1024 MB).
C.Use Amazon API Gateway with a larger payload limit to offload the image processing.
D.Use AWS Elastic Transcoder to resize images instead of Lambda.
AnswerB

More memory and timeout allow processing larger images within Lambda limits.

Why this answer

(increase memory and timeout) directly addresses the issue: increasing memory also increases CPU and network throughput, which helps process large images faster; increasing timeout gives more time. Option A (EC2 with EBS) is overkill and not serverless, losing the benefits of Lambda. Option C (API Gateway with larger payload) does not help with Lambda's internal processing limits.

Option D (Elastic Transcoder) is for video transcoding, not image resizing.

36
MCQmedium

A company runs a two-tier web application on Amazon EC2 instances behind an Application Load Balancer. The database tier is Amazon RDS for MySQL with a single Availability Zone deployment. The company requires a recovery point objective (RPO) of 5 minutes and a recovery time objective (RTO) of 2 hours. The database is 500 GB. Which solution should a solutions architect recommend to meet these requirements MOST cost-effectively?

A.Configure a Multi-AZ deployment for the RDS instance, which maintains a synchronous standby in another Availability Zone.
B.Create a read replica in a second Availability Zone and promote it manually if the primary fails.
C.Use AWS Database Migration Service (DMS) with change data capture to continuously replicate to an EC2-based MySQL instance, and fail over by repointing the application.
D.Enable automated backups with a 5-minute backup window and restore from the latest snapshot during a failover.
AnswerA

Multi-AZ for RDS maintains a synchronous standby replica in a different Availability Zone. Failover typically completes in 60–120 seconds, well within the 2-hour RTO, and because replication is synchronous, RPO is effectively zero. It also provides high availability without the cost and operational overhead of read replicas or custom replication. This is the most cost-effective solution that meets both RPO and RTO.

Why this answer

A Multi-AZ deployment maintains a synchronous standby that is automatically promoted on failure, providing near-zero RPO and failover in minutes, which satisfies both the 5-minute RPO and 2-hour RTO. Automated backups alone cannot meet the RTO because restore times are lengthy. Read replicas are asynchronous and require manual promotion, risking RPO and RTO.

DMS adds unnecessary complexity and cost.

Exam trap

The trap here is assuming that automated backups with frequent transaction logs provide fast recovery, when in fact restoring a snapshot can take hours and does not meet a tight RTO.

37
Multi-Selectmedium

A company is using Amazon S3 to store sensitive data. The security team wants to ensure that all objects uploaded to specific S3 buckets are encrypted at rest. Which TWO actions should they take? (Choose 2)

Select 2 answers
A.Use a bucket policy that denies PutObject without the x-amz-server-side-encryption header.
B.Configure default encryption on the S3 buckets to use SSE-S3 or SSE-KMS.
C.Enable S3 Cross-Region Replication.
D.Enable S3 Versioning on the buckets.
E.Enable S3 Server Access Logs.
AnswersA, B

A bucket policy denying PutObject requests that lack the x-amz-server-side-encryption header enforces encryption at upload time, satisfying the requirement that all objects in the specified buckets be encrypted at rest. This blocks unencrypted PUTs before any object is written, rather than merely detecting them afterwards.

Why this answer

Option A is correct because a bucket policy with a Deny effect on s3:PutObject conditioned on the absence of the s3:x-amz-server-side-encryption header (or Null condition on that key) forces every uploader to explicitly request server-side encryption, preventing unencrypted PUT requests. Option B is correct because configuring default bucket encryption with SSE-S3 (AES-256) or SSE-KMS automatically encrypts objects at rest even when the request does not include an encryption header, providing a baseline guarantee for all uploaded data. Option C is not relevant because Cross-Region Replication only copies objects to another bucket and does not enforce encryption at rest on the source bucket.

Option D is not relevant because Versioning preserves multiple object versions but does not itself encrypt data. Option E is not relevant because Server Access Logs record request activity for auditing, not encryption enforcement.

38
MCQeasy

A company is using Amazon ECS with Fargate launch type for a microservices application. The application experiences intermittent latency spikes. CloudWatch metrics show high CPU utilization but no obvious pattern. What should the company do to identify the cause?

A.Increase the CPU and memory for all ECS tasks.
B.Enable AWS X-Ray tracing on the ECS tasks to trace requests across microservices.
C.Set up CloudWatch Synthetics canaries to monitor the endpoints.
D.Use CloudWatch Logs Insights to analyze application logs for errors.
AnswerB

AWS X-Ray traces individual requests across microservices, exposing where latency accumulates within the call path rather than just aggregate CPU metrics. Since CloudWatch shows high CPU utilisation without an obvious pattern, X-Ray's per-request service map and segment timings pinpoint the specific downstream call or task causing intermittent spikes.

Why this answer

AWS X-Ray provides distributed tracing to pinpoint performance bottlenecks. Option A is wrong because increasing task size is a reactive fix that does not identify the cause. Option C is wrong because CloudWatch Synthetics monitors endpoint availability, not internal trace data.

Option D is wrong because CloudWatch Logs Insights is for querying logs, not for tracing requests across microservices.

39
MCQmedium

A company runs a critical web application on Amazon EC2 instances behind an Application Load Balancer. The application stores session state in a self-managed Redis cluster on a single EC2 instance. During a recent load test, the Redis instance became a bottleneck, causing session timeouts and degraded performance. The company wants to improve the scalability and availability of the session store while minimizing application changes. Which solution meets these requirements?

A.Deploy a second Redis instance on another EC2 instance and configure the application to use both instances with client-side sharding.
B.Replace the Redis cluster with an Amazon RDS for MySQL Multi-AZ database and store sessions in a table.
C.Enable sticky sessions on the Application Load Balancer and continue using the single Redis instance.
D.Migrate the session store to Amazon ElastiCache for Redis with cluster mode enabled and multiple shards.
AnswerD

ElastiCache for Redis with cluster mode enabled provides horizontal scaling and high availability across multiple shards, eliminating the single-instance bottleneck. The application can use the Redis cluster endpoint with minimal changes, as the Redis protocol remains the same. This directly addresses scalability and availability requirements.

Why this answer

ElastiCache for Redis with cluster mode enabled offers a managed, scalable, and highly available session store that integrates with existing Redis clients. It removes the single-instance bottleneck by distributing data across shards and provides automatic failover. Other options either require extensive application changes or do not address scalability and availability.

Exam trap

The trap here is assuming that sticky sessions or a second Redis instance can solve the scalability and availability issues without application changes.

40
Multi-Selectmedium

A company is running a web application on EC2 instances in an Auto Scaling group behind an ALB. The application uses an Amazon RDS for MySQL database. Recently, the application has become slow, and the operations team identifies that the database is the bottleneck due to a high number of read queries. Which TWO actions should a solutions architect take to improve read performance? (Choose two.)

Select 2 answers
A.Enable Multi-AZ for the RDS instance.
B.Implement DynamoDB Accelerator (DAX) in front of the database.
C.Scale up the RDS instance to a larger instance type.
D.Add an Amazon RDS Read Replica in the same AWS Region.
E.Implement an Amazon ElastiCache for Redis cluster to cache frequent queries.
AnswersD, E

Read Replicas offload read-only traffic from the primary RDS for MySQL instance via asynchronous replication, directly relieving the read-query bottleneck. The application can route SELECT queries to the replica endpoint, reducing contention on the primary and improving overall read throughput.

Why this answer

Option D is correct because an Amazon RDS Read Replica offloads read-only traffic from the primary MySQL instance via asynchronous replication, directly reducing the read bottleneck and improving read throughput. Option E is correct because Amazon ElastiCache for Redis caches frequent query results in memory, so repeated read queries are served from the cache instead of hitting RDS, which reduces database load and improves response times. Option A is not correct because Multi-AZ provides high availability through a synchronous standby, not additional read capacity.

Option B is not correct because DynamoDB Accelerator (DAX) is a caching layer for DynamoDB, not for Amazon RDS for MySQL. Option C is not correct because scaling up the instance increases overall capacity but does not specifically address read scaling as effectively as read replicas or caching, and it may not resolve a high-read-query bottleneck.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read replicas, assuming the standby instance can serve reads, when in fact Multi-AZ only provides failover and the standby is not accessible for read operations.

41
MCQeasy

A company runs a static website on Amazon S3 with a custom domain using Amazon Route 53. The website content is updated frequently by multiple developers. The company wants to implement a workflow where updates are automatically tested and deployed. They have existing CI/CD tools that integrate with AWS CodeCommit. The Solutions Architect needs to design a deployment pipeline that rebuilds the website only when changes are pushed to the main branch, and then invalidates the Amazon CloudFront cache if a CloudFront distribution is used. Which solution meets these requirements with the least operational overhead?

A.Use AWS CloudFormation with a custom resource that triggers a build on CodeCommit push.
B.Configure an S3 event notification to invoke an AWS Lambda function that builds and deploys the website.
C.Use AWS CodePipeline with a source stage tied to CodeCommit, a build stage using AWS CodeBuild, and a deploy stage that syncs the S3 bucket and invalidates CloudFront.
D.Use AWS Lambda@Edge to generate the website on the fly and cache at CloudFront.
AnswerC

CodePipeline orchestrates CodeCommit source, CodeBuild, and a deploy stage that syncs S3 and invalidates CloudFront, giving automatic rebuilds on main-branch pushes with minimal operational overhead. It satisfies the branch-triggered rebuild and cache invalidation requirements natively.

Why this answer

AWS CodePipeline natively integrates with CodeCommit as a source, CodeBuild as a build stage, and can deploy to S3 with a CloudFront invalidation step — all with minimal operational overhead. It supports branch-based triggers (e.g., main branch) and can be defined entirely in code, meeting the CI/CD and cache invalidation requirements.

Exam trap

The trap is over-engineering with Lambda@Edge or custom CloudFormation resources when the question explicitly states existing CI/CD tools integrate with CodeCommit — the exam rewards the managed, least-overhead pipeline service.

How to eliminate wrong answers

Option A is wrong because a CloudFormation custom resource that triggers builds on CodeCommit push is a custom, high-maintenance solution that reinvents what CodePipeline already provides. Option B is wrong because S3 event notifications trigger on object changes, not on CodeCommit pushes, and a Lambda function would need custom build logic — more operational overhead and not aligned with the existing CI/CD tooling. Option D is wrong because Lambda@Edge is for request/response manipulation at CloudFront edge locations, not for building and deploying a static site from source control.

42
MCQeasy

A company has an S3 bucket that stores sensitive data. The company wants to ensure that all objects uploaded to the bucket are encrypted at rest. Which solution should the solutions architect recommend?

A.Use a bucket policy to deny uploads that do not include the x-amz-server-side-encryption header.
B.Create an AWS Lambda function that encrypts objects after they are uploaded.
C.Configure an S3 Access Point with a policy that requires encryption.
D.Enable default encryption on the S3 bucket using SSE-S3 or SSE-KMS.
AnswerD

Correct. Default encryption on the S3 bucket automatically encrypts all objects at rest using SSE-S3 or SSE-KMS, regardless of the upload request. It is the simplest and most effective solution to ensure encryption at rest.

Why this answer

Enabling default encryption on the S3 bucket using SSE-S3 or SSE-KMS ensures that all objects are automatically encrypted at rest, regardless of whether the upload request specifies encryption. This is the simplest and most effective approach. Option A is technically viable but unnecessarily complex; a bucket policy that denies uploads without the x-amz-server-side-encryption header can enforce encryption, but it requires careful policy configuration and does not encrypt objects automatically if the header is missing—instead it rejects the upload.

Option B is inefficient and costly; using a Lambda function to encrypt objects after upload introduces latency and extra expense, whereas default encryption achieves the same result seamlessly. Option C is incorrect because S3 Access Points are designed for managing access to shared datasets, not for enforcing encryption; adding a policy there would complicate access management without providing the automatic encryption that default encryption offers.

43
MCQeasy

A company is using AWS CloudFormation to manage infrastructure. They want to ensure that any changes to a production stack are reviewed and approved before being applied. What is the BEST way to achieve this?

A.Enable termination protection on the stack.
B.Use AWS CodePipeline to automatically deploy changes.
C.Use Change Sets and require manual approval to execute them.
D.Use stack policies to prevent updates.
AnswerC

Change Sets compute the proposed modifications and surface them for inspection before any resource is altered, so a reviewer can approve or reject the diff. Requiring manual execution satisfies the stem's constraint that production changes be reviewed and approved prior to being applied.

Why this answer

AWS CloudFormation Change Sets allow you to preview how proposed changes will affect your running resources before executing them. By using Change Sets in conjunction with a manual approval process (e.g., via AWS CodePipeline or a separate review step), you can ensure changes are reviewed and approved before being applied. Option A (termination protection) only prevents stack deletion, not updates.

Option B (auto-deployment with CodePipeline) can include approval gates, but the question asks for the "BEST way" in the context of CloudFormation itself; Change Sets are the native mechanism for review. Option D (stack policies) control which resources can be updated, but do not enforce a review process.

44
MCQmedium

A company runs a batch processing workload on Amazon EC2 Spot Instances managed by an Auto Scaling group. The workload checkpoints progress to Amazon S3 every 10 minutes. Spot Instances are frequently interrupted, causing the workload to restart from the last checkpoint. The team wants to reduce the impact of interruptions and improve job completion time. Which solution should a solutions architect recommend?

A.Purchase a 1-year Reserved Instance for the batch workload and use it instead of Spot Instances.
B.Increase the checkpoint frequency to every 1 minute and continue using a single Spot capacity pool.
C.Configure the Auto Scaling group to use a mixed instances policy with multiple instance types and Availability Zones, and enable capacity rebalancing.
D.Use On-Demand Instances for the entire batch workload and enable EC2 Auto Scaling predictive scaling.
AnswerC

A mixed instances policy diversifies across instance types and Availability Zones, reducing the chance that a single Spot capacity pool interruption affects all instances. Capacity rebalancing proactively replaces instances when a Spot interruption notice is received, allowing the workload to checkpoint and migrate before termination. This directly reduces interruption impact and improves job completion time.

Why this answer

Diversifying across instance types and Availability Zones with a mixed instances policy spreads the workload over multiple Spot capacity pools, reducing correlated interruptions. Capacity rebalancing uses the two-minute Spot interruption notice to launch replacement instances and drain existing ones, so the workload can checkpoint and continue. Together they improve resilience and job completion time while retaining Spot cost savings.

Exam trap

The trap here is thinking that more frequent checkpoints or switching to On-Demand solves Spot interruptions, when the real fix is diversifying capacity pools and using capacity rebalancing to react to interruption notices.

45
MCQeasy

A company is using AWS CloudFormation to manage infrastructure. The stack creation fails with the error 'Resource handler returned message: 'User: arn:aws:sts::123456789012:assumed-role/Admin/MySession is not authorized to perform: ec2:RunInstances'. What is the MOST likely cause?

A.The IAM role used by CloudFormation does not have ec2:RunInstances permission.
B.The region specified in the template is disabled.
C.The CloudFormation template has a syntax error.
D.The AWS account is not subscribed to EC2 service.
AnswerA

CloudFormation assumes an IAM role to make API calls; if that role's policy lacks ec2:RunInstances, the stack's instance creation is denied. The error names the assumed role, confirming the missing permission is the cause of the stack creation failure.

Why this answer

The error message explicitly names the assumed role and the denied action `ec2:RunInstances`, which is the signature of an IAM permissions problem. CloudFormation uses either the user's credentials or a service role to create resources; if that principal lacks `ec2:RunInstances`, the stack fails with exactly this 'not authorized to perform' message.

Exam trap

SAP-C02 often tests whether candidates can parse the exact IAM error string and distinguish authorization failures from template syntax or region errors — the phrase 'not authorized to perform' is the giveaway.

How to eliminate wrong answers

Option B is wrong because a disabled region produces errors like 'InvalidRegion' or 'region not opted-in', not an IAM authorization failure. Option C is wrong because a template syntax error fails during validation with 'Template format error' before any resource handler runs. Option D is wrong because AWS accounts are not 'subscribed' to EC2 — EC2 is available by default in all commercial regions, and lack of subscription is not an AWS concept.

46
MCQmedium

A company has a CI/CD pipeline that builds and deploys a containerized application to Amazon ECS Fargate. The pipeline uses AWS CodeBuild to run tests and build Docker images. Recently, the pipeline has been failing intermittently with the error 'CannotPullContainerError: Error response from daemon: manifest for <image> not found'. The image is stored in Amazon ECR. The team suspects the issue is related to image tag inconsistency. The pipeline tags images with the commit hash. Which change will prevent this error?

A.Store the Docker image in Amazon S3 instead of ECR.
B.Ensure the pipeline builds and pushes the image with a unique tag, such as the commit hash, and uses that tag in the ECS task definition.
C.Use the 'latest' tag for all images.
D.Retry the failed pipeline step after a delay.
AnswerB

ECS resolves the image by the exact tag in the task definition. Pushing with the commit hash and referencing that same tag guarantees the manifest exists in Amazon ECR, eliminating the intermittent CannotPullContainerError caused by tag mismatch.

Why this answer

Ensuring that the image tag is unique and not reused prevents stale image references. Using the commit hash ensures uniqueness.

47
MCQmedium

A company runs a critical application on EC2 instances behind an Application Load Balancer. The security team requires that all traffic to the application be encrypted in transit and that the load balancer use a certificate from AWS Certificate Manager (ACM). The application currently uses HTTP. What should the company do to meet the security requirement?

A.Replace the ALB with a Network Load Balancer and associate an ACM certificate with it.
B.Change the ALB listener to TCP and use a self-signed certificate on the EC2 instances.
C.Place a CloudFront distribution in front of the ALB and configure HTTPS between viewers and CloudFront.
D.Add an HTTPS listener to the ALB using an ACM certificate, and configure the HTTP listener to redirect to HTTPS.
AnswerD

This provides encryption and uses ACM for certificate management.

Why this answer

Adding an HTTPS listener to the ALB with an ACM certificate and configuring the HTTP listener to redirect to HTTPS ensures all traffic is encrypted in transit. This meets the security requirement directly without additional components. Option A is incorrect because Network Load Balancers do not support ACM certificates for TLS termination; they require TLS termination on the backend instances.

Option B is incorrect because TCP listeners cannot terminate TLS, and self-signed certificates on EC2 instances would not provide trusted encryption for clients. Option C is incorrect because while CloudFront can provide HTTPS, it adds unnecessary complexity and cost; the requirement can be met natively with the ALB.

48
MCQhard

A company runs a containerized application on Amazon ECS with Fargate launch type. The application is deployed across multiple Availability Zones. Recently, deployments have been failing because new tasks cannot register with the Application Load Balancer (ALB) target group. The health checks are failing. What is the MOST likely cause?

A.The security group for the tasks does not allow inbound traffic from the ALB on the health check port.
B.The ECS service is configured with a desired count of zero.
C.The task definition specifies an invalid container image.
D.The ECS cluster has insufficient capacity.
AnswerA

Fargate tasks and the ALB communicate over the network, so the task security group must permit inbound traffic from the ALB security group on the health check port. Without that rule, health checks fail and tasks never register.

Why this answer

If the security group for the tasks does not allow inbound traffic from the ALB on the health check port, health checks fail and tasks cannot register. Option B is incorrect because a desired count of zero would prevent new tasks from running, but the scenario describes deployments failing due to health check failures on new tasks. Option C is incorrect: an invalid container image would cause the task to fail to start, not cause health check failures after the task is running.

Option D is incorrect because Fargate manages capacity; insufficient capacity would cause a different error (e.g., unable to provision tasks), not health check failures.

49
MCQeasy

A startup runs its application on Amazon ECS with Fargate launch type. The application uses an Application Load Balancer to distribute traffic. During a recent marketing campaign, the application experienced high latency and some requests returned 503 errors. The team suspects that the tasks are hitting resource limits. The team wants to automatically scale the tasks based on CPU utilization. Which solution should the team implement?

A.Configure Application Auto Scaling for the ECS service with a target tracking scaling policy based on average CPU utilization.
B.Create a CloudWatch alarm that triggers a Lambda function to stop idle tasks.
C.Create an Auto Scaling group for the ECS cluster and configure it to scale based on CPU utilization.
D.Use AWS Lambda to periodically check CPU utilization and update the desired count of the ECS service.
AnswerA

Application Auto Scaling target tracking adjusts the ECS service's desired task count to hold average CPU utilisation at a target, directly addressing the resource-limit saturation causing 503s and latency. It works with Fargate because scaling operates on task count, not cluster capacity.

Why this answer

Application Auto Scaling with a target tracking policy on average CPU utilization is the native, recommended way to scale ECS service tasks. It automatically adjusts the desired count of tasks to keep CPU near the target value, responding to load changes without custom code. This directly addresses the high latency and 503 errors caused by tasks hitting resource limits.

Exam trap

SAP-C02 often tests the misconception that ECS clusters are scaled with EC2 Auto Scaling groups; for Fargate, scaling is done at the service level with Application Auto Scaling, not at the cluster level.

How to eliminate wrong answers

Option B is wrong because a CloudWatch alarm that stops idle tasks only reduces capacity — it does not add tasks when CPU is high, so it cannot solve the scaling problem and may worsen 503 errors. Option C is wrong because ECS with Fargate does not use EC2 Auto Scaling groups; the cluster capacity is managed by Fargate, and scaling must target the ECS service, not an Auto Scaling group. Option D is wrong because a custom Lambda polling loop is an anti-pattern: it adds latency, requires custom code, and is less reliable than the built-in Application Auto Scaling target tracking policy.

50
MCQmedium

A CloudFormation stack is created using the template above. The stack creation fails with the error: 'The following resource(s) failed to create: [EC2Instance]'. Logs show: 'AMI 'ami-0abcdef1234567890' does not exist.' What is the most likely cause?

A.The AMI ID is not available in the region where the stack is being deployed.
B.The SQS queue name 'my-queue' is already in use.
C.The AMI ID is invalid because it contains letters.
D.The instance type t2.micro is not supported in the region.
AnswerA

AMI IDs are region-specific, so an identifier valid in one region does not resolve in another. The stem's deployment targets a region lacking that image, producing the "does not exist" failure during EC2Instance creation. Copying the AMI to the target region, or referencing a region-appropriate ID, resolves the constraint.

Why this answer

The error 'AMI ami-0abcdef1234567890 does not exist' indicates that the specified AMI ID is not available in the AWS region where the CloudFormation stack is being deployed. AMI IDs are region-specific; an AMI that exists in us-east-1 will not exist in eu-west-1 unless it was explicitly copied or is a public AMI available in that region. The most likely cause is a region mismatch between the AMI ID and the stack's deployment region.

Exam trap

SAP-C02 often tests the regional nature of AMIs, so candidates who overlook that AMI IDs are region-specific may blame the instance type or queue name instead of the region mismatch.

How to eliminate wrong answers

Option B is wrong because an SQS queue name conflict would produce an error about the queue, not about the AMI; the error explicitly references the AMI ID. Option C is wrong because AMI IDs legitimately contain hexadecimal characters including letters (a-f), so the presence of letters does not make the ID invalid. Option D is wrong because if the instance type were unsupported, the error would mention the instance type, not the AMI; t2.micro is widely supported, and the error is specifically about the AMI.

51
MCQhard

A company has a data pipeline that uses AWS Glue to process large datasets in Amazon S3. The pipeline runs daily and takes over 12 hours to complete. The company wants to reduce the processing time. Which approach would be MOST effective?

A.Increase the Glue job timeout setting to 24 hours.
B.Enable S3 Transfer Acceleration on the source bucket.
C.Increase the number of DPUs allocated to the Glue job.
D.Convert the input data from CSV to Parquet format.
AnswerC

Glue scales horizontally by adding DPUs, each providing processing capacity and memory. Increasing DPUs for this long-running job parallelises the work across more executors, cutting the 12-hour runtime, whereas other changes do not raise raw compute throughput.

Why this answer

Increasing the number of DPUs (data processing units) allocated to the Glue job allows for greater parallelism, which directly reduces processing time for CPU-bound or memory-bound workloads. Option A is incorrect because increasing the timeout does not improve performance; it only prevents the job from failing due to time limits. Option B is incorrect because S3 Transfer Acceleration speeds up data transfer to S3, not the processing within Glue.

Option D is incorrect because while converting to Parquet can improve read performance and reduce data volume, it does not address the core processing bottleneck if the job is compute-intensive; the most effective immediate step is to increase DPUs.

52
MCQeasy

A company uses AWS CloudFormation to manage its infrastructure. The operations team reports that stack updates often fail because of resource conflicts. The team wants to improve the reliability of updates without manual intervention. Which solution provides the MOST automated recovery from update failures?

A.Use CloudFormation change sets to review and approve all changes before update.
B.Write a custom AWS Lambda function that reverts changes when a stack update fails.
C.Apply a stack policy to prevent updates to critical resources.
D.Use the default CloudFormation rollback behavior that automatically reverts changes on failure.
AnswerD

CloudFormation's default rollback automatically reverts stack resources to their last known stable state when an update fails, restoring service without operator involvement. This built-in behaviour delivers the hands-off recovery the team requires, unlike manual rollback or custom remediation tooling.

Why this answer

CloudFormation's built-in rollback behavior automatically reverts all changes made during a failed stack update, restoring the stack to its last known stable state without requiring any custom code or manual intervention. This provides the most automated recovery mechanism as it is natively integrated into the CloudFormation service and requires no additional infrastructure or scripting.

Exam trap

The trap here is that candidates may overthink the solution and choose a custom Lambda function (Option B) thinking it provides more control, when in fact CloudFormation's native rollback is the most automated and reliable approach, and custom solutions often introduce additional failure points.

How to eliminate wrong answers

Option A is wrong because change sets are a review and approval mechanism that helps prevent errors before an update is executed, but they do not provide any automated recovery after a failure occurs. Option B is wrong because writing a custom Lambda function to revert changes introduces unnecessary complexity, potential for errors, and is not as reliable or automated as CloudFormation's native rollback, which handles state management and resource dependencies correctly. Option C is wrong because stack policies only prevent updates to specific critical resources during a stack update, but they do not provide any recovery mechanism if the update fails due to conflicts elsewhere.

53
MCQhard

A solutions architect is optimizing a data processing workload that runs on AWS Lambda. The function processes large JSON files stored in Amazon S3, performs CPU-intensive transformations, and writes results to Amazon DynamoDB. The function currently has 512 MB of memory and takes about 10 minutes to process each file, occasionally timing out. The architect needs to reduce processing time and avoid timeouts. Which action is MOST effective?

A.Enable Lambda provisioned concurrency to ensure the function is always warm.
B.Increase the Lambda function's memory to 3008 MB and adjust the timeout to 15 minutes.
C.Move the function to a VPC and increase the ephemeral storage to 10 GB.
D.Configure the function to use AWS X-Ray tracing and enable AWS Lambda Insights.
AnswerB

Increasing memory also proportionally increases CPU and network bandwidth for Lambda. For CPU-intensive workloads, this can significantly reduce processing time. Raising the timeout to the maximum 15 minutes provides headroom to avoid timeouts while the function completes. This is the most direct way to improve performance for this scenario.

Why this answer

For CPU-intensive Lambda functions, memory allocation directly controls CPU share. Doubling or quadrupling memory can cut execution time substantially. Raising the timeout to 15 minutes provides a safety margin.

Provisioned concurrency, VPC placement, and monitoring tools do not increase compute resources during execution, so they fail to address the core performance and timeout problem.

Exam trap

The trap here is focusing on cold starts or observability when the real bottleneck is insufficient CPU, which is tied to the memory setting in Lambda.

54
MCQmedium

A company uses AWS CloudTrail to log all API calls. The security team wants to be alerted when an IAM user creates a new access key. What is the MOST efficient way to achieve this?

A.Enable CloudTrail Insights to detect unusual key creation patterns.
B.Create a CloudWatch Events rule that matches the CreateAccessKey API call and sends an SNS notification.
C.Use CloudTrail to publish logs to CloudWatch Logs and create a metric filter to trigger an alarm.
D.Configure Amazon Athena to query CloudTrail logs and set up a scheduled query to notify.
AnswerB

CloudWatch Events (now EventBridge) pattern-matches the CreateAccessKey API call directly from CloudTrail's management event stream, triggering SNS without polling or log parsing. This satisfies the efficiency constraint: no Lambda, no CloudWatch Logs metric filters, and near-real-time alerting on the exact IAM action specified.

Why this answer

The most efficient way to alert when an IAM user creates a new access key is to use CloudWatch Events (now Amazon EventBridge) to match the CreateAccessKey API call from CloudTrail and trigger an SNS notification. Option A is incorrect because CloudTrail Insights is for detecting unusual activity patterns, not for real-time event-driven alerts. Option C is less efficient because it requires additional steps of publishing logs to CloudWatch Logs and creating a metric filter, adding complexity.

Option D is inefficient because querying CloudTrail logs with Athena is not real-time and requires custom scheduling.

55
MCQhard

A company uses AWS CloudFormation to manage infrastructure. The stack fails to update with the error: 'Resource handler returned message: The subnet 'subnet-xxx' is in use by a network interface.' The subnet is associated with a Lambda function in a VPC. The CloudFormation template is trying to delete the subnet. What should the company do to resolve this?

A.Update the Lambda function configuration to remove the VPC settings, then delete the subnet.
B.Modify the CloudFormation template to ignore the deletion failure using a DeletionPolicy attribute.
C.Use the AWS CLI to force delete the subnet.
D.Manually delete the Elastic Network Interface (ENI) from the AWS Management Console.
AnswerA

The Lambda function's elastic network interface holds the subnet, blocking deletion. Detaching the VPC configuration releases that interface, allowing CloudFormation to delete the subnet cleanly. This directly resolves the 'in use by a network interface' error the stack reports.

Why this answer

The error indicates that the subnet cannot be deleted because it is in use by a network interface, which is likely the Lambda function's elastic network interface (ENI). To resolve this, the Lambda function's VPC configuration must be removed so that the ENI is released, allowing the subnet to be deleted. This is the correct approach because CloudFormation cannot delete a subnet that still has dependencies.

Exam trap

SAP-C02 often tests the misconception that DeletionPolicy can force deletion of resources with dependencies, but it only controls retention; the real solution is to remove the dependency first.

How to eliminate wrong answers

Option B is wrong because a DeletionPolicy attribute only controls what happens to a resource when it is removed from the stack; it does not bypass the dependency check that prevents deletion of a subnet in use. Option C is wrong because AWS does not provide a force delete for subnets; the subnet must be free of dependencies. Option D is wrong because manually deleting the ENI may not be possible if it is managed by Lambda, and even if it were, it would not be a sustainable solution; the Lambda function would recreate the ENI.

56
MCQmedium

A company uses AWS CloudFormation to manage infrastructure. The operations team wants to implement a change management process where all stack updates must be reviewed and approved before execution. The team currently uses AWS CodePipeline for CI/CD. Which solution meets these requirements with the LEAST operational overhead?

A.Use CloudFormation Change Sets and require a senior engineer to execute them.
B.Write an AWS Lambda function that triggers on stack update events and requires approval via Amazon SNS.
C.Use AWS Service Catalog to govern CloudFormation templates and require approval for provisioning.
D.Store CloudFormation templates in AWS CodeCommit and use AWS CLI to execute updates after peer review.
E.Create a CodePipeline pipeline with an approval stage before the CloudFormation deployment action.
AnswerE

A manual approval stage in CodePipeline halts the pipeline until a nominated reviewer approves, then triggers the CloudFormation deploy action. This satisfies the mandated review-and-approval gate before execution while reusing the existing CI/CD tooling, so no custom change-management system is built, keeping operational overhead minimal.

Why this answer

AWS CodePipeline natively supports approval actions, allowing a manual approval stage to be inserted before the CloudFormation deployment action. This integrates directly with the existing CI/CD pipeline, requires minimal custom code, and enforces review/approval before stack updates. It is the lowest-operational-overhead solution that meets the change management requirement.

Exam trap

SAP-C02 often tests the 'least operational overhead' constraint — candidates may over-engineer with Lambda/SNS or Service Catalog when the native CodePipeline approval action is the simplest fit.

How to eliminate wrong answers

Option A is wrong because relying on a senior engineer to manually execute change sets is error-prone and does not integrate with the CI/CD pipeline, adding operational overhead. Option B is wrong because writing a Lambda function and SNS approval workflow is custom code that must be built and maintained, increasing overhead. Option C is wrong because Service Catalog governs provisioning of approved products, not approval of stack updates within a CI/CD pipeline.

Option D is wrong because storing templates in CodeCommit and using CLI after peer review is a manual process that bypasses pipeline automation and adds overhead.

57
MCQmedium

Refer to the exhibit. $ aws ec2 describe-instances --region us-east-1 --filters Name=tag:Name,Values=WebServer --query 'Reservations[].Instances[].{ID:InstanceId,State:State.Name,Type:InstanceType,LaunchTime:LaunchTime}' --output table A DevOps engineer runs the above command. The Auto Scaling group for WebServer instances has a desired count of 3, but the engineer notices that there are 5 instances with the same tag. What is the MOST likely cause?

A.The Auto Scaling group has a cooldown period that prevents immediate termination.
B.The command is filtering by the wrong tag key.
C.The other two instances are in a 'terminating' state and are not returned by the command.
D.The instances were launched with a different launch template that does not have the tag.
AnswerB

Correct. The tag key in the filter (Name) may not match the actual tag key on some instances. If the tag key is different (e.g., 'name' or 'web'), those instances would not be returned, but the engineer sees 5 instances, indicating the filter is matching instances from outside the ASG.

Why this answer

The `describe-instances` command filters by the `Name` tag with value `WebServer`, returning every instance that has that exact tag, regardless of whether it belongs to the Auto Scaling group. Seeing 5 instances while the ASG desired count is 3 indicates the filter is not specific to the ASG. The engineer is likely filtering by the wrong tag key; they should use an ASG-specific tag such as `tag:aws:autoscaling:groupName`.

Option C is incorrect because `describe-instances` returns terminating instances. Option A is less likely because cooldown periods do not affect tag-based instance queries. Option D is incorrect because instances launched without the `Name` tag would not appear in the output.

Exam trap

A common trap is to assume that instances in 'terminating' state are not returned by describe-instances, but they are. The filter's tag key must exactly match the instance tags; minor differences can lead to missing or extra results.

58
MCQhard

A company has a serverless application using AWS Lambda, API Gateway, and DynamoDB. During a traffic spike, some API requests fail with 5xx errors. The CloudWatch logs show 'ProvisionedThroughputExceededException' for DynamoDB. The team wants to handle this gracefully without losing requests. What should they do?

A.Enable auto scaling for DynamoDB and implement retry logic with exponential backoff in the Lambda function.
B.Increase the provisioned read/write capacity of the DynamoDB table to a high fixed value.
C.Use an SQS queue between API Gateway and Lambda to buffer requests.
D.Configure API Gateway to automatically retry failed requests.
AnswerA

DynamoDB's ProvisionedThroughputExceededException arises when read/write capacity units are exhausted, so on-demand or auto scaling absorbs the spike. Retries with exponential backoff and jitter let Lambda reattempt throttled calls, satisfying the requirement to handle failures gracefully without losing requests. Together they address capacity exhaustion and transient throttling.

Why this answer

The ProvisionedThroughputExceededException indicates that DynamoDB is throttling requests because the table's provisioned capacity is insufficient for the traffic spike. Enabling DynamoDB auto scaling allows the table to automatically adjust its read/write capacity based on utilization, while implementing retry logic with exponential backoff in the Lambda function ensures that throttled requests are retried after progressively longer intervals, preventing request loss and handling temporary spikes gracefully. This combination directly addresses the root cause and provides resilience.

Exam trap

SAP-C02 often tests the misconception that simply increasing provisioned capacity or adding a queue solves throttling, but the key is combining auto scaling with retry logic to handle spikes without losing requests.

How to eliminate wrong answers

Option B is wrong because setting a high fixed capacity does not handle dynamic traffic spikes efficiently and can be cost-prohibitive; it also doesn't address the need for retries when throttling still occurs. Option C is wrong because adding an SQS queue between API Gateway and Lambda changes the architecture to asynchronous processing, which may not be suitable for synchronous API requests expecting immediate responses, and it doesn't directly solve DynamoDB throttling. Option D is wrong because API Gateway does not automatically retry failed requests; retries must be implemented in the client or backend, and API Gateway's default behavior is to return the error to the caller.

59
MCQmedium

A company has a web application running on Amazon EC2 instances in an Auto Scaling group. The application writes logs to local instance storage. The operations team wants to centralize log analysis and enable real-time alerting on specific error patterns. The solution must be highly available and require minimal changes to the application. Which approach should a solutions architect recommend?

A.Modify the application to write logs directly to Amazon Kinesis Data Firehose, which delivers them to Amazon S3 for analysis.
B.Set up an AWS Lambda function that periodically connects to each instance via SSH to retrieve logs and publish them to Amazon SNS.
C.Configure the instances to mount a shared Amazon EFS file system and write logs there, then use a third-party tool to analyze the logs.
D.Install and configure the Amazon CloudWatch agent on each instance to send logs to Amazon CloudWatch Logs, then use metric filters and alarms for alerting.
AnswerD

The CloudWatch agent can collect logs from local files and send them to CloudWatch Logs without application changes. Metric filters can extract patterns and trigger alarms. This solution is highly available because CloudWatch Logs is a regional service, and it provides real-time alerting with minimal operational overhead.

Why this answer

The CloudWatch agent provides a managed way to collect logs from EC2 instances and send them to CloudWatch Logs. Metric filters can monitor for specific patterns and trigger alarms in near real-time. This requires no application changes and is highly available, making it the most efficient solution.

Exam trap

The trap here is assuming that application code changes or custom scripts are needed for log centralization, when the CloudWatch agent can handle it without modifications.

60
Multi-Selecthard

A company is running a critical microservices application on Amazon ECS with Fargate launch type. The application uses an Application Load Balancer (ALB) to distribute traffic. Recently, the team noticed that the ALB's 5xx error rate has increased. The error is HTTP 503. The team suspects the target group is unhealthy. Which THREE steps should the team take to diagnose and resolve the issue?

Select 3 answers
A.Verify the ECS service's desired count and compare with the number of healthy tasks in the target group.
B.Replace the ALB with a Network Load Balancer (NLB) to bypass health checks.
C.Check the ECS service events and task status to ensure tasks are running and passing health checks.
D.Verify that sticky sessions (session affinity) are enabled on the target group.
E.Check the ALB access logs and health check settings for the target group.
AnswersA, C, E

Comparing the ECS service's desired count against healthy targets in the target group exposes capacity shortfalls: if tasks fail ALB health checks, they are deregistered, leaving too few healthy targets and triggering HTTP 503 responses. This directly addresses the suspected unhealthy target group constraint in the stem.

Why this answer

Option A is correct because an HTTP 503 from an ALB typically means the target group has no healthy targets; comparing the ECS service's desired count against the number of healthy targets registered in the target group immediately reveals whether tasks are failing to register or failing health checks. Option C is correct because ECS service events and task status (via the ECS console, describe-services, or describe-tasks) show why tasks are stopping, restarting, or failing to reach a steady state, which is the root cause of unhealthy targets. Option E is correct because ALB access logs record the 503 responses and target health, while the target group health check settings (path, port, protocol, interval, timeout, healthy/unhealthy thresholds, and success codes) determine whether tasks are marked healthy; misconfigured health checks are a common cause of 503s.

Option B is not appropriate because replacing the ALB with an NLB does not bypass the underlying problem and NLB health checks would still mark targets unhealthy, plus it changes the architecture unnecessarily. Option D is not relevant because sticky sessions (session affinity) affect request routing to already-healthy targets and do not cause or resolve 503 errors from an unhealthy target group.

61
MCQhard

Refer to the exhibit. A CloudFormation stack was successfully created. The stack's template includes an S3 bucket and a Lambda function. A developer runs the CLI command shown but receives an error that the stack does not exist. What is the MOST likely cause?

A.The AWS CLI is configured to use a different region than where the stack was deployed.
B.The stack was deleted after creation.
C.The stack name is case-sensitive and should be 'MyApp'.
D.The '--query' parameter is incorrectly formatted.
AnswerA

Stack names are unique per region; querying the wrong region returns 'stack does not exist'.

Why this answer

The CLI command is querying the wrong region. The stack was created in us-east-1 but the CLI default region might be different. Option B is wrong because the stack was created successfully.

Option C is wrong because outputs are returned correctly. Option D is wrong because the stack name is exactly as used.

62
MCQmedium

A company runs a containerized application on Amazon ECS with Fargate. The application needs to access an Amazon S3 bucket that contains sensitive data. The security team requires that all traffic between the ECS tasks and S3 remain within the AWS network and not traverse the internet. What is the MOST secure way to meet this requirement?

A.Use an internet gateway and route traffic through a NAT gateway.
B.Enable S3 Transfer Acceleration on the bucket.
C.Create a VPC endpoint for S3 and attach it to the VPC.
D.Use a NAT gateway and update the route table to direct S3 traffic to the NAT.
AnswerC

VPC endpoint enables private connectivity to S3 without internet.

Why this answer

Using a VPC endpoint for S3 (Gateway or Interface) ensures traffic stays within the AWS network. Option A is wrong because internet traffic goes over the public internet. Option B is wrong because a NAT gateway is for outbound internet, not private access to S3.

Option D is wrong because S3 Transfer Acceleration uses the internet.

63
MCQmedium

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The application experiences intermittent high latency due to CPU spikes on some instances. The company wants to automatically replace unhealthy instances and optimize costs. What should a solutions architect do?

A.Configure a target tracking scaling policy based on average CPU utilization.
B.Use a lifecycle hook to perform a health check and terminate unhealthy instances.
C.Use an AWS Lambda function to terminate instances with high CPU.
D.Implement a scheduled scaling policy to increase instances during peak hours.
AnswerA

Target tracking adjusts desired capacity automatically to hold average CPU at a defined target, adding instances when spikes occur and removing them when demand falls. This addresses the CPU-driven latency while optimising cost, and unhealthy instances are replaced by the Auto Scaling group's health checks.

Why this answer

A target tracking scaling policy dynamically adjusts the number of instances based on average CPU utilization, which addresses intermittent CPU spikes. The Auto Scaling group also automatically replaces instances that fail ALB health checks, ensuring unhealthy instances are replaced. This optimizes costs by scaling down during low usage.

Option B is incorrect: lifecycle hooks are for custom actions during instance launch or termination, not for replacing unhealthy instances. Option C is incorrect: terminating instances with high CPU via Lambda does not integrate with Auto Scaling and could cause instability. Option D is incorrect: scheduled scaling is for predictable traffic patterns, not intermittent spikes.

Exam trap

Candidates may think that lifecycle hooks are needed to replace unhealthy instances, but ALB health checks integrated with Auto Scaling already handle this automatically.

64
MCQeasy

A company runs a static website on Amazon S3 behind Amazon CloudFront. The website uses a custom domain and SSL certificate from AWS Certificate Manager (ACM). Users report that they sometimes see an older version of the website after updates. What should the company do to ensure users always see the latest content?

A.Disable and re-enable the CloudFront distribution after each update.
B.Enable S3 bucket versioning and use version IDs in URLs.
C.Reduce the CloudFront TTL to 0 seconds for all objects.
D.Create a CloudFront invalidation for the updated files.
AnswerD

CloudFront caches objects at edge locations until the TTL expires, so updated S3 objects can still serve stale copies. An invalidation forces edge locations to fetch the current object from the origin, satisfying the requirement that users always see the latest content.

Why this answer

CloudFront caches objects at edge locations based on TTL and only re-fetches from the S3 origin when the cache expires or is explicitly invalidated. When a static site is updated, existing cached copies at edge locations continue to be served until they expire, which is why users sometimes see stale content. Creating a CloudFront invalidation for the updated files (or a wildcard path like /*) forces the edge locations to discard the cached objects and retrieve the latest versions from S3 on the next request.

Exam trap

SAP-C02 often tests the misconception that changing origin-side settings (versioning, TTLs, or toggling the distribution) refreshes edge caches, when only an explicit CloudFront invalidation purges cached objects.

How to eliminate wrong answers

Option A is wrong because disabling and re-enabling a distribution does not purge cached objects at edge locations and causes unnecessary downtime while the distribution redeploys. Option B is wrong because S3 bucket versioning only retains multiple versions of objects in the origin bucket; CloudFront still serves whatever version is cached at the edge, and version IDs in URLs would change the cache key rather than solve the staleness problem. Option C is wrong because setting TTL to 0 seconds disables caching benefits, dramatically increases origin load and latency, and is not the intended mechanism for pushing updates — invalidation is the correct, targeted tool.

65
MCQeasy

A company uses Amazon RDS for MySQL for its database. The operations team notices that read queries are slow during peak hours. The application is read-heavy and can tolerate eventual consistency. Which solution would improve read performance with minimal application changes?

A.Increase the DB instance class to a larger size.
B.Enable Multi-AZ deployment for failover support.
C.Enable RDS Proxy to pool database connections.
D.Create an RDS read replica and direct read traffic to it.
AnswerD

An RDS read replica offloads read queries to a separate MySQL instance via asynchronous replication, directly relieving the primary during peak hours. Because the application tolerates eventual consistency, replication lag is acceptable, and directing reads to the replica endpoint requires minimal application change.

Why this answer

Creating an RDS read replica and directing read traffic to it offloads read queries from the primary instance, improving read performance for a read-heavy application that tolerates eventual consistency. This requires minimal application changes—typically just updating the connection string for read operations to point to the replica endpoint. Read replicas use asynchronous replication, so they are suitable for eventually consistent reads.

Exam trap

SAP-C02 often tests the misconception that Multi-AZ improves read performance (it does not; the standby is not readable) or that RDS Proxy increases read throughput (it pools connections, not reads).

How to eliminate wrong answers

Option A is wrong because increasing the DB instance class scales both reads and writes vertically, which is more expensive and does not specifically address read-heavy load; it also requires a reboot and may not be cost-effective. Option B is wrong because Multi-AZ is for high availability and failover, not read scaling; the standby does not serve read traffic. Option C is wrong because RDS Proxy pools and shares database connections to improve scalability and resilience, but it does not increase read throughput or offload read queries from the primary.

66
MCQhard

A company runs a critical application on Amazon EC2 instances in a single AWS Region. The application uses an Amazon RDS for MySQL database with a read replica in another Availability Zone. The company wants to improve the disaster recovery posture to meet an RPO of 5 minutes and an RTO of 15 minutes in case of a regional failure. The application must be able to fail over to a secondary Region with minimal data loss. Which solution meets these requirements?

A.Configure an Amazon RDS for MySQL cross-Region read replica in the secondary Region. Promote the read replica to a standalone database in the event of a regional failure.
B.Enable multi-AZ deployment for the RDS database and use a second standby in the secondary Region with synchronous replication.
C.Use AWS Database Migration Service (AWS DMS) with change data capture (CDC) to continuously replicate data to an Amazon RDS for MySQL instance in the secondary Region.
D.Enable automated backups for the RDS database and copy the backups to the secondary Region. In the event of a failure, restore the database from the latest backup in the secondary Region.
AnswerA

A cross-Region read replica uses asynchronous replication to keep a copy of the database in another Region. The replication lag is typically under a minute, satisfying the 5-minute RPO. Promotion of a read replica to a standalone instance takes only a few minutes, well within the 15-minute RTO. This is a cost-effective and reliable DR solution for RDS for MySQL, and it meets both the RPO and RTO requirements for this scenario.

Why this answer

A cross-Region read replica provides asynchronous replication to a secondary Region, typically with sub-minute latency, which satisfies the 5-minute RPO. In a disaster, promoting the read replica to a standalone database takes only a few minutes, meeting the 15-minute RTO. This is a native RDS feature that requires minimal operational overhead and is a recommended DR pattern for RDS for MySQL.

Other options either do not meet the RTO or rely on unsupported or overly complex mechanisms.

Exam trap

The trap here is assuming that automated backups with cross-Region copy can meet a 15-minute RTO.

67
MCQmedium

A company has a web application running on Amazon EC2 instances in an Auto Scaling group. The application uses a self-signed SSL certificate on the instances, and an Application Load Balancer (ALB) terminates SSL. Users report intermittent SSL certificate errors. The security team requires that the certificate be managed and rotated automatically. Which solution should a solutions architect implement to meet these requirements?

A.Request a public certificate from AWS Certificate Manager (ACM) for the application's domain and attach it to the ALB. Remove the self-signed certificate from the instances.
B.Request a public certificate from AWS Certificate Manager (ACM) for the application's domain, attach it to the ALB, and configure the instances to use it.
C.Import the self-signed certificate into AWS Certificate Manager (ACM) and use it on the ALB. Configure ACM to rotate the certificate annually.
D.Use AWS Secrets Manager to store the self-signed certificate and configure the instances to retrieve and install it on a schedule using a Lambda function.
AnswerA

ACM provides public certificates that are automatically renewed and deployed. Attaching the ACM certificate to the ALB ensures that SSL termination uses a trusted certificate, eliminating intermittent errors caused by the self-signed certificate. Since the ALB handles SSL termination, the instances do not need certificates. This meets the requirements for managed and automatically rotated certificates.

Why this answer

The intermittent SSL errors are likely due to the self-signed certificate not being trusted by clients. Using a public ACM certificate on the ALB provides a trusted certificate that is automatically managed and renewed. Since the ALB terminates SSL, the instances do not need any certificate.

This solution meets the security team's requirement for automatic management and rotation.

Exam trap

The trap here is assuming that ACM can automatically rotate imported self-signed certificates or that instances need the certificate when an ALB terminates SSL.

68
MCQhard

A company uses AWS CloudFormation to manage infrastructure. A recent stack update failed because a resource exceeded a service quota. The team wants to be notified proactively when service limits are approaching. Which solution meets this requirement?

A.Use AWS CloudTrail to monitor API calls that indicate quota exhaustion.
B.Use AWS Config rules to check if resources are within limits.
C.Use AWS Trusted Advisor to check service limits regularly.
D.Use Amazon CloudWatch to monitor service quota usage metrics and set CloudWatch alarms.
AnswerD

CloudWatch publishes Service Quotas usage metrics, so alarms can fire before a limit is breached, satisfying the proactive notification requirement. Unlike Trusted Advisor's periodic checks, these metrics update continuously, letting the team act before a CloudFormation update fails again.

Why this answer

AWS Service Quotas publishes CloudWatch metrics (e.g., AWS/Usage namespace with Resource and Service dimensions) that track current usage against quotas. Creating CloudWatch alarms on these metrics allows proactive notification before a quota is exhausted, directly addressing the requirement. This is the AWS-recommended approach for quota monitoring at scale.

Exam trap

SAP-C02 often tests whether candidates know that Trusted Advisor is not real-time and requires paid support, while CloudWatch alarms on Service Quotas metrics provide the proactive, automated notification the question demands.

How to eliminate wrong answers

Option A is wrong because CloudTrail logs API activity after the fact — it can show a quota-exceeded error occurred, but it is reactive, not proactive, and does not emit metrics you can alarm on. Option B is wrong because AWS Config rules evaluate resource configuration compliance (e.g., 'is this bucket encrypted?'), not service quota consumption, and there are no native Config rules that track quota utilization. Option C is wrong because Trusted Advisor's service limits check is only available on Business/Enterprise Support plans, runs periodically (not real-time), and does not integrate with CloudWatch alarms for automated notification.

69
MCQhard

A company runs a microservices application on Amazon EKS. The application consists of multiple services that communicate over HTTP. The operations team wants to implement a service mesh to gain observability, traffic management, and security features without modifying application code. They also want to minimize operational overhead. Which solution should a solutions architect recommend?

A.Deploy AWS App Mesh with Envoy sidecar proxies injected into each pod.
B.Configure an Application Load Balancer (ALB) with AWS WAF to manage traffic between services.
C.Use AWS Cloud Map for service discovery and implement custom retry logic in each microservice.
D.Install and configure Istio on the EKS cluster using Helm charts.
AnswerA

AWS App Mesh is a managed service mesh that uses Envoy proxies to provide observability, traffic routing, and security without application code changes. It integrates with EKS and handles control plane management, reducing operational overhead. It supports features like circuit breaking, retries, and mutual TLS. This is the native AWS solution for service mesh on EKS, meeting all requirements.

Why this answer

AWS App Mesh is a managed service mesh that provides consistent observability, traffic control, and security for microservices on EKS. It uses Envoy sidecar proxies injected into pods, requiring no application code changes. As a managed service, it reduces operational overhead compared to self-managed options like Istio.

It supports features such as traffic routing, retries, and mutual TLS, directly addressing the requirements.

Exam trap

The trap here is assuming that any service mesh requires significant operational effort, but AWS App Mesh is managed and minimizes overhead while providing the needed features.

70
MCQmedium

A company uses AWS Lambda functions behind an Amazon API Gateway REST API. The Lambda functions query an Amazon RDS for PostgreSQL database. Recently, the company has noticed increased latency and occasional timeouts during peak hours. A solutions architect needs to improve the performance and scalability of the database layer. Which solution will meet these requirements with the LEAST operational overhead?

A.Enable Amazon DynamoDB Accelerator (DAX) on the RDS instance.
B.Add a Multi-AZ RDS Read Replica and modify Lambda to use the Read Replica for queries.
C.Increase the instance size of the RDS database to handle more concurrent connections.
D.Implement Amazon RDS Proxy to manage connection pooling between Lambda and the RDS instance.
AnswerD

Lambda opens a new database connection per invocation, exhausting PostgreSQL's connection limit during peak load. RDS Proxy pools and reuses connections, absorbing bursts without application changes. This reduces latency and timeouts with minimal operational overhead, since AWS manages the proxy infrastructure.

Why this answer

AWS Lambda functions can overwhelm an RDS database with too many concurrent connections, causing latency and timeouts. RDS Proxy pools and shares database connections, reducing the connection overhead and improving scalability without code changes. It requires minimal operational overhead because it is a managed service that integrates with IAM and Secrets Manager.

Exam trap

SAP-C02 often tests the misconception that read replicas or larger instances solve Lambda-to-RDS connection issues, when the core problem is connection pooling and the answer is RDS Proxy.

How to eliminate wrong answers

Option A is wrong because DAX is a caching service for DynamoDB, not RDS; it cannot be enabled on an RDS instance. Option B is wrong because a Multi-AZ read replica is for disaster recovery, not for offloading read queries; a read replica can offload reads but requires application changes and does not solve connection pooling, and Multi-AZ is synchronous standby. Option C is wrong because increasing instance size may handle more connections but does not address the fundamental connection management issue and can be costly and still limited by max_connections.

71
MCQeasy

A company uses Amazon CloudFront to deliver static content from an S3 bucket. They want to restrict access so that only CloudFront can access the S3 bucket. What configuration should they use?

A.Set the S3 bucket policy to allow access only from CloudFront's public IP ranges.
B.Create an origin access identity (OAI) and grant it read access to the S3 bucket.
C.Attach an IAM role to CloudFront distribution.
D.Configure CloudFront signed URLs.
AnswerB

An origin access identity (OAI) is a special CloudFront user that the distribution attaches to origin requests, letting the S3 bucket policy grant read access solely to that identity. This satisfies the requirement that only CloudFront reach the bucket, blocking direct public S3 access. Note that Microsoft Entra ID is unrelated here.

Why this answer

Create an origin access identity (OAI) and grant it read access to the S3 bucket. An OAI is a special CloudFront user that allows CloudFront to access private S3 bucket content securely. Option A is incorrect because CloudFront does not have static public IP ranges; it uses a large dynamic range.

Option C is incorrect because IAM roles are not used directly for CloudFront to S3 access; OAI is the standard method. Option D is incorrect because signed URLs control end-user access, not origin access.

72
Multi-Selecthard

A company runs a web application on Amazon ECS with Fargate launch type. The application uses an Application Load Balancer. The operations team notices that the ALB returns 503 errors during peak traffic. Which TWO actions should the solutions architect take to resolve this issue?

Select 2 answers
A.Increase the idle timeout on the ALB.
B.Enable ECS service Auto Scaling to automatically adjust the number of tasks.
C.Increase the deregistration delay on the target group.
D.Review the ECS service events for task failures or health check issues.
E.Increase the task memory allocation in the task definition.
AnswersB, D

ECS service Auto Scaling adds tasks when demand rises, so more targets register behind the ALB and absorb peak load. This directly addresses the 503s, which stem from insufficient healthy targets to serve requests, rather than from listener or health-check misconfiguration. Scaling task count restores capacity and keeps the target group populated.

Why this answer

Option B is correct because 503 errors from an ALB during peak traffic commonly indicate that the ECS service has insufficient healthy tasks to handle the load; enabling ECS service Auto Scaling (via Application Auto Scaling with target tracking on metrics like ALBRequestCountPerTarget or ECSServiceAverageCPUUtilization) allows the service to add tasks automatically to absorb the spike. Option D is correct because 503 responses can also stem from tasks failing to start, crashing, or failing ALB target group health checks; reviewing ECS service events surfaces messages such as 'service unable to place a task' or failed health checks, which is the essential diagnostic step to identify the root cause before remediation. Option A is not appropriate because the ALB idle timeout governs how long an idle connection is kept open, and increasing it does not address capacity shortages or unhealthy targets that produce 503s.

Option C is incorrect because the deregistration delay only controls how long the target group waits before completing deregistration of a draining target, which affects connection draining during scale-in or deployment, not peak-load 503 errors. Option E is incorrect because increasing task memory only helps if tasks are being killed due to out-of-memory conditions, and it does not scale capacity or diagnose the actual cause of the 503s.

Exam trap

The trap is choosing ALB-level timeout or deregistration settings for a 503 error, when 503s from an ALB almost always indicate a target health/capacity problem on the ECS side, not a load balancer configuration issue.

73
MCQmedium

A company runs a steady-state HTTP API on a fleet of Amazon EC2 instances behind an Application Load Balancer. The operations team needs to reduce cost without degrading performance or availability. They observe that CPU utilization is consistently 8–12% and memory utilization is around 40%. The workload is stateless and can tolerate a brief instance restart during deployment. Which change should a solutions architect recommend to meet the cost-reduction goal?

A.Enable detailed CloudWatch monitoring at one-minute resolution and add a dashboard for CPU and memory metrics.
B.Migrate the workload to a smaller instance type and adjust the Auto Scaling group's desired capacity to match observed peak demand.
C.Purchase a 3-year All Upfront Compute Savings Plan covering the existing instance family and run the fleet unchanged.
D.Replace the Application Load Balancer with a Network Load Balancer and route traffic directly to the instances.
AnswerB

The workload is stateless and consistently underutilized, so right-sizing to a smaller instance type and calibrating desired capacity to actual peak demand removes idle compute. Because the fleet sits behind a load balancer and tolerates brief restarts, the resize can be performed with a launch template update and an instance refresh, preserving availability while lowering the hourly run rate.

Why this answer

The fleet is stateless, load balanced, and running far below capacity, which makes right-sizing the instance type and re-calibrating desired capacity the most direct lever for cost reduction. A launch template update combined with an instance refresh can roll the change out without downtime. Savings Plans and monitoring changes reduce or report cost but never eliminate the waste from consistently idle compute.

Exam trap

The trap here is assuming that a Savings Plan or Reserved Instance discount is the fastest path to lower cost, when it only discounts capacity the workload does not actually need.

74
MCQeasy

A company runs a critical application on EC2 instances in an Auto Scaling group. They want to be notified immediately if any instance fails a status check. What is the simplest solution?

A.Configure an ELB health check and monitor the unhealthy host count.
B.Use AWS Systems Manager Automation to check instance status periodically.
C.Create a CloudWatch alarm on the StatusCheckFailed metric with an SNS action.
D.Enable AWS CloudTrail and create a metric filter for EC2 instance failures.
AnswerC

CloudWatch's StatusCheckFailed metric directly reports EC2 instance health, and an alarm on it triggers an SNS notification the moment a check fails. This satisfies the immediate-notification requirement with minimal configuration, avoiding custom scripts or Lambda polling. Auto Scaling group events alone would not surface individual instance status-check failures promptly.

Why this answer

CloudWatch can monitor the EC2 StatusCheckFailed metric and trigger an SNS notification when an instance fails a status check. This is the simplest solution because it uses built-in metrics and requires no custom scripting. Option A uses ELB health checks which monitor load balancer target health, not instance status checks.

Option B uses Systems Manager Automation which is not real-time. Option D uses CloudTrail which logs API calls, not status checks.

75
MCQmedium

A company runs a production AWS Lambda function that processes orders. Recently, the function has been timing out occasionally. The function uses a VPC with a single private subnet and has a timeout of 30 seconds. What is the MOST likely cause of the timeout?

A.The function is experiencing cold starts due to high concurrency.
B.The function is hitting the maximum concurrent execution limit.
C.The function needs to be attached to a public subnet.
D.The function does not have a NAT gateway or VPC endpoints to access external resources.
AnswerD

A Lambda function in a private subnet has no route to the internet without a NAT gateway, and no path to AWS services without VPC endpoints. Calls to external order-processing resources therefore hang until the 30-second timeout expires, producing the intermittent failures described.

Why this answer

A Lambda function attached to a VPC with only a private subnet has no route to the internet unless a NAT gateway (for outbound internet) or VPC endpoints (for AWS services) are configured. If the function calls external resources or AWS APIs without these, the calls hang until the 30-second timeout. This is the most likely cause of intermittent timeouts in this scenario.

Exam trap

SAP-C02 often tests the misconception that Lambda functions in a VPC automatically have internet access — candidates forget that VPC attachment removes default internet connectivity and that NAT or VPC endpoints are required.

How to eliminate wrong answers

Option A is wrong because cold starts add latency but typically only a few hundred milliseconds to a few seconds, not 30-second timeouts, and they would not cause consistent timeouts. Option B is wrong because hitting the concurrency limit results in throttling errors (429) or invocation failures, not timeouts of the function's own execution. Option C is wrong because Lambda functions in a VPC do not need to be in a public subnet; they need a NAT gateway or VPC endpoints for outbound access, and placing them in a public subnet does not by itself grant internet access without an internet gateway route and public IP.

Page 1 of 4 · 228 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Continuous Improvement for Existing Solutions questions.