Courseiva

SAA-C03 (SAA-C03) — Questions 1–75

935 questions total · 13pages · All types, answers revealed

Page 1 of 13

Page 2
1
Multi-Selectmedium

A production Amazon RDS database already has automated backups enabled. At 10:45 UTC, the team discovers that a faulty migration corrupted rows in a table at 10:30 UTC. The business wants the database restored to exactly the state it had at 10:30 UTC with minimal risk. Which two actions should the team take? Select two.

Select 2 answers
A.Restore the database to a new instance using point-in-time restore for 10:30 UTC.
B.Validate the restored database, then switch the application endpoint to the restored database.
C.Restore the most recent manual snapshot because it will include the 10:30 UTC state.
D.Overwrite the existing database instance in place so the application keeps the same storage volume.
E.Wait for automated backups to complete again, then replay the migration to restore the missing rows.
AnswersA, B

Correct. Point-in-time restore is the RDS recovery method for returning to a specific moment before the corruption occurred. Restoring to a new instance gives the team a clean database copy at the desired timestamp without risking the current production instance.

Why this answer

Amazon RDS Point-in-Time Restore (PITR) allows you to restore a DB instance to any second within the backup retention period, including 10:30 UTC. This uses automated backups and transaction logs to reconstruct the exact database state at that specific time, providing a precise recovery point with minimal data loss.

Exam trap

The trap here is that candidates may think manual snapshots can be used for point-in-time recovery, but they only capture a single moment and cannot roll forward to a specific time like automated backups can.

2
MCQeasy

Based on the exhibit, the web team wants the application to continue serving traffic if one Availability Zone fails. Which change best meets the requirement with the least operational overhead?

A.Increase desired capacity to 3 in the same Availability Zone so one extra instance is always available.
B.Add the unused subnet in us-east-1b to the Auto Scaling group so instances can launch in both AZs.
C.Replace the Application Load Balancer with a Network Load Balancer because it will automatically keep the app online.
D.Move the application to a larger EC2 instance type so a single server can handle the full workload.
AnswerB

Placing the Auto Scaling group in at least two Availability Zones allows AWS to distribute and replace instances across zones. Because the Application Load Balancer can route only to healthy targets, adding the second subnet is the lowest-complexity change that gives the application resilience to a full AZ outage.

Why this answer

It adds the unused subnet in us-east-1b to the Auto Scaling group, enabling EC2 instances to launch across two Availability Zones. This provides fault isolation: if one AZ fails, the ALB can route traffic to healthy instances in the other AZ. The change requires only a configuration update to the Auto Scaling group, minimizing operational overhead while meeting the high-availability requirement.

Exam trap

The trap here is that candidates often assume increasing instance count in a single AZ or using a different load balancer type alone provides high availability, but true resilience requires distributing instances across multiple Availability Zones.

How to eliminate wrong answers

Option A is wrong because increasing desired capacity to 3 in the same Availability Zone does not protect against an AZ failure; all instances remain in a single AZ, so if that AZ fails, all traffic is lost. Option C is wrong because replacing the Application Load Balancer with a Network Load Balancer does not inherently provide cross-AZ failover; the NLB still requires instances in multiple AZs to maintain availability, and the change introduces unnecessary operational overhead. Option D is wrong because moving to a larger EC2 instance type does not eliminate the single point of failure; if the AZ hosting that single instance fails, the application goes down regardless of instance size.

3
MCQmedium

You deploy a Web ACL with an AWS WAF rate-based rule intended to limit abusive traffic to your API. After the deployment, attackers still reach the backend service. ALB access logs show requests arrive at the ALB, but WAF logs indicate the Web ACL is not evaluating those requests. Which change most likely fixes the issue?

A.Associate the Web ACL with the Application Load Balancer resource ARN so WAF evaluates requests sent to that ALB.
B.Add a security group rule that drops inbound traffic from the attacker IP range at the instances' ENIs.
C.Create a target group stickiness policy so WAF can count requests consistently per client IP.
D.Enable AWS Shield Advanced but keep the Web ACL unattached because Shield automatically applies rate limiting.
AnswerA

A Web ACL only inspects traffic for resources it is explicitly associated with. Without association to the ALB's ARN, WAF never evaluates incoming requests, so the rate-based rule cannot block attackers even though requests reach the load balancer.

Why this answer

A Web ACL must be explicitly associated with a resource (such as an ALB) for AWS WAF to evaluate incoming requests. In this scenario, the Web ACL was deployed but not associated with the ALB resource ARN, so WAF never inspected the traffic. Associating the Web ACL with the ALB ensures that all requests to the ALB are evaluated by the rate-based rule before reaching the backend.

Exam trap

The trap here is that candidates assume deploying a Web ACL automatically applies it to all resources in the account, when in fact it must be explicitly associated with each resource ARN to take effect.

Why the other options are wrong

B

The issue is that the Web ACL is not evaluating requests at all, which indicates a missing association between the Web ACL and the ALB. Adding a security group rule to drop traffic from attacker IPs does not address the root cause—the Web ACL is not in the evaluation path—and security groups operate at the instance level, not at the ALB level for WAF inspection.

C

Stickiness (session affinity) ensures requests from the same client are sent to the same target, but it does not cause WAF to evaluate requests. The issue is that the Web ACL is not associated with the ALB, so WAF never inspects traffic regardless of stickiness.

D

AWS Shield Advanced does not automatically apply rate limiting; it provides DDoS protection but does not replace the need to associate a Web ACL for WAF rate-based rules. The Web ACL must be explicitly associated with a resource like an ALB to evaluate requests.

When would these options actually be correct?

B

This option would be correct in a scenario where the Web ACL is already properly associated and evaluating traffic, but attackers are still reaching the backend because the WAF rate-based rule is not effectively blocking them. In that case, adding a security group rule to drop traffic from known attacker IPs at the instance ENIs would provide an additional layer of defense.

C

A question where clients report intermittent failures or inconsistent behavior from a backend that maintains state, and the solution must ensure all requests from a client go to the same target. For example: 'Users are randomly logged out when using a stateful web application behind an ALB. Which configuration ensures session persistence?'

D

If the question described a scenario where the goal is to protect against large-scale DDoS attacks that overwhelm infrastructure, and the requirement is to get enhanced DDoS mitigation and cost protection, then enabling AWS Shield Advanced would be correct, even without a Web ACL for rate limiting.

Why candidates pick the wrong answer

B

Candidates may think that blocking attacker IPs at the instance level is a quick fix to stop abusive traffic, overlooking that the primary issue is the Web ACL not being associated with the ALB. They might also confuse the roles of security groups and WAF in traffic filtering.

C

Candidates may confuse rate-based rules (which count requests per IP) with stickiness, thinking that binding a client to one target helps WAF count requests accurately. However, WAF counts at the ALB level, not per target.

D

Candidates may think Shield Advanced includes all WAF capabilities automatically, or that it can substitute for a Web ACL, due to its name implying comprehensive protection.

4
Multi-Selectmedium

A startup runs a 24/7 web tier on Amazon EC2 with a stable baseline of 8 instances and a nightly analytics batch job that can resume from checkpoints if interrupted. The company wants to minimize monthly compute cost without hurting the always-on web tier. Which two actions should it take? Select two.

Select 2 answers
A.Buy a Compute Savings Plan for the steady web tier baseline.
B.Buy Standard Reserved Instances only for the nightly analytics batch job.
C.Run the batch job on Spot Instances and checkpoint progress frequently.
D.Move the entire workload to On-Demand Instances for maximum flexibility.
E.Use Dedicated Hosts for the batch job so the fleet is isolated.
AnswersA, C

A Compute Savings Plan reduces cost for the predictable baseline while preserving flexibility across instance families and Regions. That fits a 24/7 web tier that is expected to run continuously. It is cheaper than On-Demand for the committed portion and avoids overcommitting to a specific instance family.

Why this answer

A Compute Savings Plan offers the largest discount (up to 66%) in exchange for a 1- or 3-year hourly spend commitment, and it automatically applies to any EC2 instance family, size, or region. For the stable 8-instance web tier that runs 24/7, this plan provides significant cost savings while maintaining full flexibility to change instance types or even move to containers or Lambda, without affecting the always-on requirement.

Exam trap

The trap here is that candidates often assume Reserved Instances are always the best choice for any steady workload, but for a part-time batch job, a Savings Plan or Spot is more cost-effective, and they may overlook that Spot Instances with checkpointing are ideal for fault-tolerant, interruptible workloads.

5
MCQmedium

A read-heavy document portal repeatedly queries the same product catalogue data from DynamoDB with millisecond latency requirements. Which service can reduce read latency and table load? The architecture review board prefers a managed AWS-native control.

A.Amazon Kinesis Data Firehose
B.S3 Transfer Acceleration
C.DynamoDB Accelerator (DAX)
D.AWS Glue Data Catalog
AnswerC

DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of DynamoDB, providing microsecond latency for repeated reads while maintaining eventual consistency by default. It automatically intercepts read and write operations, populates the cache on read misses, and supports TTL-based expiration, making it ideal for read-heavy workloads with repetitive access patterns. Because the portal repeatedly queries the same product data, DAX directly addresses the latency issue without changing application code beyond the DAX endpoint.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that delivers up to 10x read performance improvement, reducing read latency to microseconds for repeated queries. It offloads read traffic from the DynamoDB table, lowering consumed read capacity units and table load, making it ideal for read-heavy workloads with millisecond latency requirements. As a fully managed, AWS-native service, DAX aligns with the architecture review board's preference for managed controls.

Exam trap

The trap here is that candidates may confuse DAX with ElastiCache (which is also a caching service but not DynamoDB-native) or assume that any AWS caching service works interchangeably, but DAX is the only managed, DynamoDB-specific cache that integrates directly with the DynamoDB API without application code changes.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose is a streaming data ingestion service for loading data into data stores and analytics tools, not a caching layer for DynamoDB reads; it cannot reduce read latency or table load for repeated queries. Option B is wrong because S3 Transfer Acceleration speeds up uploads and downloads to/from S3 over long distances using AWS edge locations, but it does not cache DynamoDB data or reduce read latency for DynamoDB queries. Option D is wrong because AWS Glue Data Catalog is a metadata repository for ETL jobs and data lake schemas, not a caching service for DynamoDB reads; it has no impact on DynamoDB read latency or table load.

6
MCQmedium

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.

A.A single-AZ Aurora cluster
B.Aurora Global Database
C.Manual snapshots copied monthly
D.An ElastiCache Redis replica
AnswerB

Aurora Global Database replicates with low latency to secondary Regions and supports faster disaster recovery than snapshot-only approaches.

Why this answer

Aurora Global Database is the correct choice because it provides a managed, cross-Region disaster recovery solution with a Recovery Point Objective (RPO) of less than 1 second and a Recovery Time Objective (RTO) of typically less than 1 minute. It uses storage-based replication to keep a secondary cluster in another AWS Region up to date with minimal latency, meeting the low RPO requirement without manual intervention.

Exam trap

The trap here is that candidates may confuse cross-Region replication with multi-AZ deployments, or assume that manual snapshots or caching solutions can meet low RPO requirements, when only a managed global database service like Aurora Global Database provides the necessary sub-second RPO and automated failover.

How to eliminate wrong answers

Option A is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, offering no disaster recovery across Regions and resulting in an unacceptably high RPO if the primary Region fails. Option C is wrong because manual snapshots copied monthly provide an RPO of up to one month, which is far too high for the low RPO requirement, and the process is not automated or managed natively for rapid recovery. Option D is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a primary data store for the trading dashboard's transactional data; it also lacks cross-Region replication for disaster recovery.

7
MCQmedium

Company A stores encrypted log files in its S3 bucket using SSE-KMS with a customer-managed KMS key. A partner application in Company B uploads objects into Company A's bucket using an IAM role in Company B. Uploads fail with an error indicating KMS access is denied (kms:Encrypt not authorized). Neither the partner IAM policy nor the S3 bucket policy currently mentions KMS. What is the most secure and correct change to allow cross-account uploads to succeed?

A.In Company A's KMS key policy, allow Company B's partner role principal to use the key for kms:Encrypt, kms:GenerateDataKey, and kms:DescribeKey, and also add a matching IAM policy in Company B that grants the partner role those same KMS actions on Company A's key ARN, constrained to the target S3 bucket context when possible.
B.In Company B's IAM policy, allow kms:Encrypt on Company A's KMS key ARN, without changing Company A's key policy.
C.Create a new KMS key in Company B and configure Company A's S3 bucket to use that key for SSE-KMS.
D.Disable key policy restrictions by setting the KMS key to enabled and removing all policy statements so that encryption automatically works for any principal.
AnswerA

Cross-account SSE-KMS requires both the KMS key policy in the key owner account and an IAM policy in the caller account to allow the required KMS actions. Scoping the permissions to the specific bucket or encryption context reduces blast radius.

Why this answer

For cross-account SSE-KMS uploads, the KMS key policy must explicitly grant the external IAM role principal the required KMS actions (kms:Encrypt, kms:GenerateDataKey, and kms:DescribeKey). Additionally, the partner account's IAM policy must also allow those same actions on the key ARN. This dual-permission model is required because KMS does not implicitly trust IAM policies in the key owner's account for cross-account access; the key policy is the authoritative gatekeeper.

Option A correctly implements both sides, and constraining to the target S3 bucket context (via kms:ViaService or kms:EncryptionContext) adds a security best practice.

Exam trap

The trap here is that candidates assume an IAM policy in the partner account alone is sufficient for cross-account KMS access, forgetting that KMS key policies are the definitive authorization mechanism for external principals.

How to eliminate wrong answers

Option B is wrong because KMS key policies are the primary access control for cross-account use; an IAM policy in Company B alone is insufficient without the key policy granting access to the external principal. Option C is wrong because using a KMS key from Company B would require Company A's S3 bucket to trust that key for SSE-KMS, which is not supported for cross-account uploads—the bucket must use its own key to decrypt. Option D is wrong because removing all policy statements from the KMS key disables all access control, making the key effectively unusable and insecure; KMS requires at least a default key policy to allow the root account, and removing it would break all encryption operations.

8
MCQmedium

A marketing site runs on x86 EC2 instances and uses open-source software with no architecture-specific licensing restriction. What should be evaluated to reduce compute cost?

A.Cross-Region data replication for all data
B.io2 Block Express volumes for all instances
C.AWS Graviton-based instances after performance testing
D.Dedicated Hosts by default
AnswerC

AWS Graviton-based instances after performance testing is the correct approach because Graviton processors (arm64 architecture) often deliver up to 20% better price-performance than comparable x86 instances for scale-out, stateless, and lighter-workload deployments like a marketing site running open-source software. The key is to first run performance testing to verify that your specific application stack—including any compiled binaries, libraries, or dependencies—is fully compatible and performs adequately on the arm64 architecture. If the workload is portable (e.g., running on Linux with open-source code), moving to Graviton instances reduces compute cost without sacrificing performance, making it a cost-effective modernization step.

Why this answer

AWS Graviton-based instances (ARM architecture) offer up to 40% better price-performance compared to comparable x86 instances for many workloads. Since the marketing site uses open-source software with no architecture-specific licensing restrictions, migrating to Graviton after performance testing can significantly reduce compute costs without sacrificing performance.

Exam trap

The trap here is that candidates assume Dedicated Hosts (Option D) always reduce costs due to 'dedicated' implying efficiency, but they actually increase costs unless specific licensing requirements (e.g., Windows Server or SQL Server) mandate physical isolation.

How to eliminate wrong answers

Option A is wrong because cross-Region data replication increases storage and data transfer costs, not reduces compute costs; it is a disaster recovery strategy, not a cost-optimization technique for compute. Option B is wrong because io2 Block Express volumes are high-performance, high-cost SSDs designed for latency-sensitive workloads, not for reducing compute costs; they would increase overall costs without addressing compute efficiency. Option D is wrong because Dedicated Hosts incur additional hourly charges for physical server isolation and are typically used for licensing or compliance requirements, not for cost reduction; they would increase compute costs compared to shared tenancy.

9
MCQeasy

An internal worker consumes messages from an Amazon SQS queue. Occasionally, a message fails validation in the worker (for example, missing required fields). Reprocessing the same bad message repeatedly wastes processing time and delays healthy messages. What is the best AWS approach to handle these poison messages without blocking the rest of the queue?

A.Configure an SQS dead-letter queue (DLQ) using a redrive policy with a maxReceiveCount.
B.Delete the SQS queue and recreate it daily to clear invalid messages.
C.Increase the consumer timeout/processing time so validation failures take longer to occur.
D.Use SNS fan-out without any DLQ and rely only on application retries.
AnswerA

With a redrive policy, SQS continues delivering the message to consumers until it has been received unsuccessfully maxReceiveCount times. After that threshold, SQS moves the poison message to a DLQ, isolating it from the main processing flow so healthy messages can continue being processed.

Why this answer

An SQS dead-letter queue (DLQ) with a redrive policy that sets a maxReceiveCount allows the worker to process a message up to a specified number of times. After that threshold is exceeded, the message is automatically moved to the DLQ, isolating the poison message and preventing it from blocking or delaying the processing of healthy messages in the main queue.

Exam trap

The trap here is that candidates may think increasing timeouts or relying on application retries alone can solve the problem, but they fail to recognize that only a DLQ with a redrive policy provides automatic, queue-level isolation of poison messages without blocking healthy message processing.

How to eliminate wrong answers

Option B is wrong because deleting and recreating the queue daily is disruptive, causes data loss of all messages (including valid ones), and does not provide a targeted mechanism to isolate only the poison messages. Option C is wrong because increasing the consumer timeout or processing time does not prevent validation failures; it only delays the retry cycle and does not remove the bad message from the queue, so it will still be reprocessed and waste resources. Option D is wrong because SNS fan-out without a DLQ and relying only on application retries means the poison message will be repeatedly delivered to all subscribers, causing infinite retries and blocking the processing of healthy messages; there is no automatic isolation mechanism.

10
MCQeasy

A media processing pipeline runs batch jobs overnight. The jobs are stateless, can be restarted from checkpoints, and can tolerate interruptions. The team wants to minimize compute cost. Which EC2 approach is the best fit?

A.Use On-Demand instances to guarantee uninterrupted capacity.
B.Use Spot Instances and design the jobs to handle interruptions by checkpointing and retrying.
C.Use a 1-year Reserved Instance for the current instance type and lock the fleet to it.
D.Use Savings Plans but still treat interruptions as failures that require manual intervention.
AnswerB

Spot is the most cost-effective EC2 option when the workload can handle interruption. Because the jobs are stateless and can resume from checkpoints, losing an instance due to a Spot interruption does not lose progress. The design aligns directly with Spot’s best-effort interruption model, minimizing compute cost while still completing the batch work.

Why this answer

Spot Instances offer significant cost savings (up to 90% off On-Demand) but can be reclaimed by AWS with a 2-minute interruption notice. Since the batch jobs are stateless, checkpointable, and interruption-tolerant, they are an ideal workload for Spot Instances. Designing the jobs to save progress to a durable checkpoint (e.g., Amazon S3) and automatically retry on interruption ensures resilience while minimizing compute cost.

Exam trap

The trap here is that candidates often assume all production workloads need On-Demand or Reserved Instances for reliability, failing to recognize that stateless, checkpointable batch jobs are the perfect use case for Spot Instances to drastically reduce costs.

How to eliminate wrong answers

Option A is wrong because On-Demand instances are the most expensive pricing model and provide no cost optimization benefit for workloads that can tolerate interruptions. Option C is wrong because a 1-year Reserved Instance locks the fleet to a specific instance type and commits to a full year of payment, which is inflexible and not cost-optimal for a batch workload that may vary in size or instance needs. Option D is wrong because Savings Plans provide a discount in exchange for a 1- or 3-year hourly spend commitment, but they do not inherently handle interruptions; treating interruptions as failures requiring manual intervention defeats the purpose of automation and increases operational overhead.

11
MCQeasy

A team runs a stateless web app on Amazon EC2 behind an Application Load Balancer. During traffic spikes, new EC2 instances take several minutes to finish bootstrapping before they can receive traffic. Which Auto Scaling configuration most directly reduces the time until additional capacity is available?

A.Increase the ALB target group deregistration delay.
B.Use an Auto Scaling warm pool so pre-initialized instances are ready to enter service.
C.Reduce the Auto Scaling group minimum size to one instance.
D.Replace the Application Load Balancer with a Network Load Balancer.
AnswerB

A warm pool keeps pre-initialised instances in a stopped or running state, so scaling out moves already-bootstrapped capacity into service instead of waiting several minutes for bootstrap, directly cutting the delay before new instances receive traffic.

Why this answer

An Auto Scaling warm pool allows you to maintain a pool of pre-initialized instances that are ready to quickly enter the target group and start serving traffic. Instead of waiting for new instances to boot and configure during a scale-out event, the warm pool provides instances that have already completed bootstrapping, drastically reducing the time to additional capacity.

Exam trap

The trap here is that candidates may confuse the deregistration delay (which handles graceful connection draining) with a mechanism to speed up instance readiness, or they may incorrectly assume that reducing the minimum size or switching to a Network Load Balancer will improve scaling speed, when neither addresses the root cause of slow bootstrapping.

Why the other options are wrong

A

Increasing the deregistration delay only keeps existing connections alive longer; it does not speed up the bootstrapping of new instances, so it does not reduce the time until additional capacity is available.

C

Reducing the minimum size to one instance does not address the bootstrapping delay; it only lowers the baseline capacity, potentially worsening performance during traffic spikes.

D

Replacing the ALB with a Network Load Balancer does not address the bootstrapping delay of EC2 instances; NLB operates at layer 4 and does not affect instance initialization time.

When would these options actually be correct?

A

This option would be correct in a scenario where the question asks how to prevent in-flight requests from being dropped during a scale-in event, such as when instances are being terminated and you need to ensure graceful connection draining.

C

If the question were about minimizing costs for a predictable, low-traffic application where over-provisioning is unnecessary, reducing the minimum size to one instance would be correct.

D

In a scenario where the application requires ultra-low latency and high throughput for TCP/UDP traffic, and the bootstrapping delay is not a concern, replacing an ALB with an NLB would be correct to reduce latency and handle millions of requests per second.

Why candidates pick the wrong answer

A

Candidates may confuse the deregistration delay with a mechanism that helps new instances become ready faster, or they might think it gives more time for bootstrapping to complete before traffic is sent.

C

Candidates may think that a smaller minimum size forces faster scaling, but it actually reduces the buffer of running instances, increasing the impact of bootstrapping delays.

D

Candidates may think that a faster load balancer (NLB) will reduce the time until new instances can receive traffic, overlooking that the bottleneck is instance bootstrapping, not load balancer performance.

12
MCQhard

Based on the exhibit, your application runs entirely in private subnets and only needs to reach Amazon S3, Amazon DynamoDB, AWS Secrets Manager, and CloudWatch Logs. The monthly bill is dominated by NAT Gateway charges. Which change most directly reduces cost while preserving private connectivity to these AWS services?

A.Replace the NAT Gateway with an Internet Gateway and keep the current private subnet routes unchanged.
B.Add a second NAT Gateway in another Availability Zone to reduce cross-AZ data transfer charges.
C.Create only interface endpoints for all four services and keep the NAT Gateway for fallback.
D.Create S3 and DynamoDB gateway endpoints, create interface endpoints for Secrets Manager and CloudWatch Logs, update route tables, and remove the NAT Gateway.
AnswerD

S3 and DynamoDB use gateway endpoints, which are the cost-effective private path for those services. Secrets Manager and CloudWatch Logs require interface endpoints for private access. Once these are in place, the NAT Gateway is no longer needed for this workload, eliminating the hourly and per-GB NAT charges while keeping traffic on the AWS network.

Why this answer

It replaces the costly NAT Gateway with free VPC Gateway Endpoints for S3 and DynamoDB, and uses AWS PrivateLink interface endpoints for Secrets Manager and CloudWatch Logs. This eliminates all internet-bound data transfer costs while keeping traffic entirely within the AWS network, directly addressing the cost concern without sacrificing private connectivity.

Exam trap

The trap here is that candidates assume all AWS services require the same type of VPC endpoint, leading them to either use only interface endpoints (costly) or keep the NAT Gateway as a safety net, missing the opportunity to use free gateway endpoints for S3 and DynamoDB.

How to eliminate wrong answers

Option A is wrong because an Internet Gateway alone does not enable private subnets to reach AWS services; private subnets still need a NAT device to route traffic through the Internet Gateway, so removing the NAT Gateway without adding endpoints would break connectivity. Option B is wrong because adding a second NAT Gateway increases costs (additional hourly charges and cross-AZ data transfer fees) rather than reducing them, and the current bill is dominated by NAT Gateway charges, not cross-AZ traffic. Option C is wrong because using interface endpoints for all four services would incur per-hour and per-GB data processing charges for S3 and DynamoDB, which are more expensive than using free gateway endpoints for those two services; keeping the NAT Gateway as a fallback also retains unnecessary costs.

13
MCQeasy

A production application uses an Amazon RDS Multi-AZ DB instance. During an unplanned failover, the database endpoint remains the same. What change should the application team make to handle the failover reliably?

A.Hard-code the new writer instance IP address after failover completes.
B.Keep using the same RDS endpoint and implement connection retry logic on failures.
C.Disable Multi-AZ and rely on manual intervention to switch endpoints.
D.Move reads to application-side caching only, and avoid reopening DB connections.
AnswerB

The RDS endpoint is DNS-based and remains constant across a Multi-AZ failover, so clients should continue using that same hostname. When failover occurs, existing active connections are dropped and in-flight transactions may fail; the application must treat those errors as transient and reconnect using retry logic with backoff or jitter. This allows the app to automatically resume writing to the new primary once DNS and the instance are ready. Do not assume an individual request succeeded; ensure retries are idempotent.

Why this answer

The RDS Multi-AZ DNS endpoint remains unchanged during a failover, automatically pointing to the new writer instance. Implementing connection retry logic with exponential backoff allows the application to handle the brief DNS propagation delay and connection interruption, ensuring reliable recovery without manual intervention.

Exam trap

The trap here is that candidates assume the endpoint changes or that Multi-AZ provides seamless failover without any application-side changes, but in reality the application must implement retry logic to handle the brief connection disruption during DNS propagation.

How to eliminate wrong answers

Option A is wrong because hard-coding the new writer instance IP address is impractical and error-prone; the IP address can change after failover, and this approach bypasses the automatic DNS update provided by Multi-AZ. Option C is wrong because disabling Multi-AZ removes high availability entirely, forcing manual endpoint switching which increases downtime and violates the goal of reliable failover handling. Option D is wrong because moving reads to application-side caching does not address the need to re-establish the database connection after failover; the application must still handle connection failures and retries for writes.

14
Multi-Selectmedium

A solutions architect is designing a cost-optimized data storage solution for a large dataset that is accessed infrequently but must be retained for compliance for 7 years. Which three actions should the architect take to minimize costs? (Choose three.)

Select 3 answers
.Store the data in Amazon S3 Glacier Deep Archive immediately after creation.
.Use Amazon S3 lifecycle policies to transition data from S3 Standard to S3 Glacier Deep Archive after 30 days.
.Enable S3 Intelligent-Tiering to automatically move data between access tiers based on usage patterns.
.Store all data in Amazon EBS gp2 volumes attached to an EC2 instance for low-latency access.
.Use S3 Object Lock in compliance mode to prevent data deletion during the retention period.
.Replicate all data to a second AWS Region using S3 Cross-Region Replication to ensure durability.

Why this answer

Amazon S3 lifecycle policies allow you to define rules that automatically transition objects to colder storage tiers like S3 Glacier Deep Archive after a specified period. This approach minimizes costs by keeping data in S3 Standard only for the initial 30 days when it might be accessed, then moving it to the lowest-cost storage class for the remaining compliance period. S3 Intelligent-Tiering automatically optimizes costs by monitoring access patterns and moving data between frequent, infrequent, and archive access tiers without manual intervention.

S3 Object Lock in compliance mode prevents any user, including the root user, from deleting or overwriting objects during the retention period, ensuring regulatory compliance.

Exam trap

The trap here is that candidates may think immediate archiving to Glacier Deep Archive is the cheapest option, but they overlook the need for lifecycle policies to balance initial access needs with long-term cost savings, and they may confuse durability (which S3 already provides) with compliance retention, leading them to select unnecessary replication.

15
MCQhard

A IoT ingestion API must ensure that only encrypted EBS volumes can be created in the account. What is the strongest preventive control?

A.Use an SCP that denies ec2:CreateVolume when the encrypted condition is false
B.Run a daily Lambda function to encrypt unencrypted volumes
C.Enable VPC Flow Logs
D.Tag encrypted volumes after creation
AnswerA

A service control policy (SCP) is the correct preventive control because it operates at the AWS Organizations root, OU, or account level and can block the ec2:CreateVolume API call when the ec2:Encrypted condition key evaluates to false. This means unauthorized or unencrypted volume creation is denied before the resource exists, across all principals in the account, regardless of their IAM permissions. SCPs do not modify resources, but they enforce policy at request time, making them an effective guardrail for encryption compliance.

Why this answer

An SCP (Service Control Policy) is the strongest preventive control because it can deny the ec2:CreateVolume API call when the encrypted condition is false, effectively blocking the creation of any unencrypted EBS volume at the account level before it happens. This is a preventive control that enforces encryption as a mandatory requirement, unlike detective or corrective measures that act after the fact.

Exam trap

The trap here is confusing preventive controls (like SCPs that block the action) with detective or corrective controls (like Lambda scripts or tagging), leading candidates to choose a reactive solution instead of the strongest preventive one.

How to eliminate wrong answers

Option B is wrong because running a daily Lambda function to encrypt unencrypted volumes is a corrective/reactive control, not a preventive one; it only fixes volumes after they have already been created unencrypted, leaving a window of non-compliance. Option C is wrong because VPC Flow Logs are a detective control that captures network traffic metadata, not a mechanism to enforce or prevent the creation of encrypted EBS volumes. Option D is wrong because tagging encrypted volumes after creation is a labeling action that provides visibility but does not prevent the creation of unencrypted volumes in the first place.

16
MCQhard

A financial services firm stores trade confirmations in an Amazon S3 bucket. Regulations require that every object be encrypted at rest with a key the firm controls and can audit independently of AWS, and that key usage be logged. The firm wants to avoid changing application code. Which encryption approach should be used?

A.SSE-S3 with bucket default encryption enabled.
B.SSE-C with the customer providing an encryption key on every PUT and GET request.
C.Client-side encryption using a key stored in an application configuration file on each server.
D.SSE-KMS with a customer managed key in AWS KMS, with CloudTrail logging of KMS API calls.
AnswerD

SSE-KMS with a customer managed key gives the firm control over the key policy, rotation, and revocation, and KMS API calls are recorded in CloudTrail for independent auditing. S3 applies the encryption server-side when objects are written, so no application code change is required. This satisfies encryption at rest, customer-controlled keys, and auditable key usage simultaneously.

Why this answer

SSE-KMS with a customer managed key provides server-side encryption that requires no application changes while giving the firm ownership of the key policy, rotation, and access. KMS integrates with CloudTrail so every use of the key can be audited independently. SSE-S3, SSE-C, and client-side encryption each fall short on either customer control, auditability, or the no-code-change constraint.

Exam trap

The trap here is equating encryption at rest with customer-controlled keys, when only a customer managed KMS key provides independent control and auditability.

17
MCQmedium

A analytics dashboard uses RDS MySQL and receives many read-only reporting queries that slow down the primary database. What should the architect add?

A.S3 lifecycle policy
B.RDS read replica and route reporting queries to it
C.Multi-AZ standby and route reads to the standby
D.A larger NAT gateway
AnswerB

An RDS read replica is a separate MySQL instance that uses asynchronous replication from the primary database. Routing reporting queries to the replica's endpoint shifts the read IOPS and CPU burden away from the primary, which continues to handle writes and can scale to support higher read throughput. This is the standard AWS pattern for offloading read-only queries in a MySQL environment, though it requires updating application data-source configuration to direct those queries appropriately.

Why this answer

Adding an RDS read replica offloads read-heavy reporting queries from the primary MySQL instance, preserving write performance. The read replica asynchronously replicates data using MySQL's native binlog replication, and routing reporting queries to its endpoint reduces contention on the primary.

Exam trap

The trap here is confusing a Multi-AZ standby (which is for high availability only and cannot serve reads) with a read replica (which is explicitly designed to offload read traffic).

How to eliminate wrong answers

Option A is wrong because S3 lifecycle policies manage object transitions and expirations in S3, not database query offloading. Option C is wrong because a Multi-AZ standby is a synchronous replica used only for failover; it does not serve read traffic (RDS does not allow direct reads from the standby). Option D is wrong because a larger NAT gateway increases outbound internet bandwidth for private subnets, which does not address database read query performance.

18
MCQmedium

A team serves static web assets (JS, CSS, images) from an Amazon S3 origin through CloudFront. Recently, the S3 origin has received a high number of requests for the same files, increasing origin data transfer costs. CloudFront access logs show many cache misses, and each request includes a unique query string used only for tracking (for example, ?utm=...). The application does not require query-string-specific content. What CloudFront change will most directly reduce origin fetches and cost?

A.Update the CloudFront cache policy to exclude query strings from the cache key so that requests differing only by tracking query parameters reuse the same cached object.
B.Lower the minimum TTL and set Cache-Control headers to no-store to force CloudFront to revalidate more often.
C.Enable Origin Shield to ensure all origin fetches go through a single regional shield with no other configuration changes.
D.Switch the S3 origin from S3 to a different storage class optimized for request rates, keeping the cache key the same.
AnswerA

CloudFront cache misses increase when the cache key includes values that vary per request. If the tracking query string is part of the cache key, each unique ?utm value generates a separate cache entry even though the underlying object (JS/CSS/image) is identical, causing repeated origin fetches. Excluding query strings from the cache key collapses those variations into a single cached object, increasing the cache hit rate and reducing origin fetches and origin data transfer.

Why this answer

CloudFront's cache policy controls which parts of a request (including query strings) are included in the cache key. By excluding the tracking query strings (e.g., `?utm=...`) from the cache key, CloudFront will treat all requests for the same file as identical, serving the cached object regardless of the query string. This directly reduces the number of origin fetches (cache misses) and lowers S3 data transfer costs, as the application does not require query-string-specific content.

Exam trap

The trap here is that candidates may think enabling Origin Shield (Option C) or changing storage classes (Option D) will solve the problem, but they overlook the fundamental issue of cache key fragmentation caused by unique query strings, which is directly addressed by adjusting the cache policy.

How to eliminate wrong answers

Option B is wrong because lowering the minimum TTL and setting `Cache-Control: no-store` would force CloudFront to revalidate or bypass the cache entirely, increasing origin fetches and costs, not reducing them. Option C is wrong because enabling Origin Shield alone, without adjusting the cache key to exclude query strings, does not address the root cause of cache misses caused by unique query strings; Origin Shield would still forward each unique query string request to the origin. Option D is wrong because switching the S3 storage class (e.g., to S3 Standard-IA or One Zone-IA) does not change the cache key behavior; the high number of unique query strings would still cause cache misses and origin fetches, and some storage classes may even incur higher per-request costs.

19
Multi-Selecthard

A image sharing application uses CloudFront in front of an S3 origin. Which two settings help keep users from bypassing CloudFront and accessing the bucket directly?

Select 2 answers
A.Enable CloudFront standard logging
B.Enable S3 static website hosting
C.Configure Origin Access Control for the S3 origin
D.Use an S3 bucket policy that allows access only from the CloudFront distribution
AnswersC, D

Origin Access Control signs CloudFront requests to the S3 origin, letting you remove public read access from the bucket policy so only the distribution's identity is granted `s3:GetObject`. This directly satisfies the requirement that users cannot bypass CloudFront, since unsigned direct requests to the bucket are denied.

Why this answer

Option C is correct because Origin Access Control (OAC) is the mechanism that lets CloudFront sign requests to the S3 origin using SigV4, so the bucket can reject any request that does not come through the distribution. Option D is correct because an S3 bucket policy that grants access only to the CloudFront distribution's service principal (or to the distribution ARN via the OAC) enforces at the bucket level that direct requests from users are denied. Together, OAC plus a restrictive bucket policy prevent users from bypassing CloudFront and hitting the S3 bucket directly.

Option A is not correct because CloudFront standard logging only records requests for auditing and does not block direct access to the origin. Option B is not correct because enabling S3 static website hosting actually makes the bucket more directly reachable via the website endpoint and does not restrict bypassing CloudFront.

Exam trap

The trap here is that candidates often confuse enabling S3 static website hosting (which creates a public endpoint) with a security control, when in fact it would undermine the goal of restricting direct access.

20
MCQeasy

Based on the exhibit, a web application must stay available if one Availability Zone fails. What is the best change to improve resilience?

A.Increase the desired capacity to 8 instances in the same subnet.
B.Add a subnet in another Availability Zone to the Auto Scaling group and keep the ALB spanning both AZs.
C.Replace the Application Load Balancer with a Network Load Balancer.
D.Move the instances to a larger instance type with more CPU and memory.
AnswerB

This places application instances across multiple Availability Zones, which protects the stateless tier from a single-AZ failure. The ALB already spans two AZs, so the missing piece is the Auto Scaling group using subnets in more than one AZ. That allows AWS to replace unhealthy instances and continue serving traffic from the surviving Zone.

Why this answer

Adding a subnet in another Availability Zone (AZ) to the Auto Scaling group and keeping the ALB spanning both AZs ensures that if one AZ fails, the ALB can route traffic to healthy instances in the other AZ. This is the standard pattern for building multi-AZ resilient architectures with Auto Scaling and ALB, as it eliminates the single point of failure at the AZ level.

Exam trap

The trap here is that candidates often think increasing instance count or size improves resilience, but without multi-AZ distribution, all instances remain vulnerable to a single AZ failure.

Why the other options are wrong

A

Increasing desired capacity within the same subnet does not protect against an Availability Zone failure; all instances remain in a single AZ, so if that AZ fails, the application becomes unavailable.

C

Replacing the ALB with a Network Load Balancer does not improve resilience across Availability Zones; it operates at a different layer and does not inherently provide cross-AZ fault tolerance for the application.

D

Increasing instance size (CPU/memory) does not address the requirement for availability zone failure resilience; it only improves performance, not fault tolerance across AZs.

When would these options actually be correct?

A

This option would be correct if the question asked for improving performance or handling increased load within a single AZ, without any requirement for multi-AZ resilience.

C

This would be correct if the question required handling millions of requests per second with ultra-low latency, or if the application needed to preserve the source IP address for backend processing, as NLB operates at Layer 4.

D

This option would be correct if the question asked for improving application performance under high load, such as 'A web application is experiencing high CPU utilization and slow response times. What change would best improve performance?'

Why candidates pick the wrong answer

A

Candidates may think that simply adding more instances improves resilience, not realizing that resilience requires distributing instances across multiple Availability Zones.

C

Candidates may think NLB is more resilient because it is simpler and faster, but they overlook that ALB already supports cross-AZ load balancing and is better suited for HTTP/HTTPS applications.

D

Candidates may think that more powerful instances inherently make the application more resilient, confusing performance improvement with high availability.

21
MCQmedium

An order-processing service consumes messages from an Amazon SQS Standard queue using a custom worker. During traffic spikes, the worker occasionally times out after performing some work but before acknowledging the message, so SQS redelivers it and it may be processed again. You also observe that a small set of “poison” messages always fail validation. What change most directly improves resilience by (1) preventing poison messages from retrying indefinitely and (2) avoiding duplicate side effects caused by legitimate retries?

A.Increase the SQS visibility timeout and, when validation fails, call DeleteMessage in the consumer to remove the message immediately.
B.Move to SNS topics with subscriptions and rely on SNS to provide exactly-once delivery to eliminate duplicates automatically.
C.Configure a dead-letter queue (DLQ) with a redrive policy that moves messages after maxReceiveCount, and implement idempotent processing in the consumer using an idempotency key.
D.Change the queue to FIFO and enable content-based deduplication, leaving the consumer logic unchanged.
AnswerC

SQS Standard is at-least-once delivery, so timeouts can cause redelivery and duplicates. A DLQ with a redrive policy prevents poison messages from retrying forever by moving them after repeated failures. Idempotent processing (for example, storing a processed marker in a database with conditional logic keyed by an idempotency key) prevents duplicate side effects when retries occur for valid messages.

Why this answer

A dead-letter queue (DLQ) with a maxReceiveCount redrive policy directly addresses the poison message problem by moving messages that repeatedly fail validation out of the main queue after a set number of retries, preventing indefinite retries. Implementing idempotent processing using an idempotency key ensures that even if a legitimate message is redelivered due to a visibility timeout, the consumer can detect and skip duplicate side effects, thus solving both requirements most directly.

Exam trap

The trap here is that candidates often confuse FIFO queues as a universal solution for both deduplication and poison message handling, but FIFO only provides exactly-once processing within a deduplication window and does not automatically handle poison messages without a DLQ, nor does it address idempotency for retries outside that window.

Why the other options are wrong

A

Increasing visibility timeout does not prevent poison messages from retrying indefinitely; they would still be redelivered until deleted manually. Also, deleting on validation failure only removes poison messages but does not address duplicate side effects from legitimate retries, as the worker may still process the same message multiple times before the timeout expires.

B

SNS does not provide exactly-once delivery; it delivers messages at least once, so duplicates can still occur. Additionally, SNS does not handle poison messages or retries, so it fails to address both requirements.

D

FIFO queues guarantee exactly-once processing but do not prevent duplicate side effects from legitimate retries (e.g., after timeout) because the same message can be redelivered with a different deduplication ID. Also, poison messages would still retry indefinitely unless a DLQ is configured, which is not mentioned.

When would these options actually be correct?

A

This option would be correct if the question asked: 'How to ensure a message is not processed by another consumer while a worker is still handling it, and how to remove invalid messages immediately?' In that scenario, increasing visibility timeout prevents premature redelivery, and deleting on validation failure removes poison messages without further retries.

B

This option would be correct if the question required decoupling message publication from processing, fanning out messages to multiple subscribers, and did not require poison message handling or duplicate prevention.

D

A question where the requirement is to eliminate duplicate messages entirely and poison messages are not a concern, or where the consumer is idempotent and the main issue is message ordering and deduplication at the queue level.

Why candidates pick the wrong answer

A

Candidates may think that increasing visibility timeout gives enough time to acknowledge, and deleting poison messages on validation failure seems like a direct fix. They overlook that poison messages would still be retried until the visibility timeout expires, and that legitimate retries due to timeouts are not handled idempotently.

B

Candidates may mistakenly believe SNS offers exactly-once delivery or confuse SNS's fan-out capability with SQS's message handling, overlooking that SNS alone cannot manage retries or poison messages.

D

Candidates may think FIFO queues provide exactly-once delivery and thus solve both problems, overlooking that redeliveries due to timeouts can still cause duplicates and that poison messages need explicit handling via DLQ.

22
MCQeasy

Based on the exhibit, what change best reduces Lambda cold-start impact for a predictable user-upload workflow?

A.Set a reserved concurrency limit for the function to protect it from throttling.
B.Enable provisioned concurrency for the function.
C.Increase the function timeout to give more time for initialization.
D.Move the function to a larger memory setting only to eliminate all initialization time.
AnswerB

Provisioned concurrency keeps a pre-initialized pool of Lambda execution environments ready to respond immediately. The exhibit shows long init duration after inactivity, which is the classic symptom of cold starts affecting user experience. Because the traffic pattern is predictable during launches, provisioned concurrency is the most direct way to reduce startup latency and smooth response times.

Why this answer

Provisioned concurrency pre-warms a specified number of execution environments so that when a user upload triggers the Lambda function, there is no cold-start latency. This is the most direct way to eliminate initialization time for a predictable workload, as it keeps instances ready to handle requests immediately.

Exam trap

The trap here is that candidates often confuse reserved concurrency (which limits concurrency) with provisioned concurrency (which pre-warms instances), or they assume that increasing memory or timeout will solve cold starts, when in fact only provisioned concurrency directly addresses initialization latency for predictable workloads.

How to eliminate wrong answers

Option A is wrong because reserved concurrency only caps the maximum number of concurrent executions to prevent throttling; it does not pre-warm instances or reduce cold-start impact. Option C is wrong because increasing the function timeout does not affect initialization time; it only extends the maximum duration a function can run, which does not address cold starts. Option D is wrong because moving to a larger memory setting can reduce initialization time by providing more CPU and resources, but it does not eliminate all initialization time, and it is not as targeted or effective as provisioned concurrency for predictable workloads.

23
MCQhard

An EC2 instance in a private subnet must access an S3 bucket that contains regulated exports for a customer analytics portal. The security team requires access to be allowed only when traffic comes through a specific VPC endpoint. What should the architect add to the bucket policy? The design must avoid adding custom operational scripts.

A.A security group rule that allows HTTPS to S3
B.A condition that matches aws:RequestedRegion to the bucket Region
C.A deny statement for all IAM users except the EC2 role
D.A condition that matches aws:sourceVpce to the endpoint ID
AnswerD

The aws:sourceVpce condition key restricts bucket access to requests originating through the named VPC endpoint, satisfying the requirement that traffic arrive only via that endpoint. It is a native policy condition, so no custom operational scripts are needed.

Why this answer

The bucket policy can use the `aws:sourceVpce` condition key to restrict access exclusively to traffic originating from a specific VPC endpoint ID. This ensures that only requests sent through that VPC endpoint are allowed, meeting the security team's requirement without requiring custom scripts or additional infrastructure.

Exam trap

The trap here is that candidates may confuse security group rules with bucket policies, or assume that restricting by IAM user or region is sufficient to enforce network-level control, when in fact only the `aws:sourceVpce` condition key directly ties access to a specific VPC endpoint.

How to eliminate wrong answers

Option A is wrong because security group rules operate at the network interface level and cannot be attached to an S3 bucket; S3 bucket policies are resource-based policies that do not support security group references. Option B is wrong because `aws:RequestedRegion` restricts the AWS Region in which the request is made, not the network path or VPC endpoint used, so it does not enforce that traffic comes through a specific VPC endpoint. Option C is wrong because denying all IAM users except the EC2 role would not restrict traffic to a specific VPC endpoint; it only controls which IAM identities can access the bucket, not the network path, and could break legitimate access from other services or users.

24
Multi-Selectmedium

A company is designing a high-performance database architecture for an e-commerce platform that experiences rapid spikes in read traffic during flash sales. The database must handle millions of reads per second with sub-millisecond latency. The data is key-value in nature, with a small number of attributes per item. Which three options should be included in the architecture? (Choose three.)

Select 3 answers
.Amazon DynamoDB as the primary database.
.Amazon RDS for MySQL with Multi-AZ and Read Replicas.
.DynamoDB Accelerator (DAX) as an in-memory cache.
.Amazon ElastiCache for Redis with cluster mode enabled.
.Amazon S3 as a primary data store accessed via Select and Range queries.
.Amazon Redshift with auto-scaling for real-time reads.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value database that delivers single-digit millisecond latency at any scale, making it ideal for high-traffic e-commerce platforms with key-value data. DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of DynamoDB, reducing read latency to microseconds for millions of reads per second. Amazon ElastiCache for Redis with cluster mode enabled provides a distributed in-memory cache that can offload read traffic from the primary database, further reducing latency and handling spikes during flash sales.

Exam trap

The trap here is that candidates often choose Amazon RDS with Read Replicas for read scaling, but they fail to recognize that relational databases cannot achieve sub-millisecond latency for millions of reads per second, and that DynamoDB with caching layers is the correct high-performance key-value solution.

25
MCQeasy

Account A hosts an IAM role (RoleInAccountA). The trust policy in Account A correctly allows a specific principal from Account B to call sts:AssumeRole. However, when Account B’s application calls sts:AssumeRole, it receives an AccessDenied error. What is the most likely missing requirement in Account B?

A.Account B’s calling principal must have an identity-based policy that allows sts:AssumeRole on RoleInAccountA’s role ARN.
B.Account A must attach an S3 bucket policy statement to allow sts:AssumeRole from Account B.
C.Account B must add kms:Decrypt permissions to the caller to satisfy AssumeRole.
D.Account B must create an SCP in the organization to allow sts:AssumeRole.
AnswerA

For a cross-account role assumption, the trust policy on RoleInAccountA is necessary but not sufficient: it lists Account B as a trusted principal, but the individual user or role in Account B must also have an attached IAM policy granting sts:AssumeRole with a Resource that includes RoleInAccountA's ARN. Without that identity-based permission, the caller has no authorization to invoke the STS API, even though the trust side is satisfied. This dual-sided authorization is the standard way to securely delegate access across accounts.

Why this answer

For an IAM role in Account A to be assumed by a principal in Account B, two conditions must be met: (1) the trust policy of the role in Account A must grant the sts:AssumeRole permission to the Account B principal, and (2) the calling principal in Account B must have an identity-based policy that explicitly allows sts:AssumeRole on the ARN of RoleInAccountA. Without this identity-based policy in Account B, the request is denied by AWS's explicit deny default, even if the trust policy in Account A is correctly configured.

Exam trap

The trap here is that candidates often assume the trust policy alone is sufficient for cross-account role assumption, forgetting that the calling principal must also have an explicit identity-based policy granting sts:AssumeRole on the target role ARN.

How to eliminate wrong answers

Option B is wrong because S3 bucket policies are used to control access to S3 resources, not to authorize sts:AssumeRole calls; sts:AssumeRole is governed by IAM policies and trust policies, not S3 bucket policies. Option C is wrong because kms:Decrypt permissions are relevant only if the role or resources accessed after assuming the role require decryption of KMS-encrypted data; they are not a prerequisite for the sts:AssumeRole API call itself. Option D is wrong because Service Control Policies (SCPs) in AWS Organizations can only deny or allow permissions for principals within the organization, but the question does not indicate that Account B is part of an organization, and even if it were, SCPs are not the missing requirement—the identity-based policy is the immediate missing element.

26
MCQmedium

A microservice reads a secret from AWS Secrets Manager using its task role (ServiceRole). The secret is configured to use a customer-managed CMK. In production, the service fails with AccessDeniedException on GetSecretValue. CloudTrail shows that Secrets Manager attempted kms:Decrypt but was denied. Which IAM policy change is most appropriate to fix the failure while keeping least privilege?

A.Add kms:Decrypt permission for the specific CMK ARN to ServiceRole, and also keep secretsmanager:GetSecretValue for the specific secret ARN.
B.Add secretsmanager:ListSecrets permission on "*" so the service can discover the secret and retry the read.
C.Add s3:GetObject permission to ServiceRole for the KMS key alias stored in an S3 bucket.
D.Add kms:Encrypt permission instead of kms:Decrypt, because the service only needs to read the secret.
AnswerA

Secrets Manager stores the secret value encrypted under a customer master key (CMK). When the microservice calls GetSecretValue, Secrets Manager must invoke KMS Decrypt to reveal the plaintext, so the role needs kms:Decrypt on that CMK in addition to secretsmanager:GetSecretValue on the secret ARN. The CloudTrail entry shows the failure occurred in KMS, not in Secrets Manager, confirming the service already passed the Secrets Manager authorization but lacked the decryption permission. Scoping both permissions to the specific ARNs maintains least privilege and fixes the AccessDenied.

Why this answer

The AccessDeniedException occurs because the task role (ServiceRole) lacks the kms:Decrypt permission for the customer-managed CMK used to encrypt the secret. Secrets Manager calls kms:Decrypt on your behalf when retrieving the secret value. Adding kms:Decrypt for the specific CMK ARN to ServiceRole, while retaining secretsmanager:GetSecretValue for the specific secret ARN, grants the minimum required permissions to decrypt and read the secret.

Exam trap

The trap here is that candidates assume secretsmanager:GetSecretValue alone is sufficient, overlooking that Secrets Manager must call kms:Decrypt with the caller's permissions when a customer-managed CMK is used.

How to eliminate wrong answers

Option B is wrong because secretsmanager:ListSecrets on "*" does not grant permission to decrypt the secret; it only lists secret metadata and does not resolve the kms:Decrypt denial. Option C is wrong because the KMS key alias is not stored in an S3 bucket in this scenario, and s3:GetObject is irrelevant to decrypting the secret; the error is about KMS decryption, not S3 access. Option D is wrong because kms:Encrypt is used to encrypt data, not to decrypt it; reading a secret requires kms:Decrypt, not kms:Encrypt.

27
MCQmedium

In AWS Organizations, a Service Control Policy (SCP) denies kms:Decrypt on a production CMK for all principals in the Finance OU. A developer in the Finance OU created/updated an IAM policy that allows secrets access, but the application still fails with AccessDenied due to the SCP. You must enable only the Finance OU to decrypt that specific CMK while keeping the SCP restrictions for other OUs. What is the correct remediation?

A.Update the developer’s IAM policy to allow kms:Decrypt on the CMK alias ARN so the request bypasses the SCP.
B.Modify the SCP so it no longer denies kms:Decrypt for that specific CMK when applied to the Finance OU, while preserving the deny behavior for other OUs.
C.Add a KMS key policy statement that allows the developer role to decrypt the CMK.
D.Attach a permissions boundary that grants kms:Decrypt so the SCP becomes irrelevant.
AnswerB

Because the SCP is what creates the Deny, the correct fix is to adjust the SCP scope/conditions so that kms:Decrypt for the specific CMK is not denied for the Finance OU. Other OUs remain under the same restrictive SCP behavior.

Why this answer

SCPs are evaluated before IAM policies and cannot be bypassed by IAM permissions. By modifying the SCP to exclude the specific CMK for the Finance OU (e.g., using a Condition key like `kms:ViaService` or a resource-level exception), you remove the explicit deny for that OU while keeping it in place for all other OUs. This ensures the developer's IAM policy can then allow `kms:Decrypt` without being blocked by the SCP.

Exam trap

The trap here is that candidates mistakenly think IAM policies or KMS key policies can override an SCP, but SCPs are a higher-order policy that always takes precedence over any allow within the account.

How to eliminate wrong answers

Option A is wrong because SCPs take precedence over IAM policies; an IAM policy allowing `kms:Decrypt` cannot bypass an SCP that explicitly denies the same action. Option C is wrong because a KMS key policy statement granting decrypt to the developer role is still subject to the SCP's explicit deny, which overrides any allow from the key policy. Option D is wrong because a permissions boundary limits the maximum permissions an IAM role can have, but it does not override an SCP; the SCP's explicit deny still applies and blocks the action.

28
MCQeasy

A development team is building a new application that stores session state in a relational database. The application experiences unpredictable read traffic, and the team wants a fully managed database that can scale read capacity automatically and provide a reader endpoint that distributes connections across multiple replicas. Which AWS service should a solutions architect recommend?

A.Amazon Aurora with Aurora Replicas and the reader endpoint.
B.Amazon Redshift with concurrency scaling enabled.
C.Amazon DynamoDB with on-demand capacity mode.
D.Amazon RDS for SQL Server with Multi-AZ deployment.
AnswerA

Aurora is a fully managed relational database compatible with MySQL and PostgreSQL, and Aurora Replicas can be added and scaled to handle read traffic. The cluster reader endpoint automatically distributes connections across the available Aurora Replicas, which is exactly the behavior described. It provides automatic storage scaling and managed replication, making it the correct relational, read-scalable choice for this scenario.

Why this answer

Aurora is a managed relational database whose Aurora Replicas serve read traffic, and the cluster reader endpoint spreads incoming read connections across those replicas automatically. That directly matches the need for automatic read scaling with a reader endpoint. DynamoDB is non-relational, RDS Multi-AZ standbys do not serve reads, and Redshift is an analytics warehouse rather than an OLTP database.

Exam trap

The trap here is assuming a Multi-AZ standby in RDS can absorb read traffic, when it exists only for failover and does not serve application reads.

29
MCQeasy

A Lambda function needs to read the current value of exactly one AWS Secrets Manager secret at startup. Which least-privilege IAM permission (action and resource scope) should you grant to the Lambda execution role?

A.secretsmanager:ListSecrets on all secrets (resource set to "*")
B.secretsmanager:GetSecretValue on only the secret’s full ARN
C.secretsmanager:UpdateSecret on the specific secret ARN
D.secretsmanager:DescribeSecret on all secrets (resource set to "*")
AnswerB

GetSecretValue is the only AWS Secrets Manager API action that returns the encrypted secret's decrypted string, which is exactly what the Lambda function must read. By restricting the Action to secretsmanager:GetSecretValue and the Resource to the secret's full ARN, the IAM policy grants access to that single secret while denying all other Secrets Manager operations. This follows least privilege because the function cannot update, list, or describe any other secret, and the full ARN is more precise than using a wildcard or partial name.

Why this answer

The Lambda function needs to read the current value of exactly one secret at startup. The least-privilege permission is `secretsmanager:GetSecretValue` scoped to that secret's full ARN. This action retrieves the secret value, and restricting the resource to the specific ARN ensures the function cannot access any other secrets.

Exam trap

The trap here is that candidates may confuse `ListSecrets` or `DescribeSecret` with `GetSecretValue`, thinking metadata retrieval is sufficient, or they may apply a broad resource scope ("*") instead of the specific ARN, violating the least-privilege principle that AWS emphasizes in the SAA-C03 exam.

Why the other options are wrong

A

The question requires reading exactly one secret's value at startup, so listing all secrets is unnecessary and violates least privilege by granting access to all secrets.

C

The question asks for reading a secret value at startup, but UpdateSecret is a write operation that modifies the secret, not reads it. It does not satisfy the requirement to read the current value.

D

The question requires reading the current secret value, but DescribeSecret only retrieves metadata (e.g., rotation date, tags) and not the secret value. Additionally, scoping to all secrets violates least privilege.

When would these options actually be correct?

A

If the Lambda function needed to discover which secrets exist (e.g., to dynamically select one based on naming pattern) before reading its value, then secretsmanager:ListSecrets on all secrets would be appropriate.

C

A Lambda function that needs to rotate a secret (e.g., update a database password) after reading it would require secretsmanager:UpdateSecret on the specific secret ARN to write the new value.

D

A question that asks for listing metadata of all secrets (e.g., to check rotation status) and least privilege is not a concern, or the requirement is to retrieve metadata for any secret in the account.

Why candidates pick the wrong answer

A

Candidates may think listing is a harmless read operation, but they overlook that least privilege demands scoping to the specific secret ARN, not all secrets.

C

Candidates may confuse the action needed for reading versus updating, or think that UpdateSecret is needed to 'refresh' the secret value, not realizing that GetSecretValue alone suffices for reading.

D

Candidates may confuse DescribeSecret with GetSecretValue, thinking it returns the secret value, or they may over-scope resources out of convenience.

30
MCQmedium

A company wants S3 access to be available only from private connectivity. They created an Interface VPC Endpoint for S3 (that provides private connectivity from their VPC to S3) and configured the application to use it from private subnets. The IAM role allows: - s3:GetObject on arn:aws:s3:::confidential-bucket/reports/* However, requests fail with AccessDenied. The S3 bucket policy includes an allow statement that permits GetObject only if: - aws:SourceVpce equals "vpce-0abc12345def6789" After redeploying the VPC endpoint, the application still uses the same IAM permissions but gets AccessDenied. What change is most likely to fix the issue?

A.Update the bucket policy to allow the new VPC endpoint ID (the vpce-* value) created by the redeployment.
B.Add internet egress via a NAT Gateway so the requests can reach S3 over the public endpoint.
C.Remove the aws:SourceVpce condition from the bucket policy to ensure the IAM permissions are sufficient.
D.Update the IAM role to add s3:PutObject permissions so the requests can be authorized.
AnswerA

The bucket policy is pinned to a specific endpoint ID using aws:SourceVpce. Redeploying or recreating the endpoint creates a new endpoint ID, so requests now present a different aws:SourceVpce value. Updating the bucket policy to match the new endpoint ID makes the condition true again while keeping access restricted to that specific private endpoint.

Why this answer

Redeploying a VPC Endpoint creates a new endpoint ID (vpce-*). The bucket policy explicitly allows access only if aws:SourceVpce matches the original endpoint ID. Since the new endpoint has a different ID, the condition fails, causing AccessDenied.

Updating the bucket policy to reference the new vpce ID restores access.

Exam trap

The trap here is that candidates assume IAM permissions alone are sufficient, overlooking that bucket policy conditions tied to a specific VPC endpoint ID become invalid after the endpoint is redeployed, causing an AccessDenied even with correct IAM roles.

How to eliminate wrong answers

Option B is wrong because adding a NAT Gateway would route traffic over the public internet, defeating the purpose of private connectivity and violating the bucket policy's SourceVpce condition. Option C is wrong because removing the condition would allow any VPC endpoint or public access to the bucket, compromising the security requirement for private-only access. Option D is wrong because the error is AccessDenied, not a missing permission; s3:PutObject is irrelevant to GetObject requests and does not address the condition mismatch.

31
MCQhard

Based on the exhibit, a public API is behind CloudFront. A single client IP is sending bursts of requests that are overwhelming the origin, and the team wants AWS to automatically mitigate the abuse at the edge without changing the application code. What should the team do?

A.Associate an AWS WAF web ACL with CloudFront and add a rate-based rule for the offending IP behavior.
B.Increase the ALB idle timeout to allow the origin to absorb more concurrent requests.
C.Add an Amazon Route 53 health check to fail over traffic to another DNS name.
D.Enable AWS Shield Advanced and rely on automatic DDoS protection for all request bursts.
AnswerA

AWS WAF is the right control at the CloudFront edge because it can inspect requests before they reach the origin and enforce a rate-based rule on abusive traffic patterns. A rate-based rule can automatically count requests by source IP and block or challenge requests that exceed the configured threshold, which directly addresses the burst traffic shown in the logs. This meets the requirement to mitigate at the edge without any application changes.

Why this answer

AWS WAF rate-based rules automatically block or rate-limit requests from a client IP when the request rate exceeds a threshold you define. By associating the web ACL with CloudFront, the rule is enforced at the edge before traffic reaches the origin, mitigating abuse without modifying application code.

Exam trap

The trap here is that candidates confuse AWS Shield Advanced's automatic DDoS mitigation (which handles network/transport layer floods) with the need for a WAF rate-based rule to stop application-layer request bursts from a single IP.

How to eliminate wrong answers

Option B is wrong because increasing the ALB idle timeout does not reduce the volume of requests hitting the origin; it only keeps idle connections open longer, which can actually worsen resource exhaustion. Option C is wrong because Route 53 health checks and failover reroute traffic to another endpoint but do not mitigate bursts from a single IP; the abusive client would simply follow the failover. Option D is wrong because AWS Shield Advanced provides enhanced DDoS protection against volumetric attacks, but it does not automatically apply per-IP rate limiting for application-layer request bursts; a rate-based rule in AWS WAF is required for that granular control.

32
MCQmedium

A healthcare provider hosts a patient-records API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. An audit finds the architecture cannot tolerate the loss of that Availability Zone. Budget is limited, and the API reads from an Amazon Aurora MySQL cluster that currently has one writer instance and no replicas. Which change most effectively addresses the audit finding?

A.Add an Aurora Replica in a second Availability Zone and register EC2 instances in that same second Availability Zone with the load balancer's target group.
B.Enable Aurora backtrack on the cluster and increase the EC2 Auto Scaling group maximum size within the existing Availability Zone.
C.Convert the cluster to Aurora Global Database and add a secondary AWS Region for the EC2 fleet.
D.Replace Aurora with a single-instance Amazon RDS for MySQL deployment and place all EC2 instances in an Auto Scaling group with a desired capacity of one.
AnswerA

Placing an Aurora Replica in a second Availability Zone gives the database a failover target, and Aurora promotes a replica automatically if the writer fails. Adding EC2 instances in that same zone removes the single-zone compute risk and lets the Application Load Balancer route to healthy targets in either zone. This directly resolves the audit finding at modest cost.

Why this answer

The audit finding is about surviving the loss of one Availability Zone, so the fix must add capacity in a second zone for both tiers. An Aurora Replica in another zone gives the database an automatic failover target, and EC2 instances in that same zone let the load balancer continue serving requests. Cross-Region designs exceed the requirement and the budget.

Exam trap

The trap here is confusing multi-Region disaster recovery with multi-Availability Zone high availability, which leads to an over-engineered and over-budget answer.

33
MCQeasy

A company runs a stateless API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. The ALB currently has a listener on port 80 only. The company wants the API to remain available if the single Availability Zone fails. What should the solutions architect do to meet this requirement with the LEAST operational overhead?

A.Enable cross-zone load balancing on the ALB and increase the desired capacity of the existing instances in the single Availability Zone.
B.Create a second ALB in another Availability Zone and use Amazon Route 53 failover routing with health checks to switch traffic when the primary zone fails.
C.Create an Amazon Machine Image (AMI) of the instances and configure an Auto Scaling group across at least two Availability Zones, registering the instances with the existing ALB target group.
D.Move the API to an Amazon S3 static website endpoint and serve the traffic through Amazon CloudFront with origin failover.
AnswerC

An Auto Scaling group spanning multiple Availability Zones automatically launches replacement instances in a healthy zone if one zone fails, and registers them with the existing target group. This provides zone-level resilience with minimal ongoing management because scaling and health replacement are handled by the service.

Why this answer

Running the stateless API in an Auto Scaling group that spans at least two Availability Zones lets the group replace instances in a surviving zone when one zone becomes unavailable. The existing ALB and target group can be reused, so the change adds zone redundancy with the least operational overhead.

Exam trap

The trap here is assuming that adding more instances inside the same Availability Zone or enabling cross-zone load balancing protects against a zone-level failure.

34
MCQhard

A Lambda-based retail API has unpredictable traffic spikes and users see latency caused by cold starts. The function must respond consistently during expected campaign windows. What should be configured? The design must avoid adding custom operational scripts.

A.A larger deployment package
B.Reserved concurrency only
C.Provisioned concurrency during campaign windows
D.CloudTrail data events
AnswerC

Provisioned concurrency is the correct choice because it instructs Lambda to pre-initialize a specified number of execution environments before traffic arrives, eliminating cold-start latency for those warm instances. During known campaign windows, you can adjust the provisioned concurrency level to align with expected demand. This allows the retail API to serve requests instantly during a traffic surge instead of waiting for environment creation. You pay for the pre-warmed capacity even when idle, but it directly solves the stated low-latency requirement.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, eliminating cold starts during campaign windows. This ensures consistent response times for the Lambda-based retail API under unpredictable traffic spikes without requiring custom scripts.

Exam trap

The trap here is confusing reserved concurrency (which prevents throttling but does not address cold starts) with provisioned concurrency (which eliminates cold starts by keeping environments warm).

How to eliminate wrong answers

Option A is wrong because a larger deployment package increases cold start duration, making latency worse. Option B is wrong because reserved concurrency only caps the maximum number of concurrent executions to prevent throttling, but does not pre-warm environments to avoid cold starts. Option D is wrong because CloudTrail data events record API activity for auditing, not performance optimization.

35
MCQmedium

A web application runs on an EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG spans three Availability Zones. After a deployment, new instances frequently fail the ALB target group health checks with HTTP 5xx responses and are quickly terminated by the ASG. What change most improves resiliency during deployments with minimal downtime by preventing premature removal of instances that are still starting?

A.Reduce the ASG health check grace period to 0 seconds so issues are detected faster.
B.Use a longer ASG health check grace period and deploy new instances using controlled replacement (for example, rolling instance refresh) so existing healthy instances continue serving while new ones warm up.
C.Restrict the ASG to a single Availability Zone so health check evaluation is simpler.
D.Disable ALB health checks so the ASG does not terminate instances on HTTP 5xx responses.
AnswerB

A longer ASG health check grace period prevents instances from being evaluated too early during normal startup time. Controlled replacement or rolling instance refresh ensures capacity is maintained while new instances warm up, so the ALB continues routing requests only to healthy targets.

Why this answer

Increasing the ASG health check grace period gives new instances more time to complete their startup and pass the ALB health checks before the ASG marks them unhealthy. A rolling instance refresh replaces instances in a controlled manner, ensuring that existing healthy instances continue serving traffic while new instances warm up, minimizing downtime and preventing premature termination.

Exam trap

The trap here is that candidates think reducing the grace period or disabling health checks will speed up recovery, when in fact it causes premature termination or serves traffic to unhealthy instances, increasing downtime.

How to eliminate wrong answers

Option A is wrong because reducing the grace period to 0 seconds would cause the ASG to terminate instances even faster when they return HTTP 5xx during startup, worsening the problem. Option C is wrong because restricting to a single Availability Zone reduces fault tolerance and does not address the root cause of premature termination during startup. Option D is wrong because disabling ALB health checks would prevent the ASG from detecting actual instance failures, leading to serving traffic to unhealthy instances and increasing downtime.

36
MCQmedium

A team serves static assets from an S3 origin through CloudFront. Cache hit ratio is low. Analytics show that requests include an Authorization header (even though the assets are public) and the cache key currently varies on that header, causing CloudFront to treat the same asset as different cache entries. What is the best change to improve cache hit ratio without breaking access controls?

A.Keep Authorization in the CloudFront cache key, but increase the origin response minimum TTL to 1 day.
B.Modify the CloudFront cache policy so the cache key does not include the Authorization header.
C.Switch the S3 origin from the current bucket to a website endpoint to enable automatic caching headers.
D.Enable CloudFront to forward all headers to S3 so origin can decide caching behavior per request.
AnswerB

CloudFront cache hit ratio depends on what constitutes a unique cache key. If Authorization is included, identical public assets requested with different Authorization values will map to different cache objects and reduce reuse. Removing Authorization from the cache key makes those requests share the same edge cache entry, improving hit ratio and reducing origin traffic. Because the scenario states the assets are public, removing Authorization from the cache key does not break access controls (access is not controlled by Authorization at the origin).

Why this answer

The low cache hit ratio is caused by the Authorization header being included in the CloudFront cache key, which creates separate cache entries for the same object even though the assets are public. By modifying the cache policy to exclude the Authorization header, CloudFront will treat all requests for the same asset as identical, dramatically improving the cache hit ratio without affecting access controls because the assets are already public.

Exam trap

The trap here is that candidates may think increasing TTL or changing the origin type will fix caching, when the real issue is the cache key composition—specifically, the Authorization header fragmenting the cache.

How to eliminate wrong answers

Option A is wrong because increasing the minimum TTL does not address the root cause—the cache key still varies on the Authorization header, so separate cache entries will persist and the cache hit ratio will remain low. Option C is wrong because switching to an S3 website endpoint does not change how CloudFront caches based on headers; the cache key is still controlled by the CloudFront cache policy, not the origin type. Option D is wrong because forwarding all headers to S3 would include the Authorization header in the cache key, making the problem worse by further fragmenting the cache.

37
MCQmedium

Based on the exhibit, the application sees several minutes of connection errors during an Aurora failover. What is the best change to reduce failover impact?

A.Change the application to use the Aurora cluster writer endpoint and retry transient connections.
B.Add an Aurora read replica and keep using the same JDBC URL.
C.Increase the EC2 instance size of the application servers.
D.Switch to a single-AZ RDS PostgreSQL instance for simpler connectivity.
AnswerA

The current configuration targets a specific instance endpoint, which becomes stale after failover. The Aurora cluster writer endpoint always resolves to the current writer, so the application can reconnect without manual endpoint changes. Adding retries with backoff helps the application survive the short DNS and connection transition during failover.

Why this answer

The Aurora cluster writer endpoint always points to the current primary instance, even after a failover. By using this endpoint and implementing retry logic for transient connection errors, the application can automatically reconnect to the new writer without manual intervention, reducing the impact of the failover from several minutes to seconds.

Exam trap

The trap here is that candidates often think adding read replicas or scaling application servers will fix failover connectivity, but the real issue is that the application must use the correct endpoint and handle transient disconnections gracefully.

Why the other options are wrong

B

Adding a read replica does not reduce failover impact because the application still uses the same JDBC URL, which points to the writer endpoint. During failover, the writer endpoint may be unavailable, and read replicas cannot handle write operations, so connection errors persist.

C

Increasing EC2 instance size addresses application-side compute capacity, not database failover connectivity issues. The connection errors stem from DNS propagation delays and endpoint changes during Aurora failover, which are unaffected by application server size.

D

Switching to a single-AZ RDS PostgreSQL instance removes high availability and does not address failover impact; it actually increases downtime risk during a failure.

When would these options actually be correct?

B

This option would be correct in a scenario where the application primarily performs read operations and needs to offload read traffic from the primary database to improve read performance or availability. For example, a reporting application that can tolerate stale reads could use read replicas to distribute load.

C

This option would be correct in a scenario where the application is experiencing performance degradation due to CPU or memory exhaustion on the application servers, and the database is healthy. For example, if an application's response times increase under load and CloudWatch shows high EC2 CPU utilization, resizing to a larger instance type would alleviate the bottleneck.

D

This option would be correct if the question asked for the simplest, lowest-cost database solution for a non-critical application that can tolerate downtime and does not require high availability or failover support.

Why candidates pick the wrong answer

B

Candidates may think that adding a read replica provides high availability or redundancy, but they overlook that the application's JDBC URL still points to the writer endpoint, which is the single point of failure during failover.

C

Candidates may assume that increasing server resources can compensate for any performance issue, including database failover delays, or they might misinterpret 'connection errors' as a symptom of insufficient application capacity rather than a database endpoint issue.

D

Candidates may think simplifying the architecture by removing Aurora's complexity will reduce failover issues, but they overlook that single-AZ lacks failover capability entirely.

38
MCQmedium

A SOC analyst needs an immutable, centralized audit record of configuration and API changes across multiple AWS accounts. Recently, an operator changed an IAM role trust policy, and investigators must determine exactly which principal made the change and which parameters were used. Your current setup sends application logs to CloudWatch Logs, but there is no organization-level API audit logging. Which approach best satisfies the requirement?

A.Enable an AWS Organizations CloudTrail organization trail that delivers management event logs (including IAM) to a centralized S3 bucket in a dedicated audit account, for all regions.
B.Use CloudWatch Logs metric filters on application logs to infer which principals changed trust policies.
C.Rely on GuardDuty alerts to provide the full request parameters for every IAM policy change.
D.Enable AWS Config only and store periodic snapshots without CloudTrail management events.
AnswerA

An organisation trail captures management events, including IAM trust policy changes, across every account and region, recording the principal identity and request parameters. Delivering to a centralised S3 bucket in a dedicated audit account provides the immutable, organisation-wide record investigators require.

Why this answer

An AWS Organizations CloudTrail organization trail captures management events (including IAM API calls like 'UpdateAssumeRolePolicy') across all accounts in the organization, delivering immutable logs to a centralized S3 bucket in a dedicated audit account. This provides the exact principal ARN, source IP, user agent, and request parameters for every API call, meeting the requirement for a centralized, immutable audit record of configuration and API changes.

Exam trap

The trap here is that candidates confuse AWS Config's resource tracking with CloudTrail's API-level auditing, failing to realize that only CloudTrail captures the 'who' and 'how' (principal and parameters) of a change, while Config only records the 'what' (state after change).

Why the other options are wrong

B

CloudWatch Logs metric filters on application logs cannot capture the full API request parameters (e.g., which principal made the change and the exact parameters used) because application logs are not authoritative for IAM changes and lack the detailed API call context required for immutable audit records.

D

AWS Config stores resource configuration changes but does not capture the full API request parameters (e.g., which principal made the change or the exact parameters used), so it cannot provide the immutable audit record of API calls required for forensic investigation.

When would these options actually be correct?

B

A SOC analyst needs to monitor application-level errors or specific patterns (e.g., failed login attempts) across multiple accounts and trigger alarms. In that scenario, CloudWatch Logs metric filters on centralized application logs would be appropriate for real-time alerting, not for immutable audit trails of API changes.

D

A question that asks for a solution to track resource configuration changes over time and detect drift, without needing to capture API caller identity or request parameters, would make AWS Config the correct answer.

Why candidates pick the wrong answer

B

Candidates may think CloudWatch Logs can serve as a centralized audit tool because it aggregates logs, but they overlook that application logs do not capture the full API request/response details needed for IAM policy changes, and they are not immutable.

D

Candidates may confuse AWS Config's configuration history with API audit logging, assuming it captures who made changes, when it only records the resulting state of resources.

39
MCQmedium

A read-heavy media archive repeatedly queries the same product catalogue data from DynamoDB with millisecond latency requirements. Which service can reduce read latency and table load? The architecture review board prefers a managed AWS-native control.

A.DynamoDB Accelerator (DAX)
B.Amazon Kinesis Data Firehose
C.AWS Glue Data Catalog
D.S3 Transfer Acceleration
AnswerA

DynamoDB Accelerator is an in-memory cache that sits in front of DynamoDB and returns items for repeated identical queries at single-digit-millisecond latency. Because this media archive repeatedly queries the same product, DAX caches those hot items and queries, reducing read latency and decreasing the read capacity units consumed by DynamoDB. It is purpose-built to accelerate DynamoDB reads while maintaining API compatibility.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that delivers microsecond read latency, directly addressing the millisecond requirement. By caching frequently accessed product catalogue data, DAX offloads read requests from the DynamoDB table, reducing table load and read capacity unit consumption. As a fully managed, AWS-native service, it aligns with the architecture review board's preference for managed controls.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration (which optimizes uploads to S3) with a caching solution for DynamoDB, or mistakenly think Glue Data Catalog or Kinesis Firehose can cache database queries, when only DAX provides in-memory acceleration for DynamoDB reads.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose is a streaming data ingestion service for loading data into data lakes or analytics tools, not a caching or read-latency reduction solution for DynamoDB. Option C is wrong because AWS Glue Data Catalog is a metadata repository for ETL and data discovery, not a cache that accelerates DynamoDB read queries. Option D is wrong because S3 Transfer Acceleration uses AWS edge locations to speed up uploads to S3 over long distances, but it does not cache DynamoDB data or reduce read latency for repeated queries.

40
MCQhard

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.

A.Instance store volumes
B.Amazon EFS with mount targets in multiple Availability Zones
C.An EBS volume attached to all instances
D.S3 mounted as a POSIX file system without a file gateway
AnswerB

Amazon EFS is a fully managed, regional NFS file system that provides shared, elastic file storage for Linux EC2 instances. By creating mount targets in each Availability Zone where instances reside, every instance can mount the same file system and read/write concurrently with POSIX semantics, automatic failover, and scalable throughput. This directly satisfies the requirement for shared file storage across multiple Linux EC2 instances in a patient portal.

Why this answer

Amazon EFS provides a fully managed, shared POSIX-compliant file system that can be mounted concurrently across multiple Linux EC2 instances. By creating mount targets in multiple Availability Zones, the file system remains accessible even if one AZ fails, meeting the high-availability requirement. This aligns with the architecture review board's preference for a managed AWS-native control.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which is limited to a single AZ and specific volume types) with the cross-AZ shared file system capability of EFS, or incorrectly assume that mounting S3 as a POSIX file system is a reliable, managed solution for shared storage.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and data is lost on instance stop or termination, so they cannot provide shared, durable storage across AZs. Option C is wrong because a single EBS volume can only be attached to one EC2 instance at a time (except for multi-attach io1/io2 volumes, which are limited to a few Nitro-based instances and still not designed for cross-AZ shared file storage). Option D is wrong because mounting an S3 bucket as a POSIX file system (e.g., via s3fs-fuse) does not provide native POSIX locking or consistency semantics, and it introduces performance and reliability issues; it is not a managed AWS-native file system service.

41
MCQeasy

You store application logs in an S3 bucket. After 30 days, the logs are rarely accessed, but you must retain them for 1 year for compliance. Which S3 feature is the best way to reduce storage cost while meeting the retention requirement?

A.Create an S3 lifecycle rule to transition older objects to a colder storage class after 30 days, then expire after 1 year
B.Keep all logs in S3 Standard and rely on lower request rates to reduce cost
C.Copy logs to EBS snapshots each week and delete the original files
D.Use S3 replication to a second bucket in another region to reduce costs
AnswerA

S3 lifecycle policies can automatically transition objects to lower-cost storage classes based on age. Transitioning after 30 days reduces ongoing storage costs because the logs are rarely accessed, while expiring after 1 year ensures you still meet the compliance retention window.

Why this answer

An S3 Lifecycle rule can automatically transition objects from S3 Standard to a colder storage class (e.g., S3 Glacier Instant Retrieval or S3 Glacier Deep Archive) after 30 days, reducing storage costs for rarely accessed logs. After 1 year, the rule can expire the objects, which permanently deletes them, meeting the compliance retention requirement without manual intervention.

Exam trap

The trap here is that candidates may think S3 Standard is always the cheapest option for infrequently accessed data, but they overlook the significant cost savings from lifecycle transitions to colder storage classes like S3 Glacier Deep Archive, which are designed for long-term archival with rare access.

Why the other options are wrong

B

S3 Standard is the most expensive storage class; relying on lower request rates does not reduce the per-GB storage cost, so it fails to minimize cost for rarely accessed logs.

C

Copying logs to EBS snapshots weekly is not cost-effective for long-term retention because EBS snapshots are designed for block-level backups of EC2 instances, not for storing log files. Additionally, this approach incurs ongoing costs for snapshot storage and does not provide native lifecycle management for infrequently accessed data.

D

S3 replication does not reduce storage costs; it increases them by storing duplicate data in another region. The question asks for cost reduction, not data redundancy or disaster recovery.

When would these options actually be correct?

B

If the question asked for the cheapest storage for data that is accessed frequently (e.g., daily) and requires low latency, S3 Standard would be the correct choice because it offers high durability and low-latency access without retrieval costs.

C

This option would be correct if the question required point-in-time recovery of an EC2 instance's root volume or data volume, with weekly backup snapshots for disaster recovery purposes, and the logs were stored on the instance's EBS volume rather than in S3.

D

A question requiring cross-region disaster recovery or compliance with data residency regulations where data must be stored in multiple geographic locations would make S3 replication the correct answer.

Why candidates pick the wrong answer

B

Candidates may think that infrequent access automatically reduces cost, overlooking that S3 Standard charges the same storage fee regardless of access frequency.

C

Candidates may think that using EBS snapshots is a valid backup strategy for any data, including logs, and that deleting the original files after copying reduces costs, without realizing that S3 lifecycle rules are more cost-effective and purpose-built for log retention.

D

Candidates may confuse replication with cost-saving mechanisms, thinking that moving data to another region automatically reduces costs, or they may overlook the additional storage and transfer fees.

42
MCQeasy

A security team requires that every object uploaded to s3://secure-bucket/uploads/ must be encrypted using SSE-KMS with a specific customer-managed KMS key. Which S3 bucket policy condition approach best enforces this requirement for PutObject requests?

A.Deny PutObject unless s3:x-amz-server-side-encryption equals "aws:kms" and s3:x-amz-server-side-encryption-aws-kms-key-id equals the required CMK ARN
B.Allow PutObject only when aws:SecureTransport is true; encryption is then guaranteed automatically
C.Deny PutObject if the request includes Content-Type other than "application/octet-stream"
D.Deny PutObject when the caller’s role is not allowed to kms:Decrypt in their IAM policy
AnswerA

This enforces the encryption choice at upload time by validating the request headers that specify SSE-KMS and the exact KMS key ID/ARN. Using a Deny condition ensures uploads that do not include the correct SSE-KMS headers (for example, unencrypted uploads or uploads using a different KMS key) are rejected immediately.

Why this answer

It uses a Deny effect with the s3:x-amz-server-side-encryption condition key set to 'aws:kms' and the s3:x-amz-server-side-encryption-aws-kms-key-id condition key set to the specific customer-managed KMS key ARN. This ensures that any PutObject request that does not include both the required encryption header and the exact KMS key identifier is denied, enforcing the encryption requirement at the bucket policy level.

Exam trap

The trap here is that candidates often confuse encryption in transit (aws:SecureTransport) with encryption at rest (SSE-KMS), or they mistakenly think that checking the caller's KMS permissions in the bucket policy is sufficient, when in fact the policy must inspect the request headers to enforce the encryption requirement.

Why the other options are wrong

B

This option only enforces SecureTransport (HTTPS), not encryption at rest. It does not require SSE-KMS or a specific KMS key, so objects could be uploaded without server-side encryption or with a different encryption method.

C

This condition restricts Content-Type, not encryption. The requirement is about SSE-KMS encryption, not the MIME type of the object.

D

Option D is wrong because the question requires encryption with a specific KMS key, not decryption permissions. Denying PutObject based on the caller's inability to decrypt does not enforce the use of the required KMS key for encryption; it only checks decryption capability, which is irrelevant for uploads.

When would these options actually be correct?

B

This would be correct if the question required enforcing encryption in transit for all S3 uploads, e.g., 'Ensure all PutObject requests use HTTPS.' Then a Deny with aws:SecureTransport false would be appropriate.

C

If the question required that only binary files (e.g., application/octet-stream) be uploaded to a specific bucket, a Deny policy on other Content-Type values would enforce that rule.

D

Option D would be correct in a scenario where the requirement is to prevent uploads by users who cannot decrypt objects in the bucket (e.g., to ensure only authorized decryptors can upload). For example: 'A security policy requires that only users with permission to decrypt objects using a specific KMS key can upload to the bucket.'

Why candidates pick the wrong answer

B

Candidates may confuse encryption in transit (HTTPS) with encryption at rest (SSE), or assume that requiring HTTPS automatically enforces server-side encryption, which is not true.

C

Candidates may confuse content restrictions with encryption requirements, or think that controlling Content-Type is a way to enforce data format standards that indirectly relate to security.

D

Candidates may confuse encryption and decryption permissions, thinking that requiring kms:Decrypt for uploads somehow enforces encryption, or they may assume that KMS key permissions are symmetric for encrypt and decrypt operations.

43
MCQmedium

A global video platform serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The team wants the control to be enforceable during normal operations.

A.A larger S3 bucket
B.Amazon CloudFront distribution with the S3 bucket as origin
C.RDS read replicas
D.An EC2 Auto Scaling group in one Region
AnswerB

Amazon CloudFront is a global content delivery network (CDN) that caches static objects from the S3 origin at edge locations geographically close to viewers. When a user requests an image or 3D asset, CloudFront serves it from the nearest edge cache, avoiding a round trip to the S3 bucket's origin Region, which dramatically reduces latency and improves throughput. CloudFront also supports features like origin shielding, signed URLs, and compression, making it the recommended service for accelerating delivery of static S3-hosted content to a worldwide audience.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content (images, JavaScript) at edge locations worldwide, reducing latency for users in distant countries. By using the S3 bucket as an origin, CloudFront serves cached copies from the nearest edge, drastically improving load times. This solution is enforceable during normal operations because CloudFront provides cache control headers and invalidation APIs to manage content freshness.

Exam trap

The trap here is that candidates may confuse scaling compute (EC2 Auto Scaling) or database (RDS read replicas) with content delivery, failing to recognize that static content performance is solved by a CDN like CloudFront, not by scaling backend resources.

How to eliminate wrong answers

Option A is wrong because increasing the S3 bucket size does not reduce latency; S3 is a regional service and does not cache content globally. Option C is wrong because RDS read replicas are for database read scaling, not for serving static files from S3. Option D is wrong because an EC2 Auto Scaling group in one Region does not address global latency; it only scales compute capacity in a single geographic area, leaving distant users unaffected.

44
MCQmedium

An application in Account B (IAM role arn:aws:iam::account-b:role/app-read) reads objects from an S3 bucket in Account A. The bucket uses SSE-KMS with a customer-managed KMS key in Account A. Object reads consistently fail with an error that includes "AccessDenied" and "kms:Decrypt". The IAM permissions in Account B for kms:Decrypt are correct, but the requests still fail. Which change will most directly fix the failure?

A.Add kms:Decrypt to the KMS key policy in Account A for the Account B role arn:aws:iam::account-b:role/app-read, and remove kms:Decrypt from the role policy in Account B.
B.Update the IAM role in Account B to use the s3:GetObject permission only, and rely on S3 to authorize KMS decrypt automatically.
C.Modify the KMS key policy in Account A to allow kms:Decrypt for the Account B role arn:aws:iam::account-b:role/app-read, using the appropriate cross-account conditions (for example, allowing the use via S3 and the expected encryption context for the bucket).
D.Switch the S3 bucket encryption from SSE-KMS to SSE-S3, keeping all existing IAM and KMS configuration unchanged.
AnswerC

For SSE-KMS, S3 must call KMS Decrypt when serving objects. KMS authorization is evaluated against the KMS key policy in Account A in addition to the identity policy in Account B. If the error includes kms:Decrypt AccessDenied in a cross-account scenario, the most direct fix is to update the KMS key policy to allow the Account B role to use the key for decrypt (often with conditions tied to S3 usage and the specific bucket/object encryption context).

Why this answer

When using SSE-KMS with a customer-managed KMS key in a cross-account scenario, the KMS key policy must explicitly grant the external IAM role (arn:aws:iam::account-b:role/app-read) permission to perform kms:Decrypt. Even if the IAM role in Account B has the correct kms:Decrypt permission, the KMS key policy in Account A acts as a resource-based policy that must also allow the cross-account principal. Without this, the KMS service denies the decrypt request, resulting in the 'AccessDenied' error.

Exam trap

The trap here is that candidates often assume IAM permissions alone are sufficient for cross-account KMS operations, forgetting that KMS key policies are resource-based and must explicitly allow external principals, even when the IAM role has the correct permissions.

Why the other options are wrong

A

The error indicates that the KMS key policy in Account A does not grant the Account B role permission to decrypt. Adding kms:Decrypt to the key policy is necessary, but removing it from the role policy in Account B is incorrect because the role still needs the permission for the request to proceed; both the key policy and the role policy must allow the action.

B

The error includes 'kms:Decrypt', indicating the KMS key policy is missing cross-account decrypt permission. Simply using s3:GetObject does not bypass KMS authorization; S3 cannot automatically authorize KMS decrypt across accounts without proper key policy.

D

Switching to SSE-S3 removes KMS involvement, but the question states that the bucket uses SSE-KMS with a customer-managed key. Changing encryption type is an indirect workaround that does not address the root cause (missing KMS key policy permissions) and may violate compliance or security requirements.

When would these options actually be correct?

A

This option would be correct if the question stated that the role in Account B had kms:Decrypt permissions but the key policy in Account A was overly permissive, and the goal was to restrict access by removing the permission from the role and relying solely on the key policy to grant cross-account access.

B

In a scenario where the S3 bucket uses SSE-S3 (not SSE-KMS) and the IAM role in Account B has s3:GetObject permission, then no KMS permissions are needed, and S3 handles decryption automatically.

D

This option would be correct if the question described a scenario where the application does not require KMS-based encryption, the bucket's encryption can be changed without affecting other dependencies, and the goal is to eliminate KMS-related permissions entirely to simplify access.

Why candidates pick the wrong answer

A

Candidates may think that removing the permission from the role simplifies the configuration or that the key policy alone is sufficient, misunderstanding that both the key policy and the IAM policy must grant the permission for cross-account access.

B

Candidates may think S3 handles all authorization automatically, overlooking that cross-account KMS access requires explicit key policy grants, not just IAM permissions.

D

Candidates may think that switching to SSE-S3 removes the KMS dependency and thus the error, without considering that the bucket is already configured with SSE-KMS and changing encryption type is a significant change that may not be allowed or desired.

45
MCQhard

A healthcare analytics platform stores derived datasets in an Amazon S3 bucket. Regulatory rules require that every object remain recoverable for 90 days after creation even if an application bug issues a delete, and that no object version be permanently destroyed during that window. The team wants the strongest protection with the least custom code. Which S3 feature should the solutions architect enable?

A.S3 Object Lock in compliance mode with a 90-day retention period.
B.S3 Lifecycle rules that transition objects to S3 Glacier Deep Archive after 1 day and expire them after 90 days.
C.S3 Object Lock in governance mode with a 90-day retention period.
D.S3 Versioning combined with a bucket policy that denies s3:DeleteObject.
AnswerA

Compliance mode prevents any principal, including the AWS account root user, from deleting or overwriting a protected object version until the retention period expires. Setting a 90-day retention satisfies the recoverability window with no custom code, giving the strongest immutability guarantee available for S3 objects.

Why this answer

S3 Object Lock in compliance mode enforces write-once-read-many protection for a defined retention period and cannot be bypassed by any user, including the root user. A 90-day retention period meets the recoverability requirement without custom code, making it the strongest and least-effort control for preventing permanent destruction of object versions.

Exam trap

The trap here is treating governance mode as equivalent to compliance mode, when governance mode can be bypassed by users holding the s3:BypassGovernanceRetention permission.

46
MCQmedium

A financial services company runs a high-traffic REST API on Amazon EC2 instances behind an Application Load Balancer. The API retrieves user session data from an Amazon DynamoDB table for every request. During peak hours, DynamoDB read capacity is exhausted, causing throttling and increased latency. The workload is read-heavy and the session data is accessed frequently but changes infrequently. The solutions architect needs to reduce DynamoDB read load and improve API response times with minimal application changes. Which solution meets these requirements?

A.Migrate the session data to Amazon ElastiCache for Redis and modify the application to read from Redis instead of DynamoDB.
B.Create a read replica of the DynamoDB table in another region and direct all reads to the replica.
C.Increase the DynamoDB read capacity to a higher provisioned level and enable auto scaling.
D.Enable DynamoDB Accelerator (DAX) for the table and modify the application to use the DAX client for session reads.
AnswerD

DAX is an in-memory cache for DynamoDB that provides microsecond read latency and reduces the number of read capacity units consumed. It requires minimal code changes: the application uses the DAX client instead of the standard DynamoDB client. Since session data is read frequently and changes infrequently, DAX is ideal. It offloads read traffic from the table, preventing throttling and improving API response times.

Why this answer

DynamoDB Accelerator (DAX) is a fully managed, highly available in-memory cache for DynamoDB that delivers up to a 10x performance improvement, from milliseconds to microseconds. It is API-compatible with DynamoDB, so the application only needs to use the DAX client. For read-heavy workloads with infrequent updates, DAX reduces read capacity consumption and improves response times without requiring a major rewrite or data migration.

Exam trap

The trap here is assuming that simply increasing read capacity or adding a read replica will reduce latency, when the real need is to offload repeated reads from the table with an in-memory cache.

47
MCQhard

A risk simulation workload generates analytics files that are accessed unpredictably. Some files become hot again months later. The team wants automatic storage cost optimisation without retrieval delays. What should be used?

A.Manual monthly review and object copying
B.S3 Glacier Flexible Retrieval for all files
C.S3 Intelligent-Tiering
D.EFS One Zone for analytics files
AnswerC

S3 Intelligent-Tiering automatically monitors and moves objects between four access tiers—frequent, infrequent, archive instant, and archive access—based on changing usage patterns, with zero retrieval fees and no latency impact. It allows the risk simulation workload to pay only for the storage class each file actually needs, while keeping every object immediately available for analysis.

Why this answer

S3 Intelligent-Tiering is the correct choice because it automatically moves objects between access tiers (frequent, infrequent, and archive instant access) based on changing access patterns, without any retrieval delays. This handles the unpredictable access described—files that become hot again months later—by keeping them in the archive instant access tier until access resumes, then promoting them instantly. It optimizes storage costs automatically without manual intervention or cold retrieval waits.

Exam trap

The trap here is that candidates often choose S3 Glacier Flexible Retrieval for cost savings, overlooking the 'without retrieval delays' requirement, which disqualifies any cold storage option that requires restoration time.

How to eliminate wrong answers

Option A is wrong because manual monthly review and object copying is not automatic, introduces operational overhead, and risks cost inefficiency or retrieval delays if the review cycle misses changing access patterns. Option B is wrong because S3 Glacier Flexible Retrieval has retrieval delays (minutes to hours) and is not suitable for files that may become hot again unpredictably, as it would cause unacceptable wait times for immediate access. Option D is wrong because EFS One Zone is a file system, not an object storage service, and does not provide the cost optimization or automatic tiering needed for unpredictable access patterns on analytics files.

48
MCQhard

Based on the exhibit, a central deployment role in Account A is assumed by several CI/CD pipelines from Account B. The role must remain reusable, but the team wants the TeamA pipeline to upload artifacts only to s3://artifact-bucket/teamA/prod/ without creating a separate IAM role. What is the best approach?

A.Use an IAM user in Account B and hard-code the narrower S3 path in its access key policy.
B.Add a bucket ACL that grants write access only to the TeamA pipeline session name.
C.Attach a permissions boundary to the central role so every pipeline session inherits the narrower prefix automatically.
D.Pass an STS session policy when TeamA assumes the role to further restrict the temporary credentials to the teamA/prod prefix.
AnswerD

An STS session policy is specifically designed to reduce the permissions of temporary credentials for a single assume-role session. The reusable base role can remain broad enough for multiple pipelines, while TeamA can pass a session policy that limits effective permissions to the teamA/prod prefix. This preserves the shared role model and achieves least privilege without creating a separate IAM role.

Why this answer

When the TeamA pipeline assumes the central IAM role in Account A, it can pass an STS session policy that further restricts the temporary credentials to only allow actions on the s3://artifact-bucket/teamA/prod/ prefix. This approach keeps the role reusable for other pipelines while enforcing a narrower permission scope at the session level, without requiring a separate IAM role.

Exam trap

The trap here is that candidates often think a permissions boundary (Option C) can dynamically restrict individual sessions, but permissions boundaries set a hard limit on the role's overall permissions and cannot be applied per-session like an STS session policy can.

How to eliminate wrong answers

Option A is wrong because IAM users in Account B cannot directly access resources in Account A via hard-coded access keys; cross-account access requires IAM roles and trust policies, and hard-coding keys violates security best practices. Option B is wrong because S3 bucket ACLs do not support restricting access based on an IAM role session name; ACLs are legacy and cannot filter by session tags or names. Option C is wrong because a permissions boundary sets the maximum permissions for the role itself, not for individual sessions; it would apply to all pipelines assuming the role, not just TeamA, and cannot dynamically restrict to a specific prefix per session.

49
MCQhard

A financial services company runs a three-tier web application on AWS. The application servers in a private subnet must retrieve database credentials from AWS Secrets Manager at startup. The security team requires that the credentials never be stored on disk and that access be granted only to the specific IAM role attached to the instances. Which solution meets these requirements with the LEAST operational overhead?

A.Retrieve the secret from AWS Secrets Manager using the AWS SDK with the instance's IAM role, and configure automatic rotation for the secret.
B.Use AWS Systems Manager Parameter Store SecureString parameters and retrieve them with the AWS CLI during instance bootstrap.
C.Embed the credentials in the AWS Lambda environment variables and have the application servers retrieve them through an API Gateway endpoint.
D.Store the credentials in an encrypted Amazon S3 object and have the application download and decrypt them at startup using the instance role.
AnswerA

Secrets Manager integrates natively with IAM roles, so the application can call GetSecretValue using temporary credentials from the instance profile without storing anything on disk. Built-in rotation for supported databases updates the secret and the database password automatically, minimizing operational effort. IAM policies can restrict access to the specific secret and role, satisfying least privilege. This is the intended AWS pattern for this scenario.

Why this answer

AWS Secrets Manager is purpose-built for storing and rotating database credentials. Using the instance's IAM role to call GetSecretValue keeps credentials in memory, avoids disk storage, and supports least privilege through IAM policies scoped to the specific secret. Automatic rotation removes manual effort.

The other options either require custom rotation logic, expose secrets in less secure locations, or add unnecessary components.

Exam trap

The trap here is treating Parameter Store SecureString as equivalent to Secrets Manager, ignoring that native automatic rotation for database credentials is a Secrets Manager feature.

50
MCQmedium

A company runs a critical two-tier web application on AWS. The web tier consists of Amazon EC2 instances behind an Application Load Balancer (ALB) in a single Availability Zone. The database tier is an Amazon RDS for MySQL DB instance in the same Availability Zone. A recent power outage in that Availability Zone caused a full application outage. The company wants to redesign the architecture to survive an Availability Zone failure with minimal operational overhead. Which solution meets these requirements?

A.Use an Application Load Balancer with cross-zone load balancing enabled and enable Multi-AZ on the RDS instance.
B.Enable RDS automated backups with a 30-day retention period and create a read replica in a second Availability Zone.
C.Place the EC2 instances in an Auto Scaling group spanning two Availability Zones, and take nightly snapshots of the RDS instance to restore in the second AZ.
D.Deploy the web tier across two Availability Zones with an ALB, and convert the database to an RDS Multi-AZ DB instance.
AnswerD

Spreading EC2 instances across two AZs behind an ALB removes the single point of failure for the web tier, and an RDS Multi-AZ DB instance maintains a synchronous standby in a different AZ. On failure, RDS automatically fails over to the standby, and the ALB routes traffic to healthy instances. This design survives an AZ outage with minimal operational overhead.

Why this answer

The correct design distributes the stateless web tier across multiple Availability Zones and uses an RDS Multi-AZ DB instance for automatic database failover. This eliminates single points of failure at both tiers and requires little ongoing operational effort. Other options either leave the web tier in one AZ or rely on manual recovery for the database, failing the resilience and minimal-overhead goals.

Exam trap

The trap here is assuming that enabling Multi-AZ on RDS alone provides full application resilience while ignoring the single-AZ web tier.

51
MCQmedium

An orders service publishes payment instructions to an Amazon SQS Standard queue. The downstream processor sometimes times out after it has already applied the payment, but before it can delete the message from the queue. As a result, the same payment instruction can be processed more than once. The team wants the strongest way to prevent duplicate side effects while keeping the system decoupled. What should they implement?

A.Keep the queue as SQS Standard but increase the visibility timeout so duplicates are less likely to reappear during timeouts.
B.Change the queue to an SQS FIFO queue and use a stable deduplication ID derived from the payment instruction ID.
C.Make the downstream processor idempotent by recording processed payment instruction IDs in a durable datastore and ignoring repeats.
D.Use an ALB health check to restart the downstream processor when timeouts occur.
AnswerC

SQS Standard is at-least-once delivery, so the same message can be delivered more than once if the consumer times out before deleting it. Idempotent processing is the strongest protection against duplicate side effects because it prevents repeat application of the payment even when the message is redelivered.

Why this answer

Making the downstream processor idempotent ensures that duplicate payment instructions are safely ignored, even if the same message is delivered more than once. This approach provides the strongest guarantee against duplicate side effects without requiring changes to the queue type or increasing visibility timeouts, and it keeps the system fully decoupled.

Exam trap

The trap here is that candidates often assume that switching to a FIFO queue or increasing visibility timeout fully solves duplicate processing, but they overlook that the downstream processor's timeout after applying the payment is the root cause, which idempotency directly addresses.

How to eliminate wrong answers

Option A is wrong because increasing the visibility timeout only reduces the likelihood of duplicates but does not eliminate them; a timeout can still occur after processing, leading to the same duplicate issue. Option B is wrong because switching to an SQS FIFO queue with a deduplication ID prevents duplicate messages from being delivered, but it does not prevent the downstream processor from timing out after applying the payment and before deleting the message, so the same message could be redelivered and processed again. Option D is wrong because an ALB health check only restarts the downstream processor when timeouts occur, but it does not prevent duplicate processing of the same payment instruction.

52
MCQmedium

Your order-processing system uses EventBridge rules to send events to a Lambda function that updates order status. Over the last week, some events fail with a transient database timeout, and the Lambda retries intermittently but then the events are lost (no alerts after failures). You want at-least-once processing, bounded retries, and a way to inspect unprocessable events for later reprocessing. Which architecture change best meets these requirements?

A.Send EventBridge events to an SQS queue, configure a redrive policy to move messages to a dead-letter queue (DLQ) after a defined receive count, and make the Lambda processing idempotent.
B.Invoke Lambda directly from EventBridge in asynchronous mode, and increase the Lambda timeout to reduce failures.
C.Use SNS topics with Lambda subscriptions, but remove all retry and DLQ configuration to minimize duplicate events.
D.Store failed events only in CloudWatch logs, and have operators manually copy log entries back into the database for reprocessing.
AnswerA

Routing events through SQS decouples delivery from Lambda, so the redrive policy bounds retries and diverts unprocessable messages to a DLQ for later inspection and reprocessing, while idempotent handlers preserve at-least-once semantics across duplicate deliveries.

Why this answer

It introduces an SQS queue between EventBridge and Lambda, which provides a durable buffer for events. The redrive policy moves events to a dead-letter queue (DLQ) after a defined number of failed processing attempts, ensuring bounded retries and preserving unprocessable events for later inspection and reprocessing. Making the Lambda idempotent guarantees at-least-once processing even if duplicate events occur.

Exam trap

The trap here is that candidates may think increasing Lambda timeout or relying on asynchronous invocation retries alone is sufficient, but they overlook the need for a DLQ to capture and inspect events that persistently fail, which is a key requirement for operational visibility and reprocessing.

Why the other options are wrong

B

Asynchronous Lambda invocation from EventBridge has limited retry (0-2 attempts) and no DLQ support, so events lost after transient failures cannot be inspected or reprocessed, failing the requirement for bounded retries and inspectability.

C

Removing retry and DLQ configuration prevents at-least-once processing and makes it impossible to inspect unprocessable events, directly contradicting the requirements.

D

Storing failed events only in CloudWatch logs and manually reprocessing them does not provide automated retries, bounded retries, or a systematic way to inspect and reprocess unprocessable events, violating the requirements for at-least-once processing and automated reprocessing.

When would these options actually be correct?

B

This option would be correct if the question required minimal cost and complexity for a non-critical system where occasional event loss is acceptable, and there was no need for DLQ or reprocessing.

C

In a scenario where the requirement is exactly-once processing with no duplicates and no need for retries or inspection of failed events, and the system can tolerate event loss on failure.

D

This option would be correct in a scenario where the requirement is to have a simple, low-cost solution for debugging and manual intervention, with no need for automated retries or DLQ, and where the volume of failures is very low and acceptable to handle manually.

Why candidates pick the wrong answer

B

Candidates may think asynchronous invocation automatically handles retries and durability, overlooking that EventBridge async targets have no DLQ and limited retry, and that increasing timeout doesn't prevent loss from other transient failures.

C

Candidates may think SNS with Lambda is simpler and that removing retries reduces duplicates, but they overlook the need for reliability and failure inspection.

D

Candidates may think that logging failures is sufficient for auditing and that manual reprocessing is acceptable, underestimating the need for automated retry and DLQ mechanisms to ensure reliability and reduce operational overhead.

53
MCQeasy

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The architecture review board prefers a managed AWS-native control.

A.CloudFront caching with appropriate TTLs
B.AWS Backup Vault Lock
C.IAM Access Analyzer
D.S3 Select
AnswerA

CloudFront can serve cached content from edge locations when the origin is temporarily unavailable.

Why this answer

CloudFront caching with appropriate TTLs allows cached responses to be served to users even when the S3 origin is temporarily unavailable. By setting a minimum TTL (e.g., 0 seconds for fresh content, but a higher default or maximum TTL for stale content), CloudFront can continue delivering previously cached pages from edge locations during an S3 outage, ensuring high availability and resilience. This is a managed AWS-native feature that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates may confuse data protection features (like Backup Vault Lock) or data retrieval tools (like S3 Select) with caching and origin resilience, overlooking that CloudFront's TTL-based caching is the direct AWS-managed solution for serving content during origin outages.

How to eliminate wrong answers

Option B (AWS Backup Vault Lock) is wrong because it is a data protection feature for backup vaults that prevents deletion of backups, not a mechanism to serve cached content during an origin outage. Option C (IAM Access Analyzer) is wrong because it analyzes resource-based policies to identify unintended public access, not to cache or serve static content. Option D (S3 Select) is wrong because it is a query-in-place feature that retrieves subsets of data from objects using SQL expressions, and it does not provide caching or resilience against origin outages.

54
MCQhard

A company runs an internal API on Amazon EC2 instances in a private subnet. The API must call AWS Systems Manager Parameter Store to read configuration values. The security team wants to avoid long-lived credentials on the instances and avoid routing traffic over the public internet. Which combination of steps should be taken?

A.Attach an IAM role to the instances and route Parameter Store calls through a NAT gateway.
B.Use an EC2 instance profile with a gateway VPC endpoint for Systems Manager.
C.Attach an IAM role to the instances and create an interface VPC endpoint for Systems Manager.
D.Store IAM user access keys in AWS Secrets Manager and retrieve them from the instance at boot.
AnswerC

An attached IAM role provides temporary credentials through instance metadata, removing long-lived keys. An interface VPC endpoint powered by AWS PrivateLink keeps Parameter Store traffic inside the VPC, so no internet or NAT gateway is needed. Together they meet both the credential and the private connectivity requirements.

Why this answer

Temporary credentials come from an IAM role attached to the instances, eliminating static access keys. Private access to Parameter Store is provided by an interface VPC endpoint, which uses AWS PrivateLink to keep traffic within the AWS network. Combining the role and the interface endpoint satisfies both the credential hygiene and the no-public-internet constraints.

Exam trap

The trap here is reaching for a gateway VPC endpoint for Systems Manager, when gateway endpoints exist only for Amazon S3 and Amazon DynamoDB.

55
MCQmedium

A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?

A.S3 lifecycle transition to Glacier Flexible Retrieval
B.An EBS snapshot schedule
C.S3 Cross-Region Replication with versioning enabled
D.A CloudFront distribution
AnswerC

S3 Cross-Region Replication (CRR) asynchronously replicates every new object to a destination bucket in a different AWS Region, and it requires versioning to be enabled on both the source and destination buckets. This maintains a separate, durable copy of the uploaded documents that can be promoted to production during a regional outage, satisfying disaster recovery and compliance requirements. Replication is automatic, can be filtered by prefix or tags, and preserves object metadata and versions.

Why this answer

S3 Cross-Region Replication (CRR) with versioning enabled automatically copies objects from a source bucket in one AWS Region to a destination bucket in another Region, meeting the disaster recovery requirement for a geographically separate copy. Versioning must be enabled on both buckets to support replication of all object versions, ensuring consistency and recoverability. This is the native S3 feature designed for cross-region data redundancy without custom scripting or third-party tools.

Exam trap

The trap here is that candidates confuse S3 Cross-Region Replication with S3 lifecycle policies or other storage services like EBS snapshots, failing to recognize that CRR is the only option that directly creates a second copy of S3 objects in a different AWS Region for disaster recovery.

How to eliminate wrong answers

Option A is wrong because S3 lifecycle transition to Glacier Flexible Retrieval only moves objects to a lower-cost storage class within the same bucket and region; it does not create a copy in another AWS Region. Option B is wrong because EBS snapshots are for Amazon Elastic Block Store volumes attached to EC2 instances, not for S3 objects, and they cannot replicate data across regions automatically without additional configuration like copying snapshots manually. Option D is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations for low-latency delivery; it does not provide persistent cross-region storage replication for disaster recovery.

56
Multi-Selectmedium

A company runs a stateless web application on Amazon EC2 instances behind an Application Load Balancer. The application experiences predictable traffic patterns: low traffic at night and high traffic during business hours. The company wants to optimize costs without compromising availability. Which two actions should be taken? (Choose two.)

Select 2 answers
A.Configure an Auto Scaling group with scheduled scaling to match the predictable traffic patterns.
B.Use a Network Load Balancer instead of an Application Load Balancer to reduce costs.
C.Enable detailed monitoring for the EC2 instances to improve scaling responsiveness.
D.Purchase Reserved Instances or Savings Plans for the baseline capacity.
E.Use Spot Instances for all instances in the Auto Scaling group to reduce compute costs.
AnswersA, D

Scheduled scaling allows the Auto Scaling group to adjust capacity based on known traffic patterns, such as scaling out before business hours and scaling in at night. This ensures sufficient capacity during peaks and reduces costs during low-traffic periods by terminating unused instances. It directly addresses the predictable nature of the workload.

Why this answer

Scheduled scaling aligns capacity with predictable traffic, scaling out during business hours and in at night to reduce costs. Purchasing Reserved Instances or Savings Plans for the baseline covers the continuously running capacity at a discount. Together, these actions optimize cost while maintaining availability.

Other options either risk availability, add cost, or do not address the predictable pattern effectively.

Exam trap

The trap here is assuming that Spot Instances are always the best cost-saving measure, but using them for all instances can jeopardize availability in a stateless web application, and detailed monitoring adds cost without addressing the predictable scaling need.

57
MCQeasy

A team stores important documents in Amazon S3. They want to recover earlier versions if someone overwrites or deletes a file by mistake. What should they enable?

A.Amazon S3 Versioning
B.Amazon EBS snapshots
C.Amazon CloudWatch logs
D.VPC flow logs
AnswerA

Amazon S3 Versioning is the correct solution as it automatically retains multiple variants of an object in the same bucket, each with a unique version ID. This crucial feature enables recovery from both accidental overwrites and deletions, ensuring the integrity and availability of important documents. When an object is modified or deleted, S3 does not remove the previous version, but rather stores it as a non-current version, allowing for easy restoration to any prior state.

Why this answer

Amazon S3 Versioning is the correct choice because it allows you to preserve, retrieve, and restore every version of every object stored in an S3 bucket. When enabled, S3 automatically maintains a unique version ID for each object, so if a file is overwritten or deleted, the previous version remains accessible. This directly addresses the requirement to recover earlier versions after accidental modification or deletion.

Exam trap

The trap here is that candidates may confuse S3 Versioning with backup services like EBS snapshots, but versioning is an S3-native feature for object-level recovery, not a volume-level backup mechanism.

Why the other options are wrong

B

Amazon EBS snapshots are used for backing up Amazon Elastic Block Store volumes attached to EC2 instances, not for S3 object versioning. They do not provide the ability to recover earlier versions of S3 objects.

C

Amazon CloudWatch logs capture log data from AWS resources, not file versions. They cannot recover overwritten or deleted S3 objects.

D

VPC flow logs capture IP traffic information for network interfaces in a VPC, not file version history in S3. They cannot recover overwritten or deleted S3 objects.

When would these options actually be correct?

B

When a question asks for a backup solution for EC2 instance volumes (e.g., to recover from accidental data loss or corruption), enabling EBS snapshots would be the correct answer.

C

A question asks: 'A company needs to monitor API calls to an S3 bucket for security analysis. What should they enable?' CloudWatch Logs would be correct if the question specified logging API activity via CloudTrail and storing logs in CloudWatch.

D

A question asks: 'A company needs to analyze network traffic patterns and troubleshoot connectivity issues between EC2 instances in a VPC. What should they enable?' VPC flow logs would be the correct answer.

Why candidates pick the wrong answer

B

Candidates may confuse EBS snapshots with versioning because both involve creating point-in-time backups, but they apply to different AWS services (EBS vs. S3).

C

Candidates may confuse logging (CloudWatch) with versioning, thinking logs can track changes and enable recovery, but logs only record events, not object versions.

D

Candidates may confuse 'logs' with version tracking, or think that any logging feature can help recover data, not understanding that VPC flow logs are for network metadata only.

58
MCQmedium

A global video platform serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The design must avoid adding custom operational scripts.

A.A larger S3 bucket
B.Amazon CloudFront distribution with the S3 bucket as origin
C.RDS read replicas
D.An EC2 Auto Scaling group in one Region
AnswerB

Amazon CloudFront distribution with the S3 bucket as origin is the correct solution because CloudFront caches static images and video at edge locations around the world, dramatically reducing the distance data must travel to reach users. When a user requests content, CloudFront serves it from the nearest edge cache, and only fetches from the S3 origin on a cache miss, which also offloads throughput from S3 and lowers recurring costs. CloudFront integrates securely with S3 via Origin Access Control (OAC), ensuring that the bucket stays private while the CDN delivers content globally.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content (images, JavaScript files) at edge locations worldwide. By distributing content closer to users, it reduces latency and improves load times for distant countries without requiring any custom operational scripts or changes to the S3 bucket.

Exam trap

The trap here is that candidates may think increasing S3 bucket size or adding compute resources (EC2, RDS) can solve latency issues, but the correct solution is a CDN like CloudFront that brings content physically closer to users.

How to eliminate wrong answers

Option A is wrong because increasing the size of an S3 bucket does not improve performance; S3 bucket size has no impact on latency or throughput for static content delivery. Option C is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to accelerate delivery of static files stored in S3. Option D is wrong because an EC2 Auto Scaling group in a single Region does not reduce latency for users in distant countries; it only provides compute scaling within one geographic area, not global edge caching.

59
MCQhard

Based on the exhibit, a batch platform in Account B must assume a role in Account A. Only the specific role arn:aws:iam::222233334444:role/BatchRunner should be allowed to assume it, and the design must prevent any other role in Account B from reusing the same external ID. Which change best meets the requirement?

A.Add an identity-based policy to the BatchRunner role that allows sts:AssumeRole on the target role.
B.Change the trust policy principal from account root to arn:aws:iam::222233334444:role/BatchRunner and keep the ExternalId condition.
C.Replace the ExternalId condition with a role session name condition so only BatchRunner sessions are accepted.
D.Attach an SCP to Account B that denies sts:AssumeRole unless the request comes from BatchRunner.
AnswerB

Restricting the trust policy principal from the entire 222233334444 account root to the exact ARN arn:aws:iam::222233334444:role/BatchRunner is the correct fix because it follows least privilege by allowing only that specific role to assume the target role. Keeping the ExternalId condition preserves protection against the confused deputy problem, requiring the caller to present the correct ExternalId in the sts:AssumeRole request before the trust policy authorizes the assumption.

Why this answer

The trust policy on the target role in Account A must restrict the principal to the exact BatchRunner role ARN (arn:aws:iam::222233334444:role/BatchRunner) rather than the entire Account B root. This ensures that only that specific role can assume the target role. Keeping the ExternalId condition adds an additional layer of security by requiring a unique identifier that only BatchRunner knows, preventing any other role in Account B from reusing the same external ID.

Exam trap

The trap here is that candidates often think an identity-based policy on the assuming role (Option A) is sufficient, but the trust policy on the target role must explicitly restrict the principal to the specific role ARN, not just the account root.

How to eliminate wrong answers

Option A is wrong because identity-based policies on the BatchRunner role cannot grant it permission to assume a role in another account; the trust policy on the target role must explicitly allow the BatchRunner principal, and the BatchRunner role also needs an sts:AssumeRole permission, but the key missing change is the principal restriction. Option C is wrong because a role session name condition (sts:RoleSessionName) is set by the assuming entity and can be spoofed by any role in Account B, so it does not prevent other roles from reusing the same external ID. Option D is wrong because Service Control Policies (SCPs) are applied at the organization or OU level in AWS Organizations, not to individual accounts, and they cannot restrict based on a specific role ARN within the same account; they also cannot enforce the external ID requirement.

60
MCQhard

Based on the exhibit, a batch-processing service runs on Amazon EC2. The workload is Linux-based, can run on ARM64, and is CPU-bound during its nightly processing window. The team wants the best throughput per dollar without changing the application logic. Which EC2 instance family should the solutions architect recommend?

A.C7g instances based on AWS Graviton processors
B.R7i instances because more memory will improve CPU-bound job throughput.
C.M7a instances because general-purpose families are always the safest performance choice.
D.T3 instances because burstable instances can handle occasional nighttime spikes at lower cost.
AnswerA

C7g instances are compute optimized and use Graviton processors, which often deliver strong price-performance for CPU-bound Linux workloads that can run on ARM64. The exhibit shows the application is compatible and even benchmarks faster on ARM.

Why this answer

The C7g instances are based on AWS Graviton processors (ARM64 architecture), which offer up to 25% better performance per dollar compared to x86-based instances for CPU-bound workloads. Since the workload is Linux-based, can run on ARM64, and is CPU-bound, the C7g family provides the best throughput per dollar without requiring any application logic changes.

Exam trap

The trap here is that candidates may choose memory-optimized or general-purpose instances (like R7i or M7a) thinking they are safer, or burstable instances (T3) assuming they handle spikes cheaply, without recognizing that compute-optimized ARM64 instances (C7g) provide the best throughput per dollar for CPU-bound, ARM64-compatible workloads.

How to eliminate wrong answers

Option B is wrong because R7i instances are memory-optimized, designed for workloads that require large amounts of memory, not for CPU-bound jobs where additional memory does not improve throughput. Option C is wrong because M7a instances are general-purpose and balance compute, memory, and networking, but they are not optimized for CPU-bound workloads and use x86 architecture, which typically offers lower performance per dollar compared to ARM64-based instances for this specific scenario. Option D is wrong because T3 instances are burstable and designed for workloads with low baseline CPU usage and occasional spikes, but they are not suitable for sustained CPU-bound processing during a nightly window, as they would exhaust CPU credits and incur performance throttling or additional costs.

61
MCQmedium

A media analytics company ingests a continuous stream of JSON clickstream events, roughly 20,000 records per second, into an Amazon Kinesis Data Streams stream with 32 shards. Downstream consumers must be able to re-read the same records up to 7 days later to rebuild a reporting index. Which combination of settings should the team use to maximize the number of records each consumer can read per second while preserving this replay capability?

A.Switch the stream to on-demand capacity mode, keep the retention period at 24 hours, and have consumers read with the Kinesis Client Library because it partitions reads across shards with no per-shard limit.
B.Increase the shard count to 64, set the data retention period to 168 hours, and have each consumer use enhanced fan-out with its own dedicated 2 MiB/s read throughput per shard.
C.Increase the shard count to 64, set the retention period to 168 hours, and rely on the shared GetRecords polling model because it automatically scales to 10 MiB/s per shard when many consumers register.
D.Keep 32 shards, set the retention period to 24 hours, and have consumers poll with the GetRecords API using a single shared throughput budget of 2 MiB/s per shard.
AnswerB

Adding shards raises the aggregate write and read ceiling, the 168-hour retention period keeps records replayable for a full week, and enhanced fan-out gives each registered consumer a dedicated 2 MiB/s per-shard pipe rather than sharing the 2 MiB/s per-shard limit. This directly satisfies both the throughput and the 7-day replay requirements.

Why this answer

The scenario needs higher aggregate read throughput plus a full week of replayable data. Enhanced fan-out is the only mechanism that gives each consumer a dedicated 2 MiB/s per-shard read pipe instead of sharing the 2 MiB/s per-shard polling budget, and the 168-hour retention setting preserves records for the required rebuild window. Adding shards raises the aggregate ceiling further.

Exam trap

The trap here is assuming that adding more registered consumers to a standard Kinesis Data Streams stream increases the shared 2 MiB/s per-shard read throughput limit.

62
MCQhard

Based on the exhibit, a CI pipeline assumes a shared deployment role in Account A. The role can access several artifact prefixes, but this pipeline must only upload to teamA/prod/ and decrypt using a single KMS key for this execution. Changing the shared role would affect other pipelines. Which approach should the pipeline use?

A.Attach a permission boundary to the pipeline's assumed session so the temporary credentials cannot exceed the shared role permissions.
B.Pass an inline session policy in the AssumeRole request that further restricts the temporary credentials to teamA/prod/ and the approved KMS key.
C.Add an SCP to Account A that forces all roles to use the same S3 prefix and key whenever they are assumed.
D.Change the role trust policy to allow only the teamA/prod/ prefix and the key ARN because trust policies can scope S3 object paths directly.
AnswerB

STS session policies are designed to further restrict the permissions of temporary credentials issued by AssumeRole. In this case, the shared role can remain reusable for other pipelines, while this one execution is narrowed to the exact S3 prefix and KMS key required. The effective permissions become the intersection of the role permissions and the session policy, which preserves least privilege without changing the shared role itself.

Why this answer

An inline session policy passed in the AssumeRole request allows you to further restrict the temporary credentials' permissions without modifying the shared role itself. This ensures the pipeline can only upload to teamA/prod/ and decrypt using the specified KMS key, while other pipelines using the same role remain unaffected.

Exam trap

The trap here is that candidates confuse permission boundaries (which set a maximum limit) with session policies (which further restrict a specific session), or mistakenly think trust policies can scope resource-level permissions like S3 prefixes or KMS keys.

How to eliminate wrong answers

Option A is wrong because a permission boundary sets the maximum permissions for the role but does not dynamically restrict the session to specific prefixes or keys; it would still allow access to all prefixes the role can access. Option C is wrong because SCPs apply to all principals in the account and cannot be scoped to a single pipeline's session without affecting other roles and users. Option D is wrong because trust policies control who can assume the role, not what actions the assumed session can perform; S3 object paths cannot be scoped in trust policies.

63
Multi-Selectmedium

An Aurora PostgreSQL application has an OLTP writer and a reporting dashboard that issues many read-only queries. The writer is healthy, but read latency rises noticeably during reporting windows. Which two changes should you make? Select two.

Select 2 answers
A.Add Aurora Replicas to scale out the read workload.
B.Send read-only application traffic to the reader endpoint.
C.Scale up only the writer instance and keep all queries on it.
D.Replace the cluster with a single-AZ RDS instance to reduce replication overhead.
E.Move the dashboard to DynamoDB without changing the query model.
AnswersA, B

Aurora Replicas are independent compute instances in the same Aurora cluster that share the underlying storage volume. By adding one or more replicas, you create additional read endpoints that can absorb dashboard queries and other read-only traffic, directly offloading the writer instance. Because Aurora's storage is distributed and replicated separately, adding replicas does not cause significant write overhead, making horizontal read scaling the most efficient and cost-effective solution for an OLTP workload with heavy reads.

Why this answer

Adding Aurora Replicas (Option A) is correct because Aurora Replicas are dedicated read-only instances that share the same underlying storage volume as the writer, allowing you to scale read capacity linearly without impacting write performance. Sending read-only traffic to the reader endpoint (Option B) is correct because the reader endpoint automatically load-balances connections across all available Aurora Replicas, ensuring that dashboard queries are distributed and do not overload a single instance.

Exam trap

The trap here is that candidates may think scaling up the writer instance (Option C) is sufficient, but they overlook that read-heavy workloads require horizontal read scaling via replicas, not just vertical scaling of the writer.

Why the other options are wrong

C

Scaling up the writer instance does not offload read queries; the writer still handles all traffic, so read latency remains high during reporting windows. Aurora Replicas are needed to distribute read-only queries.

E

Moving the dashboard to DynamoDB without changing the query model is wrong because DynamoDB is a NoSQL database with a different query model (key-value and document), so existing SQL queries from the reporting dashboard would not work without significant application changes.

When would these options actually be correct?

C

In a scenario where the database is CPU-bound on the writer due to write-heavy workload and read latency is acceptable, scaling up the writer instance (vertical scaling) would be correct to increase write throughput.

E

This option would be correct in a scenario where the reporting dashboard requires extremely low-latency access to a simple, high-traffic dataset (e.g., user session data or IoT events) and the application can be refactored to use DynamoDB's query patterns, with the original Aurora cluster handling only transactional writes.

Why candidates pick the wrong answer

C

Candidates may think that a more powerful writer can handle both reads and writes faster, overlooking that Aurora's architecture separates read scaling via replicas.

E

Candidates may think DynamoDB is always a good choice for read-heavy workloads due to its scalability and low latency, overlooking the fact that the existing query model (SQL) is incompatible with DynamoDB's NoSQL interface.

64
MCQeasy

A company runs EC2 instances in private subnets and needs to access Amazon S3 objects without using a NAT gateway. They want the traffic to stay within AWS private networking as much as possible (no internet egress). Which VPC endpoint type should they create for Amazon S3?

A.Create an Interface VPC endpoint for S3 and point the instances to it
B.Create a Gateway VPC endpoint for S3 and update the route tables to use it
C.Create a NAT gateway and allow outbound HTTPS to S3
D.Create a VPC endpoint service and manually register S3 as a provider endpoint
AnswerB

Gateway VPC endpoints for S3 are the supported way to send S3 traffic from private subnets without NAT. They add routes in the relevant route tables (via S3 prefix lists) so requests to S3 go through the AWS network. This avoids internet egress and keeps the path private to the extent intended by VPC endpoint routing.

Why this answer

A Gateway VPC endpoint for S3 is the correct choice because it uses prefix lists and route table entries to send S3 traffic directly through AWS's private network without leaving the AWS backbone or requiring a NAT gateway. This endpoint type supports S3 and DynamoDB only, and it does not incur hourly charges, making it cost-effective for private subnet instances to access S3 objects securely.

Exam trap

The trap here is that candidates often confuse Gateway endpoints (for S3/DynamoDB) with Interface endpoints (for other AWS services), or incorrectly assume that a NAT gateway is required for private subnet egress, missing that Gateway endpoints provide a free, private alternative for S3 access.

How to eliminate wrong answers

Option A is wrong because an Interface VPC endpoint for S3 uses an Elastic Network Interface (ENI) with a private IP, but it still requires a NAT gateway or internet gateway for private subnet instances to reach it unless the endpoint is in the same subnet; more importantly, Gateway endpoints are the recommended and simpler option for S3. Option B is the correct answer. Option C is wrong because a NAT gateway allows outbound internet traffic, which violates the requirement to keep traffic within AWS private networking and avoid internet egress.

Option D is wrong because a VPC endpoint service is used to expose your own services to other VPCs via AWS PrivateLink, not to access AWS services like S3; you cannot manually register S3 as a provider endpoint.

65
MCQhard

Based on the exhibit, a media company serves versioned JavaScript and CSS files from an Amazon S3 origin through CloudFront. After a frontend release, the cache hit ratio dropped sharply even though the file names are versioned. The application team says the browser requests include the same Authorization header on every asset request because the frontend and API share one domain. What should the solutions architect do to improve CloudFront cache hit ratio without changing the application authentication model for the API?

A.Enable S3 Transfer Acceleration on the bucket so CloudFront fetches objects faster from the origin.
B.Create a CloudFront cache policy that excludes Authorization, cookies, and unnecessary query strings from the cache key.
C.Switch the origin from S3 to an Application Load Balancer so CloudFront can cache dynamic responses more effectively.
D.Configure CloudFront to forward every viewer header to the origin so the origin can decide whether the content is cacheable.
AnswerB

This reduces cache fragmentation because CloudFront can reuse the same cached object for many viewers. Since the assets are immutable and versioned, the Authorization header is not needed to vary the cache for these files. Keeping API authentication separate preserves the application model while improving hit ratio.

Why this answer

The sharp drop in cache hit ratio is caused by the Authorization header being included in the cache key, which makes each request unique even though the file names are versioned. By creating a CloudFront cache policy that excludes the Authorization header (and unnecessary cookies/query strings) from the cache key, CloudFront can serve cached responses to requests with different Authorization headers, restoring the cache hit ratio without altering the application's authentication model for the API.

Exam trap

The trap here is that candidates may think the Authorization header is required for caching or that forwarding all headers is safe, but in reality, including it in the cache key destroys cache efficiency for static assets, and the correct solution is to exclude it via a cache policy.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration improves upload/download speed over long distances but does not affect CloudFront's cache key or hit ratio. Option C is wrong because switching to an Application Load Balancer would not solve the cache key issue; ALB is for dynamic content and would not improve caching for static versioned files served from S3. Option D is wrong because forwarding every viewer header to the origin would include the Authorization header in the cache key, making each request unique and further reducing the cache hit ratio, which is the opposite of what is needed.

66
MCQeasy

A DynamoDB-backed multi-tenant app experiences throttling during a promotion. Most writes and reads target tenant "ACME" and use the same partition key value, causing a hot partition. Which design change most directly improves performance?

A.Add a "shard" component to the partition key (for example, tenantId + hashed bucket) to spread traffic across partitions
B.Increase the table’s read capacity without changing the partition key
C.Switch all reads to strongly consistent reads to guarantee faster results
D.Store ACME data in S3 and query it directly to avoid DynamoDB throttling
AnswerA

DynamoDB throughput is distributed across physical partitions. If one partition key value receives most traffic, that partition throttles. Adding a shard component to the partition key increases the number of partition key values being used, spreading requests across more partitions and reducing hot-partition throttling.

Why this answer

Adding a shard component to the partition key (e.g., appending a random or hash-based suffix to the tenant ID) distributes writes and reads for the same tenant across multiple physical partitions. This directly alleviates the hot partition caused by all ACME traffic hitting a single partition key value, allowing DynamoDB to utilize its full provisioned throughput across partitions.

Exam trap

The trap here is that candidates may think increasing total table capacity (Option B) solves throttling, but they overlook that DynamoDB throttles at the partition level, not the table level, so a single hot partition remains constrained regardless of total capacity.

How to eliminate wrong answers

Option B is wrong because increasing the table’s read capacity does not fix the hot partition issue—DynamoDB distributes throughput evenly across partitions, so a single partition can still throttle even if total table capacity is high. Option C is wrong because strongly consistent reads do not improve performance; they are slower and consume more read capacity units than eventually consistent reads, and they do not spread traffic across partitions. Option D is wrong because storing ACME data in S3 and querying it directly bypasses DynamoDB’s low-latency access patterns and introduces additional complexity (e.g., S3 eventual consistency, lack of native querying), making it an inefficient and indirect solution for a hot partition problem.

67
MCQmedium

A production application writes to an Amazon Aurora PostgreSQL cluster. Users report that during business-hour reporting runs, write latency increases. The application team wants to keep the writer focused on OLTP writes while still providing low-latency reads for reporting queries. What architectural approach should the solutions architect recommend?

A.Create Aurora read replicas and direct reporting read-only connections to the cluster reader endpoint.
B.Resize the writer instance to a larger class so it can handle both writes and reads with fewer slowdowns.
C.Enable cross-region replication for the entire cluster so reporting always runs in the secondary Region.
D.Disable read replicas and use caching only in the application layer, keeping all queries connected to the writer endpoint.
AnswerA

Aurora read replicas are separate DB instances that share the same underlying storage volume, so they serve read-only traffic without adding load to the writer. The cluster reader endpoint automatically load-balances connections across all replicas, allowing reporting queries to run in parallel with production writes. This decouples read and write workloads, reducing contention on the writer and improving overall responsiveness. It also lets you scale read capacity independently by adding or resizing replicas.

Why this answer

A is correct because creating Aurora read replicas and directing reporting read-only connections to the cluster reader endpoint offloads read traffic from the writer instance. This allows the writer to focus on OLTP writes, while the reader endpoint load-balances read-only queries across replicas, providing low-latency reads for reporting without impacting write performance.

Exam trap

The trap here is that candidates may think resizing the writer instance (Option B) is sufficient, but the exam tests the architectural principle of separating read and write workloads to avoid resource contention, not just scaling vertically.

Why the other options are wrong

B

Resizing the writer instance to a larger class does not offload read traffic from the writer; reporting queries still compete with OLTP writes on the same instance, failing to isolate workloads and reduce write latency.

C

Cross-region replication does not reduce latency for reporting queries in the primary region; it creates a separate cluster in another region, which would not help with low-latency reads for local reporting and introduces additional cost and complexity.

D

Disabling read replicas and using only application-layer caching forces all reporting queries through the writer endpoint, increasing write latency during reporting runs. This contradicts the goal of offloading reads to keep the writer focused on OLTP writes.

When would these options actually be correct?

B

A solutions architect needs to improve overall database performance for a single-instance Aurora PostgreSQL that experiences high CPU and memory pressure from both reads and writes, with no requirement to separate read traffic. Scaling up the instance class provides more resources to handle the combined load.

C

When the requirement is disaster recovery across regions and reporting can tolerate higher latency, or when the primary region is heavily loaded and reporting queries can be offloaded to a different geographic location with acceptable latency.

D

In a scenario where the application requires strong read-after-write consistency and cannot tolerate even eventual consistency, and the reporting workload is small enough that caching handles most reads, directing all queries to the writer endpoint ensures immediate consistency.

Why candidates pick the wrong answer

B

Candidates may think that a larger instance can simply 'power through' the workload, overlooking the architectural best practice of separating read and write traffic to avoid contention.

C

Candidates may think that replicating to another region automatically distributes read load, but they overlook that cross-region replication is for disaster recovery, not for reducing read latency in the same region.

D

Candidates may think caching alone can solve read performance without understanding that reporting queries often involve complex aggregations that caching cannot fully address, and they may underestimate the impact of mixing read and write workloads on the same instance.

68
MCQmedium

An S3 bucket stores user-uploaded images. Access patterns are unpredictable: some objects are never read again, while others are occasionally retrieved months later. The team wants to reduce storage cost without having to manually track access frequency or run periodic analyses. Which S3 storage and lifecycle approach is the best fit?

A.Enable S3 Intelligent-Tiering so objects can automatically move between access tiers based on observed access patterns.
B.Use S3 Glacier Instant Retrieval for all objects immediately to minimize storage cost.
C.Create a lifecycle rule that transitions objects to Standard-IA after a fixed 30 days, regardless of access.
D.Keep all objects in S3 Standard and reduce costs by enabling server access logging compression.
AnswerA

S3 Intelligent-Tiering is designed for unknown or changing access patterns. It monitors access and automatically moves objects between tiers (for example, between frequent-access and infrequent-access tiers) based on actual usage, which avoids the need to manually decide transition schedules. This directly meets the requirement to reduce storage cost while eliminating ongoing manual tracking or periodic analysis.

Why this answer

S3 Intelligent-Tiering is the best fit because it automatically moves objects between access tiers (frequent, infrequent, archive instant, archive) based on changing access patterns, eliminating the need for manual tracking or lifecycle rules. This optimizes storage costs for unpredictable access patterns without requiring you to define fixed time-based transitions or perform periodic analyses.

Exam trap

The trap here is that candidates often choose a fixed lifecycle rule (Option C) thinking it is simpler, but they overlook the retrieval fees and inefficiency of applying a rigid time-based policy to unpredictable access patterns, whereas Intelligent-Tiering adapts dynamically without manual tuning.

How to eliminate wrong answers

Option B is wrong because storing all objects immediately in S3 Glacier Instant Retrieval incurs higher retrieval costs and minimum storage charges (90 days) for objects that may never be accessed again, and it does not adapt to unpredictable patterns. Option C is wrong because a fixed 30-day transition to Standard-IA does not account for objects that are accessed frequently after 30 days, leading to retrieval fees, and it fails to optimize for objects that are never accessed again. Option D is wrong because enabling server access logging compression does not reduce storage costs for the objects themselves; it only reduces log storage size, and keeping all objects in S3 Standard is more expensive than using Intelligent-Tiering for unpredictable access.

69
MCQhard

Based on the exhibit, the security team wants centralized detection and alerting for both successful and failed attempts to change S3 bucket policies and KMS key policies across multiple accounts. Which approach best meets the requirement?

A.Enable S3 server access logging on each bucket and archive the logs in the security account.
B.Use AWS Config rules only, because Config records every successful and failed API call automatically.
C.Create an organization CloudTrail trail for management events and add EventBridge rules in the security account to alert on PutBucketPolicy and PutKeyPolicy events, including failed calls.
D.Enable GuardDuty in every account and use its findings as the main source for policy change notifications.
AnswerC

An organization CloudTrail trail delivers read/write management events from all accounts in the AWS Organization to a single S3 bucket (and optionally CloudWatch Logs) in the security account, creating a centralized audit trail. CloudTrail records both successful and failed API calls, including PutBucketPolicy and PutKeyPolicy, with event details such as caller identity, source IP, and request parameters. By adding Amazon EventBridge rules that match these specific event names—including `errorCode` fields for failed calls—the security team can trigger near-real-time alerts or automated remediation, making this the most direct and complete solution.

Why this answer

An organization CloudTrail trail captures management events (including PutBucketPolicy and PutKeyPolicy) across all accounts in the organization, and EventBridge rules in the security account can filter for both successful and failed API calls (using the `errorCode` field) to trigger centralized alerts. This provides the required centralized detection and alerting for policy changes across multiple accounts.

Exam trap

The trap here is that candidates may confuse S3 server access logging (which logs object-level access) with CloudTrail (which logs management API calls), or assume AWS Config automatically records all API calls, when in fact Config only tracks configuration changes and not failed API attempts.

How to eliminate wrong answers

Option A is wrong because S3 server access logging logs object-level access requests, not management API calls like PutBucketPolicy, and it does not capture KMS key policy changes at all. Option B is wrong because AWS Config rules evaluate resource configurations and compliance, but they do not automatically record every API call; they rely on configuration changes and cannot directly alert on failed API calls. Option D is wrong because GuardDuty focuses on threat detection (e.g., anomalous behavior, compromised credentials) and does not natively provide detailed alerting for specific management API calls like PutBucketPolicy or PutKeyPolicy, especially for failed attempts.

70
MCQeasy

A small analytics team runs a nightly batch job on a single Amazon EC2 instance. The job starts at 2:00 AM and finishes by 4:00 AM. The instance is idle for the rest of the day. The team wants to reduce EC2 costs and is willing to accept that the instance may be stopped and started. Which action will reduce costs MOST effectively?

A.Move the job to a larger instance type to finish faster.
B.Enable detailed monitoring and create a CloudWatch alarm to reboot the instance if CPU utilization is low.
C.Stop the instance when the job completes and start it again before the next run using an automated schedule.
D.Purchase a 1-year All Upfront Reserved Instance for the instance.
AnswerC

Stopping an instance when it is not needed means you do not pay for compute hours while it is idle. Since the job runs for only about two hours each night, stopping the instance for the remaining 22 hours drastically reduces cost. Automated start and stop can be implemented with instance scheduler solutions or EventBridge and Lambda.

Why this answer

When an instance is needed only for a short, predictable window, stopping it during idle hours removes the compute charges for those hours. Automated start and stop scheduling is a simple and effective cost optimization for intermittent workloads. Reserved Instances and larger instance types do not eliminate the cost of idle time.

Exam trap

The trap here is assuming that a Reserved Instance or a larger instance type will reduce cost, when the biggest saving comes from not paying for the instance at all during the long idle period.

71
Multi-Selectmedium

A retail API runs on Amazon EC2 instances behind an Application Load Balancer and stores orders in an Amazon RDS for PostgreSQL database. A test that stopped one Availability Zone caused the API to return errors because all application servers were in the same AZ and the database was single-AZ. Which two changes should the architect make to continue serving traffic during a single-AZ failure? Select two.

Select 2 answers
A.Increase the EC2 instance size and keep all application servers in the same subnet.
B.Configure the Auto Scaling group to launch instances across private subnets in at least two Availability Zones.
C.Replace the Application Load Balancer with a Network Load Balancer in a single Availability Zone.
D.Convert the RDS for PostgreSQL database to a Multi-AZ deployment.
E.Add an Amazon RDS read replica and point the application to the replica endpoint.
AnswersB, D

Spreading instances across private subnets in at least two Availability Zones lets the Auto Scaling group replace capacity in a surviving AZ, satisfying the requirement to keep serving traffic when one AZ fails. The load balancer then routes only to healthy targets.

Why this answer

Option B is correct because an Auto Scaling group that spans private subnets in at least two Availability Zones ensures application instances remain available if one AZ fails, and the ALB can route to healthy targets in the surviving AZ. Option D is correct because converting RDS for PostgreSQL to a Multi-AZ deployment creates a synchronous standby in a different AZ and automatically fails over the database endpoint, eliminating the single-AZ database as a point of failure. Option A is wrong because increasing instance size does not address the single-AZ placement of the application servers.

Option C is wrong because a single-AZ Network Load Balancer still fails when that AZ fails and does not improve database resilience. Option E is wrong because a read replica is asynchronous, is not an automatic failover target for writes, and pointing the application at the replica endpoint does not provide a highly available primary database.

Exam trap

The trap here is that candidates often think a read replica can serve as a high-availability solution for writes, but read replicas are asynchronous and do not support automatic failover for the primary database.

Why the other options are wrong

A

Increasing EC2 instance size and keeping all servers in one subnet does not provide fault tolerance across Availability Zones; a single AZ failure would still take down all application servers.

C

A Network Load Balancer (NLB) in a single AZ cannot provide cross-AZ failover; the question requires serving traffic during a single-AZ failure, which demands multi-AZ architecture. An NLB alone does not address the lack of application server redundancy.

E

A read replica does not provide automatic failover; the application would need to manually switch to the replica endpoint, which does not address the single-AZ failure of the primary database. The question requires continued serving traffic during a single-AZ failure, which Multi-AZ provides by automatic failover.

When would these options actually be correct?

A

If the question described a performance bottleneck (e.g., CPU or memory saturation) and the goal was to improve throughput for a single-AZ workload, increasing instance size would be correct.

C

When the requirement is to handle extremely high throughput with low latency for TCP/UDP traffic, and the application is already deployed across multiple AZs with its own failover logic. An NLB in a single AZ could be correct if the question explicitly states that only one AZ is used and the goal is to maximize performance within that AZ.

E

A read replica would be correct if the question asked for offloading read traffic from the primary database to improve read performance, or if the requirement was to have a standby for disaster recovery in a different region (cross-region read replica) without automatic failover.

Why candidates pick the wrong answer

A

Candidates may think bigger instances alone can handle failures, confusing vertical scaling with high availability.

C

Candidates may confuse load balancer types, thinking an NLB provides better availability than an ALB, or they may overlook that the NLB is still confined to a single AZ, which does not solve the multi-AZ failure requirement.

E

Candidates may confuse read replicas with Multi-AZ deployments, thinking that a read replica provides high availability, but it does not offer automatic failover and is primarily for read scaling.

72
MCQeasy

A company runs EC2 workloads in one region with somewhat steady overall demand. Over time, the team frequently changes instance families (for performance/optimization) and sometimes changes instance size, but wants predictable cost discounts. Which purchase option provides the best balance of cost savings and flexibility?

A.Standard Reserved Instances for a specific instance family and size only.
B.Savings Plans (Compute Savings Plans), scoped for flexible EC2 usage in the region.
C.Spot Instances for all workloads, assuming interruptions will never happen.
D.On-Demand only, because it avoids the complexity of purchase option scopes.
AnswerB

Compute Savings Plans provide discounted pricing for steady usage while allowing flexibility across instance families, OS, and sizes within the selected scope (for example, region). That matches the scenario: demand is steady enough for discounts, but the underlying instance type choices change frequently to meet performance needs.

Why this answer

Compute Savings Plans offer the best balance of cost savings and flexibility because they provide up to 66% discount in exchange for a commitment to a consistent amount of compute usage (measured per hour) in a region, but they automatically apply to any EC2 instance family, size, OS, or tenancy, as well as AWS Fargate and Lambda. This matches the team's need to frequently change instance families and sizes while still getting predictable discounts, unlike Standard RIs which lock you to a specific family and size.

Exam trap

The trap here is that candidates often confuse Standard Reserved Instances (which lock family/size) with Convertible RIs (which allow family changes but require a 1:1 exchange and still have restrictions), or they assume Savings Plans only apply to EC2, missing that Compute Savings Plans also cover Fargate and Lambda, making them the most flexible option for compute cost optimization.

How to eliminate wrong answers

Option A is wrong because Standard Reserved Instances require a commitment to a specific instance family and size (e.g., m5.large), which prevents the team from freely changing instance families for performance optimization without incurring modification fees or losing the discount. Option C is wrong because Spot Instances can be interrupted with a 2-minute warning when AWS needs capacity back, making them unsuitable for steady workloads where interruptions are assumed to never happen—this violates the fundamental design of Spot as a cost-optimization tool for fault-tolerant or flexible workloads. Option D is wrong because On-Demand pricing offers no discount (0% savings) and avoids complexity only by paying full price, which fails to meet the requirement for predictable cost savings.

73
Matchingmedium

Match the disaster recovery strategy to the recovery posture it best fits for a Regional outage.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Lowest cost option where the environment is rebuilt from backups and hours of downtime are acceptable.

Keep only the critical core running in the secondary Region, then scale out after failover.

Run a scaled-down but functional environment in another Region for faster cutover.

Serve production traffic from more than one Region at the same time for the fastest recovery.

Why these pairings

These pairs match disaster recovery strategies to their recovery postures, aligning with AWS DR strategies where RTO and RPO define the recovery objectives.

74
MCQmedium

A payments service receives payment orders by consuming messages from an Amazon SQS Standard queue. The downstream processor occasionally exceeds its processing timeout. As a result, some messages reappear in the queue and may be processed more than once. The team wants to prevent duplicate side effects (for example, double-charging) and also ensure poison messages do not repeatedly consume processing capacity. What approach best satisfies both goals?

A.Implement idempotent processing (for example, store processed payment IDs in DynamoDB) and configure an SQS dead-letter queue (DLQ) using a redrive policy with an appropriate maxReceiveCount.
B.Rely only on increasing the SQS visibility timeout so duplicates rarely occur, without adding idempotency checks or a DLQ.
C.Switch to a FIFO queue and delete messages immediately upon receipt to avoid duplicates.
D.Move the workload to SNS and use synchronous HTTP endpoints so the sender retries until the receiver confirms success.
AnswerA

With SQS Standard’s at-least-once delivery, duplicates can occur. Idempotency ensures repeated processing of the same payment ID does not create duplicate side effects. A DLQ with redrive policy isolates poison messages: after a message is received and fails processing more than maxReceiveCount times, SQS moves it to the DLQ instead of cycling it back to the main queue indefinitely.

Why this answer

It addresses both requirements: idempotent processing (e.g., storing processed payment IDs in DynamoDB) ensures that even if a message is processed more than once, duplicate side effects like double-charging are prevented. Configuring an SQS dead-letter queue (DLQ) with a redrive policy and an appropriate maxReceiveCount (e.g., 3 or 5) automatically moves messages that exceed the maximum number of receives to the DLQ, preventing poison messages from repeatedly consuming processing capacity.

Exam trap

The trap here is that candidates often confuse 'exactly-once delivery' (FIFO queues) with 'exactly-once processing,' failing to realize that idempotency is still required to handle failures after message receipt, and that a DLQ is necessary to manage poison messages regardless of queue type.

Why the other options are wrong

B

Increasing visibility timeout reduces duplicates but does not guarantee idempotency; messages can still be processed multiple times if the timeout is exceeded. It also fails to handle poison messages that repeatedly fail processing.

C

FIFO queues guarantee exactly-once processing, but the question states messages reappear due to processing timeout; deleting immediately upon receipt would lose messages that fail processing, and FIFO does not prevent duplicate side effects if processing is not idempotent.

D

SNS with synchronous HTTP endpoints does not guarantee exactly-once processing; the sender may still retry, and the receiver could process duplicates. It also lacks a mechanism to handle poison messages that repeatedly fail, as there is no dead-letter queue.

When would these options actually be correct?

B

This option would be correct in a scenario where the processing time is consistently predictable and the only concern is to avoid temporary overlaps, with no requirement for duplicate prevention or poison message handling.

C

A question where the requirement is to process messages in strict order without duplicates, and the processing is idempotent or the message is deleted only after successful processing (e.g., using a FIFO queue with a consumer that deletes after processing and a DLQ for failures).

D

This option would be correct in a scenario where the team needs to fan out messages to multiple subscribers and requires immediate, synchronous confirmation of processing, with no concern for duplicate prevention or poison message handling.

Why candidates pick the wrong answer

B

Candidates may think that increasing visibility timeout is a simple fix to prevent duplicates, overlooking the need for idempotency and poison message management.

C

Candidates may think FIFO queues eliminate duplicates entirely, but they only prevent duplicates during delivery, not during processing; they also overlook the need for idempotency and poison message handling.

D

Candidates may think synchronous processing eliminates duplicates because the sender waits for a response, but they overlook that retries can still cause duplicates, and there is no built-in poison message handling.

75
MCQmedium

A telemetry pipeline uses RDS MySQL and receives many read-only reporting queries that slow down the primary database. What should the architect add? The architecture review board prefers a managed AWS-native control.

A.Multi-AZ standby and route reads to the standby
B.RDS read replica and route reporting queries to it
C.S3 lifecycle policy
D.A larger NAT gateway
AnswerB

Creating an RDS read replica is the correct approach because it provisions a separate, read-only MySQL instance that asynchronously replicates all changes from the primary database. Reporting and analytics queries can be pointed to the read replica's own endpoint, offloading read-heavy telemetry workloads from the primary and freeing its compute/IO capacity for writes. This pattern is purpose-built for scaling read throughput and matches the stated requirement to support read-only reporting traffic.

Why this answer

RDS read replicas are designed specifically to offload read-heavy workloads like reporting queries from the primary database. They provide an asynchronous read-only copy of the database that can handle SELECT statements without impacting the primary's write performance. This is a fully managed AWS-native solution that aligns with the architecture review board's preference.

Exam trap

The trap here is confusing Multi-AZ standby (which is for failover only) with read replicas (which are for read scaling), leading candidates to incorrectly choose Option A.

How to eliminate wrong answers

Option A is wrong because a Multi-AZ standby is a synchronous replica used for high availability and failover, not for read traffic; it cannot serve read queries directly. Option C is wrong because an S3 lifecycle policy manages object storage transitions and expiration, which is unrelated to offloading database read queries. Option D is wrong because a larger NAT gateway increases outbound internet capacity for private subnets, which does not address read query load on an RDS database.

Page 1 of 13

Page 2