Courseiva

SAA-C03 (SAA-C03) — Questions 526–600

935 questions total · 13pages · All types, answers revealed

Page 7

Page 8 of 13

Page 9
526
MCQhard

A solutions architect is designing a high-performance computing (HPC) workload on AWS that requires a shared POSIX-compliant file system with high throughput and low latency for thousands of concurrent compute instances. The workload is temporary, running for a few hours each week, and the team wants to minimize cost. Which storage solution should the architect recommend?

A.Amazon FSx for Windows File Server with Multi-AZ deployment and SSD storage, shared across all compute instances.
B.Amazon S3 with S3 Transfer Acceleration and a custom FUSE client mounted on each compute instance to provide POSIX access.
C.Amazon EFS with Provisioned Throughput and mount targets in multiple Availability Zones, mounted on all compute instances.
D.Amazon FSx for Lustre with a scratch file system, linked to an Amazon S3 bucket for data import and export.
AnswerD

FSx for Lustre is designed for HPC workloads, providing a POSIX-compliant file system with sub-millisecond latencies and hundreds of GB/s throughput. A scratch file system is cost-effective for temporary workloads and does not replicate data, which suits the few-hours-per-week pattern. Linking to S3 allows seamless data import and export, reducing manual data movement.

Why this answer

FSx for Lustre is purpose-built for HPC, offering a POSIX-compliant, high-throughput, low-latency file system. A scratch deployment is cost-effective for temporary workloads and can be linked to S3 for data staging. EFS, S3 with FUSE, and FSx for Windows File Server do not provide the required performance or protocol compatibility for this scenario.

Exam trap

The trap here is assuming that Amazon EFS with Provisioned Throughput can match the performance of FSx for Lustre for HPC, when Lustre is specifically optimized for high-throughput, low-latency parallel file access.

527
MCQmedium

A serverless API built with AWS Lambda serves latency-sensitive requests. The team observes intermittent slow responses during traffic ramp-ups and expects some users to hit the API immediately after a period of inactivity. Which configuration best reduces cold-start latency during these ramp-ups?

A.Enable Lambda provisioned concurrency on a published alias used by the API, and set a minimum provisioned concurrency greater than zero.
B.Increase the Lambda function’s memory setting; cold starts will always be eliminated regardless of traffic patterns.
C.Switch the Lambda runtime to a newer language version and remove any VPC configuration so the function never cold starts.
D.Set an API Gateway stage variable to "warm" the function at request time, which forces immediate initialization.
AnswerA

Provisioned concurrency keeps a defined number of Lambda execution environments initialized and ready behind a specific alias. When traffic ramps up—especially after inactivity—invocations can use pre-initialized environments, reducing or eliminating cold starts for those requests.

Why this answer

Lambda provisioned concurrency keeps a specified number of execution environments initialized and ready to respond immediately, eliminating cold starts for those invocations. By setting a minimum provisioned concurrency greater than zero on the alias used by API Gateway, the function remains warm even after periods of inactivity, ensuring consistent low latency during traffic ramp-ups.

Exam trap

The trap here is that candidates confuse provisioned concurrency with reserved concurrency, or assume that increasing memory or changing runtime settings can fully eliminate cold starts, when only provisioned concurrency guarantees pre-warmed execution environments for latency-sensitive workloads.

How to eliminate wrong answers

Option B is wrong because increasing memory reduces cold-start duration but does not eliminate cold starts; they still occur after inactivity. Option C is wrong because switching runtimes or removing VPC configuration does not prevent cold starts; VPC-enabled functions have additional cold-start overhead, but all Lambda functions can cold start regardless of runtime or VPC settings. Option D is wrong because API Gateway stage variables are static configuration values, not mechanisms to warm functions; they cannot force initialization at request time.

528
MCQmedium

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.

A.Subnets in at least two Availability Zones with health checks enabled
B.All instances in one larger subnet
C.A Network Load Balancer in one subnet
D.A single EC2 instance with detailed monitoring
AnswerA

Spreading the Auto Scaling group across subnets in at least two Availability Zones with health checks enabled ensures instances in a surviving AZ absorb traffic when one fails. This satisfies the requirement to tolerate a single AZ failure using AWS-native controls.

Why this answer

An Auto Scaling group configured with subnets in at least two Availability Zones ensures that if one AZ fails, the remaining AZ(s) can continue to serve traffic. Health checks on the EC2 instances allow the Auto Scaling group to detect and replace unhealthy instances, maintaining the desired capacity across the surviving AZs. This aligns with the requirement for a managed AWS-native control to tolerate an AZ failure.

Exam trap

The trap here is that candidates might think a single large subnet or a Network Load Balancer provides AZ resilience, but subnets are AZ-scoped and an NLB is a separate load-balancing component, not an Auto Scaling group configuration setting.

How to eliminate wrong answers

Option B is wrong because placing all instances in one larger subnet, even if it spans multiple AZs (which is not possible as subnets are AZ-specific), does not provide AZ failure tolerance; a single AZ failure would take down all instances. Option C is wrong because a Network Load Balancer (NLB) is not a component of an Auto Scaling group configuration; the question asks what the Auto Scaling group should include, and an NLB is a separate resource, not a configuration setting within the group. Option D is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of one Availability Zone; if that instance resides in the failed AZ, the application becomes unavailable, and detailed monitoring does not provide redundancy.

529
MCQmedium

A solutions architect is designing an S3 bucket for a order processing API. The objects must never be publicly accessible, even if a developer later adds an overly broad bucket policy. What should the architect configure?

A.Enable S3 Block Public Access at the account or bucket level
B.Enable server access logging on the bucket
C.Create an IAM policy that denies s3:GetObject to anonymous users
D.Enable S3 Transfer Acceleration
AnswerA

Enabling S3 Block Public Access at either the account or bucket level is the most robust and recommended control for preventing unintended public exposure of S3 buckets. This feature provides four distinct settings that can be applied to block public access granted through new or existing bucket policies, access control lists (ACLs), or any combination thereof. By enforcing these settings, S3 Block Public Access acts as a comprehensive safeguard, overriding any conflicting permissions that might otherwise inadvertently grant public read or write access to objects.

Why this answer

S3 Block Public Access provides a definitive override that prevents any public access to objects, regardless of bucket policies or ACLs. By enabling this setting at the account or bucket level, the architect ensures that even if a developer later adds an overly broad bucket policy, the objects remain inaccessible to anonymous users. This is the only option that guarantees no public access can be inadvertently granted.

Exam trap

The trap here is that candidates may think an IAM policy denying anonymous access is sufficient, but they miss that bucket policies can override IAM policies when both are evaluated, making S3 Block Public Access the only foolproof solution.

How to eliminate wrong answers

Option B is wrong because server access logging only records requests made to the bucket; it does not enforce any access restrictions. Option C is wrong because an IAM policy that denies s3:GetObject to anonymous users can be overridden by a later bucket policy that grants public access, as IAM and bucket policies are evaluated together and a bucket policy can explicitly allow what an IAM policy denies. Option D is wrong because S3 Transfer Acceleration is a performance feature that speeds up uploads over long distances; it has no effect on access control or public accessibility.

530
MCQmedium

A solutions architect is designing an S3 bucket for a IoT ingestion API. The objects must never be publicly accessible, even if a developer later adds an overly broad bucket policy. What should the architect configure? The design must avoid adding custom operational scripts.

A.Enable S3 Transfer Acceleration
B.Create an IAM policy that denies s3:GetObject to anonymous users
C.Enable S3 Block Public Access at the account or bucket level
D.Enable server access logging on the bucket
AnswerC

S3 Block Public Access provides four independent settings—BlockPublicAcls, IgnorePublicAcls, BlockPublicPolicy, and RestrictPublicBuckets—that prevent users from making objects or buckets public through ACLs or bucket policies. When enabled at the account level, these settings act as a guardrail for all current and future buckets, and they take precedence over any conflicting public ACL or bucket policy. For an IoT ingestion API where objects must remain private, this is the decisive control.

Why this answer

S3 Block Public Access provides a definitive override that prevents any public access to objects, regardless of bucket policies or object ACLs. This setting, when enabled at the account or bucket level, ensures that even if a developer later attaches an overly permissive bucket policy, the public access is blocked. It meets the requirement of avoiding custom operational scripts by being a native, configurable S3 feature.

Exam trap

The trap here is that candidates may think an IAM policy denying s3:GetObject to anonymous users is sufficient, but anonymous users are not IAM principals, so such a policy has no effect on anonymous access granted by a bucket policy.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is a performance feature that speeds up uploads over long distances using edge locations; it does not control access permissions or prevent public access. Option B is wrong because an IAM policy that denies s3:GetObject to anonymous users only applies to IAM principals, not to anonymous requests; anonymous users are not IAM entities, so this policy would not block public access granted by a bucket policy. Option D is wrong because server access logging records requests to the bucket for auditing purposes but does not enforce any access restrictions or prevent public access.

531
MCQeasy

A retail company hosts a product catalogue API on Amazon EC2 instances behind an Application Load Balancer. The API serves mostly small JSON responses and is read-heavy. Users in a distant continent report slow response times even though the origin servers are not heavily loaded. The company cannot change the application and wants the lowest-latency read experience globally. Which service should they use?

A.An Application Load Balancer with cross-zone load balancing enabled and a larger instance fleet.
B.An Amazon Route 53 latency-based routing policy with additional EC2 instances in a second Region.
C.Amazon CloudFront with the Application Load Balancer as the origin, configuring cache behaviors for the API paths.
D.AWS Global Accelerator with an accelerator in front of the Application Load Balancer.
AnswerC

CloudFront caches responses at edge locations close to users, so repeated reads of the same catalogue JSON are served from the edge with low latency. It supports the ALB as a custom origin, honors cache headers, and can forward dynamic requests to the origin while caching cacheable GET responses, which directly solves the distant-user latency for a read-heavy API.

Why this answer

A read-heavy API returning cacheable JSON is a strong fit for a CDN. CloudFront terminates user connections at nearby edge locations and serves cached responses, while forwarding cache misses to the ALB origin. This reduces round-trip time for distant users without changing the application or standing up a second Region.

Exam trap

The trap here is confusing Global Accelerator's network-path optimization with actual content caching at the edge.

532
Multi-Selecteasy

A service processes messages from an Amazon SQS queue. Sometimes the worker finishes the business logic but does not delete the message before the visibility timeout expires, so the message is delivered again. Which two changes improve resilience and reduce the impact of duplicate processing? Select two.

Select 2 answers
A.Make the message handler idempotent.
B.Set the SQS visibility timeout long enough for normal processing to complete.
C.Switch from SQS to Amazon SNS for reliable buffering.
D.Shorten the queue retention period so messages expire quickly.
E.Disable retries in the consumer application.
AnswersA, B

SQS provides at-least-once delivery, meaning the same message can be delivered to a consumer more than once, especially during network timeouts or consumer crashes. An idempotent handler design ensures that processing a duplicate message does not cause duplicate side effects, such as creating duplicate database records or processing the same payment twice. This is typically achieved by storing a unique message identifier or business key and checking for prior processing before executing the business logic. Idempotency is the fundamental corrective control for SQS's inherent lack of exactly-once semantics.

Why this answer

Making the message handler idempotent ensures that even if a message is processed multiple times (due to visibility timeout expiry), the business outcome remains the same. Idempotency is a key design pattern for resilient architectures when using at-least-once delivery systems like SQS. Option B is correct because setting the visibility timeout long enough for normal processing prevents premature redelivery, reducing the chance of duplicate processing in the first place.

Exam trap

The trap here is that candidates often think disabling retries or switching to SNS will solve the duplicate processing issue, but they fail to recognize that SQS's at-least-once delivery model inherently requires idempotent consumers and proper visibility timeout configuration.

533
MCQmedium

A analytics dashboard uses RDS MySQL and receives many read-only reporting queries that slow down the primary database. What should the architect add? The team wants the control to be enforceable during normal operations.

A.S3 lifecycle policy
B.RDS read replica and route reporting queries to it
C.Multi-AZ standby and route reads to the standby
D.A larger NAT gateway
AnswerB

An RDS read replica offloads read-only reporting traffic from the primary MySQL instance via asynchronous replication, directly relieving the contention described. Because the replica is a distinct endpoint, the team can enforce routing of reporting queries to it during normal operations, satisfying the stem's enforceability constraint.

Why this answer

B is correct because an RDS read replica is designed to offload read-heavy workloads from the primary database instance. By routing reporting queries to the read replica, the primary database is freed from processing these read-only requests, improving overall performance. This solution is enforceable during normal operations as the read replica is always available for reads, unlike a Multi-AZ standby which is not accessible for reads.

Exam trap

The trap here is confusing a Multi-AZ standby (which is not readable) with a read replica (which is readable), leading candidates to incorrectly choose C thinking the standby can serve reads.

How to eliminate wrong answers

Option A is wrong because an S3 lifecycle policy manages object transitions and expirations in S3, not database query routing or read offloading. Option C is wrong because a Multi-AZ standby is a synchronous replica used only for failover and is not accessible for read queries during normal operations; routing reads to it would fail. Option D is wrong because a larger NAT gateway increases outbound internet capacity for private subnets, which does not address database read performance or query routing.

534
MCQmedium

Based on the exhibit, the team wants to stop poison messages from consuming worker capacity and also prevent duplicate side effects if the same message is delivered more than once. Which design change best meets the requirement?

A.Increase the SQS queue batch size so each worker processes more messages per request.
B.Replace SQS with Amazon SNS and let each worker subscribe directly to the topic.
C.Configure a dead-letter queue and make the handler idempotent by storing a durable processed-message key.
D.Disable retries and shorten the visibility timeout so failed messages disappear sooner.
AnswerC

A dead-letter queue isolates messages that repeatedly fail so they stop wasting worker capacity. Idempotency ensures a message processed more than once does not create duplicate side effects, which is essential when visibility timeouts expire or retries occur. Together, these controls address both poison-message handling and at-least-once delivery behavior.

Why this answer

A dead-letter queue isolates poison messages that repeatedly fail processing, preventing them from consuming worker capacity. Making the handler idempotent by storing a durable processed-message key (e.g., using DynamoDB or a database) ensures that even if the same message is delivered more than once, duplicate side effects are avoided. This combination directly addresses both requirements: stopping poison messages from wasting resources and preventing duplicate processing.

Exam trap

The trap here is that candidates often think disabling retries or increasing batch size solves poison messages, but they fail to realize that only a dead-letter queue isolates problematic messages, and idempotency is required to handle duplicate deliveries inherent in SQS's at-least-once delivery model.

How to eliminate wrong answers

Option A is wrong because increasing the SQS batch size does not prevent poison messages from consuming worker capacity; it only makes each worker process more messages per request, which could actually increase the impact of poison messages. Option B is wrong because replacing SQS with Amazon SNS and having workers subscribe directly to the topic removes the ability to decouple producers and consumers, and SNS does not provide message retention, retries, or a dead-letter queue mechanism, so poison messages would still be delivered and could cause duplicate side effects. Option D is wrong because disabling retries and shortening the visibility timeout would cause failed messages to disappear sooner, but this does not prevent duplicate side effects (messages could still be redelivered before being deleted) and does not isolate poison messages—they would simply be lost, not handled.

535
MCQhard

Based on the exhibit, a company stores sensitive PDFs in S3 and serves them through CloudFront. Direct requests to the S3 object URL must fail, but CloudFront should still be able to fetch the files securely. Which solution best satisfies the requirement?

A.Leave the bucket public but require CloudFront signed cookies for all users.
B.Use an S3 access point and give it a public policy so CloudFront can reach the objects.
C.Configure CloudFront Origin Access Control for the S3 origin and update the bucket policy to allow only that distribution.
D.Use S3 object ACLs to grant read access only to users behind CloudFront.
AnswerC

Origin Access Control (OAC) enables CloudFront to authenticate every origin request to S3 using SigV4, so S3 can distinguish requests made by your distribution from all other traffic. You then write a bucket policy that allows s3:GetObject only when the principal is cloudfront.amazonaws.com and the aws:SourceArn matches your distribution's ID, thereby denying direct S3 bucket URL access. This is the standard secure pattern for a private S3 origin and is preferred over the older Origin Access Identity. It ensures that if someone finds a direct S3 URL, the bucket policy rejects it.

Why this answer

CloudFront Origin Access Control (OAC) allows CloudFront to authenticate requests to an S3 origin using a specific identity, and the bucket policy can be configured to grant access only to that CloudFront distribution. This ensures that direct S3 object URL requests fail (since they lack the CloudFront signature), while CloudFront can still fetch the files securely using the OAC mechanism.

Exam trap

The trap here is that candidates often confuse signed URLs/cookies (which control user access) with origin access controls (which control how CloudFront fetches from S3), leading them to pick options that still allow direct S3 access.

How to eliminate wrong answers

Option A is wrong because leaving the bucket public would allow anyone with the S3 object URL to access the files directly, violating the requirement that direct requests must fail. Option B is wrong because an S3 access point with a public policy would still allow direct public access to the objects, bypassing CloudFront. Option D is wrong because S3 object ACLs cannot restrict access based on the requester being behind CloudFront; they only grant permissions to specific AWS accounts or canonical users, not to CloudFront distributions.

536
MCQeasy

A latency-sensitive API is implemented with AWS Lambda. During traffic ramp-ups, users sometimes experience slow responses due to cold starts. The team wants to ensure fast initialization for a baseline level of concurrent requests. Which AWS feature should they use?

A.Lambda provisioned concurrency
B.Increase reserved instances for EC2
C.Enable S3 event notifications for every request to the API
D.Decrease the function timeout to reduce execution variability
AnswerA

Lambda provisioned concurrency pre-initializes a specified number of execution environments and holds them in a ready state, so requests are served immediately without the initialization delay that causes cold starts. By configuring a baseline concurrency, you ensure that latency-sensitive APIs experience consistent response times even during traffic spikes, as new environments are warmed ahead of actual invocations.

Why this answer

Lambda Provisioned Concurrency keeps a specified number of execution environments initialized and ready to respond immediately, eliminating cold starts for those concurrent requests. This directly addresses the latency-sensitive API requirement during traffic ramp-ups by ensuring fast initialization for a baseline level of concurrency.

Exam trap

The trap here is that candidates may confuse 'provisioned concurrency' with 'reserved concurrency' (which only caps concurrency, not pre-warms) or think that reducing the function timeout or adding S3 triggers can somehow mitigate cold starts.

How to eliminate wrong answers

Option B is wrong because reserved instances for EC2 apply to EC2 compute capacity, not to Lambda functions, and would not address Lambda cold starts. Option C is wrong because S3 event notifications are used to trigger Lambda functions on S3 object events, not to pre-warm Lambda execution environments, and adding them for every API request would introduce unnecessary complexity and latency. Option D is wrong because decreasing the function timeout does not reduce cold start latency; it only limits the maximum execution duration, and may actually increase execution variability by forcing premature terminations.

537
MCQmedium

A caching layer uses Amazon ElastiCache for Redis in front of a stateless web service. The service must continue to read cached responses during maintenance events and should automatically fail over to another node if one AZ becomes impaired. Which design change best satisfies this requirement?

A.Deploy a single-node Redis cluster and rely on application-level retries when cache misses occur.
B.Configure an ElastiCache Redis replication group with automatic failover across multiple Availability Zones.
C.Move the cache into the VPC but keep it in one Availability Zone to reduce network latency.
D.Use a Memcached cluster and configure only client-side connection pooling without failover support.
AnswerB

A Redis replication group with automatic failover maintains a primary and replicas across Availability Zones; if the primary's AZ is impaired, ElastiCache promotes a replica automatically, so cached reads continue during maintenance events as the stem requires.

Why this answer

An ElastiCache Redis replication group with automatic failover across multiple Availability Zones ensures that if the primary node or its AZ becomes impaired, a read-replica in another AZ is automatically promoted to primary. This allows the stateless web service to continue reading cached responses without interruption, satisfying both the maintenance and AZ impairment requirements.

Exam trap

The trap here is that candidates often confuse Memcached's simplicity with Redis's replication capabilities, assuming that client-side connection pooling alone can handle failover, when in fact Memcached lacks any built-in replication or automatic failover mechanism.

Why the other options are wrong

A

A single-node Redis cluster lacks automatic failover; if the node or its AZ becomes impaired, the service cannot read cached responses until the node is restored, violating the requirement for continued reads during maintenance and AZ impairment.

C

Keeping the cache in one Availability Zone does not provide automatic failover to another node if that AZ becomes impaired, failing the requirement for high availability during maintenance events.

D

Memcached does not support automatic failover or multi-AZ replication; if an AZ becomes impaired, the cache becomes unavailable, violating the requirement for continued reads during maintenance.

When would these options actually be correct?

A

This option would be correct in a scenario where the application can tolerate cache unavailability (e.g., reads from a slower database are acceptable) and the primary goal is cost minimization, with no requirement for high availability or automatic failover.

C

If the requirement was to minimize latency for a single-AZ application with no high availability needs, and the question explicitly stated that AZ impairment is not a concern, then deploying in one AZ would be appropriate.

D

If the requirement was for a simple, low-latency cache that can tolerate data loss and does not need automatic failover (e.g., caching non-critical, recomputable data in a single AZ), Memcached with client-side pooling would be appropriate.

Why candidates pick the wrong answer

A

Candidates may think a single-node cluster is simpler and cheaper, and assume application-level retries are sufficient to handle failures, underestimating the need for automatic failover to maintain cache availability during AZ impairments.

C

Candidates may think that reducing network latency by keeping the cache in one AZ is more important than high availability, or they may overlook the failover requirement and focus solely on performance.

D

Candidates may confuse Memcached's simplicity and speed with high availability, or assume that client-side connection pooling alone provides failover, not realizing Memcached lacks built-in replication and automatic failover.

538
MCQeasy

A company stores millions of small, rarely accessed backup objects in Amazon S3 Standard. The objects must remain immediately retrievable within milliseconds and be retained for at least five years, but the company wants to reduce storage cost. Which action should the company take?

A.Convert the bucket to S3 One Zone-IA to save on storage and retrieval.
B.Add a lifecycle rule to transition the objects to S3 Standard-IA after 30 days.
C.Add a lifecycle rule to transition the objects to S3 Glacier Deep Archive after 30 days.
D.Enable S3 Intelligent-Tiering and rely on it to move objects automatically.
AnswerB

S3 Standard-IA offers lower per-gigabyte storage pricing than S3 Standard while still providing millisecond latency and immediate access. It is intended for long-lived, infrequently accessed data, so it fits the requirement to keep objects instantly retrievable for five years while cutting storage cost. A lifecycle rule automates the transition.

Why this answer

S3 Standard-IA keeps objects immediately accessible with millisecond latency while charging less per gigabyte than S3 Standard, which matches the requirement to retain rarely accessed backups for five years without retrieval delays. A lifecycle rule can transition the objects automatically after 30 days. Glacier Deep Archive is cheaper but too slow, Intelligent-Tiering adds monitoring fees on small objects, and One Zone-IA sacrifices AZ resilience.

Exam trap

The trap here is chasing the absolute cheapest storage class and overlooking that archival tiers cannot deliver millisecond retrieval.

539
MCQmedium

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

A.Set the ECS deployment configuration to maximum percent 100 so tasks replace instances faster during rollouts.
B.Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.
C.Enable ALB connection draining longer than expected so existing connections survive longer during an AZ event.
D.Reduce task memory reservations to pack both tasks onto a single EC2 instance.
AnswerB

Multi-AZ instance capacity ensures tasks have eligible compute in another AZ when one AZ loses instances.

Why this answer

The current architecture has a single point of failure: the ASG spans only one subnet (one AZ), so if that AZ fails, all EC2 instances are lost, and the ECS service cannot run any tasks. By increasing the ASG min size to at least 2 and ensuring it uses subnets in at least two AZs, the ASG will maintain at least one healthy instance in each AZ, allowing the ECS service to survive a single-AZ outage. This aligns with the AWS Well-Architected Framework's principle of deploying across multiple AZs for high availability.

Exam trap

The trap here is that candidates may focus on the ALB's multi-AZ configuration and overlook that the compute layer (ASG/EC2) is the actual bottleneck, leading them to choose connection draining or deployment settings that do not address the fundamental lack of cross-AZ capacity.

How to eliminate wrong answers

Option A is wrong because setting the ECS deployment configuration maximum percent to 100 controls how many tasks are replaced during a rolling update, not the ability to survive an AZ failure; it does not address the underlying lack of EC2 capacity across AZs. Option C is wrong because ALB connection draining only gracefully terminates existing connections during deregistration or health check failures; it does not prevent service interruption when all EC2 instances in the single AZ become unavailable. Option D is wrong because reducing task memory reservations to pack both tasks onto a single EC2 instance actually increases risk—if that single instance or its AZ fails, both tasks are lost, violating the resilience requirement.

540
MCQmedium

A video platform uses Amazon Aurora. The workload has many short-lived database connections from Lambda functions, causing connection storms. What should be added? The design must avoid adding custom operational scripts.

A.S3 Select
B.An internet gateway
C.A larger Route 53 hosted zone
D.RDS Proxy
AnswerD

RDS Proxy is a fully managed, highly available database proxy that sits between the application and Amazon Aurora, maintaining a pool of established connections and reusing them for multiple app sessions. It dramatically reduces the overhead of opening and closing connections, which is particularly valuable for serverless applications like AWS Lambda that scale quickly and create connection storms. By multiplexing client connections through a small set of persistent connections, RDS Proxy keeps Aurora within its max connection limits and also improves failover resilience by preserving connection paths.

Why this answer

RDS Proxy is the correct choice because it sits between Lambda functions and the Aurora database, pooling and reusing database connections. This prevents connection storms by reducing the overhead of establishing new connections for each short-lived Lambda invocation, without requiring custom scripts or application changes.

Exam trap

The trap here is that candidates may think scaling the database (e.g., using Aurora Auto Scaling) or adding more compute resources solves connection storms, but the real bottleneck is the connection overhead itself, which RDS Proxy directly addresses without custom scripts.

How to eliminate wrong answers

Option A is wrong because S3 Select is a service for retrieving subsets of data from objects in Amazon S3 using SQL expressions; it does not manage database connections or address connection storms. Option B is wrong because an internet gateway enables VPC-to-internet communication for public subnets; it has no role in database connection pooling or reducing connection overhead. Option C is wrong because a larger Route 53 hosted zone increases the number of DNS records you can host but does not affect database connection management or mitigate connection storms.

541
MCQhard

A financial services company must store audit logs in S3 for 7 years and ensure that no one — including the AWS account root user — can delete or overwrite the logs during the retention period. Which S3 Object Lock configuration should a solutions architect use?

A.Object Lock in Compliance mode with a 7-year retention period
B.Object Lock in Governance mode with a 7-year retention period
C.S3 Versioning with a lifecycle rule to transition objects to Glacier after 7 years
D.A bucket policy with Deny for s3:DeleteObject applied to all principals including root
AnswerA

S3 Object Lock in Compliance mode establishes an unalterable Write Once, Read Many (WORM) state for objects. This mode prevents any user, including the AWS account root user, from deleting or overwriting objects until the specified retention period, in this case, 7 years, has expired. The retention period cannot be shortened or removed by anyone, ensuring the highest level of data immutability required for strict financial regulatory compliance.

Why this answer

S3 Object Lock in Compliance mode prevents ALL users — including the root account — from deleting or overwriting objects before the retention period expires. The retention period itself cannot be shortened once set in Compliance mode.

Governance mode also prevents most deletions, but users with s3:BypassGovernanceRetention permission (and the root account) can delete objects or shorten the retention period. For regulatory requirements where not even root can override, Compliance mode is mandatory.

Exam trap

Candidates choose Governance mode because 'governance' sounds strict. In AWS terminology, Governance is the LESS strict option — it can be bypassed by privileged users. Compliance mode is immutable: no one can remove the retention until the period expires.

This distinction is critical for financial regulations like SEC Rule 17a-4 and FINRA requirements.

Why the other options are wrong

B

Governance mode can be bypassed by the root account and users with s3:BypassGovernanceRetention permission. This does NOT meet the requirement that no one including root can delete the logs.

C

S3 Versioning prevents accidental deletion by keeping previous versions, but a privileged user can permanently delete all versions. Lifecycle rules manage storage class transitions — they do not prevent deletion. Compliance mode is required.

D

Bucket policies cannot restrict the root account. IAM policies (including resource-based policies) cannot override root user permissions. Only AWS Organizations SCPs and S3 Object Lock Compliance mode can restrict root's ability to delete S3 objects.

542
MCQeasy

A small e-commerce company runs a web application on a single Amazon EC2 instance in one Availability Zone. The instance stores session state locally and the database runs on the same instance. The company wants the application to survive an Availability Zone failure with minimal changes and no data loss for committed orders. Which combination of changes should the architect recommend?

A.Create an Amazon Machine Image of the instance and configure a weekly cron job to launch a replacement instance in a second Availability Zone if the first fails.
B.Attach an additional Amazon EBS volume to the instance and take hourly snapshots to Amazon S3 for disaster recovery.
C.Move the database to Amazon RDS with a Multi-AZ deployment, place the web tier in an Auto Scaling group across multiple Availability Zones, and store session state in Amazon ElastiCache for Redis.
D.Enable detailed monitoring on the EC2 instance and create an Amazon CloudWatch alarm that reboots the instance when the status check fails.
AnswerC

Amazon RDS Multi-AZ provides a synchronous standby in another Availability Zone with automatic failover, protecting committed orders. An Auto Scaling group across AZs keeps the web tier available, and ElastiCache for Redis externalizes session state so users are not tied to one instance, delivering resilience with minimal application redesign.

Why this answer

RDS Multi-AZ keeps a synchronous standby in a second Availability Zone and fails over automatically, protecting committed orders. An Auto Scaling group across AZs maintains the web tier, and externalizing sessions to ElastiCache for Redis removes dependence on a single instance, so the application survives an AZ failure with minimal redesign.

Exam trap

The trap here is treating EBS snapshots or instance reboots as high availability, when they only provide backup or recover the same failed Availability Zone.

543
MCQeasy

Based on the exhibit, which Amazon EFS performance mode is the best fit for this workload?

A.Use General Purpose performance mode for low-latency access.
B.Use Max I/O performance mode to optimize for the highest possible latency tolerance.
C.Use One Zone storage class to increase metadata speed.
D.Use Provisioned Throughput mode because it is the only performance mode available.
AnswerA

General Purpose is the best EFS performance mode when the priority is low latency for small file operations. The exhibit describes a moderate number of clients and latency-sensitive metadata access, which matches the strengths of General Purpose. It is the usual choice for most applications unless the workload specifically needs very large-scale parallel throughput.

Why this answer

The General Purpose performance mode is the best fit for this workload because it provides the lowest latency for file operations, which is critical for latency-sensitive applications such as web serving, content management, and development environments. EFS General Purpose mode is optimized for workloads where consistent low-latency access is required, making it the default and recommended choice for most use cases.

Exam trap

The trap here is that candidates confuse performance modes (General Purpose vs. Max I/O) with throughput modes (Bursting vs. Provisioned) or storage classes (Standard vs.

One Zone), leading them to select options that address throughput or availability rather than latency requirements.

Why the other options are wrong

B

Max I/O performance mode is designed for high throughput and can handle high I/O, but it does not optimize for latency tolerance; it actually has higher latency variability compared to General Purpose mode, making it unsuitable for a workload requiring low-latency access.

C

One Zone storage class is a storage class, not a performance mode; it does not affect metadata speed. Metadata performance is determined by the performance mode (General Purpose or Max I/O), not the storage class.

D

Provisioned Throughput mode is not a performance mode; it is a throughput setting available within General Purpose or Max I/O performance modes. The question asks for a performance mode, and Provisioned Throughput is not one of the two performance modes (General Purpose and Max I/O).

When would these options actually be correct?

B

A question describing a workload with high throughput requirements, such as big data processing or media transcoding, where latency is less critical and the application can tolerate higher variability, would make Max I/O the correct choice.

C

A question asks which EFS storage class to use for a workload that can tolerate data loss in the event of an Availability Zone failure and requires lower storage costs. In that scenario, One Zone storage class would be correct.

D

This option would be correct if the question asked: 'Which throughput mode should be used for a workload that requires a consistent, high throughput regardless of the amount of data stored?' In that scenario, Provisioned Throughput mode is the correct choice.

Why candidates pick the wrong answer

B

Candidates may mistakenly think 'Max I/O' implies better performance across all metrics, including latency, or they may confuse it with Provisioned Throughput mode, assuming it offers more control over performance.

C

Candidates may confuse storage classes with performance modes, or incorrectly believe that One Zone storage class improves metadata performance due to its local nature.

D

Candidates may confuse 'performance mode' with 'throughput mode' because both terms involve optimizing file system performance, leading them to select Provisioned Throughput as a performance mode.

544
MCQhard

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The architecture review board prefers a managed AWS-native control.

A.Use CloudFront signed URLs
B.Use Amazon SQS standard queue and design consumers to be idempotent
C.Use UDP messages sent directly to workers
D.Use an in-memory queue on one EC2 instance
AnswerB

An SQS standard queue is designed for decoupling components and provides at-least-once delivery with high throughput, ensuring that every message placed in the queue is eventually delivered to a consumer. Because at-least-once delivery can produce duplicate messages, consumers must be built to process each event idempotently so that repeated handling does not corrupt state or create duplicate outcomes. This combination satisfies the requirement that every event is processed.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, which guarantees that every message is processed at least once, meeting the requirement that every event must be processed. Duplicate processing is acceptable because the consumer can be designed to handle idempotency. This is a managed, AWS-native service that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates may confuse 'at-least-once' delivery with 'exactly-once' delivery and incorrectly choose a solution like FIFO queues (not listed) or dismiss SQS standard queues due to the duplicate processing allowance, but the question explicitly states duplicates are acceptable if idempotency is handled, making SQS standard the correct choice.

How to eliminate wrong answers

Option A is wrong because CloudFront signed URLs are used for securing content delivery, not for event processing or messaging. Option C is wrong because UDP is a connectionless, unreliable protocol that does not guarantee delivery, so it cannot ensure at-least-once processing. Option D is wrong because an in-memory queue on a single EC2 instance is not managed, not AWS-native, and introduces a single point of failure, violating the requirement for a resilient, managed service.

545
Multi-Selectmedium

An order-processing worker consumes messages from Amazon SQS. Occasionally, the worker times out after successfully creating a payment record but before deleting the message, which causes duplicate charges during retries. Some messages also fail validation repeatedly because required fields are missing. Which two changes should the team make? Select two.

Select 2 answers
A.Make the payment step idempotent using a unique transaction identifier.
B.Configure an SQS dead-letter queue with a redrive policy.
C.Reduce the visibility timeout so failed messages return to the queue faster.
D.Run only one long-lived worker instance so the queue can never be processed twice.
E.Switch from a standard queue to a FIFO queue and remove all other changes.
AnswersA, B

Correct. SQS provides at-least-once delivery, so the same message can be processed more than once if the worker times out, retries, or crashes after partially completing the work. An idempotency key lets the application recognize that the payment was already created and prevents duplicate charges.

Why this answer

Making the payment step idempotent using a unique transaction identifier ensures that if the same message is processed multiple times due to a timeout, the payment is only charged once. This is a common pattern for handling at-least-once delivery semantics in Amazon SQS, where the worker must be designed to handle duplicate messages safely.

Exam trap

The trap here is that candidates often think reducing the visibility timeout will speed up recovery, but it actually increases the chance of duplicate processing, and they may also overlook that a FIFO queue alone does not fix the worker's failure to delete the message after processing.

546
MCQmedium

A analytics dashboard uses an Application Load Balancer in one Region. Global users need lower network latency to the application without caching dynamic responses. What should be considered?

A.AWS Global Accelerator
B.S3 Cross-Region Replication
C.AWS Backup cross-Region copy
D.CloudFront only with long TTLs
AnswerA

AWS Global Accelerator is correct because it uses static anycast IP addresses at AWS edge locations to route inbound TCP/UDP traffic over the high-quality AWS global network, delivering it to the Application Load Balancer without relying on content caching. This directly reduces latency, jitter, and packet loss for dynamic API requests from global users, making it ideal for an analytics dashboard that repeatedly fetches live data.

Why this answer

AWS Global Accelerator uses the AWS global network and Anycast IPs to route traffic to the optimal Regional endpoint, reducing latency for global users without caching dynamic responses. It does not cache content, so dynamic data is always fetched from the origin, meeting the requirement of no caching while improving network performance via the AWS backbone.

Exam trap

The trap here is that candidates often choose CloudFront for any global latency improvement, but the requirement of 'no caching dynamic responses' disqualifies CloudFront unless TTL=0 is used, which still incurs edge request overhead, whereas Global Accelerator is purpose-built for non-cached dynamic traffic.

How to eliminate wrong answers

Option B (S3 Cross-Region Replication) is wrong because it replicates static objects across S3 buckets in different Regions, which does not reduce latency for dynamic application responses served by an ALB. Option C (AWS Backup cross-Region copy) is wrong because it is a backup and disaster recovery feature for copying backup data across Regions, not a mechanism to lower network latency for live application traffic. Option D (CloudFront only with long TTLs) is wrong because CloudFront caches content at edge locations, and using long TTLs would serve stale cached responses, violating the requirement of no caching for dynamic responses.

547
MCQhard

A DynamoDB table for a retail API has a partition key based only on the current date. Write throttling occurs during business hours. What is the best design change? The architecture review board prefers a managed AWS-native control.

A.Use a higher-cardinality partition key that distributes writes across partitions
B.Create a global secondary index with the same date key
C.Reduce the table's write capacity
D.Move the table to S3 Glacier Instant Retrieval
AnswerA

With only three possible partition key values, all writes for a given value are routed to the same partition, which is hard-capped at 1,000 write capacity units per second. A higher-cardinality partition key, such as a composite key combining customer ID with a timestamp, spreads write traffic across many partitions and uses the table's provisioned throughput more effectively, preventing hot-partition throttling.

Why this answer

Using only the current date as a partition key creates a hot partition because all writes for the day target a single partition, leading to throttling. A higher-cardinality partition key, such as a composite key combining date with a unique attribute like user ID or order ID, distributes writes evenly across multiple partitions, fully utilizing DynamoDB's provisioned throughput. This is the best managed-native solution to resolve write throttling without changing the table's capacity or moving data.

Exam trap

The trap here is that candidates often think adding a GSI or adjusting capacity solves throttling, but the root cause is the partition key's low cardinality, which only a higher-cardinality key can fix by distributing writes across partitions.

How to eliminate wrong answers

Option B is wrong because a global secondary index (GSI) with the same date key does not solve the hot partition issue; GSIs have their own throughput and inherit the same write distribution problem, potentially causing throttling on the index. Option C is wrong because reducing the table's write capacity would worsen throttling during business hours, not resolve the underlying hot partition caused by the poor key design. Option D is wrong because S3 Glacier Instant Retrieval is an object storage class for infrequently accessed data with millisecond retrieval, not a replacement for DynamoDB's low-latency, high-throughput key-value access, and moving the table would break the API's real-time requirements.

548
MCQhard

A SaaS provider runs a multi-tenant application on Amazon RDS for PostgreSQL. The database is 2 TB and experiences steady read-heavy traffic during business hours. The provider wants to offload read traffic to reduce load on the primary instance and lower cost compared to scaling up the primary. The application can tolerate slightly stale reads for reporting queries. Which solution is MOST cost-effective?

A.Enable Multi-AZ deployment and route reporting queries to the standby instance.
B.Migrate the database to Amazon DynamoDB with on-demand capacity mode.
C.Create a read replica and route reporting queries to it.
D.Increase the size of the primary RDS instance to a larger instance class.
AnswerC

A read replica is a separate RDS instance that receives asynchronous replication from the primary. It offloads read-only reporting queries, reducing load on the primary and allowing the primary to remain at its current size. Because the application tolerates slightly stale reads, asynchronous replication lag is acceptable. The read replica can be sized independently and stopped or scaled down when not needed, making it more cost-effective than scaling up the primary.

Why this answer

A read replica offloads read-only reporting queries from the primary instance, reducing contention and allowing the primary to remain appropriately sized. Because asynchronous replication introduces slight lag, the application's tolerance for stale reads makes this acceptable. This approach is more cost-effective than scaling up the primary or enabling Multi-AZ, which does not provide readable standby capacity.

Exam trap

The trap here is assuming that a Multi-AZ standby instance can be used to serve read traffic, when it is inaccessible for normal operations.

549
MCQhard

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The design must avoid adding custom operational scripts.

A.A FIFO queue without a redrive policy
B.A dead-letter queue with an appropriate maxReceiveCount
C.A larger message retention period only
D.Short polling instead of long polling
AnswerB

A dead-letter queue with a configured maxReceiveCount is the standard SQS mechanism for poison messages. When a message is received that many times without being deleted, SQS automatically moves it to a separate DLQ, isolating the problematic message from the main queue. This lets you inspect and debug the payload while healthy traffic continues unaffected. This also protects downstream consumers from continuous failure loops.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount allows messages that repeatedly fail processing to be moved out of the source queue after a specified number of receive attempts. This prevents poison messages from blocking retries and consuming processing resources, without requiring custom operational scripts.

Exam trap

The trap here is that candidates may think increasing retention or switching polling modes solves poison messages, but only a DLQ with maxReceiveCount directly addresses repeated failures without custom scripts.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; it still requires a DLQ configuration to move failing messages out. Option C is wrong because increasing the message retention period only keeps messages longer but does not prevent poison messages from repeatedly failing and blocking retries. Option D is wrong because short polling (immediate return with fewer messages) does not address poison message handling; it only affects message availability and latency, not failure management.

550
MCQeasy

A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.

A.IAM Access Analyzer
B.AWS Backup Vault Lock
C.CloudFront caching with appropriate TTLs
D.S3 Select
AnswerC

CloudFront caching with appropriate TTL values lets the CDN store static objects at edge locations and continue serving those cached copies even when the S3 origin is temporarily unreachable, as long as the cached object has not expired. Setting longer TTLs for immutable content (e.g., versioned images, scripts, and stylesheets) reduces the frequency of origin fetches and widens the window of resilience during an S3 outage. This is the only proposed solution that keeps content available to end users during an origin failure, directly addressing the requirement to tolerate an S3 outage.

Why this answer

CloudFront caches responses from the S3 origin based on configured TTLs (Cache-Control or Expires headers). If the S3 origin experiences a short outage, CloudFront can still serve cached content to users as long as the TTL has not expired, ensuring availability without custom scripts. This is the most direct and resilient feature for this use case.

Exam trap

The trap here is that candidates may confuse backup or access control features (like Backup Vault Lock or IAM Access Analyzer) with availability mechanisms, or think S3 Select provides caching, when the correct answer is simply leveraging CloudFront's built-in caching TTLs to serve stale content during origin outages.

How to eliminate wrong answers

Option A is wrong because IAM Access Analyzer helps identify unintended access to resources, not caching or origin resilience. Option B is wrong because AWS Backup Vault Lock prevents deletion of backups, not caching or serving stale content during origin outages. Option D is wrong because S3 Select is a feature to retrieve subsets of data from objects using SQL queries, not related to caching or origin resilience.

551
MCQmedium

A marketing site stores logs in S3. Logs are queried for 30 days, rarely accessed for one year, and then retained for compliance. What should reduce storage cost?

A.S3 lifecycle policy that transitions objects to lower-cost storage classes over time
B.Keep all logs in S3 Standard indefinitely
C.Use EBS snapshots for the logs
D.Move all logs immediately to S3 Glacier Deep Archive
AnswerA

An S3 lifecycle policy automates object transitions based on age, aligning storage cost with actual access patterns. Since logs are queried only for 30 days, rules can move older logs to S3 Standard-IA after 30 days, then to S3 Glacier Instant Retrieval or S3 Glacier Flexible Retrieval, and eventually to S3 Glacier Deep Archive for archival. This approach minimizes storage cost while keeping recent logs immediately accessible, and it requires no manual intervention. Lifecycle rules are evaluated daily per object, making them the ideal cost optimization mechanism for time-decaying log data.

Why this answer

An S3 Lifecycle policy automates the transition of objects from S3 Standard (frequently accessed) to lower-cost storage classes like S3 Standard-IA (infrequent access) after 30 days, then to S3 Glacier Deep Archive for long-term compliance retention. This matches the access pattern: frequent queries for 30 days, rare access for a year, then archival storage, minimizing cost without manual intervention.

Exam trap

The trap here is that candidates might choose immediate archiving (Option D) to minimize storage cost, overlooking the 30-day query requirement and the retrieval latency/cost of Glacier Deep Archive, or mistakenly think EBS snapshots (Option C) are a valid alternative for log storage.

How to eliminate wrong answers

Option B is wrong because keeping all logs in S3 Standard indefinitely incurs the highest per-GB storage cost, ignoring the significant cost savings from transitioning to lower-cost tiers for rarely accessed and archived data. Option C is wrong because EBS snapshots are block-level backups for EC2 volumes, not designed for object storage of logs; using them would require an EC2 instance to manage the logs, adding compute and management overhead. Option D is wrong because immediately moving all logs to S3 Glacier Deep Archive would incur retrieval costs and delays (hours) for the 30-day query period, violating the requirement for frequent queries during that time.

552
MCQmedium

A marketing team runs a report-generation process that must execute once per day at 02:00 UTC. It usually completes in 10315 minutes, but sometimes takes up to 45 minutes due to varying data volumes. They currently run the workload on an EC2 instance that is always on, which wastes money during off-hours. The team wants to minimize operational overhead and pay mainly for actual execution time. What is the best architecture choice?

A.Use a scheduled Amazon EC2 Auto Scaling group that keeps a minimum of one instance running at all times.
B.Use an EventBridge schedule to run the report as an Amazon ECS task on AWS Fargate and write results to S3.
C.Use AWS Lambda triggered by an EventBridge schedule at 02:00 UTC and write results to S3.
D.Use an EMR cluster provisioned daily with manual teardown to ensure the instance is always available before 02:00.
AnswerB

Fargate runs the ECS task only when EventBridge triggers the schedule, so billing covers actual task runtime rather than idle hours. It removes server patching and capacity management, and the 45-minute maximum fits comfortably within a single scheduled task run.

Why this answer

Amazon ECS on AWS Fargate is the best choice because it eliminates the need to manage servers, scales automatically, and charges only for the vCPU and memory resources consumed during task execution. The EventBridge schedule triggers the Fargate task at 02:00 UTC, and the report is written to S3, which provides durable, cost-effective storage. This architecture minimizes operational overhead and cost by avoiding an always-on EC2 instance.

Exam trap

The trap here is that candidates may choose AWS Lambda without considering its 15-minute execution timeout, which cannot handle the 45-minute maximum runtime of this report-generation process.

Why the other options are wrong

A

This option keeps an instance running at all times, which wastes money during off-hours and does not minimize operational overhead or pay mainly for actual execution time.

C

AWS Lambda has a maximum execution timeout of 15 minutes, but the report-generation process can take up to 45 minutes, so Lambda cannot handle the entire workload.

D

Provisioning an EMR cluster daily with manual teardown introduces significant operational overhead, which contradicts the requirement to minimize operational overhead. Additionally, EMR is designed for big data processing (e.g., Spark, Hive) and is overkill for a simple report-generation task, leading to higher costs and complexity.

When would these options actually be correct?

A

A scheduled Auto Scaling group with a minimum of one instance is correct when the workload requires a persistent server (e.g., for real-time processing or stateful applications) and cost optimization is not the primary concern.

C

A question where the task completes in under 15 minutes, requires no containerization, and benefits from Lambda's serverless, pay-per-execution model with minimal operational overhead.

D

This option would be correct if the workload involved large-scale data processing (e.g., running Spark or Hive jobs on terabytes of data) that requires a distributed cluster, and the team already has operational processes in place to manage cluster lifecycle automatically (e.g., using AWS Step Functions or Lambda for teardown). The question would emphasize cost savings by running only when needed, but not prioritize minimizing overhead.

Why candidates pick the wrong answer

A

Candidates may think Auto Scaling is cost-effective, but the 'minimum one instance' requirement contradicts the goal of paying only for execution time.

C

Candidates see 'EventBridge schedule' and 'serverless' and assume Lambda is the simplest choice, overlooking the 15-minute timeout limit.

D

Candidates may think EMR is suitable for any 'report generation' task, especially if they associate it with data processing. The manual teardown might seem like a simple way to save costs, but they overlook the operational burden and the fact that simpler services (like ECS Fargate) can handle the workload more efficiently.

553
MCQeasy

A startup runs a nightly batch job on a single EC2 instance that reads a large dataset from Amazon S3, performs transformations, and writes results back to S3. The job takes about two hours, and the team wants the job to restart automatically if the instance fails or is terminated by AWS. The job is idempotent and can safely resume from the beginning. What is the MOST operationally efficient way to meet this requirement?

A.Use EC2 Auto Recovery to recover the instance onto new hardware and attach an Amazon EBS volume that persists the job's progress.
B.Place the instance in an Auto Scaling group with a minimum and desired capacity of one, use a launch template that installs and starts the job at boot, and configure a lifecycle hook to keep the instance in service until the job completes.
C.Create an Amazon CloudWatch alarm on the StatusCheckFailed metric that invokes an AWS Lambda function to call the EC2 RebootInstances API.
D.Convert the job into an AWS Lambda function with a 15-minute timeout and trigger it on a nightly Amazon EventBridge schedule.
AnswerB

An Auto Scaling group with a desired capacity of one continuously maintains a single healthy instance, so if AWS terminates or the instance fails, the group replaces it and the launch template's user data reruns the idempotent job. A lifecycle hook can delay termination long enough for a clean shutdown, and this is fully managed with no custom monitoring code.

Why this answer

A single-instance Auto Scaling group is the managed way to keep exactly one healthy instance running and to replace it automatically when it is terminated or fails. Supplying the job through a launch template's user data means every replacement instance starts the job again, and because the job is idempotent, restarting from the beginning is acceptable. No custom alarm or recovery scripting is required.

Exam trap

The trap here is reaching for CloudWatch alarms and Lambda automation when the built-in Auto Scaling group replacement behavior already satisfies the restart requirement.

554
MCQhard

Based on the exhibit, downstream payment timeouts cause EventBridge deliveries to back up and some events are retried until they age out. What change best improves resilience and preserves events during downstream outages?

A.Increase the Lambda timeout so each invocation can wait longer for the payment API.
B.Put an Amazon SQS queue between EventBridge and the consumer, and have workers drain the queue with a DLQ for poison messages.
C.Switch the target to a Lambda function with reserved concurrency of zero during outages.
D.Replace EventBridge with CloudWatch Logs subscriptions so the consumer can poll the log stream later.
AnswerB

SQS is the right durability and buffering layer for this requirement. EventBridge can publish orders.checkout events to a queue, and workers can consume them at a controlled rate even when the payment API is unavailable. This decouples event ingestion from downstream processing, absorbs bursts, and preserves events until the outage ends. A DLQ provides a safe landing zone for messages that continue to fail after retries so they are not silently dropped.

Why this answer

Introducing an SQS queue between EventBridge and the consumer decouples the event delivery from the downstream payment API. During outages, events are stored durably in SQS and can be processed later without being lost. A Dead Letter Queue (DLQ) captures events that fail repeatedly, preventing poison messages from blocking the queue and ensuring no events age out due to retry exhaustion.

Exam trap

The trap here is that candidates often assume increasing timeouts or concurrency adjustments can fix backpressure issues, but they fail to recognize that decoupling with a durable queue is the only way to preserve events during extended downstream outages without losing them to retry expiration.

Why the other options are wrong

A

Increasing Lambda timeout does not address the root cause of downstream payment API timeouts; it only makes the Lambda wait longer, potentially exacerbating backpressure and event aging without improving resilience.

C

Setting reserved concurrency to zero during outages would stop all invocations, causing all events to be lost or retried until they age out, rather than improving resilience or preserving events.

D

CloudWatch Logs subscriptions deliver log data to a consumer in near-real-time but do not provide a durable buffer or retry mechanism for downstream failures; events would still be lost if the consumer is unavailable, and there is no built-in DLQ for poison messages.

When would these options actually be correct?

A

This option would be correct if the question were about a Lambda function that fails due to a slow but eventually successful downstream API, and the goal is to avoid premature timeouts without changing the architecture.

C

This option would be correct in a scenario where you need to temporarily throttle or stop processing from a specific event source to protect a downstream system from overload, while using a DLQ or retry mechanism to preserve events for later processing.

D

This option would be correct if the requirement was to archive all events for long-term storage and allow a consumer to process them on its own schedule, with no need for real-time delivery or automatic retries during outages.

Why candidates pick the wrong answer

A

Candidates may think that giving the Lambda more time will allow it to succeed eventually, overlooking that the downstream API is failing and that waiting longer does not prevent event loss or backpressure.

C

Candidates may think that reducing concurrency can prevent overload, but setting it to zero halts all processing, which contradicts the goal of preserving events during outages.

D

Candidates may think that using CloudWatch Logs provides a persistent log that can be replayed later, overlooking that EventBridge already offers retries and a DLQ-like mechanism, and that log subscriptions do not buffer events during consumer downtime.

555
MCQeasy

You need to run batch jobs on EC2. The jobs can tolerate interruptions: if an instance is terminated, the job can restart from checkpoints. To reduce compute cost as much as possible, what is the best choice?

A.EC2 On-Demand Instances to avoid interruptions
B.EC2 Spot Instances with checkpoint-based interruption handling
C.Savings Plans to guarantee capacity for the entire year
D.Reserved Instances with no interruption handling
AnswerB

Spot Instances are priced lower because AWS can reclaim capacity. When your workload can be interrupted and later restarted from checkpoints, the interruption model is compatible with Spot, making it the most cost-optimized option among the choices.

Why this answer

Spot Instances offer significant cost savings (up to 90% compared to On-Demand) but can be reclaimed by AWS with a two-minute warning. Since the batch jobs can tolerate interruptions and restart from checkpoints, Spot Instances are the most cost-effective choice. This aligns with the requirement to reduce compute cost as much as possible while handling interruptions gracefully.

Exam trap

The trap here is that candidates often choose On-Demand or Reserved Instances because they fear interruptions, but the question explicitly states the jobs can tolerate interruptions, so the most cost-effective option is Spot Instances, not a more expensive but stable alternative.

Why the other options are wrong

A

On-Demand Instances are more expensive than Spot Instances and do not offer cost savings. Since the job can tolerate interruptions via checkpoints, Spot Instances provide the lowest cost.

C

Savings Plans do not provide interruption handling or cost reduction for batch jobs that can tolerate interruptions; they offer discounted rates in exchange for a commitment but do not address the need for the lowest cost with interruption tolerance.

D

Reserved Instances require a 1- or 3-year commitment and do not inherently handle interruptions; they are designed for steady-state workloads, not fault-tolerant batch jobs where cost reduction is the priority.

When would these options actually be correct?

A

If the batch jobs are time-sensitive and cannot tolerate any interruptions, such as real-time data processing or critical financial calculations, On-Demand Instances would be the correct choice to ensure continuous availability.

C

A question where the workload requires consistent, long-term compute capacity (e.g., a production web server running 24/7) and the goal is to reduce costs with a usage commitment, not to handle interruptions.

D

For a long-term, steady-state workload (e.g., a web server running 24/7 for 3 years) where cost savings over On-Demand are needed and interruptions are unacceptable, Reserved Instances would be the best choice.

Why candidates pick the wrong answer

A

Candidates may choose On-Demand because they are familiar with its reliability and want to avoid the complexity of handling interruptions, overlooking that the question explicitly states the jobs can tolerate interruptions.

C

Candidates may think Savings Plans always reduce costs significantly, overlooking that Spot Instances offer deeper discounts and are better suited for fault-tolerant workloads like batch jobs.

D

Candidates may think Reserved Instances always reduce costs, but they overlook that the workload is batch and can tolerate interruptions, making Spot Instances cheaper and more appropriate.

556
MCQmedium

Account Y provides a role named AnalyticsReadOnly to engineers in Account X. The role trust policy currently allows sts:AssumeRole from the Account X principal. A new security requirement states that only STS sessions created with MFA are allowed to assume the role. Which trust policy condition is the best choice to enforce MFA for sts:AssumeRole?

A.Add a condition "Bool": { "aws:MultiFactorAuthPresent": "true" } in the role trust policy for the sts:AssumeRole action.
B.Add a condition "StringEquals": { "aws:username": "mfa-user" } in the IAM policy attached to the role.
C.Add a condition requiring "sts:ExternalId" to equal a fixed value in the trust policy.
D.Add a condition "Bool": { "aws:SecureTransport": "true" } in the trust policy to require HTTPS.
AnswerA

aws:MultiFactorAuthPresent is a condition key designed to reflect whether MFA was used when establishing the STS session. By requiring it to be true in the trust policy, STS denies AssumeRole when the caller did not authenticate with MFA.

Why this answer

The `aws:MultiFactorAuthPresent` condition key checks whether the principal used MFA to obtain the session credentials. By adding a `Bool` condition set to `"true"` in the trust policy for the `sts:AssumeRole` action, only STS sessions that were created after MFA authentication will be allowed to assume the role. This directly enforces the security requirement without affecting other authentication methods.

Exam trap

The trap here is that candidates may confuse `aws:MultiFactorAuthPresent` with other condition keys like `aws:SecureTransport` or `aws:username`, or think that `sts:ExternalId` can enforce MFA, when in fact only the `Bool` condition on `aws:MultiFactorAuthPresent` directly checks MFA status.

How to eliminate wrong answers

Option B is wrong because `aws:username` refers to the IAM user name, not the MFA status; requiring a specific username does not enforce MFA and can be bypassed if that user does not use MFA. Option C is wrong because `sts:ExternalId` is used to prevent the confused deputy problem in cross-account access, not to enforce MFA. Option D is wrong because `aws:SecureTransport` only ensures the session uses HTTPS/TLS, which is already required for AWS API calls, and does not verify MFA usage.

557
MCQhard

A company has a critical application running on Amazon EC2 instances that must access an Amazon RDS for MySQL database. The security team requires that the database credentials are never stored on the EC2 instances and that access to the database is auditable. The database is in a private subnet and only accepts connections from the application's security group. The company wants to implement a solution that automatically rotates the database password every 90 days. Which solution meets these requirements?

A.Store the database credentials in AWS Systems Manager Parameter Store as a SecureString parameter. Use a custom Lambda function to rotate the password every 90 days and update the parameter.
B.Store the database credentials in AWS Secrets Manager with automatic rotation enabled. Configure the EC2 instance role to allow secretsmanager:GetSecretValue, and have the application retrieve the secret at runtime.
C.Use IAM database authentication for Amazon RDS for MySQL. Configure the EC2 instance role with rds-db:connect permission, and have the application generate an authentication token.
D.Store the database credentials in an encrypted Amazon S3 bucket. Use an AWS Lambda function triggered by Amazon EventBridge to rotate the password every 90 days and update the S3 object.
AnswerB

AWS Secrets Manager stores the credentials securely and supports automatic rotation for Amazon RDS for MySQL. The EC2 instance role grants permission to retrieve the secret, so credentials are never stored on the instances. Access to the secret can be audited using AWS CloudTrail. This meets all requirements: no hardcoded credentials, automatic rotation, and auditability.

Why this answer

AWS Secrets Manager is the managed service designed for storing and rotating database credentials. It provides automatic rotation for Amazon RDS for MySQL, integrates with IAM for access control, and logs access via CloudTrail. The EC2 instances retrieve the secret at runtime, so credentials are never stored locally.

This meets the no-hardcoded-credentials, automatic rotation, and auditability requirements.

Exam trap

The trap here is assuming that IAM database authentication can satisfy a requirement for rotating a database password, when in fact it eliminates passwords entirely and requires application code changes.

558
MCQmedium

A trading analytics system deploys 10 EC2 instances that exchange very frequent, low-latency messages over the network. The instances must be placed as close together as possible to minimize network hop count and inter-node jitter. Which deployment choice best matches this requirement?

A.Use a spread placement group to distribute instances across multiple underlying hardware to improve overall availability.
B.Use a cluster placement group so the instances are placed close together to reduce latency and jitter.
C.Use no placement group and rely on the Auto Scaling group to balance instance placement automatically.
D.Use a partition placement group so each instance is assigned to separate failure domains for low variance.
AnswerB

A cluster placement group is the optimal choice because it launches instances in a single Availability Zone on the same underlying hardware and low-latency network backbone. This physical proximity minimizes network hop count and inter-instance communication latency, while the consistent network path reduces jitter—variation in packet delivery times—which is critical for a jitter-sensitive trading analytics system. The group is designed specifically for tightly coupled, high-performance workloads that require high-throughput, low-latency inter-node traffic, and when combined with enhanced networking, it can achieve microsecond-level latency and high bandwidth. Note that this comes at the cost of fault isolation, as a single hardware failure can affect all instances, but for this performance-driven use case, that trade-off is acceptable.

Why this answer

A cluster placement group is designed for low-latency, high-throughput scenarios by placing all instances in a single Availability Zone within the same rack or logical cluster, minimizing network hop count and inter-node jitter. This directly meets the requirement for very frequent, low-latency messaging between 10 EC2 instances.

Exam trap

The trap here is that candidates often confuse 'low latency' with 'high availability' and choose a spread placement group (Option A) thinking it reduces jitter, when in fact it increases network distance and latency by distributing instances across hardware.

How to eliminate wrong answers

Option A is wrong because a spread placement group distributes instances across distinct underlying hardware to maximize availability, which increases network distance and latency, opposite to the requirement. Option C is wrong because relying on an Auto Scaling group without a placement group does not guarantee close physical proximity; instances may be placed across different racks or AZs, increasing jitter. Option D is wrong because a partition placement group isolates instances into separate failure domains (partitions) to reduce correlated failures, but this increases network hops between partitions, not minimizing latency.

559
MCQmedium

A analytics dashboard uses an Application Load Balancer in one Region. Global users need lower network latency to the application without caching dynamic responses. What should be considered? The design must avoid adding custom operational scripts.

A.AWS Global Accelerator
B.S3 Cross-Region Replication
C.AWS Backup cross-Region copy
D.CloudFront only with long TTLs
AnswerA

Global Accelerator routes user traffic over the AWS global network to the nearest edge location, reducing latency for dynamic content without caching. It satisfies the no-custom-scripts constraint because it is a managed service configured declaratively, unlike Route 53 latency routing with custom health-check logic.

Why this answer

AWS Global Accelerator uses the AWS global network to route traffic from edge locations to the optimal regional endpoint, reducing latency and jitter for global users. It does not cache content, making it ideal for dynamic responses that cannot be cached. The service requires no custom scripts, as it integrates directly with the Application Load Balancer via a static IP address or DNS name.

Exam trap

The trap here is that candidates often confuse Global Accelerator with CloudFront, assuming both are for caching, but Global Accelerator does not cache content and is specifically designed for non-cacheable, dynamic traffic requiring low latency and fast failover.

How to eliminate wrong answers

Option B (S3 Cross-Region Replication) is wrong because it replicates objects across S3 buckets in different regions, but it does not reduce network latency for dynamic application traffic; it is designed for data redundancy and disaster recovery, not for real-time request routing. Option C (AWS Backup cross-Region copy) is wrong because it copies backup data across regions for compliance or disaster recovery, and it has no impact on live application latency or traffic routing. Option D (CloudFront only with long TTLs) is wrong because CloudFront caches content at edge locations, which violates the requirement to avoid caching dynamic responses; long TTLs would serve stale data, and disabling caching would negate the latency benefit, while custom scripts would be needed to bypass caching for dynamic content.

560
MCQhard

A company runs a stateful web application on a single Amazon EC2 instance in a public subnet. The application stores session data on the instance's root volume. The company wants to make the application highly available across two Availability Zones and ensure that session data is preserved if an instance fails. Which solution should a solutions architect recommend?

A.Create an Amazon Machine Image (AMI) of the instance, launch two new instances in two Availability Zones, and configure an Application Load Balancer with sticky sessions.
B.Move the session data to an Amazon ElastiCache for Redis cluster with Multi-AZ enabled, and configure an Auto Scaling group across two Availability Zones behind an Application Load Balancer.
C.Attach an Amazon Elastic Block Store (EBS) volume to the instance and create a snapshot schedule. Launch a second instance in another Availability Zone and attach the same EBS volume.
D.Enable termination protection on the instance and create a scheduled AWS Lambda function to reboot the instance if it becomes unhealthy.
AnswerB

Storing session data in ElastiCache for Redis with Multi-AZ enabled externalizes the state, allowing any instance in the Auto Scaling group to access it. The Auto Scaling group across two Availability Zones ensures high availability, and the Application Load Balancer distributes traffic. If an instance fails, a new one launches and retrieves session data from ElastiCache.

Why this answer

The recommended solution is to externalize session data to an ElastiCache for Redis cluster with Multi-AZ enabled and run the application in an Auto Scaling group across two Availability Zones behind an Application Load Balancer. This design ensures that session data survives instance failures and the application remains available if one AZ fails. The load balancer distributes traffic to healthy instances.

Exam trap

The trap here is assuming that sticky sessions or EBS snapshots can preserve session state across instance failures, when session data stored on an instance's root volume is lost if the instance fails.

561
MCQhard

Based on the exhibit, what is the best change to improve read performance without increasing write latency on the primary database?

A.Create an RDS read replica and direct the reporting queries to the replica endpoint.
B.Convert the DB instance to Multi-AZ so the primary can serve more reads.
C.Increase the primary instance class to a larger size and keep all traffic on one writer.
D.Migrate the reporting workload to DynamoDB to gain faster reads.
AnswerA

A read replica offloads the long-running read-only reports from the primary database, which preserves write performance and reduces read latency for the reporting workload. Because the business accepts slightly stale report data, the asynchronous replication delay is acceptable. This is the most direct and AWS-native way to separate read pressure from writes.

Why this answer

Creating an RDS read replica offloads read-heavy reporting queries from the primary database instance, improving read performance without increasing write latency on the primary. The replica operates asynchronously, so writes on the primary are not blocked or delayed by the reporting workload.

Exam trap

The trap here is confusing Multi-AZ (which only provides failover redundancy) with read replicas (which provide read scaling), leading candidates to incorrectly select Multi-AZ as a performance solution.

Why the other options are wrong

B

Multi-AZ is designed for high availability and failover, not for read scaling. It does not serve reads from the standby; the standby is only used for failover, so it does not improve read performance.

C

Increasing the primary instance class improves overall performance but does not specifically improve read performance without increasing write latency; it also increases cost and does not offload read traffic from the primary.

D

Migrating to DynamoDB would require re-architecting the application and data model, which is not a simple change to improve read performance on an existing RDS primary database. It also does not address the goal of not increasing write latency on the primary, as DynamoDB is a different database service.

When would these options actually be correct?

B

If the question asked for the best change to improve database availability and minimize downtime during a planned maintenance or failure, Multi-AZ would be correct. For example: 'What change provides automatic failover in case of an AZ outage?'

C

If the question asked for the best way to improve overall database performance for both reads and writes without adding a read replica or changing architecture, scaling up the instance class would be correct.

D

This option would be correct if the question asked for a fully managed NoSQL solution to handle high-scale, low-latency reads for a new application with flexible schema requirements, and the existing relational database is not suitable for the workload.

Why candidates pick the wrong answer

B

Candidates may mistakenly think Multi-AZ provides read scaling because it involves a second instance, confusing it with read replicas.

C

Candidates may think that a larger instance handles more reads inherently, overlooking that read replicas offload reads without affecting write latency.

D

Candidates may think DynamoDB offers faster reads due to its managed nature and low-latency performance, but they overlook the significant migration effort and the fact that the question specifically targets improving read performance on the existing primary database without increasing write latency.

562
MCQmedium

A healthcare company runs a critical patient-records API on Amazon EC2 instances behind an Application Load Balancer in a single AWS Region. The compliance team mandates that the API remain available even if an entire AWS Region becomes unavailable. The company wants a cost-effective solution that avoids running full production capacity in a second Region at all times. Which approach BEST meets these requirements?

A.Deploy the API in a second Region with a minimal warm standby environment, and use Amazon Route 53 failover routing with health checks to shift traffic when the primary Region fails.
B.Create an Amazon CloudFront distribution with the existing ALB as the origin, and enable origin failover to a second ALB in another Region.
C.Take regular Amazon EBS snapshots of the EC2 instances and copy them to a second Region, then restore the instances manually if the primary Region fails.
D.Enable Multi-AZ deployment for the EC2 instances and the Application Load Balancer, and rely on the existing single-Region architecture to survive a Regional failure.
AnswerA

A warm standby keeps a scaled-down but functional stack in the secondary Region, which can be scaled up during a failover. Route 53 failover routing with health checks automatically redirects traffic when the primary endpoint becomes unhealthy. This balances cost and recovery time, meeting the cross-Region availability requirement without paying for full duplicate capacity.

Why this answer

A warm standby in a second Region with Route 53 failover routing provides cross-Region resilience while controlling cost. The standby runs at reduced capacity and can be scaled up during a Regional failure. Health checks trigger automatic DNS failover, so clients reach the surviving Region without manual intervention.

This design directly addresses the need to survive a full Region outage without duplicating full production capacity at all times.

Exam trap

The trap here is assuming that Multi-AZ deployment provides protection against a complete AWS Region failure, when it only protects against Availability Zone failures within a single Region.

563
MCQmedium

A company runs a stateful analytics workload on EC2 instances that use EBS volumes. The data must be restorable in another Region after a major outage, with frequent point-in-time recovery. Which approach provides the most suitable replication mechanism for the EBS-backed data?

A.Create scheduled EBS snapshots and copy them to another Region, then restore the volumes from those snapshots during recovery.
B.Enable EBS multi-attach to spread the workload across AZs and replicate snapshots automatically between Regions.
C.Use RDS read replicas in another Region and keep the analytics dataset in an RDS instance only.
D.Rely on instance store for durability and copy only AMIs across Regions.
AnswerA

Scheduled EBS snapshots copied cross-Region provide point-in-time recovery and durability in a second Region, satisfying the restore-elsewhere requirement. Snapshots capture only changed blocks after the first, keeping replication efficient, and restoration creates fresh volumes from any snapshot in the target Region.

Why this answer

Scheduled EBS snapshots provide point-in-time backups of EBS volumes, which can be copied to another Region using the cross-Region snapshot copy feature. During recovery, you restore volumes from those snapshots in the target Region, ensuring the data is restorable after a major outage. This approach meets the requirements for frequent point-in-time recovery and cross-Region durability.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which is for high availability within a single AZ) with cross-Region replication, or mistakenly think instance store provides durability for long-term data recovery.

Why the other options are wrong

B

EBS multi-attach allows attaching a volume to multiple EC2 instances in the same AZ, but it does not replicate snapshots across Regions or provide cross-Region disaster recovery. It is designed for clustered applications within a single AZ, not for multi-Region replication.

C

RDS read replicas are for relational databases, not for analytics workloads on EC2 with EBS volumes. The question specifies EBS-backed data, not RDS-managed data, so using RDS would require migrating the dataset and does not replicate EBS snapshots.

D

Instance store volumes are ephemeral and lose data on instance stop/termination, making them unsuitable for durable, restorable data. Copying AMIs does not replicate the analytics data itself.

When would these options actually be correct?

B

An exam question requiring high availability for a clustered application (e.g., a shared file system) within a single AZ, where multiple EC2 instances need concurrent read/write access to the same EBS volume, and the question asks for the feature that enables this.

C

A company runs a web application using Amazon RDS for MySQL and needs to offload read traffic to a secondary Region for low-latency queries. RDS read replicas in another Region would be the correct answer for read scaling and cross-Region disaster recovery.

D

For a stateless workload where only the AMI (OS and application) needs to be available in another Region for disaster recovery, and the data is stored externally (e.g., in S3 or a database).

Why candidates pick the wrong answer

B

Candidates may confuse 'multi-attach' with cross-Region replication, or think that spreading across AZs implies automatic cross-Region backup, not realizing multi-attach is limited to one AZ and does not handle snapshots.

C

Candidates may confuse cross-Region replication capabilities of RDS with the need to replicate EBS data, assuming RDS can handle any analytics workload, or they may overlook that the question explicitly mentions EBS volumes and EC2 instances.

D

Candidates may confuse instance store with EBS, or think that AMI copying provides data replication, overlooking the ephemeral nature of instance store.

564
MCQmedium

A company hosts a web application on EC2 instances behind an Application Load Balancer (ALB) in us-east-1. A static failover site is hosted in an S3 bucket with static website hosting enabled. The company needs automatic DNS failover to the S3 bucket if the primary ALB becomes unhealthy. Which Route 53 configuration achieves this?

A.Configure Route 53 Failover routing with a health check on the ALB as PRIMARY and the S3 bucket website endpoint as SECONDARY
B.Configure Route 53 Weighted routing with 100% weight on the ALB and 0% on the S3 bucket
C.Configure Route 53 Latency routing with records in both regions to route to the healthiest endpoint
D.Configure Route 53 Geolocation routing with North American users directed to the ALB and all others to S3
AnswerA

Failover routing is the correct Route 53 policy for active-passive architecture, where the ALB is the primary endpoint and the S3 static website endpoint is the secondary. You configure a primary record pointing to the ALB, associate it with a Route 53 health check that monitors the ALB, and create a secondary record pointing to the S3 website endpoint. When the health check fails, Route 53 automatically returns the S3 endpoint, providing DNS-level failover. The S3 bucket must be configured for static website hosting with a publicly accessible endpoint, and this setup requires no manual intervention to switch traffic.

Why this answer

Route 53 Failover routing uses health checks to route traffic to a primary resource and automatically switch to a secondary when the primary health check fails.

Configuration: Create a Route 53 health check targeting the ALB endpoint. Create a PRIMARY alias A record pointing to the ALB with the health check associated. Create a SECONDARY alias A record pointing to the S3 static website endpoint. When the ALB health check fails, Route 53 returns the S3 endpoint automatically.

Exam trap

Route 53 offers multiple routing policies. Failover routing is active-passive — one primary resource, one standby. Weighted routing splits traffic percentages (active-active).

Latency routing picks the lowest-latency endpoint. Geolocation routes by user geography. Only Failover routing provides automatic primary/secondary switchover based on health checks.

Weighted routing at 100%/0% does NOT failover when the 100% target fails.

Why the other options are wrong

B

Weighted routing at 100%/0% does not failover. When the 100% target (ALB) is unhealthy, Route 53 does not automatically redirect to the 0% target (S3). Weighted routing splits traffic by percentage without health-check-based switching.

C

Latency routing routes to the lowest-latency endpoint for each client. It does not implement primary/secondary logic. If us-east-1 is unhealthy, some clients may still be routed there unless combined with health checks (but even then this is not a defined primary/secondary failover).

D

Geolocation routing directs traffic by user geography — North American users always go to the ALB even when it fails. S3 only receives other-region traffic. This is not a failover configuration.

565
MCQhard

A DynamoDB table for a travel booking site has a partition key based only on the current date. Write throttling occurs during business hours. What is the best design change? The architecture review board prefers a managed AWS-native control.

A.Create a global secondary index with the same date key
B.Move the table to S3 Glacier Instant Retrieval
C.Reduce the table's write capacity
D.Use a higher-cardinality partition key that distributes writes across partitions
AnswerD

A date-only partition key funnels every write into a single partition, hitting its 1,000 WCU ceiling regardless of table capacity. Switching to a high-cardinality key such as booking ID or customer ID spreads writes across many partitions, eliminating the hot partition and the throttling, while remaining fully AWS-native.

Why this answer

Using a low-cardinality partition key like the current date causes all writes to land on a single partition, leading to throttling. By choosing a higher-cardinality partition key (e.g., combining date with a user ID or booking ID), writes are distributed evenly across multiple partitions, leveraging DynamoDB's internal partitioning to handle the throughput. This is a managed, AWS-native design change that resolves hot partition issues without additional services.

Exam trap

The trap here is that candidates often confuse a GSI as a solution for write performance, when in fact GSIs only help with read query patterns and do not alleviate write hot spots on the base table.

How to eliminate wrong answers

Option A is wrong because creating a global secondary index (GSI) with the same date key does not solve the write throttling; GSIs have their own write capacity and inherit the same hot partition problem from the base table's partition key. Option B is wrong because moving the table to S3 Glacier Instant Retrieval is not a managed AWS-native control for DynamoDB write throttling; S3 is a different storage service and cannot replace DynamoDB's real-time transactional write capabilities. Option C is wrong because reducing the table's write capacity would worsen throttling during business hours, as it lowers the maximum allowed writes per second, directly contradicting the need to handle high write demand.

566
MCQeasy

A company stores several petabytes of archived regulatory records in Amazon S3. The records must be retained for seven years and are almost never accessed, but if an auditor requests a record, it must be retrievable within 12 hours. The company wants the lowest storage cost that still meets the retrieval requirement. Which S3 storage class should the solutions architect choose?

A.S3 Glacier Deep Archive
B.S3 Standard-Infrequent Access
C.S3 Glacier Flexible Retrieval
D.S3 One Zone-Infrequent Access
AnswerA

S3 Glacier Deep Archive is the lowest-cost S3 storage class and is designed for long-term retention of data accessed rarely, with a standard retrieval time of 12 hours and a bulk retrieval option of 48 hours. Because the auditor requirement allows up to 12 hours, Deep Archive satisfies the retrieval window while minimizing storage cost for petabytes held over seven years.

Why this answer

The lowest-cost storage class that still returns data within the allowed 12-hour window is S3 Glacier Deep Archive, whose standard retrieval completes within 12 hours. Glacier Flexible Retrieval and the Infrequent Access classes cost more per gigabyte, and paying for faster retrieval or millisecond access provides no benefit when the retrieval requirement is so permissive.

Exam trap

The trap here is choosing a faster retrieval class out of caution, when the stated 12-hour window is exactly what Glacier Deep Archive's standard retrieval is designed to satisfy.

567
MCQmedium

A company hosts a public website on Amazon EC2 instances behind an Application Load Balancer. The site is static HTML, CSS, and images, and the same content is served to all visitors. Origin CPU is high because every request is forwarded to the instances, and visitors in remote Regions see slow page loads. The team wants to reduce origin load and improve global latency with the least operational effort. Which solution should be used?

A.Create an Amazon CloudFront distribution with the ALB as the origin and configure cache behaviors for the static paths
B.Enable AWS Global Accelerator with the ALB as the endpoint
C.Configure the ALB with cross-zone load balancing and enable sticky sessions
D.Move the static assets to an Amazon S3 bucket and serve them directly with S3 website hosting
AnswerA

CloudFront caches the static HTML, CSS, and images at edge locations, so repeat requests are served from the edge instead of hitting the ALB and EC2 instances. This lowers origin CPU because far fewer requests reach the instances, and global visitors get content from a nearby edge. It requires only a distribution and cache behavior configuration, matching the low-effort requirement.

Why this answer

For a static site served identically to all visitors, edge caching removes the majority of requests from the origin. CloudFront caches the static paths at edge locations, reducing EC2 CPU and improving latency for remote visitors, and it can be added in front of the existing ALB with only configuration changes, satisfying the low-effort requirement.

Exam trap

The trap here is choosing Global Accelerator for latency, when it optimizes the network path but never caches content, leaving origin load unchanged.

568
MCQmedium

A financial analytics company runs a nightly batch job that reads 4 TB of compressed log data from an Amazon S3 bucket and writes aggregated results to another S3 bucket. The job runs on a fleet of 8 Amazon EC2 instances in a single AWS Region, and the team wants the highest possible aggregate read throughput while minimizing request costs. The data is already stored in S3 Standard, and the team does not want to change the storage class. Which solution best meets these requirements?

A.Enable S3 Cross-Region Replication to a second bucket and have half the instances read from the replica.
B.Use S3 Transfer Acceleration on the source bucket and have the EC2 instances read through the accelerated endpoint.
C.Mount the source bucket as a file system on each EC2 instance using an S3 File Gateway and read the objects through the mounted share.
D.Have the EC2 instances issue ranged GET requests against many distinct prefixes in parallel to spread the load across S3 partitions.
AnswerD

S3 automatically partitions a bucket by key prefix, and each partition supports a limited request rate. By issuing ranged GETs and spreading reads across many prefixes, the fleet drives parallel requests to multiple partitions, raising aggregate throughput. This is the standard approach for maximizing read performance from EC2 in the same Region without changing storage class or paying for acceleration.

Why this answer

S3 scales read throughput by partitioning data across multiple internal partitions, each with its own request rate limit. Spreading parallel ranged GET requests across many prefixes lets the fleet hit multiple partitions simultaneously, maximizing aggregate throughput. Transfer Acceleration, S3 File Gateway, and Cross-Region Replication all add cost or complexity without addressing the per-prefix request limit for in-Region batch reads.

Exam trap

The trap here is assuming that a single S3 bucket can only serve a fixed amount of read throughput regardless of how requests are distributed across key prefixes.

569
Multi-Selectmedium

A company needs to give an external auditing firm read-only access to specific objects in an Amazon S3 bucket for 30 days. The firm has its own AWS account and should not receive long-term credentials. The company wants to minimize the blast radius if the firm's account is compromised. Which two steps should the company take? (Choose two.)

Select 2 answers
A.Create an IAM user in the company account for the auditing firm and email the access key and secret key to the firm's security contact.
B.Create an IAM role in the company account with a trust policy that allows the auditing firm's AWS account to assume it, and attach a policy granting s3:GetObject on the specific object prefix.
C.Attach an S3 bucket policy that grants s3:GetObject to the auditing firm's IAM user ARN, and disable S3 Block Public Access to allow the cross-account policy to work.
D.Make the S3 bucket public and share the object URLs with the auditing firm so no AWS credentials are needed.
E.Require the auditing firm to call sts:AssumeRole and use the resulting temporary credentials, and set a maximum session duration and an external ID in the trust policy.
AnswersB, E

Cross-account access without long-term credentials is achieved with an IAM role that the external account assumes. Scoping the role policy to s3:GetObject on a specific prefix limits the blast radius if the external account is compromised. This is the standard AWS pattern for temporary cross-account access and meets the read-only requirement.

Why this answer

Cross-account temporary access is best implemented with an IAM role in the resource account that the external account assumes. Scoping the role policy to the specific prefix and requiring sts:AssumeRole with a maximum session duration and an external ID limits what a compromised external account can do and prevents the confused deputy problem. Long-term IAM users and public buckets fail the security goals.

Exam trap

The trap here is thinking that cross-account S3 access requires making the bucket public or disabling Block Public Access, when role assumption with a scoped policy is the intended mechanism.

570
MCQhard

A company runs a critical application on Amazon EC2 instances in a single Availability Zone. The application writes data to an Amazon RDS for MySQL DB instance that is not Multi-AZ. The company wants to improve the resilience of the database tier so that it can survive an Availability Zone failure with minimal downtime and no data loss. The application uses the database endpoint from the RDS console. Which solution meets these requirements?

A.Create a read replica in another Availability Zone and manually promote it if the primary fails.
B.Convert the DB instance to a Multi-AZ deployment and update the application to use the same endpoint.
C.Enable automated backups with a retention period of 35 days and copy the backups to another Region.
D.Migrate the database to Amazon DynamoDB with global tables and update the application to use the DynamoDB endpoint.
AnswerB

Converting to Multi-AZ creates a synchronous standby in a different AZ. RDS automatically fails over to the standby in the event of an AZ failure, and the application can continue using the same endpoint. This provides minimal downtime and no data loss because replication is synchronous.

Why this answer

Multi-AZ deployments for Amazon RDS provide a synchronous standby replica in a different Availability Zone. In the event of an AZ failure, RDS automatically fails over to the standby, and the application continues to use the same endpoint. Because replication is synchronous, there is no data loss.

This meets the requirements of minimal downtime and no data loss without application changes.

Exam trap

The trap here is confusing read replicas with Multi-AZ standbys. Read replicas are for offloading reads and can be promoted manually, but they use asynchronous replication and do not provide automatic failover.

571
Multi-Selecthard

A logistics company runs an Amazon RDS for MySQL database that supports a shipment tracking application. Read replicas are already in use, but the primary instance's CPU is saturated by a small number of long-running analytical queries that the reporting team runs directly against the primary. The company wants to offload these analytical queries and improve primary performance while keeping the application's transactional writes fast. (Choose two.)

Select 2 answers
A.Enable Multi-AZ on the primary instance so the standby can serve the analytical queries.
B.Increase the primary instance size and add more Provisioned IOPS to the storage volume to absorb the analytical queries.
C.Move the analytical workload to Amazon Redshift and load data from RDS using AWS Database Migration Service or zero-ETL integration.
D.Take frequent manual snapshots of the primary during reporting runs so the queries read from the snapshot.
E.Create a read replica dedicated to analytics and point the reporting queries to that replica's endpoint.
AnswersC, E

Redshift is a columnar data warehouse built for large analytical scans, so moving reporting there removes that load from the primary entirely. Loading via AWS DMS or RDS zero-ETL integration keeps the warehouse refreshed, and the primary no longer spends CPU on long analytical queries, which improves transactional write performance.

Why this answer

The goal is workload isolation. A dedicated read replica for reporting moves analytical reads off the primary without changing the application, while a purpose-built warehouse such as Amazon Redshift removes heavy scans entirely and keeps the warehouse current through DMS or zero-ETL. Together they protect transactional write latency.

Exam trap

The trap here is assuming a Multi-AZ standby can serve read traffic, when it is only a failover target.

572
MCQeasy

Your team hosts versioned static assets (for example, /static/app-<buildHash>.js). Each build hash never changes, but you release new files on new URLs. To maximize cache hit rate and reduce origin load using CloudFront, what should you do when generating HTTP responses for these assets?

A.Set Cache-Control: no-cache so CloudFront always revalidates with the origin
B.Set Cache-Control: public, max-age=31536000, immutable for the versioned assets
C.Set Cache-Control: max-age=0 and rely on CloudFront to cache by default
D.Disable CloudFront caching and forward all headers and query strings to the origin
AnswerB

For content-addressed/versioned URLs, a long max-age lets CloudFront treat the object as fresh for a long period. Adding the immutable directive tells clients not to revalidate while the max-age is still valid, supporting high cache hit rates and fewer origin fetches for repeat requests.

Why this answer

Setting `Cache-Control: public, max-age=31536000, immutable` tells CloudFront and browsers to cache the versioned asset for one year (31536000 seconds) and never revalidate, since the URL changes with each new build. The `immutable` directive (RFC 8246) signals that the content will never change on that URL, eliminating conditional revalidation requests and maximizing cache hits, which reduces origin load.

Exam trap

The trap here is that candidates confuse 'no-cache' (which still allows caching but forces revalidation) with 'no-store' (which forbids caching entirely), or they assume that `max-age=0` is acceptable for versioned assets, not realizing it forces revalidation and reduces cache efficiency.

Why the other options are wrong

A

Setting Cache-Control: no-cache forces CloudFront to revalidate with the origin on every request, defeating the purpose of caching versioned static assets that never change. This increases origin load and reduces cache hit rate.

C

Setting max-age=0 forces CloudFront to revalidate with the origin on every request, defeating the purpose of caching immutable versioned assets and increasing origin load.

D

Disabling CloudFront caching and forwarding all headers and query strings defeats the purpose of using a CDN for static assets, increasing origin load and reducing performance. Versioned assets with immutable hashes should be cached aggressively.

When would these options actually be correct?

A

This option would be correct for dynamic content that must always be fresh, such as real-time user-specific data or API responses where you want CloudFront to validate with the origin on each request to ensure the latest version is served.

C

For dynamic content that changes frequently and must always be fresh (e.g., real-time stock prices), setting max-age=0 ensures CloudFront revalidates with the origin on each request, providing the latest data.

D

This option would be correct if the question involved dynamic content that must be personalized per user (e.g., based on cookies or authorization headers) and cannot be cached at the edge, requiring every request to go to the origin.

Why candidates pick the wrong answer

A

Candidates may think 'no-cache' is a safe default to avoid serving stale content, not realizing that for immutable versioned assets, it unnecessarily adds revalidation overhead and misses the opportunity to cache aggressively.

C

Candidates may mistakenly believe that CloudFront automatically caches responses with max-age=0, or they confuse 'no-cache' behavior with 'max-age=0', thinking it still allows caching.

D

Candidates may think that forwarding all headers and query strings ensures freshness and correctness, not realizing that for versioned static assets, caching is safe and beneficial.

573
MCQhard

Based on the exhibit, which storage design best supports the application servers' shared working directory requirement?

A.Mount Amazon EFS on every EC2 instance and use it as the shared workspace.
B.Attach one gp3 EBS volume to each instance and synchronize the files with cron jobs.
C.Store the artifacts in S3 and have each node read them directly from S3 as a filesystem.
D.Use instance store on each instance because it provides the fastest local file access.
AnswerA

EFS provides shared, persistent, POSIX-compliant file access across multiple EC2 instances and Availability Zones. That matches the requirement that all nodes see the same workspace immediately and that files survive instance replacement. It is the right choice when the application needs a common filesystem rather than an object store or local-only disk.

Why this answer

Amazon EFS provides a fully managed, NFS-based shared file system that can be mounted concurrently on multiple EC2 instances across multiple Availability Zones. This directly satisfies the requirement for a shared working directory where all application servers can read and write files simultaneously without additional synchronization overhead.

Exam trap

The trap here is that candidates often confuse object storage (S3) with shared file storage, assuming S3 can serve as a drop-in replacement for a POSIX filesystem, but S3 lacks file locking, atomic renames, and low-latency metadata operations required for a shared working directory.

Why the other options are wrong

B

Synchronizing files with cron jobs introduces latency and inconsistency, failing to provide a real-time shared working directory as required by the application servers.

C

Using S3 as a filesystem (e.g., via s3fs) introduces latency and consistency issues that are unsuitable for a shared working directory requiring low-latency, POSIX-compliant file operations across multiple EC2 instances.

D

Instance store provides temporary, block-level storage that is physically attached to the host, but data is lost if the instance is stopped or terminated. It cannot serve as a persistent shared working directory across multiple EC2 instances.

When would these options actually be correct?

B

This option would be correct if the requirement was for each instance to have its own independent working directory with periodic backups or synchronization for disaster recovery, not a shared real-time workspace.

C

When the application servers need to read static artifacts (e.g., configuration files, binaries) that are rarely updated, and the design prioritizes cost savings over low-latency file operations, S3 with direct reads (e.g., via SDK) would be a correct choice.

D

For a scenario requiring the highest I/O performance for temporary, non-persistent data that is unique to each instance (e.g., scratch space for large-scale data processing), where data loss on instance stop is acceptable.

Why candidates pick the wrong answer

B

Candidates may think that using EBS volumes with cron-based sync is a cost-effective way to achieve file sharing, underestimating the complexity and inconsistency of distributed synchronization.

C

Candidates may think S3 is a universal storage solution and overlook its lack of POSIX semantics and higher latency compared to EFS, especially when the question mentions 'shared working directory' which implies frequent, concurrent file operations.

D

Candidates may assume that the fastest local storage (instance store) is always the best choice, overlooking the requirement for persistence and sharing across instances.

574
MCQeasy

A web application runs on an Amazon EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ALB is configured to use at least two Availability Zones (AZs), but the ASG currently uses subnets in only one AZ. If that AZ becomes unavailable, the application stops serving requests. Which change most directly improves resilience to an AZ outage?

A.Keep the ASG in one Availability Zone, but reduce ALB health check intervals.
B.Place the ASG across multiple Availability Zones by configuring it with subnets in at least two AZs.
C.Switch the load balancer from an ALB to an NLB to remove HTTP health check dependency.
D.Add an Amazon SQS queue to buffer requests during failures.
AnswerB

An ASG launches instances into the AZs of the subnets you specify. By placing the ASG in at least two AZs, the ALB can route traffic to healthy targets in the remaining AZ(s) if one AZ fails, enabling recovery as new instances maintain desired capacity.

Why this answer

Distributing an Auto Scaling group across multiple Availability Zones (AZs) ensures that if one AZ fails, the remaining AZs continue to serve traffic. The Application Load Balancer (ALB) is already configured for at least two AZs, but the ASG’s single-AZ subnet placement creates a single point of failure. By adding subnets in at least two AZs to the ASG, the application becomes resilient to an AZ outage without any other architectural changes.

Exam trap

The trap here is that candidates assume the ALB’s multi-AZ configuration automatically protects the application, overlooking that the ASG must also span multiple AZs to provide compute redundancy.

How to eliminate wrong answers

Option A is wrong because reducing health check intervals only detects failures faster but does not eliminate the single point of failure; if the sole AZ becomes unavailable, no healthy instances exist to serve traffic. Option C is wrong because switching from an ALB to an NLB does not address the root cause—the ASG is still in one AZ—and HTTP health checks are not the issue; the ALB can already perform health checks across AZs. Option D is wrong because adding an SQS queue buffers requests but does not provide compute capacity in another AZ; without instances in a second AZ, the queue cannot process requests during an AZ outage.

575
MCQhard

A healthcare analytics platform ingests records into an Amazon Aurora MySQL cluster. Compliance rules require that the cluster remain writable even if an entire Availability Zone is lost, and that recovery happen without operator action. The team also wants read traffic to scale independently of the writer. Which configuration should a solutions architect choose?

A.Create an Aurora MySQL cluster with one writer instance and two Aurora Replicas distributed across three Availability Zones, and connect the application to the cluster reader endpoint for read queries and the cluster endpoint for writes.
B.Create an Aurora MySQL cluster with one writer instance and enable Aurora Global Database with a secondary cluster in a different Region, directing all reads to the secondary Region's reader endpoint.
C.Create an Aurora MySQL cluster with one writer instance, and configure an Amazon Route 53 failover record that points to a standby cluster in the same Region when health checks fail.
D.Create an Aurora MySQL cluster with a single writer instance and one reader instance in the same Availability Zone, and enable automated backups with a 35-day retention period.
AnswerA

Aurora stores six copies of data across three Availability Zones, and distributing the writer plus two Aurora Replicas across those zones means a zone failure leaves surviving replicas available. Aurora automatically promotes a replica to writer and updates the cluster endpoint, so the application reconnects without manual steps, while the reader endpoint load-balances read queries across the replicas.

Why this answer

Aurora's storage layer replicates each data volume six ways across three Availability Zones, and placing the writer plus two Aurora Replicas in separate zones means a zone failure does not remove all compute endpoints. Aurora automatically promotes a surviving replica and repoints the cluster endpoint, so writes resume with no operator involvement, while the reader endpoint spreads read queries across the replicas for independent read scaling.

Exam trap

The trap here is treating automated backups or a Route 53 failover record as Availability Zone resilience, when both depend on a manual or external recovery path rather than Aurora's own automatic replica promotion.

576
MCQmedium

A media company distributes on-demand video to viewers in North America, Europe, and Asia. The videos are stored in a single Amazon S3 bucket in us-east-1 and are served directly from S3. Viewers in Asia report slow start times and frequent buffering. The company wants to reduce latency for all viewers with minimal operational overhead. Which solution meets these requirements?

A.Enable S3 Transfer Acceleration on the bucket and update the application to use the accelerated endpoint.
B.Configure an S3 Cross-Region Replication rule to copy objects into buckets in eu-west-1 and ap-southeast-1, and have the application redirect viewers to the nearest bucket.
C.Create an Amazon CloudFront distribution with the S3 bucket as the origin, and configure origin access control (OAC).
D.Move the bucket to ap-southeast-1 and use S3 Same-Region Replication to keep a secondary copy in us-east-1.
AnswerC

CloudFront caches video segments at edge locations close to viewers on every continent, so Asian viewers retrieve content from a nearby POP instead of crossing the Pacific to us-east-1. OAC lets the distribution access the private bucket securely without making objects public. This directly addresses global latency with minimal operational effort because AWS manages the edge network.

Why this answer

CloudFront is the managed content delivery service that caches objects at edge locations worldwide, which removes the long round trip to a single Region for viewers on other continents. Pairing it with origin access control keeps the bucket private while allowing only the distribution to fetch objects, so no objects need to be made public and no viewer-redirection logic is required.

Exam trap

The trap here is assuming that accelerating the path to a single origin Region is equivalent to caching content near viewers, when only a CDN places copies of the objects at the edge.

577
MCQmedium

A test environment runs on x86 EC2 instances and uses open-source software with no architecture-specific licensing restriction. What should be evaluated to reduce compute cost? The design must avoid adding custom operational scripts.

A.Cross-Region data replication for all data
B.AWS Graviton-based instances after performance testing
C.io2 Block Express volumes for all instances
D.Dedicated Hosts by default
AnswerB

AWS Graviton processors use Arm64 architecture, delivering better price-performance than comparable x86 instances. Since the software is open-source with no architecture-specific licensing restriction, and no custom operational scripts are added, migrating after performance testing satisfies the cost-reduction requirement without introducing operational overhead.

Why this answer

AWS Graviton-based instances (ARM architecture) offer up to 40% better price-performance compared to x86 instances for many workloads. Since the environment uses open-source software with no architecture-specific licensing restrictions, migrating to Graviton after performance testing can significantly reduce compute costs without requiring custom operational scripts, as AWS provides native support for ARM-based instances.

Exam trap

The trap here is that candidates may confuse cost optimization with performance improvement or licensing requirements, leading them to select Dedicated Hosts or high-performance storage options that actually increase costs.

How to eliminate wrong answers

Option A is wrong because cross-region data replication increases data transfer and storage costs, and it does not directly address compute cost reduction. Option C is wrong because io2 Block Express volumes are high-performance, high-cost EBS volumes designed for I/O-intensive workloads, not for reducing compute costs, and they would increase storage costs unnecessarily. Option D is wrong because Dedicated Hosts incur additional per-host charges and are used for licensing or compliance requirements, not for cost optimization; they would increase compute costs rather than reduce them.

578
MCQeasy

An order-processing system publishes an event whenever a payment succeeds. Three downstream services (inventory, shipping, and analytics) must react independently. Analytics sometimes has high latency, but order processing must not be blocked. What is the best AWS approach to decouple these consumers?

A.Have order processing call each service synchronously via HTTPS and retry on failures.
B.Publish payment events to SNS (or EventBridge) and let each downstream service consume independently (for example, via SQS queues or other async targets).
C.Store events in a single relational database table and let consumers poll continuously for new rows.
D.Send events directly from the producer to each consumer EC2 instance using SSH tunnels.
AnswerB

Using pub/sub decouples the producer from consumers. Order processing publishes once and can complete without waiting for each downstream service. Each consumer receives events independently, so analytics latency does not directly block inventory or shipping processing.

Why this answer

Amazon SNS (or EventBridge) enables asynchronous, fan-out messaging where a single payment-success event is published once and delivered independently to multiple downstream services (inventory, shipping, analytics) via SQS queues or other targets. This decouples the producer from consumer latency—analytics can take its time without blocking order processing—and ensures each consumer processes the event at its own pace, meeting the requirement for independent, non-blocking reactions.

Exam trap

The trap here is that candidates may choose synchronous integration (Option A) because it seems simpler, failing to recognize that the requirement 'must not be blocked' explicitly demands asynchronous decoupling, not just retries.

How to eliminate wrong answers

Option A is wrong because synchronous HTTPS calls with retries tightly couple the producer to all consumers; if analytics has high latency, order processing is blocked waiting for responses, violating the requirement that it must not be blocked. Option C is wrong because storing events in a single relational database table introduces a single point of failure, creates a polling bottleneck, and tightly couples consumers to a shared schema and table, which is not a decoupled, scalable architecture. Option D is wrong because sending events directly via SSH tunnels requires direct network connectivity to each EC2 instance, introduces security risks, and tightly couples the producer to consumer instances, making it brittle and unscalable.

579
MCQmedium

An internal worker consumes messages from an Amazon SQS Standard queue. Recently, some messages fail validation in the worker (for example, missing required fields), causing the worker to crash before it can successfully process those messages. Those messages keep getting retried repeatedly, slowing down processing of valid messages. The team wants a resilient mechanism to quarantine bad messages after a limited number of receive attempts. What should they implement?

A.Increase the SQS visibility timeout to several hours so the worker does not retry too quickly.
B.Configure a redrive policy with a Dead-Letter Queue (DLQ) and set maxReceiveCount so poison messages are moved to the DLQ after repeated failures.
C.Switch the queue to an SNS topic and subscribe the worker directly, eliminating message retries.
D.Enable KMS encryption with a new CMK to ensure validation errors stop occurring.
AnswerB

An SQS DLQ with a redrive policy is specifically designed for poison-message handling. When a message exceeds maxReceiveCount without successful processing (for example, the worker crashes before deletion), SQS moves the message to the DLQ. This quarantines bad messages and protects throughput for valid messages.

Why this answer

Amazon SQS supports configuring a redrive policy with a Dead-Letter Queue (DLQ) that automatically moves messages after a specified number of receive attempts (maxReceiveCount). This isolates poison messages that fail validation and cause crashes, preventing them from being retried indefinitely and slowing down valid message processing. The worker can then focus on valid messages while the DLQ stores the problematic ones for later analysis or manual intervention.

Exam trap

The trap here is that candidates may think increasing the visibility timeout (Option A) solves the retry problem, but it only delays retries without eliminating the root cause, while the DLQ mechanism (Option B) provides a proper quarantine by moving messages after a configurable number of receive attempts.

How to eliminate wrong answers

Option A is wrong because increasing the visibility timeout to several hours would only delay retries, not prevent them; the worker would still crash repeatedly on the same invalid messages after each timeout expires, and valid messages would be blocked for hours. Option C is wrong because switching to an SNS topic eliminates message retries entirely, but the worker would still crash on invalid messages without any retry mechanism or quarantine, and SNS does not provide a built-in DLQ for consumer-side failures. Option D is wrong because enabling KMS encryption with a new CMK addresses data encryption at rest and in transit, but has no effect on message content validation errors or crash handling; encryption does not fix missing required fields or prevent retries.

580
MCQmedium

A company is deploying a web application on AWS. The application runs on Amazon EC2 instances behind an Application Load Balancer. The company wants to protect the application from common web exploits such as SQL injection and cross-site scripting, and also wants to restrict access to specific geographic regions. Which combination of AWS services should a solutions architect use to meet these requirements?

A.AWS Network Firewall with stateful rule groups and domain lists.
B.AWS Shield Advanced with a web ACL and IP set rules.
C.AWS WAF with managed rule groups and geographic match conditions.
D.Amazon GuardDuty with threat intelligence and custom findings.
AnswerC

AWS WAF can be attached to an Application Load Balancer and provides managed rule groups that protect against common exploits like SQL injection and cross-site scripting. It also supports geographic match conditions to allow or block requests based on the country of origin. This directly addresses both requirements.

Why this answer

AWS WAF is the appropriate service for protecting web applications behind an Application Load Balancer from application-layer attacks. Its managed rule groups cover common threats like SQL injection and cross-site scripting, and geographic match conditions can allow or block traffic by country. The other services either operate at different layers or provide detection rather than inline prevention.

Exam trap

The trap here is confusing network-layer firewalls with application-layer web application firewalls; only AWS WAF inspects HTTP requests for exploits and supports geographic matching.

581
MCQhard

A logistics company stores shipment events in an Amazon S3 bucket. An analytics team must be able to recover any object version that is accidentally overwritten or deleted for at least 90 days, and objects must be protected from permanent deletion by any user, including the root user, during that window. Which combination of S3 features meets these requirements with the LEAST operational overhead?

A.Enable S3 Versioning and configure S3 Cross-Region Replication to a second bucket in another Region with a 90-day lifecycle expiration.
B.Enable S3 Versioning and configure an S3 Lifecycle rule to transition noncurrent versions to S3 Glacier Deep Archive after 90 days.
C.Enable S3 Versioning and apply an S3 Object Lock retention period of 90 days in compliance mode on the bucket.
D.Enable S3 Versioning and add a bucket policy that denies s3:DeleteObject for all principals for 90 days.
AnswerC

S3 Object Lock with a 90-day compliance-mode retention prevents any principal, including the root user, from deleting or overwriting the protected object versions until the retention expires. Versioning is required for Object Lock, and compliance mode provides the strongest immutability, directly satisfying both recovery and permanent-deletion-prevention requirements with minimal ongoing effort.

Why this answer

S3 Object Lock requires versioning and enforces a retention period during which object versions cannot be deleted or overwritten by anyone, including the root user. Compliance mode cannot be shortened or bypassed, so it uniquely satisfies the requirement to prevent permanent deletion for at least 90 days while still allowing recovery of prior versions.

Exam trap

The trap here is thinking that versioning plus a restrictive bucket policy or replication equals immutability, when only Object Lock retention actually blocks deletion by the root user.

582
MCQmedium

A web application for a claims portal is behind an Application Load Balancer. The application must be protected from common SQL injection and cross-site scripting attacks with minimum operational overhead. What should the architect deploy?

A.AWS WAF associated with the Application Load Balancer
B.AWS Shield Advanced only
C.Network ACLs on the public subnets
D.Security groups on the application instances
AnswerA

AWS WAF associated with the Application Load Balancer is the correct answer because it operates at Layer 7, inspecting HTTP/HTTPS requests, headers, body, and query strings for signatures of SQL injection and cross-site scripting (XSS). Managed rule groups such as the Core Rule Set (CRS) and OWASP Top 10 rules specifically detect and block these application-layer exploits. Attaching AWS WAF to an ALB enables centralized, content-aware filtering of all traffic destined to your application, which is exactly the capability required to protect a claims portal against these attack types.

Why this answer

AWS WAF is a web application firewall that integrates directly with an Application Load Balancer to filter and monitor HTTP/HTTPS requests. It provides managed rules specifically designed to block common attack patterns like SQL injection and cross-site scripting (XSS) with minimal operational overhead, as the rules are pre-configured and automatically updated by AWS.

Exam trap

The trap here is that candidates often confuse network-layer controls (like security groups or NACLs) with application-layer protection, assuming they can block SQL injection or XSS, but these operate at Layer 3/4 and cannot inspect HTTP request bodies or headers for malicious content.

How to eliminate wrong answers

Option B is wrong because AWS Shield Advanced provides DDoS protection, not application-layer attack filtering for SQL injection or XSS. Option C is wrong because Network ACLs operate at the subnet level (Layer 3/4) and cannot inspect application-layer payloads for SQL injection or XSS patterns. Option D is wrong because security groups act as stateful firewalls at the instance level (Layer 3/4) and cannot perform deep packet inspection for application-layer attacks.

583
MCQmedium

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The design must avoid adding custom operational scripts.

A.Subnets in at least two Availability Zones with health checks enabled
B.All instances in one larger subnet
C.A Network Load Balancer in one subnet
D.A single EC2 instance with detailed monitoring
AnswerA

An Auto Scaling group configured with subnets in at least two Availability Zones ensures that if an entire AZ becomes unavailable, the ASG can launch replacement instances in another AZ, preserving capacity. Health checks on the load balancer and EC2 allow the ASG to detect failed instances and replace them automatically. This architecture eliminates the single point of failure that comes with a single AZ and provides fault tolerance for both planned maintenance and unexpected outages.

Why this answer

Placing subnets in at least two Availability Zones ensures that if one AZ fails, the Auto Scaling group can launch instances in the remaining healthy AZ, maintaining application availability. Health checks integrated with the Application Load Balancer allow the Auto Scaling group to automatically replace unhealthy instances without custom scripts, aligning with the requirement to avoid operational overhead.

Exam trap

The trap here is that candidates often assume a single larger subnet or a Network Load Balancer provides AZ resilience, but they fail to recognize that without multiple subnets in distinct AZs, the architecture cannot survive an AZ failure, and custom scripts would be needed for health checks without ELB integration.

How to eliminate wrong answers

Option B is wrong because placing all instances in one larger subnet confines them to a single Availability Zone, violating the requirement to tolerate the failure of one AZ. Option C is wrong because a Network Load Balancer operates at Layer 4 and does not provide the health check integration needed for Auto Scaling group instance replacement; additionally, placing it in one subnet creates a single point of failure. Option D is wrong because a single EC2 instance, even with detailed monitoring, cannot survive an AZ failure and does not leverage Auto Scaling for automatic recovery.

584
MCQmedium

A company uses an Amazon Aurora DB cluster in a Multi-AZ configuration. During a planned failover of the writer instance, the database endpoints in the application are updated incorrectly. After failover, reads work but writes fail with connection errors and timeouts for several minutes. The team currently uses the instance endpoint for the writer. What should they change to improve write resilience during failovers?

A.Continue using the instance endpoint, but increase application retry count so the writer changes are handled more quickly.
B.Use the Aurora cluster writer endpoint for all write operations.
C.Use a read replica endpoint for writes because it is typically stable across failovers.
D.Disable Multi-AZ failover so the writer instance never changes and writes remain consistent.
AnswerB

Aurora provides a writer endpoint designed specifically for write traffic. During failover, Aurora updates where the writer endpoint points, so the same DNS name continues to resolve to the current writer instance without requiring manual endpoint changes in the application.

Why this answer

The Aurora cluster writer endpoint always points to the current primary (writer) instance, even after a failover. By using this endpoint instead of a static instance endpoint, the application automatically resolves to the new writer without manual updates, eliminating connection errors and timeouts during failover transitions.

Exam trap

The trap here is that candidates confuse the instance endpoint (which is static and tied to a specific instance) with the cluster endpoint (which is dynamic and always points to the current writer), assuming any endpoint will automatically follow failover.

How to eliminate wrong answers

Option A is wrong because increasing the retry count does not fix the root cause—the application is still pointing to the old (now read-only) instance endpoint, so writes will continue to fail until the endpoint is manually corrected. Option C is wrong because read replica endpoints point to read-only instances; writes to a read replica will always fail with an error, regardless of failover state. Option D is wrong because disabling Multi-AZ failover removes high availability entirely, making the database vulnerable to a single point of failure, which contradicts the goal of improving write resilience.

585
Multi-Selecthard

A claims workflow requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.

Select 2 answers
A.Point-in-time recovery
B.DAX
C.Deletion protection or tightly controlled delete permissions
D.Global secondary indexes
AnswersA, C

PITR allows restoration to a specific second within the supported recovery window.

Why this answer

Point-in-time recovery (PITR) for DynamoDB enables continuous backups with 35-day granularity, allowing restoration to any second within that window. This directly satisfies the point-in-time recovery requirement without custom scripts, as it is a native AWS feature.

Exam trap

The trap here is that candidates often confuse DAX with a data protection feature, but DAX only accelerates reads and has no role in backup or deletion prevention.

586
MCQmedium

A web application uses an Amazon Aurora DB cluster for a read-heavy workload. The application team needs higher read throughput but cannot change the database schema. They want to avoid blocking writes and are willing to route read traffic separately. What is the most appropriate architecture change?

A.Create Aurora read replicas and route SELECT queries to an Aurora reader endpoint.
B.Scale up the writer instance storage only; read capacity will automatically increase without using a reader endpoint.
C.Move the Aurora cluster to Multi-AZ deployment mode only; read scaling is handled automatically without replicas.
D.Replace the cluster with a single RDS instance because it offers consistent performance for both reads and writes.
AnswerA

Aurora Replicas are independent compute nodes attached to the same distributed storage volume as the writer, so they can serve SELECT traffic without competing for the writer's CPU or memory. The Aurora reader endpoint automatically load-balances connections across all healthy replicas, enabling the cluster to scale read throughput far beyond a single writer instance. Routing only SELECT queries to the reader endpoint preserves write consistency because all write operations continue through the writer endpoint, and replica lag is typically low enough for most use cases.

Why this answer

Creating Aurora read replicas and routing SELECT queries to the Aurora reader endpoint is the most appropriate architecture change because Aurora's reader endpoint distributes read traffic across up to 15 low-latency read replicas, providing higher aggregate read throughput without blocking writes. This approach requires no schema changes and allows the application to separate read and write traffic, directly addressing the read-heavy workload requirement.

Exam trap

The trap here is that candidates often confuse Multi-AZ deployment (which provides failover only) with read replica scaling, or mistakenly believe that scaling storage or using a single instance can improve read throughput without schema changes.

How to eliminate wrong answers

Option B is wrong because scaling up the writer instance storage does not increase read throughput; Aurora's read capacity is tied to compute resources (e.g., instance class) and the number of replicas, not storage size. Option C is wrong because Multi-AZ deployment in Aurora is for high availability and failover, not for scaling read throughput; read scaling requires dedicated read replicas with a reader endpoint. Option D is wrong because replacing the cluster with a single RDS instance would eliminate read scaling capabilities and introduce a single point of failure, making it unsuitable for a read-heavy workload that needs higher throughput without blocking writes.

587
MCQeasy

An S3 bucket stores application logs. After 30 days, the team rarely accesses the logs, but compliance requires keeping them for 18 months. Which setup most directly reduces storage cost while maintaining compliance?

A.Configure an S3 Lifecycle policy to transition objects to a colder storage class after 30 days and expire (delete) them after 18 months.
B.Enable S3 Versioning and rely on deleting old versions after 30 days to reduce storage costs while keeping the latest data.
C.Move the bucket to a different AWS region farther from the users to reduce the likelihood of accidental reads and thereby lower storage costs.
D.Switch all objects to S3 Glacier Instant Retrieval immediately, regardless of object age, to minimize storage charges.
AnswerA

This lifecycle policy is the right approach because it automatically transitions 30-day-old log objects to a lower-cost storage class — such as S3 Standard-IA or S3 Glacier Instant Retrieval — which directly matches the noted shift in access pattern after the first 30 days. The separate expiration action after 18 months enforces the compliance requirement to retain logs for exactly that period and then delete them, preventing both premature deletion and runaway storage growth. Lifecycle transitions are managed by S3, so the team does not need custom code to move or expire objects.

Why this answer

S3 Lifecycle policies allow you to automatically transition objects to cheaper storage classes (e.g., S3 Standard-IA or S3 Glacier Deep Archive) after 30 days, reducing storage costs for rarely accessed logs. The policy also sets an expiration action to delete objects after 18 months, meeting the compliance requirement without manual intervention.

Exam trap

The trap here is that candidates may think moving to a different region or using versioning reduces costs, but the core concept is that S3 Lifecycle policies directly automate cost optimization by transitioning to colder storage classes and expiring data, which is the most direct and compliant approach.

How to eliminate wrong answers

Option B is wrong because enabling S3 Versioning and deleting old versions does not address the need to keep logs for 18 months; it only manages versions, not the primary objects, and can increase costs due to storing multiple versions. Option C is wrong because moving the bucket to a different region does not reduce storage costs; it may increase data transfer costs and does not change the storage class or lifecycle management. Option D is wrong because switching all objects to S3 Glacier Instant Retrieval immediately, regardless of age, would likely increase costs for frequently accessed logs in the first 30 days, as this storage class has higher retrieval costs and is not optimal for data that is still being accessed.

588
Multi-Selectmedium

A company is running a production web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The workload has predictable traffic spikes during business hours and low traffic at night. The current architecture uses On-Demand EC2 instances, leading to high costs. The company wants to reduce costs without sacrificing availability or performance. Which three of the following strategies would help achieve this goal? (Choose three.)

Select 3 answers
.Purchase Reserved Instances for the baseline capacity that runs 24/7.
.Add Spot Instances for the entire workload during peak hours.
.Use Auto Scaling with a mixed instances policy that includes On-Demand and Spot Instances.
.Migrate to AWS Lambda for all web application traffic.
.Implement a scheduled scaling action to increase capacity before business hours and decrease after.
.Consolidate all instances into a single larger instance to reduce overhead.

Why this answer

Purchasing Reserved Instances for the baseline 24/7 capacity provides a significant discount (up to 72%) compared to On-Demand pricing, directly reducing costs for the always-running portion of the workload. This strategy is correct because it matches the predictable, steady-state traffic component without sacrificing availability or performance.

Exam trap

The trap here is that candidates may think Spot Instances can be used for the entire workload during peak hours, but they overlook the interruption risk and the requirement for the workload to be fault-tolerant, which a production web application behind an ALB typically is not without careful design.

589
MCQmedium

A web application for a IoT ingestion API is behind an Application Load Balancer. The application must be protected from common SQL injection and cross-site scripting attacks with minimum operational overhead. What should the architect deploy? The design must avoid adding custom operational scripts.

A.AWS WAF associated with the Application Load Balancer
B.Network ACLs on the public subnets
C.Security groups on the application instances
D.AWS Shield Advanced only
AnswerA

AWS WAF integrated with an Application Load Balancer operates at Layer 7 to inspect each HTTP/HTTPS request for malicious patterns, including SQL injection and cross-site scripting (XSS). It uses managed rule groups and custom rules to filter traffic before it reaches the application instances. WAF supports real-time visibility and rate-based controls, making it the only listed option that can inspect application payloads and block these attack types.

Why this answer

AWS WAF is a web application firewall that helps protect web applications from common web exploits like SQL injection and cross-site scripting (XSS). By associating an AWS WAF web ACL with the Application Load Balancer, you can filter and monitor HTTP(S) requests based on rules that match malicious patterns, with no custom scripts or operational overhead. This is the most efficient and managed solution for the stated requirements.

Exam trap

The trap here is that candidates often confuse network-layer controls (NACLs, security groups) with application-layer protection, or assume Shield Advanced alone covers all web threats, when in fact WAF is specifically designed for Layer 7 attack mitigation like SQLi and XSS.

How to eliminate wrong answers

Option B is wrong because Network ACLs are stateless packet filters that operate at the subnet level (Layer 3/4) and cannot inspect application-layer payloads like SQL or XSS patterns. Option C is wrong because security groups are stateful virtual firewalls that control traffic based on IP addresses, ports, and protocols (Layer 3/4), and they lack the ability to perform deep packet inspection for web application attacks. Option D is wrong because AWS Shield Advanced provides DDoS protection and enhanced detection, but it does not include the rule-based filtering needed to block SQL injection or XSS attacks.

590
MCQhard

A DynamoDB table for a retail API has a partition key based only on the current date. Write throttling occurs during business hours. What is the best design change?

A.Use a higher-cardinality partition key that distributes writes across partitions
B.Create a global secondary index with the same date key
C.Reduce the table's write capacity
D.Move the table to S3 Glacier Instant Retrieval
AnswerA

A date-only partition key concentrates every write into a single partition, hitting its throughput ceiling. A higher-cardinality key spreads items across many partitions, so write capacity is consumed in parallel and throttling during business-hour peaks is eliminated.

Why this answer

Using a partition key based solely on the current date creates a 'hot partition' because all writes for that day target the same partition, leading to throttling. A higher-cardinality partition key (e.g., combining date with a unique attribute like user ID or order ID) distributes write traffic evenly across multiple partitions, allowing DynamoDB to utilize its full throughput capacity and eliminating throttling.

Exam trap

The trap here is that candidates may think a GSI can solve write throttling, but GSIs only help with read patterns and do not redistribute write load on the base table.

How to eliminate wrong answers

Option B is wrong because creating a global secondary index (GSI) with the same date key does not change the base table's partition key; writes still target the same hot partition, so throttling persists. Option C is wrong because reducing the table's write capacity would worsen throttling, not solve it, as the issue is uneven distribution, not insufficient total capacity. Option D is wrong because S3 Glacier Instant Retrieval is an object storage service for archival data, not a transactional database; it cannot support DynamoDB's low-latency read/write operations or query patterns.

591
MCQmedium

A document portal requires consistent high IOPS for a transactional database on EC2. Which EBS volume type is most suitable? The architecture review board prefers a managed AWS-native control.

A.sc1 Cold HDD
B.Instance store only
C.Provisioned IOPS SSD such as io2
D.st1 Throughput Optimized HDD
AnswerC

Provisioned IOPS SSD volumes like io2 are purpose-built for business-critical workloads that require consistent, predictable high IOPS and low latency. They allow you to provision specific IOPS independent of capacity, and io2 offers 99.999% durability, making it the correct choice for a document portal where database performance is essential.

Why this answer

The scenario requires consistent high IOPS for a transactional database, which demands low-latency, predictable performance. Provisioned IOPS SSD (io2) is the only EBS volume type that allows you to specify a guaranteed IOPS rate independent of volume size, making it ideal for latency-sensitive transactional workloads. It is also a managed AWS-native service, satisfying the architecture review board's preference.

Exam trap

The trap here is that candidates may confuse 'high IOPS' with throughput-optimized HDDs (st1) or mistakenly think instance store offers managed persistence, but the key differentiator is the need for consistent, provisioned IOPS and managed durability that only io2 provides.

How to eliminate wrong answers

Option A is wrong because sc1 Cold HDD is designed for infrequently accessed, cold data with low cost, offering burstable throughput but very low IOPS, making it unsuitable for transactional databases requiring consistent high IOPS. Option B is wrong because instance store provides temporary, block-level storage that is physically attached to the host, but it is not managed (data is lost on instance stop/termination) and does not qualify as a managed AWS-native control. Option D is wrong because st1 Throughput Optimized HDD is optimized for large, sequential workloads like big data and log processing, not for random I/O patterns of transactional databases, and it cannot guarantee high IOPS.

592
MCQhard

A media processing workflow in private subnets downloads large amounts of data from S3 through a NAT gateway. NAT data processing charges are high. What should the architect use to reduce cost? The design must avoid adding custom operational scripts.

A.S3 Object Lambda
B.AWS Shield Advanced
C.Gateway VPC endpoint for Amazon S3
D.A larger NAT gateway
AnswerC

A gateway VPC endpoint for Amazon S3 is a free resource that you attach to a subnet's route table, enabling instances to reach S3 via private IP addresses. Traffic to S3 through a gateway endpoint does not traverse an internet gateway or NAT gateway, thereby eliminating the NAT data processing charge. It is the recommended and most cost-effective way to access S3 from private subnets when no internet connectivity is required.

Why this answer

A Gateway VPC endpoint for Amazon S3 allows instances in private subnets to access S3 directly over the AWS network without traversing a NAT gateway. This eliminates NAT data processing charges because traffic stays within the AWS backbone, reducing costs significantly for large data downloads.

Exam trap

The trap here is that candidates often confuse Gateway VPC endpoints with Interface VPC endpoints, assuming both incur costs, or mistakenly think NAT gateways are required for all private subnet outbound traffic, missing the S3-specific optimization.

How to eliminate wrong answers

Option A is wrong because S3 Object Lambda is used to transform data on the fly during retrieval, not to reduce data transfer costs or bypass NAT gateways. Option B is wrong because AWS Shield Advanced provides DDoS protection, not cost optimization for S3 data transfer. Option D is wrong because a larger NAT gateway would increase, not reduce, costs due to higher hourly and data processing charges.

593
MCQmedium

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The design must avoid adding custom operational scripts.

A.An EBS snapshot schedule
B.S3 Cross-Region Replication with versioning enabled
C.S3 lifecycle transition to Glacier Flexible Retrieval
D.A CloudFront distribution
AnswerB

Cross-Region Replication automatically copies objects to a destination bucket in another Region, and versioning is a prerequisite for it. This meets the disaster-recovery copy requirement without custom operational scripts, since replication is handled natively by S3.

Why this answer

S3 Cross-Region Replication (CRR) automatically replicates objects to a destination bucket in a different AWS Region, meeting the disaster recovery requirement without custom scripts. Versioning must be enabled on both source and destination buckets for CRR to function, as it tracks object versions and ensures consistency during replication.

Exam trap

The trap here is that candidates may confuse S3 Cross-Region Replication with S3 lifecycle policies or CloudFront, thinking they provide cross-region replication, but only CRR with versioning enabled meets the DR requirement without custom scripts.

How to eliminate wrong answers

Option A is wrong because EBS snapshots are for Amazon Elastic Block Store volumes attached to EC2 instances, not for S3 objects; they cannot replicate S3 data across regions. Option C is wrong because S3 lifecycle transitions to Glacier Flexible Retrieval only change storage class within the same region for cost optimization, not replicate data to another region. Option D is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations for low-latency access, not a replication mechanism for disaster recovery across regions.

594
MCQeasy

A stateless web application runs on Amazon EC2 instances across two Availability Zones. The team wants unhealthy instances to be removed automatically and replaced without manual action. What is the best solution?

A.Place the instances in a single subnet and increase the instance size.
B.Use an Application Load Balancer with an Auto Scaling group and configure health checks.
C.Use a network ACL to detect failed instances and restart them.
D.Store the web servers on EBS volumes so the data survives failures.
AnswerB

An Application Load Balancer distributes traffic across healthy targets, and an Auto Scaling group can replace instances that fail health checks. Together, they provide automatic recovery from instance failure and keep the application available across multiple Availability Zones. This is the standard resilient design for stateless EC2 web tiers.

Why this answer

An Application Load Balancer (ALB) with an Auto Scaling group provides automated health checks and instance replacement. The ALB performs HTTP/HTTPS health checks against the instances, and when an instance fails the health check, the Auto Scaling group automatically terminates the unhealthy instance and launches a new one to maintain the desired capacity. This ensures the stateless web application remains available across both Availability Zones without manual intervention.

Exam trap

The trap here is that candidates confuse network ACLs (stateless packet filters) with health check mechanisms, or assume that persistent storage (EBS) alone provides high availability without an orchestration layer like Auto Scaling.

Why the other options are wrong

A

Placing instances in a single subnet and increasing instance size does not provide automatic unhealthy instance replacement; it only increases capacity within one Availability Zone, offering no fault tolerance or self-healing.

C

Network ACLs are stateless packet filters at the subnet level; they cannot detect failed instances or trigger instance replacement. They do not perform health checks or automate recovery.

D

Storing web servers on EBS volumes does not automate instance replacement or health checks; it only preserves data across failures, but manual action is still required to replace failed instances.

When would these options actually be correct?

A

This option would be correct if the question asked for a solution to handle increased load for a single-instance application that does not require high availability, and the goal was to vertically scale the instance to improve performance.

C

A question asks for a security layer to block specific IP ranges from accessing a subnet, or to allow/deny traffic at the subnet boundary based on source/destination IP and port. In that context, a network ACL is the correct answer.

D

If the question required ensuring data persistence for stateful applications after instance failure, such as a database server where data must survive termination, then using EBS volumes with termination protection would be correct.

Why candidates pick the wrong answer

A

Candidates may think that a larger instance can handle failures better, or they confuse vertical scaling with the automatic recovery provided by Auto Scaling groups.

C

Candidates may confuse network ACLs (which filter traffic) with health check mechanisms, or mistakenly think ACLs can monitor instance health and trigger actions.

D

Candidates may think that preserving data on EBS volumes automatically handles failures, confusing data durability with instance recovery automation.

595
MCQmedium

A media archive requires consistent high IOPS for a transactional database on EC2. Which EBS volume type is most suitable? The team wants the control to be enforceable during normal operations.

A.Provisioned IOPS SSD such as io2
B.st1 Throughput Optimized HDD
C.Instance store only
D.sc1 Cold HDD
AnswerA

io2 is a Provisioned IOPS SSD EBS volume engineered for business-critical workloads that demand consistent, single-digit-millisecond latency and sustained IOPS. For a 4 TB media archive requiring durable, high-performance storage, io2 delivers 99.999% durability and allows you to provision up to 256,000 IOPS (with Block Express), ensuring predictable performance for transactional or indexing workloads rather than relying on burst credits.

Why this answer

The io2 Provisioned IOPS SSD volume type is designed for latency-sensitive transactional database workloads that require consistent high IOPS. It allows you to specify a guaranteed IOPS rate (up to 256,000 IOPS for io2 Block Express) and provides 99.999% durability, making it ideal for enforcing performance control during normal operations.

Exam trap

The trap here is that candidates often confuse throughput-optimized HDD (st1) with IOPS-focused SSD, assuming 'high throughput' implies high IOPS, but st1 is designed for sequential access and cannot provide the low-latency random I/O that transactional databases require.

How to eliminate wrong answers

Option B (st1 Throughput Optimized HDD) is wrong because it is a throughput-optimized HDD volume designed for large, sequential workloads like big data and log processing, not for transactional databases requiring consistent low-latency IOPS. Option C (Instance store only) is wrong because instance store volumes are ephemeral and data is lost on instance stop or termination, making them unsuitable for persistent database storage. Option D (sc1 Cold HDD) is wrong because it is a cold HDD volume optimized for infrequently accessed data with the lowest cost, offering very low IOPS and throughput that cannot meet the demands of a transactional database.

596
MCQmedium

An order-quote Lambda function is invoked directly by API Gateway. Traffic is predictable during the business day, and the first request after scaling from zero causes unacceptable latency. The team wants to keep the current architecture and reduce cold-start impact. Which configuration should they use?

A.Increase the function timeout so the first invocation has more time to finish.
B.Enable provisioned concurrency for the Lambda function.
C.Set reserved concurrency to a fixed number and leave the rest unchanged.
D.Increase the memory size only to eliminate cold starts.
AnswerB

Provisioned concurrency keeps a set number of Lambda execution environments initialized and ready to serve traffic. That directly reduces or removes cold starts for predictable workloads such as business-hours APIs. It is the most appropriate choice when the team wants to preserve serverless architecture while delivering consistent response times for the first request and subsequent requests.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, so when the first request arrives after scaling from zero, it is served by a pre-warmed instance instead of incurring a cold start. This directly addresses the unacceptable latency without changing the architecture or requiring code modifications.

Exam trap

The trap here is that candidates confuse reserved concurrency (which caps concurrent executions) with provisioned concurrency (which pre-warms instances), or mistakenly believe that increasing memory or timeout can eliminate the cold-start initialization delay.

Why the other options are wrong

A

Increasing the function timeout does not reduce cold-start latency; it only allows the function to run longer, but the initial cold-start delay remains.

C

Reserved concurrency limits the maximum concurrent executions for a function but does not pre-warm instances, so it does not reduce cold-start latency for the first request after scaling from zero.

D

Increasing memory size can reduce cold start duration but does not eliminate cold starts; the first request after scaling from zero still incurs cold start latency.

When would these options actually be correct?

A

In a scenario where a Lambda function consistently times out due to long processing times (e.g., processing large files) and the team needs to ensure completion without changing architecture, increasing the timeout would be correct.

C

A question where a function must not exceed a certain concurrency limit to avoid throttling downstream resources (e.g., a database with limited connections) would make reserved concurrency the correct answer.

D

A question where the goal is to reduce execution time of a Lambda function that is consistently hitting the maximum timeout, and the function is CPU-bound, so more memory (and thus more CPU) speeds up execution.

Why candidates pick the wrong answer

A

Candidates may think that giving the function more time compensates for the cold-start delay, misunderstanding that cold-start is about initialization time, not execution duration.

C

Candidates may confuse reserved concurrency with provisioned concurrency, thinking that reserving capacity eliminates cold starts, but reserved concurrency only caps concurrency, not pre-initializes environments.

D

Candidates may believe that more memory eliminates cold starts entirely, or they confuse memory allocation with keeping the function warm.

597
MCQmedium

An order processing workflow uses Amazon SQS as the decoupling layer between a producer and a consumer Lambda function. The consumer intermittently fails due to a downstream dependency. The team has observed that certain “poison” messages keep being retried repeatedly and prevent other messages from being processed efficiently. Which SQS configuration most directly addresses this issue?

A.Set the SQS queue’s retention period to 10 years and rely on application retries to eventually succeed.
B.Increase visibility timeout to a very large value and avoid dead-letter queues to keep ordering stable.
C.Configure a redrive policy with a dead-letter queue (DLQ) and set an appropriate visibility timeout greater than the maximum processing time.
D.Switch the queue to FIFO and remove retries in the Lambda event source mapping entirely.
AnswerC

A redrive policy defines a dead-letter queue (DLQ) and a maxReceiveCount; once a message is received that many times without being deleted, SQS moves it to the DLQ, quarantining poison messages for inspection or manual redrive. Setting the visibility timeout longer than the worst-case processing time prevents the message from becoming visible again while a consumer is still working, which would otherwise cause duplicate deliveries. Together, these settings bound both the retry window and the queue depth, allowing transient failures to retry while isolating permanent failures without losing data.

Why this answer

Configuring a redrive policy with a dead-letter queue (DLQ) allows messages that repeatedly fail processing to be moved out of the main queue after a specified number of receive attempts. Setting an appropriate visibility timeout greater than the maximum processing time ensures that messages are not made visible again before the consumer finishes processing, preventing premature retries. This directly isolates poison messages so they no longer block the processing of other messages in the queue.

Exam trap

The trap here is that candidates may think increasing visibility timeout or switching to FIFO alone will handle failed messages, but without a DLQ, poison messages remain in the queue and continue to block other messages, which is the core issue described.

Why the other options are wrong

A

Setting the retention period to 10 years does not address poison messages; it only keeps messages longer. Relying on application retries without a DLQ allows poison messages to be retried indefinitely, blocking other messages.

B

Increasing visibility timeout to a very large value does not prevent poison messages from blocking the queue; they will still be retried indefinitely, and without a DLQ, failed messages cannot be isolated for analysis or skipped.

D

Switching to FIFO and removing retries does not address poison messages; FIFO ensures strict ordering but does not prevent problematic messages from blocking the queue, and removing retries would cause immediate failures without handling the root cause.

When would these options actually be correct?

A

If the question asked for ensuring message durability for long-term processing where failures are transient and no poison messages exist, setting a long retention period would be correct.

B

In a scenario where message ordering is critical and you cannot tolerate any message loss or reordering, and you have a mechanism to ensure all messages eventually succeed (e.g., idempotent processing with exponential backoff), you might avoid DLQs and use a very large visibility timeout to prevent premature retries.

D

This option would be correct in a scenario where the requirement is to process messages in strict order and any failed message must be discarded immediately to avoid blocking subsequent messages, with no tolerance for retries or reordering.

Why candidates pick the wrong answer

A

Candidates may think increasing retention gives more time for retries to succeed, overlooking that poison messages never succeed and need isolation via a DLQ.

B

Candidates may think that a large visibility timeout gives more time for processing and avoids reordering, but they overlook that poison messages will still block the queue and cause repeated failures without a DLQ to divert them.

D

Candidates may think FIFO queues solve all ordering issues and that removing retries simplifies error handling, overlooking that poison messages still need isolation via DLQs and that retries are essential for transient failures.

598
MCQeasy

A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application must remain available if the instance fails or if its Availability Zone becomes unavailable. The startup wants a managed solution that requires minimal operational overhead. Which solution should a solutions architect recommend?

A.Launch a second EC2 instance in the same Availability Zone and configure an Elastic IP address for failover.
B.Create an Auto Scaling group spanning at least two Availability Zones and attach it to an Application Load Balancer.
C.Migrate the application to AWS Lambda and expose it through an API Gateway REST API.
D.Create an Auto Scaling group with a desired capacity of one instance and place it in a single Availability Zone.
AnswerB

An Auto Scaling group spanning multiple Availability Zones automatically replaces failed instances and can launch replacements in healthy zones when one zone fails. Attaching an Application Load Balancer distributes traffic across healthy instances and provides a single stable endpoint. This is a managed, low-overhead solution that meets both instance and zone failure requirements.

Why this answer

An Auto Scaling group that spans at least two Availability Zones can replace a failed instance and can also launch instances in a healthy zone when one zone becomes unavailable. Attaching an Application Load Balancer provides a stable endpoint and health-based routing, delivering a managed, resilient architecture with minimal operational effort.

Exam trap

The trap here is believing that a single-zone Auto Scaling group provides Availability Zone resilience, when it can only replace instances within the same zone.

599
MCQeasy

Your application runs in private subnets with no NAT gateway. It needs to call AWS Secrets Manager to retrieve secrets. For private connectivity without internet egress, which VPC endpoint type should you create for AWS Secrets Manager?

A.An Interface VPC endpoint (AWS PrivateLink) for secretsmanager in your Region
B.A Gateway VPC endpoint for secretsmanager
C.A NAT gateway in the private subnet route table
D.A VPC peering connection to the AWS public network hosting Secrets Manager
AnswerA

Secrets Manager supports Interface VPC endpoints powered by AWS PrivateLink. An interface endpoint attaches an elastic network interface (ENI) with a private IP address to your chosen subnets, allowing your instances to call the Secrets Manager API without traversing the internet, a NAT gateway, or an internet gateway. You can also enable private DNS so standard secretsmanager.region.amazonaws.com resolves to the private IP, and traffic stays entirely inside the VPC.

Why this answer

An Interface VPC endpoint (AWS PrivateLink) creates an elastic network interface in your subnet with a private IP address, allowing your instances to communicate with AWS Secrets Manager over the AWS network without traversing the internet. Since your application runs in private subnets with no NAT gateway, this is the only supported endpoint type for Secrets Manager, as Gateway endpoints are only available for S3 and DynamoDB.

Exam trap

The trap here is that candidates often confuse Gateway endpoints (which are free and only for S3/DynamoDB) with Interface endpoints (which incur hourly charges and support many services like Secrets Manager), leading them to incorrectly select option B.

How to eliminate wrong answers

Option B is wrong because Gateway VPC endpoints are only supported for Amazon S3 and DynamoDB, not for AWS Secrets Manager. Option C is wrong because a NAT gateway requires an internet gateway and public subnet, and the question explicitly states there is no NAT gateway and no internet egress allowed. Option D is wrong because VPC peering connects two VPCs within the AWS network, not to the AWS public network hosting Secrets Manager; Secrets Manager is accessed via service endpoints, not through peering.

600
MCQmedium

A dev sandbox has unpredictable DynamoDB traffic with long idle periods and occasional spikes. Which capacity mode should minimize operational overhead and avoid paying for idle provisioned capacity?

A.Reserved capacity for maximum daily traffic
B.Provisioned capacity set for peak traffic
C.DynamoDB on-demand capacity mode
D.Global tables in every Region
AnswerC

DynamoDB on-demand capacity mode charges you only for the actual read and write requests your application makes, with no minimum capacity and no need to forecast traffic patterns. It instantly accommodates spikes from unpredictable developer workloads, and when the table is idle, you pay nothing beyond storage costs. This directly solves the problem of paying for unused capacity during long idle periods, which is exactly why on-demand is the correct choice for this dev sandbox.

Why this answer

DynamoDB on-demand capacity mode (Option C) is ideal for unpredictable workloads with long idle periods and occasional spikes because it automatically scales to handle traffic without requiring any capacity planning. You pay only for the reads and writes you actually perform, eliminating the cost of idle provisioned capacity and the operational overhead of managing scaling.

Exam trap

The trap here is that candidates confuse 'reserved capacity' (an EC2/RDS concept) with DynamoDB capacity modes, or think that provisioned capacity with auto-scaling is always cheaper, ignoring the cost of idle capacity during long idle periods.

How to eliminate wrong answers

Option A is wrong because Reserved capacity is not a DynamoDB capacity mode; it is a pricing model for EC2 or RDS, and DynamoDB does not offer reserved capacity. Option B is wrong because Provisioned capacity set for peak traffic would incur costs for idle capacity during long idle periods, and you would still need to manually adjust or use auto-scaling to handle spikes, increasing operational overhead. Option D is wrong because Global tables are a replication feature for multi-Region active-active setups, not a capacity mode, and they do not address cost or overhead from idle capacity or traffic spikes.

Page 7

Page 8 of 13

Page 9